Files
msd-core/CHANGELOG.md
2026-08-19 13:51:02 +00:00

660 KiB
Raw Blame History

Changelog

All notable changes to GSD will be documented in this file.

Format follows Keep a Changelog.

Unreleased

[1.11.0] - 2026-08-19

Added

  • resolve-execution now tells the truth about what the agent will run at — the query reported only the config-cascade effort, which is not what an installed agent uses when its effort: frontmatter was hand-stripped or drifted. --json adds effort_effective (read from the installed agent frontmatter for the claude runtime; "inherit" when the key is absent) and effort_effective_source (frontmatter | frontmatter-absent | resolved). All existing fields, including --pick effort, are unchanged. (#3534) (#3542)

  • The install manifest now records which runtime and scope wrote it — a global and a project-local install used to write two gsd-file-manifest.json files that neither named their own runtime nor their own scope, so nothing could answer "which GSD surfaces are installed, where". The manifest gains manifestVersion, runtime and scope, and a new read-only Installed Surface Resolver reads both scopes at once. Manifests written by earlier versions are read without error and need no reinstall. (#2872) (#3323)

  • Opt-in .git/hooks/pre-commit guard for commit_docs — gsd-tools commit-docs-guard enable/disable writes (or removes) a pre-commit hook that shells out to the existing check-commit verb, refusing a commit that stages .planning/ files while commit_docs resolves to false. Closes the one bypass earlier phases of epic #2292 could not reach: a plain git add -A && git commit run by hand or by a script outside GSD's own tooling. Fully opt-in by maintainer narrowing — no install path wires it in by default (regression-locked by tests/commands.test.cjs's E2 row); enable refuses rather than overwrites an existing foreign pre-commit hook, refuses when core.hooksPath would make the written hook inert, and resolves the real hooks directory via git rev-parse --git-path hooks so a linked worktree or submodule (where .git is a file) is handled correctly rather than assuming a literal .git/hooks path. The hook is identified by a stable # gsd-core:commit-docs-guard marker line, checked by presence rather than byte-equality. (#3588) (#3609)

  • installRuntimeArtifacts() now returns the plan it executed — per kind, per scope, including on the combined OpenCode/Kilo family path that previously returned nothing — so an install's correctness is a value a caller can assert, not something only re-readable from disk afterward. Install IO routes through a new injectable fs seam (install-fs-adapter.cts), letting a full install run end-to-end against a fake adapter with no real destination filesystem contact; failures still throw rather than becoming a value, and a best-effort cleanup that fails is now visible in the return instead of silently swallowed. Writes on disk are unchanged. Completes ADR-58's never-landed cleanup rollout step. (#2874) (#3568)

  • Capability skills are now named at the install consent prompt — installing a third-party capability whose only contribution was skills printed "ships no executable surfaces (declarative only)" and listed nothing, even though each SKILL.md body lands verbatim in your agent's instruction context. The pre-install disclosure now names every contributed skill in its own section and states plainly that the bodies are not content-scanned. Values interpolated into the prompt are escaped across every disclosed surface, so a crafted name can no longer forge additional lines of disclosure text. No stored consent is disturbed and no re-consent prompt fires. (#3248) (#3253)

  • validate agents now reports Codex .toml model posture, not just presence — on a codex install it flags any agent whose .toml pins a GSD tier alias or a claude-* id (which Codex rejects with a 400, so the agent never spawns) or carries a model_reasoning_effort with no model. Previously the check confirmed only that agent files existed, so a stale install from before the passive-model posture reported healthy right up until a typed agent failed to start. Read-only — it names the offending agent and value and never edits your files. Reports not_codex and reads nothing on other runtimes. (#3242) (#3290)

  • effort sync now repairs stale Codex .toml files without a reinstall — on a codex install it strips a model pin that Codex rejects (a tier alias or a claude-* id) and an orphaned model_reasoning_effort, so agents fall back to the always-available session model. An explicit real-Codex pin is left alone. It is a dry run by default — pass --apply to write — and only the offending lines are removed: line endings, BOM, comments, key order, and any keys you added by hand are preserved byte-for-byte, so a repair is a two-line diff rather than a reformatted file. A file that cannot be parsed is refused and reported, never partially rewritten, and writes are atomic. Pairs with validate agents, which detects the same drift. The claude path is unchanged. (#3243) (#3296)

  • ~/.gsd/defaults.json shadowing is now diagnosed instead of silent — in any project with a .planning/config.json, global model-side keys (model_profile, model_overrides, models, dynamic_routing, runtime, …) were silently ignored for model resolution; a file named defaults.json applied to no real project with no signal. GSD now prints a one-time stderr warning naming the shadowed keys. Resolution precedence is unchanged; global effort keeps working via effort sync and never warns. (#3532) (#3540)

  • /gsd-review now records which model each reviewer actually used — REVIEWS.md frontmatter gains models: and model_sources:, so an unpinned lane's verdict is no longer attributable to an unknown model. (#2295) (#3649)

  • The read-injection scanner now reports which rules fired as structured data — its PostToolUse output carries a findings array of {ruleId, match} records alongside the human-readable advisory, so consumers no longer have to parse the advisory sentence to learn what was detected (the advisory text itself is unchanged). (#3523) (#3548)

  • Plans can now opt into a specialist executor via a per-plan agent_hint: frontmatter field — execute-phase dispatches the named subagent instead of gsd-executor when it resolves on the active runtime, and falls back to gsd-executor when the field is absent, blank, or the named agent does not resolve (byte-identical to today). Resolution consults the active runtime's agent directory (project-local and user-global, across filename variants) via a new gsd-tools resolve-agent query, and the hint flows through phase-plan-index as plan_json.agent_hint. Default-on via workflow.agent_hint_routing (set false to disable); covers the Agent()-based dispatch (harness-worktree and sequential). (#1689) (#3417)

  • resolveTriggerSurface (Runtime Artifact Layout Module) resolves the /gsd- trigger surface — winner, shadowedBy, and nested-router registration — per runtime/scope, and a new runtime.triggerPrecedence descriptor axis (required-with-default) decides same-trigger collisions; agents and kimi-agents are never trigger-bearing. (#3291)

  • Complexity-triggered refactor proposals — after a phase runs, GSD can now measure the complexity of the code that phase touched and surface a scoped refactor proposal when a function crosses a threshold or drifts past its recorded anchor, so entropy gets caught while it is still one function instead of a rewrite. Advisory and off by default; enable with gsd config-set refactor.trigger_enabled true. (#1953) (#3261)

  • check:contract-drift — a machine-enforced agent-contract registry — sentinel markers, read-tag gates, and deleted-file test references can no longer drift silently: the Agent Registry table in gsd-core/references/agent-contracts.md is now linted against what agents emit and what workflows consume, and lint-removed-but-needed catches tests that pin files your PR deleted. (#3565) (#3571)

  • runtime-homes now exports its non-registry config-home descriptors — KIMI_HOOKS_TOML_DESCRIPTOR, NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, GSD_LOCATION_ENV_KEYS, and the ConfigHomeDescriptor type are public, so consumers that need the set of config-location env vars (rather than a single resolved path) can derive it instead of hand-maintaining a copy. resolveKimiHooksTomlDir() behaviour is unchanged; its descriptor is simply named rather than inline (#3156).

    The test-instrumentation scripts no longer ship in the npm package — scripts/run-tests.cjs, scripts/live-config-guard.cjs, scripts/affected-tests-lib.cjs, and scripts/run-affected-tests.cjs are now excluded from the tarball (they are one closed require chain of repo-only test tooling). npm test in an installed package was already inoperable (tests/ has never shipped); a deep import of scripts/run-tests.cjs from the published package — an unsupported surface — will now be MODULE_NOT_FOUND (#3156). (#2677)

  • Per-phase commit_docs override — set phase_commit_docs.<phase-id> to commit one phase's .planning/ artifacts (e.g. an architecture phase) while keeping other phases local, without flipping the project-wide commit_docs switch. (#3587) (#3601)

  • audit-open acknowledge now suppresses open audit items at future milestone closes — deferring an item via /gsd-complete-milestone previously only wrote a human-readable note; the item resurfaced at every later close with no way to silence it short of resolving it for real. The new audit-open acknowledge --category <cat> --milestone <ver> [--at <date>] ... CLI verb writes a verdict-preserving audit_acknowledged marker that suppresses the item starting at the next audit scan, without ever touching the artifact's own status: field, and self-invalidates the moment the artifact's observed state changes again. query audit-open --json now also reports an acknowledged count per category alongside counts, so a clean close can be told apart from one that is clean only because prior items are still suppressed. (#3458) (#3555)

  • Effort now supports inherit — "follow the session" is a first-class, declarable choice — effort.agent_overrides, routing_tier_defaults, and effort.default accept inherit; the install-time writer omits the effort: frontmatter key for agents resolving to it (Codex omits the model_reasoning_effort pin), and effort sync --apply no longer re-adds a hand-stripped key — an absent key under inherit is in-sync, and a present one is stripped. An explicit inherit never escalates on failed attempts. (#3533) (#3541)

  • A lint rule now keeps Windows binary resolution in one place — re-implementing PATH/PATHEXT lookup outside the platform seam is rejected at lint time, so the four divergent resolvers epic #3411 removed cannot quietly come back. No change to how GSD behaves at runtime. (#3619) (#3636)

  • The EoS Registry now lists GSD for Reasonix — discover the independently maintained onionviolet/gsd-reasonix protocol-v1 host integration for Reasonix, including exact install and uninstall commands, supported interface points, and negotiated host axes. (#3403)

  • Windows binary resolution now has one owner — GSD resolves a command name to the file Windows can actually start, in the single platform seam, instead of four divergent copies. Reviewer lanes, execTool, and the capability spawn path all share it, so a .cmd/.bat shim resolves and runs where it previously failed with spawn ENOENT. macOS and Linux behavior is unchanged. (#3411) (#3621)

  • GSD now warns when .planning/ is gitignored but still tracked by git — adding .planning/ to .gitignore has no effect on files git already tracks, so planning docs kept landing in commits while commit_docs reported false. validate health now reports this as W029 with the git rm -r --cached remedy. (#3586) (#3598)

  • Diagnostic rules for .planning/ health checks now have a single parsed subject to read from — src/planning-snapshot.cts composes the already-consolidated milestone, phase, and plan derivations into one scope-carrying projection, so a rule can no longer re-derive a field's location from raw document text the way three now-inert validate health predicates once did (#3162). No command output changes yet — validate health migrates onto it in a follow-up phase. (#3308) (#3402)

  • Quick tasks can now be archived at milestone close-out. /gsd-complete-milestone offers an opt-in prompt to sweep .planning/quick/ into .planning/milestones/<version>-quick/ with a generated README.md index and a reset Quick Tasks Completed table, and /gsd-cleanup offers the same archival retroactively for milestones that were already closed. (#2142) (#3592)

Changed

  • The plan-phase AI-integration capability gate no longer lists substring-collidable keywords — bare eval (a substring of ordinary phase-goal words like evaluation and retrieval) is replaced by llm eval, and the under-specified ai system is dropped, per maintainer triage on the linked issue. The gate is a capability prompt, not a hard block, so this is a precision improvement: phase goals like "add evaluation metrics" or "build the retrieval layer" no longer invite a spurious AI-SPEC branch, and genuinely AI-flavored goals still match on the precise framework and technique names. (#2115) (#3431)

  • /gsd-explore research passes now disposition each surfaced claim three ways — admit (survives a prompted-to-refute pass and is grounded in a source, shown with the source), refute (a source contradicts it, dropped or corrected), or abstain (unverifiable, or a source-vs-prior conflict). Abstained claims go to a separate Unresolved ledger instead of being smoothed into confident prose, so you can see what the research could not stand behind. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim — a blog post contradicting your engines field is an abstain, the engines field itself is a refute — and your own prior belief is never authoritative alone. A finding that comes back with no disposition at all is ledgered as an abstain rather than silently dropped or asserted as prose. Two guards ship with it: conflict-abstention (a source-vs-prior conflict routes to the ledger, not a silent pick-a-side) and a tier floor (a would-be admit is presented as an abstain when the researcher's resolved tier is budget-level or could not be determined, because an under-tiered or unverified researcher over-defers to whatever source it was handed; corrections are unaffected). Keying the floor on the resolved tier rather than the model id keeps it working on non-Claude installs, where the model id is often blank or substituted by the runtime. The floor narrows this gap rather than closing it — a config that deliberately repoints one tier at another tier's model can still report a higher tier than what actually runs. Claims-side analogue of the honest verifier. (#2229) (#2543)

  • STATE.md now records the commit it was written against — a new state_head frontmatter stamp lets /gsd-health and smart-entry report how far the codebase has moved since STATE.md was last written, so a long-stale STATE.md can be discounted rather than read at face value. Health adds advisory W024 once the gap reaches 20 commits. This is a freshness proxy, not a drift measurement: the count includes commits that never touched anything STATE.md describes, and the stamp refreshes on any state write — so it is always worded as approximate and never gates anything. The stamp is omitted entirely when the commit cannot be resolved to the project's own repository — a project nested inside an unrelated checkout reports unknown rather than borrowing that repo's freshness. (#2573) (#2622)

  • Live-plan counting now has one owner, so superseded plans stop being scheduled and nested-layout phases stop reporting zero — scanPhasePlans is the sole source of which plans exist and which are outstanding. Twenty-one call sites that re-derived it from filenames now route through it, so a plan marked status: superseded is no longer scheduled into an execute-phase wave, phases using the nested plans/ layout no longer report zero plans, and stray summaries no longer inflate completion. (#3183) (#3199)

  • A percentage is now withheld everywhere its scope is not COMPLETE, not just at the sites Phase 3 reached — closing ADR-3180 §7.6 rule 4 at the two remaining gaps an isolated review caught: state json's buildStateFrontmatter no longer hardcodes SCOPE.COMPLETE when deriving progress.percent (it now threads the real listMilestonePhaseDirs scope through _diskScanCache, including its prose-fallback path, so a genuinely unreadable .planning/phases directory can no longer surface a stale or falsely-earned number there while every other surface withholds), and roadmap analyze --json now exposes the scope that actually gates progress_percent as its own progress_scope field — distinct from the top-level scope (heading-windowing identity) — so a consumer can tell why progress_percent is null from the JSON alone instead of seeing scope: "complete" next to an unexplained null. state update-progress also now writes a [gsd-tools] WARNING: line to stderr when it silently no-ops on a non-COMPLETE scope, so the skip is not visible only to a JSON reason field most callers never read. state sync now also withholds: it no longer hardcodes SCOPE.COMPLETE when deriving the percentage it writes into STATE.md's body — a non-COMPLETE scope (confirmed reproducible on TRUNCATED and UNSCOPED fixtures, not just the previously-checked UNREADABLE case) skips the Progress: write entirely and records a Progress: skipped — …(#3217) entry in changes, instead of persisting a fabricated percentage that could disagree with the same write's own (already-scoped) frontmatter progress: block. 0 under a genuinely COMPLETE scope is unaffected and still renders. Tier-2: progress_percent, percent, and plan_percent are number | null; computeProgressPercent requires a scope argument; roadmap analyze --json gains a new progress_scope field; state sync --raw's changes array can now contain a scope-skip entry and correspondingly withhold a Progress: body write it would previously have made. (#3217) (#3318)

  • gsd-plan-checker now flags same-wave plans that are coupled but don't say so — two plans in the same wave that share mutable state (a config key, table, migration, env var, singleton) or depend on each other's execution order, with no depends_on edge between them, are reported as an advisory Dimension 3 finding. The coupling gets settled at plan time instead of surfacing as an intermittent failure during parallel execution. docs/AGENTS.md's plan-checker entry, which claimed eight verification dimensions and listed eight names matching none of the agent's actual fifteen, is corrected to the real list in the same change. (#1954) (#3237)

  • state validate now runs its drift scan for STATE.md files whose phase lives only in frontmatter, instead of silently skipping the scan and reporting a false-clean result (#3162); it also no longer lets a frontmatter status: key shadow the body Status field. Its output gains a scope field (complete/truncated/unscoped/unreadable) reporting whether the check could actually run — valid still means no drift was found, and is not derived from scope. (#3187)

    state complete-phase's idempotency guard now consults frontmatter current_phase (via the same fallback chain as state validate), so a STATE.md whose phase lives only in frontmatter is no longer silently rolled back on a re-run of state complete-phase --phase N. It also gains a new refusal path: when the frontmatter cannot be parsed, the command now errors out ("Unable to read STATE.md frontmatter; refusing to run complete-phase to avoid a destructive rollback") instead of guessing. (#3187)

    workstream list/status/progress's per-workstream state projection (status, current_phase, last_activity) now resolves those fields from frontmatter when the body has no corresponding field, instead of reporting them absent — a frontmatter-only STATE.md's workstream inventory output changes accordingly. (#3187) (#3283)

  • effort.routing_tier_defaults now merges over the built-in tier defaults instead of replacing them — previously, creating an effort block without routing_tier_defaults silently disabled the built-in tier ladder (light:low / standard:high / heavy:xhigh), collapsing every non-overridden agent to high; one agent_overrides entry could reshape 20+ agents you never named. A partial block now fills gaps from the built-ins, and an invalid value falls back to that tier's built-in. (#3531) (#3539)

  • milestone complete no longer lets a stale STATE.md body line overwrite fresher frontmatter — it wrote through a path that re-derived frontmatter from the body with no preservation pass, so a stale Stopped-at line silently replaced a newer curated value, exactly as phase complete did before it was fixed. It now runs the same preservation the rest of the write path uses, and reports each field it protected in a new preservation_warnings array instead of staying silent about the divergence. (#3469) (#3501)

  • GSD now requires Node 24 or newer — the engines.node floor moves from 22 to 24, and the Node 22 test lane is retired. Node 22 entered Maintenance LTS and this project tracks the Active LTS line; the change is what lets regex escaping delegate to the built-in RegExp.escape instead of a hand-rolled implementation. If you are on Node 22, upgrade before updating GSD. (#3416)

  • Worktree-wave merges now warn when a plan branch committed outside its declared scope — the execute-phase cleanup gauntlet compares each branch's actual committed diff against the files_modified the plan declared and reports every path outside it. Advisory only: the merge still proceeds and the exit status is unchanged. (#2596) (#3264)

  • Progress percentages now come from one owner — every .planning/ completion percentage the CLI reports is computed by a single shared function instead of six hand-inlined copies, so a rounding or ceiling fix can no longer land on one command and silently miss the others. Reported values are unchanged. (#3180) (#3223)

  • Fallow binary resolution now shares the platform seam — resolving the fallow binary uses the same PATH/PATHEXT logic as every other spawn, so on Windows a fallow.cmd shim resolves correctly and an extensionless npm shim is no longer picked up in its place. node_modules/.bin is still searched before PATH, and the POSIX executable-bit check is unchanged. (#3618) (#3633)

  • validate health splits two previously-conflated warning codes into their own codes — W021 now covers only the phase-id-convention mismatch it originally meant; the STATE-vs-ROADMAP milestone-complete mismatch it used to also report moves to the new W026. Likewise W017 now covers only orphan worktrees; the stale-worktree case moves to the new W027. (#3405)

  • validate consistency's warnings are now coded diagnostics — each entry is a {code, message, fix, repairable} object instead of a bare string. Findings that overlap with validate health (a phase in ROADMAP.md with no directory on disk, or vice versa) now carry the exact same W006/W007 codes validate health already uses for them, so there's one vocabulary for that finding, not two. The four subjects unique to this command (phase/plan numbering gaps, orphan summaries, plans missing wave frontmatter) get a new C001-C004 code range. (#3407)

  • Digit-leading phase names now resolve consistently by bare number — phases such as "24/7 Autonomy", "80/20 Cleanup", and "12-Factor Refactor" now resolve across every phase verb instead of appearing missing; ambiguous directory collisions now fail loudly with their candidate paths instead of silently selecting the first match. /gsd and /gsd:progress also stop under-reporting: their verify-failed check shares the same directory selection, so a failed verification in one of these phases is surfaced rather than read as a healthy phase, and phase directories carrying a project-code prefix (MEM-05-…) are no longer skipped by that check entirely. The same selection now backs every remaining consumer that had resolved directories on its own, so phases list, phase remove, phase next-decimal, the schema-drift gate, the init-manager overview, roadmap analyze, and the milestone-completion and health consistency checks stop reporting these phases as having no directory. /gsd-health no longer reports one of these phases as both missing from disk and absent from the roadmap at the same time (W006 + W007), and phase remove now refuses — without deleting or renumbering anything — when two directories claim the same bare phase number. phase remove also stops writing a phase count one too high into STATE.md when the phase it just deleted was one of these digit-leading directories (#2528). (#2559)

  • state validate's warnings are now coded diagnostics, and the drift field is gone — each entry is a {code, severity, message, remedy} object (seven codes, S001-S007) naming exactly what STATE.md disagrees with the filesystem about and how to fix it, instead of a bare string. The separate drift object every response used to carry is removed entirely; every condition it used to report (a conflicting phase reference, a missing phases directory, a plan-count mismatch, a stale executing status) is now one of the seven coded warnings, so no information is lost, it's just structured. valid and scope are unchanged. (#3407)

  • /gsd-progress and /gsd-execute-plan stop counting superseded plans as outstanding work — seven prompt-layer sites across execute-plan.md, plan-phase.md, plan-review-convergence.md and progress.md counted plans with a raw ls *-PLAN.md | wc -l, so a plan marked status: superseded was still counted as outstanding, a phase on the nested plans/ layout (#3139) reported zero plans it actually had, and loosely-named plan files were missed entirely. Every site now calls phase find, which gains three additive fields — plan_count/summary_count (live, superseded excluded — 'how much is left') and plan_count_all (physical, every plan on disk — 'what did the planner write') — so what a workflow shows and what phase find reports for the same phase are now the same number. This also fixes a dead route: progress.md's Route 0 resume-incomplete-phase check read .plans/.summaries arrays that its producer, roadmap.analyze, never emitted (it emits plan_count/summary_count scalars), so both counts were always 0 and the check had never fired at all — it now fires correctly. This is a behavior change you'll notice: plan/summary counts shown by these workflows will move — toward being correct. (#3218) (#3327)

  • validate health --repair no longer resets config.json or regenerates STATE.md automatically — these two repairs are destructive (they lose custom settings or session history), so they're now reported with their fix described but never auto-applied; run the suggested command yourself to apply them. (#3405)

  • Gap-closure planning no longer documents a completion marker nothing reads — the planner emitted ## GAP CLOSURE PLANS CREATED but no workflow had a dispatch branch for it, so completion was always detected via the gap_closure: true fix-plan artifacts anyway; the dead marker is retired and the artifact route (verify-work --gaps spawn → plans → execute-phase --gaps-only) is now the documented contract. (#3440) (#3443)

  • Progress, stats, and phase listings now stay within the current milestone. progress, stats, and phases list no longer count backlog (999.*) or pre-milestone (0-*) directories as current-milestone phases, and phases clear / milestone complete no longer delete or archive those directories. phases list --phase and --include-archived are unaffected, since they intentionally look up or list beyond the current milestone. (#3185) (#3222)

  • Codex agents now inherit the session model instead of getting a pinned per-tier model — if you install for codex with a runtime set and any model_profile other than inherit, GSD no longer writes a model (or model_reasoning_effort) line into ~/.codex/agents/<agent>.toml. This fixes typed agents failing to spawn with 400 invalid_request_error: "The 'sonnet' model is not supported when using Codex with a ChatGPT account", which degraded the whole plan/execute flow to a generic-agent fallback. To keep pinning a model, set an explicit real-Codex id in model_overrides (e.g. {"model_overrides": {"gsd-planner": "gpt-5.6-sol"}}) — that path is unchanged. The installer prints a one-time notice when it drops a pin. Codex-only; all other runtimes are untouched. (#3241) (#3276)

  • gsd-verifier now says why a verified truth holds, not just that it does — a truth that reaches ✓ VERIFIED is additionally classified against three incidental-reliance patterns (an undeclared precondition, an ordering or side effect nothing enforces, a truth that is only true under the test fixture) and, when one matches, is reported as ✓ VERIFIED (coincidental-reliance) with an entry in the new coincidental_reliance_items frontmatter list naming what to harden. Purely advisory: the base ✓ VERIFIED token is unchanged, the truth still counts toward the score, the overall status is unaffected, and no human-verification item is emitted — a passing phase still passes. Only a consumer matching the truth-row verdict cell for exact equality (rather than as a substring) needs to tolerate the suffix. Two limits stated up front: the check is endogenous, and so measurably weaker than the exogenous backstop tag gsd-core/references/honest-verifier.md routes on — advisory status is the consequence, and its precision is unmeasured; and gsd-core/workflows/verify-phase.md is not edited, receiving the rule through its eager import of the verification-report template rather than a second inline copy, because it sits 29 bytes under its size hard cap. (#1955) (#3250)

  • Install scope is now resolved once, as a value — the installer and the modules downstream of it no longer each re-derive whether an install is global or local from a bare string. One module owns the scope axis and reports its config home, its per-scope settings file, and whether it requires a consent record. No behavior changes for any install. (#2870) (#3278)

  • runtime.hostBehaviors is now a closed vocabulary — the capability-manifest field that carries per-host install and adaptation switches was validated by nothing, so a typo'd or invented key was silently ignored forever. Its 59 keys are now enumerated, and a key outside the vocabulary is ignored with a non-fatal warning naming the capability and the key. It is never a validation error: a manifest authored against a newer GSD degrades visibly rather than failing the build, and an out-of-tree runtime descriptor carrying a bespoke key keeps installing. No shipped capability is affected. (#2801) (#3272)

  • The ADR gate now resolves documentation links and checks H1 status brackets — a link in docs/adr/ that pointed nowhere, and an H1 whose trailing [Status] bracket contradicted its own Status: field, both passed CI green; readers and agents following those citations hit dead ends the build had already blessed. gen-adr-index.cjs --check now fails on either, naming the file, the line, and the unresolved target. Links inside fenced or inline code are left alone, and resolution is case-exact on every platform. A new --json flag reports the same findings as a structured document with stable reason codes, so tooling never has to pattern-match an error message. (#2704) (#3266)

  • The claude reviewer in /gsd:review no longer inherits your CLAUDE.md or auto-memory — the lane now declares CLAUDE_CODE_DISABLE_CLAUDE_MDS=1 CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 (CLAUDE.md loading and auto-memory are independently-toggled mechanisms, so each gets its own variable), merged into that one spawn's environment, so it reviews the same self-contained prompt the gemini and codex reviewers already receive. It was previously the only reviewer additionally seeing your global CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory — a context asymmetry against the workflow's own independent-review premise, and a measured ~4k extra input tokens per spawn. Carried as declared lane data (invoke.env, ADR-2782), not a bespoke handler; nothing reaches the orchestrating session or any other lane in the run. Affects /gsd:review (and the convergence flow that reuses it) invoked from a non-Claude-Code runtime; inside Claude Code the claude reviewer already self-skips for independence. (#2483) (#2493)

  • Agent required-reading enforcement now actually fires — spawner workflows and commands emitted <files_to_read> while agents gate on <required_reading>, so the "you MUST Read every listed file" clause never triggered; the canonical tag is now <required_reading> everywhere (46 spawn blocks across 24 workflows), with a repo guard banning the legacy tag so the two vocabularies can never drift apart again. (#3423) (#3432)

  • The installer module no longer re-exports internals it does not own. bin/install.js exported 197 names, 70 of which were either dead or plain pass-throughs to the modules that actually implement them — kept for "existing consumers" that turned out not to exist, since no production code has ever required the file. Those 70 are gone and their tests now import the owning modules directly. No installed output changes. (#2876) (#3615)

  • A truncated milestone window is no longer reported as an empty milestone — roadmap analyze now emits a scope field (complete/truncated/unscoped/unreadable) so phase_count: 0 from a genuinely fresh milestone is distinguishable from a window that closed before reaching the roadmap's phase sections, and milestone complete refuses to archive on a truncated window instead of moving every phase directory in the project. (#3184) (#3209)

  • init now reports the host runtime it is actually running under — inside a Codex session GSD reported agent_runtime: claude, and checked the wrong directory for installed agents, because runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json. A detection rung now sits beneath both explicit sources, resolving codex from Codex's own session environment. Explicit settings still win, no shared defaults are written, and model resolution is untouched. (#3245) (#3307)

  • The plan drift guard now flags the same fact stated two ways — when ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md contradict each other about a phase status, a success criterion, a requirement ID or a domain term, plan review reports it in REVIEWS.md naming both locations and which one is authoritative, instead of letting a fresh-context agent act on the stale copy. The phase-status axis is decided deterministically rather than by judgment, so a STATE/ROADMAP contradiction is caught the same way every time — and a disagreement about whether a phase is complete is always reported, never written off as one document lagging the other. Advisory only; it never blocks convergence, and the judgment axes key on contradicting knowledge rather than similar-looking text. Runs under the existing plan_review.source_grounding switch — no new setting. (#1956) (#3259)

  • Milestone names are no longer truncated at a parenthesis, and a phase heading is never mistaken for the milestone — a ROADMAP whose ### Phase N heading mentioned a version could cause a wrong milestone: to be written to STATE.md, and a milestone named v3.3 — Portability (Windows) was recorded and rendered as Portability. Milestone identity now has one implementation; when it cannot be determined it is reported as absent instead of defaulting to a plausible-looking v1.0/milestone. (#3216) (#3226)

  • /gsd-progress no longer implies re-execution when only the verification report is missing — the routing message now explains that running /gsd-execute-phase on a historical phase resumes at the verification gates and does not re-run already-summarized plans, and softens the unrecognized-status message to acknowledge an intentional non-standard marker. (#1762) (#3439)

  • Phase completion is now decided by a single disk-strict predicate — a ticked ROADMAP checkbox no longer carries machine authority. isPhaseComplete (src/verification.cts) is the one owner: a phase is complete exactly when its *-VERIFICATION.md reads passed, read unconditionally — plan count is never a precondition. This changes four observable surfaces: init manager now reports a zero-plan phase with a passing verification as complete instead of the retired not_required sentinel (#3168); roadmap analyze's checkbox override is removed, so a ticked checkbox with outstanding plans or no passing verification now reports incomplete instead of complete; roadmap update-plan-progress routes through the same owner (and, unchanged, still refuses to write a completion checkbox/date while any plan lacks a *-SUMMARY.md); and gsd-core/workflows/mvp-phase.md stops ORing a checkbox-derived PHASE_COMPLETE into its completion decision, deciding on disk status alone. A ticked checkbox is not deleted — only its authority over these commands is removed.

    workstream list/workstream status's per-phase complete status (via buildWorkstreamInventory) is now routed through the same owner instead of its own summaryCount >= planCount-plus-verification-verdict rule — a zero-plan phase with a passing verification now reports complete there too, and a phase whose *-VERIFICATION.md is absent no longer reports complete on summary count alone (the pre-existing "verifier-disabled projects still complete" tolerance is retired under disk-strict). That tolerance was #2645's deliberate boundary — missing/unknown/stale verdicts counted as non-failing so a project that never runs the verifier would not report 0% forever. Disk-strict retires it and closes #2645's Goodhart hole from the other side: deleting a *-VERIFICATION.md now lowers the reported completion instead of raising it. A project that does not run the verifier will report its phases incomplete. (#3186) (#3306)

  • state json and the state-mutating commands now agree with what is actually on disk — a stale body annotation could beat a fresher curated frontmatter value in state json output, and commands reported fields as updated that the write pipeline had already discarded while staying silent about fields it restored. Preservation is now enforced in one place across every path, each command reconciles its report against the persisted file, and a value dropped because this write deliberately removed its body line is reported rather than silently lost. (#3471) (#3519)

  • must_haves.key_links[].pattern now uses RE2 syntax — backreferences and look-around are no longer supported in a key-links pattern, because they are the constructs that require a backtracking engine and cannot be evaluated in guaranteed linear time. A pattern using them is reported as pattern_neutralized: "unsupported" with the link marked unverified, rather than being silently matched as literal text. Ordinary patterns, including every example shipped in the docs, are unaffected. (#3477) (#3496)

  • Agent files now install identically whether you run a full install or apply a surface. Every runtime materializes its agents from its capability descriptor, so /gsd-surface --materialize no longer skips agent files for Cline, Codex, Hermes, Kilo, OpenCode and Kimi Code — previously it wrote none for those runtimes, leaving an install missing the agents a fresh install would have created. Installed output is byte-identical to before for every runtime. (#2866) (#3600)

Removed

  • The undocumented runtime.hostBehaviors.reviewerCli capability field has been removed — it was superseded by the declared reviewer body in 1.9.0 and kept working for one release as a derived alias. A manifest that still sets it contributes no reviewer lane and now reports a non-fatal warning naming the capability, at build time on stderr and at install time through the overlay loader; nothing crashes and no other behavior changes. Every shipped reviewer lane already declares a reviewer body, so the roster is unchanged — if you maintain an out-of-tree runtime descriptor that relied on the flag, declare a reviewer body to restore the lane. (#2801) (#3272)
  • Two workflow files that shipped to every runtime but were never loaded are gone — discovery-phase.md and plan-milestone-gaps.md had no command, agent, or skill referencing them, and docs/INVENTORY.md claimed callers for one that did not exist. A new lint rule now fails the build if any shipped workflow becomes unreachable again. (#3560) (#3564)
  • Removed the orphaned verify-phase workflow (~40 KB shipped to every runtime, never loaded) — its still-live verification gates (decision-coverage validation, test-quality audit, infrastructure-phase human-verification scoping) moved to a reference the verifier agent actually loads, so they run again instead of shipping as dead prose; installs are ~40 KB lighter and PRs to the verifier no longer mirror a dead twin to keep lockstep tests green. (#1891) (#3422)

Fixed

  • withPlanningLock no longer reports a phantom "held by a live process" timeout when .planning/ cannot be created — a best-effort try { platformEnsureDir(...) } catch { /* ok */ } swallowed the real mkdir failure (EACCES/ENOSPC/EROFS), so the subsequent lock write failed with ENOENT (parent missing), and because ENOENT is in the lock's retry set (added for a Docker overlay-fs race) the loop spun the full 10 s budget before throwing a misattributed contention error that pointed operators at a nonexistent lock-holder. The mkdir failure now propagates immediately with its real filesystem errno and message, so an unwritable or full disk is reported as itself, not as concurrent-writer contention. The Docker overlay-fs ENOENT lock-write race (directory present) is still retried as before, and every code path where .planning/ already exists or can be created is unchanged. Part of epic #1879 (distinguish "absent" from "corrupt/permission-denied" across engine read paths). (#1884) (#3472)

  • The idle/staleness detector now fires when last_activity carries a description — a last_activity written in the shape templates/state.md prescribes ([YYYY-MM-DD] — [What happened]) parsed to NaN, and because the detector treats an unparseable value as "not stale" it failed open to false. Any project whose last_activity kept its description was never reported idle, no matter how long it had sat. The leading date is now parsed out of the value, so the description no longer blinds the only staleness signal in the front door. An impossible calendar date such as 2026-02-30 is now rejected outright rather than silently rolling forward to a real — and wrong — date. (#2570) (#2571)

  • /gsd-ship now detects and recovers a PR wedged by the ship-note commit — when the [ci skip] ship note leaves required checks unstarted, ship re-triggers CI instead of leaving the PR unmergeable. (#2783)

    Note: This introduces a latency tradeoff. All /gsd-ship invocations now poll GitHub PR state for up to 15 seconds to ensure the commit was processed and check if recovery is needed, even for repositories without required checks. (#2818)

  • spec-phase Step 5.5 now surfaces the edge-probe's proposed edges to the resolution loop instead of discarding them — the deterministic coverage report was computed, validated, then reduced to a single applicable-count, so the resolution loop re-derived edge categories from requirement prose and the engine's proposals never reached it. The report is now rendered into context and its rows are consumed as a floor the model unions with its own classification (still adding any category the classifier missed), so the written ## Edge Coverage reflects the engine's deterministic taxonomy rather than model-invented categories; --auto gets the same floor. (#3102) (#3391)

  • /gsd-quick --validate no longer trusts a verification result it cannot actually read — quick parsed the verifier's status by grepping the whole report rather than its frontmatter, so a status: line in the report's prose could be picked up alongside or instead of the real one, staleness was never detected at all, and a range of valid and malformed reports alike resolved to a value no routing arm matched — leaving the orchestrator to improvise at the moment the pipeline had failed. Quick now reads the same frontmatter-anchored, staleness-aware verification.status query that execute-phase, verify-work and progress already use, and routes missing / unknown / stale through an explicit arm instead of falling through. (#3174) (#3205)

  • The verifier's non-inferable (backstop) abstention rule now defines "explicit evidence" where the verifier is guaranteed to read it. Step 3 item 5b used the term undefined — its definition was stranded in gsd-core/references/honest-verifier.md behind a stale references/ cite that does not resolve, so the term fell back to the verifier's default notion of evidence (symbol presence + wiring), the exact false-pass the #1154 abstention protocol exists to refuse. 5b now carries the definition inline (a passing wired held-out/property-based test or directly observed behavior; presence + wiring never qualifies), the AFK never-silent/never-halt completion line and the insufficient_spec-vs-manual-UAT distinction ship in the eagerly-loaded verifier-phase-gates.md reference, and the agent file's three stale bare references/ cites are gone: the two at 5c and the MVP-mode section now resolve under the gsd-core/ prefix, and 5b's is superseded by the inline definition itself. (#3206) (#3435)

  • Autonomous/auto-mode no longer auto-approves unmet <precondition> checkpoints, and the blocker loop now halts needs_human instead of retrying forever — the checkpoint an executor returns when a task's <precondition> is unmet (an unmet user_setup step, a missing env var, an absent prior-phase artifact) now carries gate="blocking-human", which both auto-mode bypass layers (executor checkpoint protocol and execute-phase checkpoint handling) honor, so it always stops for a human instead of being silently approved with a synthetic "approved" and then failing <verify> on the still-missing prerequisite. Independently, /gsd:autonomous's blocker handler now counts "Fix and retry" attempts per phase step and, after 3 failed attempts, escalates to a terminal needs_human halt that surfaces the unmet items and records a ## Needs Human STATE.md row, ending the observed multi-hour retry loops on operator-gated plans. (#3210) (#3528)

  • state add-roadmap-evolution and state add-decision no longer persist a literal Phase ? when --phase is omitted — both commands built their entry from the raw CLI flag's ? fallback instead of the phase already recorded in STATE.md, even with current_phase: 3 present in frontmatter. Both now resolve the phase through a strict write-path ladder (frontmatter current_phase → body Current Phase → Phase: X of Y scoped strictly to ## Current Position), leaving ? only when genuinely unresolvable; an explicit --phase still wins. The resolver deliberately does not reuse the read-path resolveStatePhase, whose document-wide fallback could adopt a stale historical | Phase | N | table row. A guard test now sweeps src/*.cts for any new raw phase || '?' call site. (#3481) (#3522)

  • Frontmatter round-trips no longer double backslashes on every state write — escapeDoubleQuoted escaped \, ", and control characters on each serialize while the parser only stripped the outer quote delimiters, so every read-modify-write cycle doubled existing escapes (2ⁿ−1 backslashes after n cycles). syncStateFrontmatter carries last_activity_desc through that seam on every state command, growing STATE.md unboundedly — the reported 134 MB file OOMed state.record-session after 26 writes. Double-quoted scalars are now un-escaped on parse via the exact inverse of the escaper, making serialize→parse a fixed point; unrecognized escapes are kept literally so hand-authored files parse unchanged. (#3497) (#3521)

  • /gsd:code-review now derives the phase diff base from GSD's own commit scopes instead of a prose phrase, ending silently wrong review scopes — the diff base fed to the reviewer file-list fallback, the SUMMARY↔diff cross-check union, the reviewer agent's diff_base, and the fallow --changed-since structural pass was greped from commit messages for the literal "Phase N" and kept the oldest match, so any prose mention anywhere in history (a planning commit deferring work "to Phase N per D-09", a doc commit using "### Phase N" as a format example) silently set the base months before the phase existed — on a real repo ~4 phases too early, inflating the reviewer's reading list ~78% with no warning — while GSD's own commits (docs(phase-N):, feat(N-MM):, docs(N):), which never contain the literal phrase, were never matched at all. All three derivations now anchor on the subject-line conventional-commit phase scope (both padded 06 and unpadded 6 spellings, since workflows emit the unpadded roadmap number), commit bodies can no longer capture the base, and a history with no scope-style commits fails loudly with the existing no-base warning and --files escape hatch instead of silently picking an arbitrary commit. (#3503) (#3526)

  • uat_path is now pinned to the phase's own UAT artifact instead of being picked by unsorted directory-listing order — both uat_path projections (init plan-phase and init phase-op) selected the phase's *-UAT.md with a bare first-match .find() that had no phase-membership check and no ordering, so a stray or cross-phase 04-UAT.md sitting in phase 03's directory could become phase 03's uat_path, and which file won was filesystem-dependent (creation order on APFS, hash order on ext4/XFS) — meaning two machines on the same commit could emit different uat_path values for the same phase, sending downstream workflows to read another phase's UAT state. Both sites now route through a shared phase-pinned resolver (resolveUatFile, sibling of the resolveVerificationFile rule from #3357/#3492): the phase's own <token>-UAT.md always wins, otherwise the alphabetically-first dashed candidate, deterministically on every machine. (#3518) (#3525)

  • MemPalace sub-features whose defaults are enabled now run when their config keys are absent — the earlier capture_artifacts absent-key fix (#2982) had been applied to only one of six hand-written config gates; the remaining gates for mempalace.mirror_kg (knowledge-graph mirroring in the capture and recall skills, their command mirrors, and the curator agent) and mempalace.diary_journal (per-agent diary entries at ship) still required the key to be explicitly present and true, so a project that enabled MemPalace without writing every sub-toggle silently never mirrored KG facts or wrote diary entries, with no warning. All six gates now treat an absent key as enabled (matching the capability registry's declared default: true) and disable the behavior only on an explicit false; default-off switches (mempalace.enabled, cross_project_tunnels) still require explicit opt-in, and a registry-parity regression test keeps future default-true keys from reintroducing the inversion. (#3479) (#3527)

  • Roadmap Plans: lines keep their hand-written text instead of being overwritten with a plan count — roadmap update-plan-progress replaced everything after the Plans: label whenever the line did not already begin with a canonical N/N plans token, silently destroying freeform prose, a TBD note, or a hand-written annotation. A sentence that wrapped onto a second line lost only its first line, leaving the continuation stranded so the roadmap asserted something nobody wrote — at exit 0, in a diff that read as a routine count bump. The count is now written only over a real count token or the fresh-template placeholder, and a single-plan phase (1 plan) is recognized rather than frozen. (#3584) (#3635)

  • Running a capability's own test suite no longer silently deactivates it — bundleContentHash digested every entry under a capability bundle with no exclusions, so ordinary Python bytecode caching (__pycache__/*.pyc, written by any plain python3 run) changed the consent-binding hash. The capability then reported inactive with no error and no warning, and loop render-hooks quietly dropped its step and gate — indistinguishable from never having installed it. An empty __pycache__ directory was enough to trigger it, since the digest binds directory existence. Only a *.pyc/*.pyo file sitting directly inside a __pycache__ directory is now excluded from the digest; a .pyc/.pyo file anywhere else stays bound, since a sourceless legacy .pyc there is still importable and executable. A __pycache__/.pytest_cache directory has only its own marker suppressed — its contents still bind the digest normally. node_modules and other executable content stay bound, excluded entries still count toward the walk's caps, and the filter runs after the symlink rejection so it cannot smuggle one past. (#3631) (#3650)

  • Corrected model-profiles.md: model and effort do not resolve through one shared precedence ladder — the reference previously claimed a models[phase_type] or dynamic_routing override flips both, and that an effort config change takes effect like a model change. In reality effort (claude runtime) is baked into agent frontmatter at install time and requires node gsd-tools.cjs effort sync --apply to change; Codex agents pin model_reasoning_effort in generated .toml files. (#3530) (#3536)

  • requirements mark-complete now flips the traceability row when ## Traceability holds more than one table — updateTableCell no longer binds to the first table in the section; it scans for the table that actually carries the requested column. A section with a phase-summary table above the requirement rows previously made the Status write silently bail (table_unmatched) while the checkbox still flipped, leaving the row at Pending indefinitely. Single-table sections are unchanged. (#3255) (#3377)

  • init execute-phase no longer hands a directory slug to the phase-start flow as the phase display name. When a phase's working directory already exists on disk, the disk-lookup path derived phase_name from the directory-name remainder — itself an already-slugified value (phase.add writes ${num}-${slug} dirs) — so phase_name and phase_slug came out byte-identical. The execute-phase workflow forwards phase_name into state begin-phase --name, which wrote that raw slug into STATE.md's current_phase_name on every phase start (loop-termination-and-baseline-correctness instead of Loop-Termination and Baseline Correctness). init execute-phase now prefers the ROADMAP's curated display name (### Phase N: <Name>) for phase_name, matching the no-disk fallback path that already did this correctly; phase_slug is unchanged so branch-name construction is unaffected. The state begin-phase override mechanism (#2821/#2736) is untouched. (#3171) (#3429)

  • Two GSD workflows told agents that a Claude Code Agent() spawn blocks until the subagent finishes — Claude Code backgrounds subagents by default, so /gsd-execute-phase could treat a wave as returned when it had not, and /gsd-debug lost its session-manager handoff in exactly the way #2196 was filed to fix. The dispatch notes now match this package's own shipped capability matrix, and both debug spawns carry the run_in_background: false opt-out they always needed. (#3177) (#3281)

  • Executor dispatch prompts now state checkpoint gate semantics: gate="blocking" (the default) is auto-approvable in auto-mode, only gate="blocking-human" always surfaces to a human. The phase-level and single-plan-level orchestrators no longer leave room to compose dispatch text that refuses auto-approval, which stalled autonomous runs at ordinary blocking checkpoints. (#3478)

  • gsd-tools validate health no longer flags .planning/WINDOWS.md as an unrecognized file — the broken-windows ledger that gsd-core's own windows command writes is now registered as a canonical .planning/ artifact. Previously the W019 warning advised archiving or deleting a file that, with workflow.windows_enforce on, gates /gsd-ship. (#3224) (#3369)

  • gap-analysis check gap-analysis.plan-post no longer reports prose trailing the requirement ID list as missing requirements. ROADMAP Requirements lines routinely carry locked-decision annotations, ambiguity scores, and prohibition notes after the ID list; passing that value verbatim into --phase-req-ids previously caused every prose word to be reported as an individually-missing requirement, drowning the real coverage signal. Tokens that cannot be requirement IDs (prose, punctuation, dates) are now dropped after range expansion. (#3438)

  • state.patch now reports a field as updated only when its post-write on-disk value matches the requested value; fields the write pipeline re-derives away (e.g. current_phase, current_phase_name) are reported as failed instead of phantom updated (#3487)

  • User profile and dev-preferences files are no longer lost when an install or uninstall is interrupted. These files were held only in memory while GSD deleted and rebuilt the directory containing them, so pressing Ctrl-C — or any crash during the copy — destroyed them permanently. On the main install path that window spanned the entire gsd-core tree rebuild. They are now staged to disk before anything is deleted, and any copy orphaned by an interrupted run is restored automatically on the next install or uninstall. (#1874) (#3600)

  • /gsd:code-review-fix <phase> --auto now commits the converged REVIEW.md alongside REVIEW-FIX.md and reliably commits REVIEW-FIX.md at all — the --auto re-review loop overwrote REVIEW.md every iteration but the workflow's single docs commit staged only REVIEW-FIX.md, so the committed REVIEW.md stayed at iteration 1 and contradicted the committed REVIEW-FIX.md (and the converged REVIEW.md plus .iterN.md backups survived only as uncommitted working-tree state). Separately, the two inline frontmatter validators exported REVIEW_PATH into a node -e body that reads process.env.FIX_REPORT_PATH, so the status check was always empty and REVIEW-FIX.md was never committed (the user was wrongly told the agent produced malformed output). The validators now export FIX_REPORT_PATH, the --auto commit stages REVIEW.md too, and spent .iterN.md backups are removed on successful convergence (retained on degradation). Non-auto single-pass runs are unchanged. (#3190) (#3434)

  • Three shell guards that could never fire now do — the planner's Walking Skeleton mode never activated on any project, phase planning recorded an empty requirement list instead of TBD, and completing a milestone with no phase summaries could hang instead of finishing. Each read a value that came back empty on success, so the fallback written to handle it was unreachable. (#3409) (#3558)

  • Shipped workflow/agent citations resolve again — 43 backticked references/<name>.md cites across 19 shipped files were dead pointers from every install location; all repaired to the canonical gsd-core/references/<name>.md form, and a new sweep gate fails the build on any future bare cite across the runtime-loaded trees. (#3576) (#3596)

  • A genuinely milestone-sectioned ROADMAP whose STATE.md asserts a milestone token matching no heading no longer has progress.total_phases clobbered to the on-disk phase-directory count (e.g. 25 -> 4) on every state-mutating command. The stored total is preserved (or the key omitted when nothing is stored), a stderr warning names the unbounded milestone token, and progress.percent stays withheld as before. (#3480)

  • Phase-directory collisions in .planning/phases/ (two in-scope dirs normalizing to the same phase number) no longer resolve by filesystem mtime — a checkout-order signal that made progress.total_plans and completed_plans differ across clones of the same commit. The survivor is now chosen deterministically by lexicographic directory name, and the collision is surfaced as a stderr warning naming both directories. (#3486)

  • code-review: every phase diff-base derivation now uses the same anchored, POSIX-portable phase-mention grep. Fixes wrong review scope from /gsd:code-review when a phase has no SUMMARY artifacts: the reviewer diff_base and the fallow --changed-since base no longer resolve to old unrelated commits whose messages merely contain the phase digits, and the anchored search now actually matches on macOS (the previous \b word boundary is not POSIX ERE and silently matched nothing there). (#3437)

  • phase add no longer files new phases inside archived roadmap history — the insertion point used the file's last horizontal rule, which on a long roadmap sits deep in shipped/archive content, so new phases landed under an unrelated archived phase's heading instead of at the end of the active phase list. Insertion is now scoped to the current milestone. (#3163) (#3400)

  • Agent isolation guard enforces on multi-runtime machines — the isolation guard (and Cursor's subagent-start fallback) resolved the project runtime from the host-wide ~/.gsd/defaults.json, which names whichever runtime installed last; on machines with two runtimes this confidently picked the wrong runtime and silently disabled executor worktree policing. Both now read the per-install .gsd-runtime marker above that file. (#3566) (#3589)

  • state update-progress no longer writes two different completion percentages in one call — the verb printed plan throughput (summaries/plans) to stdout and into the body Progress: bar, while the same write independently derived the frontmatter progress.percent as the deliberate min(plan, phase) cap. Mid-phase, when plan throughput runs ahead of phase completion, STATE.md contradicted itself and state json disagreed with the command that had just written it — silently, at exit 0. All surfaces now derive from the single canonical computation, and its reported plan counts come from the same milestone window as the percent, so the verb's own output can no longer disagree with itself. When that computation withholds a percent, the verb withholds too rather than substituting a different metric. The min-cap definition is unchanged. (#3583) (#3634)

  • GSD skills no longer override the caller's effort level (#3151) — invoking /gsd-plan-phase, /gsd-execute-phase, /gsd-autonomous, /gsd-next, /gsd-progress, or /gsd-stats previously set output_config.effort to a static value baked into the skill frontmatter; when that differed from the session's effort (which it did ~76% of the time), it invalidated the entire prompt cache at both scope boundaries (skill entry and exit). These skills now run at the session's existing effort level (no effort: emitted into SKILL.md). The elevated-effort intent is preserved on the source command files; only the skill-frontmatter emission is dropped. The separate agent-effort surface is unaffected. (#3425)

  • Global OpenCode/Kilo installs no longer pin a tier-default model over your session selection — a project's model_profile: "inherit" was invisible to the install-time resolver on global installs (it probes from the install dir and never reaches the project), so the balanced default silently baked e.g. anthropic/claude-opus-4-8 into the agent frontmatter, which those runtimes use over the live /model selection — producing "Model not found" on providers without that exact id. A profile that cannot be verified now bakes no model: line, so subagents follow the session model as documented; declare model_profile in ~/.gsd/defaults.json to pin tiers machine-wide. (#3543) (#3563)

  • phase remove no longer corrupts STATE.md after removing an inserted (decimal) phase — the removed-phase write prepended a second, partially-wrong frontmatter block (and left the phase's ROADMAP heading behind, so total_phases kept counting it); removal now updates STATE.md in place as a single block, drops the heading, and clamps phase counts at zero. (#3572) (#3594)

  • validate health --backfill now works without also passing --repair — previously it silently did nothing unless --repair was also set, due to an unreachable internal gate. (#3405)

  • gsd-health's STATE/ROADMAP staleness warning (W011) now reads the current phase from the YAML frontmatter format gsd-tools itself writes (current_phase), in addition to the legacy prose, canonical body, and pipe-table forms, and suppresses the warning when the recorded status reports completion in the state writer's own vocabulary (status: completed). The stale-worktree warning (W027) no longer advises unconditional forced removal: its remediation now directs checking for uncommitted work first (git -C status --porcelain), removing non-destructively when clean, with --force presented as an explicit opt-in to discard changes. (#3452)

  • Completing one phase no longer marks the whole milestone done — state complete-phase wrote the body prose Phase N complete, and the status normalizer matches complete as a substring, so finishing phase 2 of 4 collapsed the milestone-level STATE.md frontmatter to status: completed while the very same call correctly recorded completed_phases: 2 of total_phases: 4. Downstream automation that gates on milestone status — auto-advance, archival, ship gating — was told a half-open milestone was finished. Milestone status is now derived from those counters instead of from phase-level prose. (#3578) (#3614)

  • progress, stats, and query progress now report a real percentage inside a workstream — under --ws, these commands counted the workstream's own phases and plans but read the milestone window from the project root, which workstream create has already migrated away. The scope resolved as unreadable and the percentage was withheld, so a fully-complete workstream reported no progress at all. milestone complete no longer archives every phase directory when its milestone window is unreadable — it previously fell back to moving everything on disk in that case; it now declines to archive and reports why, leaving the phase directories in place. (#3597) (#3607)

  • gsd-tools stats no longer counts phantom phases from inline code — prose mentioning ### Phase N: inside an inline code span (e.g. a roadmap explaining its own numbering) inflated phases_total with a never-completing Not-Started row and deflated completion percent; stats now requires the same digit-bearing phase id shape roadmap analyze uses, so the two agree. (#3569) (#3591)

  • parseDeferredItems now counts heading-delimited deferred items as ONE entry (a heading plus its descriptive sub-bullets) instead of one per bullet, across flat, container-heading, and mixed-depth files; headless one-bullet-per-item files are unchanged. A bolded - **Status:** resolved marker now resolves its item instead of surfacing as a bogus unresolved entry. (#3488)

  • state.record-session no longer shrinks your phase count — a project whose ROADMAP declares more phases than it has directories on disk (phases 5 and 6 planned but not started yet) had progress.total_phases silently overwritten with the directory count, converging on the right number only once the last phase directory happened to exist. A flat roadmap carrying an ordinary heading like ## Progress was being misread as milestone-sectioned. Known limit: two milestone sections carrying no version token, no status marker and not the word "Milestone" are still not detected as sectioning. (#3204) (#3230)

  • Parallel phases running in the same working tree no longer corrupt STATE.md silently: state.begin-phase, state.advance-plan and phase.complete now consult a milestone claim (.planning/milestone.lock) keyed by phase + session id, and surface a visible milestone_conflict warning (stderr plus a typed JSON field, and phase.complete's warnings[]) when another live session holds a different phase — instead of silently overwriting the single Current Position slot. (#3455)

  • /gsd-map-codebase --fast now actually runs the fast scan — the flag routed to "the scan workflow" in prose but named no path any runtime could resolve, and the command loaded only the full four-agent map workflow, so scan.md was never read and the single-agent scan was improvised rather than executed. (#3561) (#3562)

  • State validation properly detects drift — Resolved an issue where state validation would silently fail to detect drift because it skipped scanning entirely when the shipped template lacked a specific field. (#3208)

  • Phase writes now guard the current milestone's scope. phase add/add-batch/insert reject a description containing a level 1-3 heading with a milestone marker (version token, status marker, or the word Milestone) before anything is written, and the edit-phase workflow captures roadmap milestone-scope (new read-only probe) around its in-place section write and rolls the edit back with an explicit error if the milestone window's scope or phase set changed. (#3446)

  • /gsd-quick and the UAT-diagnosis step no longer abort with a FATAL on a non-Claude runtime that can actually isolate — both dispatch sites resolved worktree isolation from a hardcoded RUNTIME != "claude" test, so every non-Claude host was refused regardless of what it could actually do. They now read the negotiated dispatch.isolation capability (#2584), and installs for runtimes that declare worktree support no longer stamp workflow.use_worktrees to false, which had pre-empted that negotiation. A runtime is judged by what it declares rather than by its name. A host that declares no isolation primitive at all still fails closed when worktrees are explicitly enabled — that FATAL is the fail-closed contract, not the bug — and a host whose isolation model the single-agent sites cannot express degrades to sequential, one agent at a time, on the main working tree. (#2728)

  • Reviewer lanes that declare source-grounded evidence are now verified at run time: a review citing zero file:line source evidence is stamped [reviewed-without-source-citations] and down-weighted in the Consensus Summary, instead of silently riding its declared evidence class at full weight (gemini plan-only reviews were measured doing exactly this). (#3436)

  • Managed /gsd:debug auto-resume no longer stalls after an answered checkpoint: the respawned session manager now receives the recorded next action and checkpoint status, plus the disposition that prior checkpoints were already answered, so the debug loop proceeds on the persisted next step instead of stopping behind the no-progress guard. (#3476)

  • Concurrent Claude Code sessions no longer share one active-workstream pointer — Claude Code exports its session id as CLAUDE_CODE_SESSION_ID, but the session-identity probe only listened for CLAUDE_SESSION_ID, so session-scoped workstream isolation never engaged on Claude Code: every session in a working tree resolved through the single shared .planning/active-workstream pointer, and a STATE.md update belonging to one workstream could be written silently into another's directory. The probe now accepts CLAUDE_CODE_SESSION_ID (no other key's precedence changed); concurrent sessions each keep their own session-scoped pointer again. (#3557) (#3570)

  • plan-phase: a completed --gaps planning run's Next Up handoff now recommends /gsd:execute-phase --gaps-only (matching the gap-closure scope just planned) instead of the whole-phase /gsd:execute-phase . Standard and --reviews runs are unchanged. (#3453)

  • GSD no longer commits .planning/ files you told it to ignore — several workflow steps staged planning artifacts with raw git add, bypassing the commit_docs setting and the .gitignore auto-detect entirely, so planning docs reached shared history anyway. (#3585) (#3590)

  • Global Claude installs load skill content correctly again — the installer rewrote @~/.claude/ file references to @$HOME/.claude/, which Claude Code does not expand, silently leaving every GSD skill with an empty execution_context (the model got scaffolding but never the workflow body). @-references now stay on the tilde form Claude resolves. (#3133) (#3393)

  • Twenty-five folded test suites no longer run twice on every CI lane — three consolidated install suites each carried a verbatim second copy of a contiguous run of folded regression blocks (~5,800 lines), left behind by a stale-base re-application during the test-consolidation epic. Every duplicated block registered and passed twice, so nothing reported it, and a contributor fixing one of those regressions could edit one copy and leave the other asserting the old behavior with the suite still green. The duplicates are deleted, and a new local/no-duplicate-fold-marker ESLint rule fails the build if a folded suite ever appears twice in one host file again. (#3271) (#3285)

  • Items left unresolved when a milestone closes are no longer invisible to every later audit — query audit-open's four phase-scoped scanners read only .planning/phases/, so once a milestone closed and its phase directories moved to .planning/milestones/vX.Y-phases/, any UAT gap, verification gap, context question or deferred item still open at that moment vanished from the pre-close audit permanently. In a fully-archived project the scanners returned nothing at all, which is indistinguishable from a clean tree — and because the audit sums every category into one has_open_items boolean, that could report a clean close it had not verified. All four now scan the archived milestone directories as well, and each item says which milestone it came from. (#3458) (#3555)

  • roadmap tools recognize table-style phase listings — a ROADMAP whose current-milestone phases are declared as markdown table rows (| 20 | … |) reported phase_count: 0 and found: false across roadmap.analyze, roadmap.get-phase, init.phase-op, and the milestone filter; all four surfaces now resolve table-declared phases (progress tables and fenced examples excluded). (#3577) (#3599)

  • A terminal session now follows the workstream your repo says is active — with .planning/active-workstream naming a workstream, any invocation that had never run workstream use silently resolved the flat .planning/ tree instead: it misreported milestone, phase and progress on reads, and wrote to the superseded flat STATE.md. Because the stale tree is well-formed, nothing warned, and the documented workaround was to prepend GSD_WORKSTREAM= or --ws to every command. A session that has never set its own pointer now inherits the repo marker. Session isolation is unchanged — a session that owns a pointer is never repointed — and the two workstream-mode fail-safe guards now say whether a marker exists but failed to resolve, instead of claiming none is set. Note that clearing a session's pointer returns it to inheriting the marker rather than forcing flat mode. (#3579) (#3616)

  • branching_strategy: "phase"/"milestone" once again lands the first strategy-scoped commit on the strategy branch — gsd-tools query commit now creates and switches to a brand-new phase/milestone branch (restoring the #1278 intent), instead of creating it without switching and leaving the commit on the base branch. The #3079 protection is preserved: an already-existing strategy branch is still never silently switched to (it warns and commits on the current branch). The first fresh create is now logged to stderr instead of being silent, and the misleading "already exists" warning no longer recurs on every subsequent commit once HEAD is on the strategy branch. (#3207) (#3363)

  • A file belonging to another phase no longer blocks the phase you are in — sixteen scans (plus the single-pick fallback inside resolveVerificationFile) collected verification and UAT artifacts from a phase directory without checking they belonged to that phase, so a stray or copied file such as 04-VERIFICATION.md sitting in phase 03's directory contributed its status to phase 03. The worst case was not cosmetic: a stray file carrying gaps_found or human_needed pushed a blocker that flipped the UAT-passed predicate to false, and transition gates on that — so a leftover file could refuse to let a phase advance. Some scans could also claim the opposite, reporting verification passed on the strength of a file the phase does not own. All of them now check phase membership. Where a directory's own phase cannot be determined from its name, every file is still included, so no scan silently loses a phase's real blockers; where it can, a phase holding only another phase's report now correctly reports having none of its own rather than adopting it. (#3511) (#3535)

  • The build no longer requires Node 24: escapeRegex falls back to an in-file metachar escape when RegExp.escape is absent (#3498) — RegExp.escape is ES2026 (Node 24+), and src/pattern.cts called it unconditionally, so npm run build itself failed on Node 22 (gen-loop-host-contract consumes the module), breaking the gsd-test linux-node22 verification lane. The seam now prefers the built-in when present and falls back otherwise — still the single owner of escaping (#3212 invariant preserved). Behavior on Node 24+ is unchanged; match behavior below Node 24 is verified equivalent by regression tests that neuter RegExp.escape in a child process. (#3499)

  • Full-line # comments in .planning/STATE.md (and every frontmatter surface) now survive a mutating write — parseYamlRegion carries column-0 comments through to reconstructFrontmatter via a Symbol-keyed channel, and syncStateFrontmatter propagates that channel across its fresh-rebuild of the frontmatter object, so a comment like # NOTE: current_phase is hand-maintained is no longer silently destroyed on the next state verb. Comment-less frontmatter is unchanged; data identity (keys/values/arrays/nested) is preserved alongside the comments. (#3257) (#3387)

  • A plan SUMMARY whose frontmatter declares status: blocked is no longer counted as a completed plan. Previously both the progress counters written to STATE.md (state planned-phase / begin-phase / record-session) and the phase-plan-index read path paired PLAN and SUMMARY files by filename existence alone, so a blocked plan counted as done and was omitted from the incomplete list. Filename existence remains the fallback when a SUMMARY carries no status field, and status: halted summaries still count as completion records (a designed stop), so untouched projects are unaffected. (#3459)

  • Auto-chain phase completion now runs the same post-processing as a normal transition (#1526) — completing a phase via /gsd:execute-phase (auto-chain) previously skipped the transition workflow's graduation scan, session-continuity, project-reference, accumulated-context, and current-position updates, leaving project state different from a normal transition. execute-phase now delegates post-completion processing to the transition workflow (post-completion mode: skips re-verify + re-running phase.complete to avoid a double-write). Identity/standalone transition behavior is unchanged. (#3419)

  • Capability skill bodies are now documented as an instruction surface — the capability trust model previously grouped skills with inert assets as "non-executable" surfaces whose consent is lighter because they do not execute code. A skill body does not execute code; it instructs the agent that does. The docs now state that a capability's SKILL.md bodies reach your agent's instruction context verbatim and are not content-scanned, and capability authors are told the same on the authoring side. No behavior changed and no existing consent was invalidated. (#3247) (#3249)

  • A capability's ship:pre gate now actually blocks the ship — the ship preflight resolved every declared ship:pre gate but enforced only the built-in security and broken-windows capabilities, so any other capability's blocking: true gate was resolved, evaluable, and then silently dropped: a phase shipped past its own declared failing condition with nothing evaluated, nothing warned, and nothing logged. Preflight now dispatches every active gate generically — honoring each gate's own blocking and onError — matching the contract execute:wave:post, execute:post and plan:post already implement. (#3559) (#3608)

  • package-legitimacy-gate tests now locate the executor RULE 3 section by its heading inside the deviation rules block, so unrelated prompt edits can no longer redirect or silently defeat the package-install guardrail assertions (#3489)

  • Known-defect warnings that lived only in docs are now enforced checks — six failure modes that CONTEXT.md merely described are now caught automatically, including unbounded subprocesses that could hang a run indefinitely and an unscoped frontmatter read that could pick up a body line. Writing the checks surfaced nine live instances, all fixed. (#2896) (#3325)

  • The init phase-op, init plan-phase and init execute-phase queries no longer hand consumers a fully-formed absolute path for a REQUIREMENTS.md/STATE.md/ROADMAP.md that does not exist. Those three fields were built with a bare path.join and no existence check, so a non-null value was indistinguishable from the file actually being there — even as the conditional sibling fields in the same payload (patterns_path, context_path, ...) already returned null for absent files, and ultraplan-phase.md explicitly gates its REQUIREMENTS.md read on requirements_path is not null. Each of the three reading sites now returns null when the file is absent and its absolute path when present. The project/milestone-bootstrap and doc-ingest emitters that use these paths as write-targets for not-yet-created files are intentionally unchanged. (#3188) (#3430)

  • phase.complete no longer rewrites STATE.md frontmatter stopped_at with a stale body 'Stopped at:' line: the completion now refreshes the session continuity line it implies ('Phase N complete, ready to plan Phase N+1') and applies the standard field-preservation policy on its atomic commit path. state record-session no longer reports 'Stopped At' as updated when the value is already current. (#3491)

  • Cross-AI reviewer lanes now run on Windows: the declared bare CLI name is resolved through one shared PATH+PATHEXT lookup before spawning, so npm-installed .cmd/.bat shims start via cmd.exe mediation instead of failing with spawn ENOENT (#3275). (#3445)

  • Six STATE.md frontmatter fields whose preservation policy is declared in the field-classification table were not honored by the table-driven preservation pass — last_activity_desc, paused_at, current_phase, current_plan (preserve-when-unchanged) and milestone, milestone_name (preserve-if-placeholder). The pass now implements every declared row, so editing a preservation row is a one-row table edit as the table's own contract documents. Curated frontmatter values for paused_at / current_phase / current_plan now survive a body-only write (e.g. state update) even when the body carries a stale-but-present derived value — previously only an absent derived value triggered the fallback, so a stale body value silently overwrote the curated frontmatter value. last_activity_desc is now governed by a single rule (the table row) rather than a separate date-comparison guard that could disagree with it. (#3258) (#3447)

  • state writes no longer shrink progress.total_phases to the started-phase count when ROADMAP.md is absent — with no readable roadmap, every state command persisted the on-disk phase-directory count as the declared total (only phases that had started counted, so a 5-phase project read 50-100% complete with 3-4 phases unstarted); the stored frontmatter total now wins, with a warning, and state json reports the same preserved value. (#3573) (#3595)

  • ui-plan-gate no longer blocks planning on a UI-token match alone: the gate now requires static frontend evidence (a package.json UI-framework dependency or a component-framework file in the tree) before blocking, so a phase section naming a hyphenated repo like dashboard-financeiro no longer trips the gate in a repo with no frontend. The gate result also surfaces matchedToken/matchedLine so operators can see what triggered the flag. (#3451)

  • Windows/Claude Code: /gsd-update and re-running the installer now migrate stale .sh managed hook commands (gsd-validate-commit, gsd-graphify-update, gsd-session-state, gsd-phase-boundary) in settings.json/settings.local.json to the current bash-runner-omission format — removing the redundant nested bash that the pre-#580/#3393 shape spawns on every hook fire. Custom user hooks are never touched. (#3460)

  • gsd-tools query state update-progress no longer rewrites Progress to 0% after a milestone close — when the current-milestone phase scan finds zero plans (the post-archive state, where .planning/phases/ is empty), the command is now a no-op that leaves STATE.md unchanged, instead of mapping 0/0 to 0% and destroying the shipped [██████████] 100% record. The legitimate 0% case (plans exist, none summarized) still writes 0%. (#3233) (#3375)

  • STATE.md preservation now enforces every policy its own table declares — a field could be declared preserve-when-unchanged and quietly go unenforced, because the executor branched on field names rather than on the declared policy, so four of eight rows were honored by a weaker mechanism elsewhere and two policies had no implementation at all. Preservation is now dispatched from the classification table, a declared row nothing enforces fails loudly instead of silently, and a whitespace-only curated value is no longer treated as a real one. (#3468) (#3495)

  • state planned-phase now refreshes the Current Position Phase: line (the body source current_phase is re-derived from) instead of leaving a stale previous-phase line behind, so STATE.md frontmatter, body prose, and state json stay coherent; the --name argument is persisted into the Phase line and current_phase_name instead of being silently dropped. (#3490)

  • Executor dispatch prompts no longer list companion files as raw @-include lines that Claude Code never expands inside an Agent() prompt string. The orchestrator now build-time embeds execute-plan.md and its companion references (summary template, checkpoints, tdd, worktree-path-safety, executor-examples) into the dispatched gsd-executor prompt, so execute-plan-only steps (segment_execution, previous_phase_check, verification_failure_gate, update_codebase_map) actually reach executors instead of silently never running. (#3462)

  • roadmap validate now emits a V005 warning and exits non-zero when the active milestone's window is truncated — phase entries exist in ROADMAP.md but are excluded from the milestone's resolved section (e.g. an intervening version-bearing heading closes the window before its own Phase sections). Previously this passed silently with {"warnings":[]}. (#3444)

  • /gsd-plan-phase's §13a Decision Coverage Gate no longer reports false total-coverage failures when a decision's own body contains a bulleted cross-reference to a sibling decision — a bullet nested (indented) under an already-open decision, elaborating on how it relates to another decision, was previously indistinguishable from a malformed top-level declaration attempt. A single such bullet forced the whole coverage analysis to could-not-parse, discarding every decision that DID parse correctly and reporting covered: 0 even when every decision was, in fact, fully covered by the phase's plans. (#3169) (#3424)

  • ZCode installs now strip mcp__* tool grants from installed GSD subagents at install time. ZCode's dispatcher treats every mcp____* entry in an agent's tools: frontmatter as a required MCP server and hard-fails the subagent spawn (CONFIGURATION_ERROR) when it is not connected, so /gsd-quick --full and plan/execute-phase flows failed out of the box with zero MCP servers configured. Installed ZCode agents now declare only core tools; MCP tools remain available when servers are connected. Claude Code installs are unchanged. (#3483)

  • /gsd-sync-skills now refuses cross-runtime skill sync — skill content and directory layout are runtime-specific (the installer applies per-runtime converters/adapter headers/brand swaps/layout rules), and two runtimes alias another runtime's skills root, so a verbatim cross-runtime copy silently corrupted destination skills and could overwrite a runtime the user never named. sync now refuses any --to that differs from --from and points at the installer, keeping identity sync (--from == --to) as a no-op. (#3025) (#3404)

  • milestone.complete no longer records the wrong line as a release's accomplishment — the one-liner was extracted from the first bold text under the SUMMARY's first heading, so an incidental first heading (a rule list, deviation notes) could contribute Rule 1 - Bug or NeutralPath as the milestone's permanent accomplishment in MILESTONES.md. Extraction now anchors to a Summary/Overview/Accomplishments heading and falls back to empty when none is present. (#3170) (#3401)

  • ESLint now actually runs on 56 previously-unlinted source files — a file matching no files: glob was not linted-and-clean, it was skipped entirely while eslint . still exited 0. All of hooks/ and eslint-rules/ sat in that blind spot. A new drift guard fails the build if any tracked source file resolves to zero rules without a recorded reason, so the class cannot silently regrow. (#3059) (#3277)

  • gsd-tools validate health and validate consistency no longer flag sentinel phase directories (999.x backlog/interim, 0.x drafts) — the disk-vs-roadmap comparison now applies the isSentinelPhaseId guard that the phase commands already had. Sentinel ids are defined as never-on-roadmap, so a 999-interim directory previously produced a permanent spurious W007 ("Phase 999 exists on disk but not in ROADMAP.md", advice to add it to the roadmap or delete it — both wrong) and a spurious "Gap in phase numbering: N → 999". Real (non-sentinel) orphans and genuine numbering gaps still warn. (#3225) (#3371)

  • roadmap.analyze now reports the real phase count instead of a silent phase_count: 0 when a CLOSED milestone heading sits between the active milestone heading and its own phase-detail sections. A prior refactor (#3184) already added a scope discriminator so the empty result was distinguishable from a genuinely empty milestone; this closes the other half of the issue — the consuming resume gate (workflows/next.md Route 0) iterates .phases[], so an empty array silently disarmed the safety invariant regardless of the scope field. When the scoped window comes back empty, is non-COMPLETE scope, and phase directories exist on disk, the query re-scans the shipped-milestone-stripped document and populates the phase list while keeping scope non-COMPLETE so the result remains flagged as best-effort. (#3165) (#3428)

  • The 1.4.0 changelog entry for Cursor slash commands now credits the PR that shipped it — the entry describing gsd install --cursor writing .cursor/commands/ cited #803 (the Cline PR, which the adjacent entry cites correctly) instead of #805, so anyone tracing the Cursor commands surface landed in an unrelated change. (#2359) (#3252)

  • Settings no longer offer worktree isolation on runtimes that cannot honor it, and health warns before execution fails closed — previously /gsd-settings recommended "Yes" and persisted workflow.use_worktrees: true on every runtime, handing installs whose declared dispatch.isolation capability is none the exact value /gsd-execute-phase and /gsd-quick fail closed on. On those runtimes the Worktrees question now offers "No (Recommended)" / "Leave unchanged" (never an enabling option), warns when the config carries an inherited explicit true, and /gsd-health surfaces such a config as new warning W025 with a DEGRADED status before execution-time failure. Runtimes that declare harness-worktree or orchestrator-worktree are unaffected — the gate is the declared capability, never the runtime name. Both surfaces resolve isolation through the new inspect-dispatch-isolation query, a sentinel-free sibling of dispatch-isolation: the dispatch verb records its decision to the executor-isolation sentinel by design, which a read-only diagnostic must never trigger. The inspection verb rejects --force-isolation, --phase and --plan as usage errors rather than accepting and ignoring them — the recording verb applies --force-isolation after resolution, so silently ignoring it would hand the same argv two different answers. Both surfaces also distinguish "this runtime declares no isolation primitive" from "the capability could not be resolved", and say which one happened instead of reporting a resolver failure as a capability verdict. (#2486) (#2531)

  • The /gsd slash command in Pi now visibly renders its output (progress, errors) via Pi's ctx.ui.notify mechanism instead of a return value Pi silently discards. (#3485)

  • The optional pre-commit hook now actually checks command-alias drift — every guard in .githooks/pre-commit was inert: nine matched paths under the retired sdk/ tree and invoked npm scripts that no longer exist, and the tenth watched gitignored build outputs that git can never stage. Staging src/command-aliases.cts now runs check:alias-drift instead of passing silently. (#2725) (#3273)

  • init.progress no longer infers the next phase from stray out-of-order artifacts — a phase directory created out of order (e.g. a phase-9 UAT evidence file while roadmap phase 8 was still pending and unscaffolded) dragged the reported frontier forward, making init.progress skip Phase 8 and disagree with roadmap.analyze; the frontier is now derived from roadmap order, with artifacts as corroborating evidence only. (#3581) (#3603)

  • phase complete now selects the lowest genuinely-outstanding lower-numbered phase as next_phase instead of a merely-positionally-next higher phase heading, and keeps STATE.md frontmatter current_phase and current_phase_name paired (both describe the same phase) even for narrative-prose STATE.md files (#3482)

  • planning-config.md documented "light" as an allowed workflow.code_review_depth value, but config-set only accepts quick/standard/deep — the reference now matches the validator, pinned by a doc↔capability-registry parity test. The agent_skills row now also documents the array-of-strings form for assigning multiple skill sets to one agent type. (#3449)

  • Documentation no longer points at files that were renamed or deleted — docs/INVENTORY.md claimed its roster was anchored by six drift-control tests when five had been deleted, and the four translations named a seventh that the English file had already dropped. CONTEXT.md, VERSIONING.md, docs/CONFIGURATION.md and docs/skills/discovery-contract.md pointed at issue-NNN- test filenames and sdk/ paths that no longer exist, and VERSIONING.md described an SDK bundling step the release workflow does not perform. Most consequentially, docs/TESTING-SUITES.md instructed contributors to add drift acknowledgments to a file CONTRIBUTING.md says to never use — following it put the entry where the contributing guide forbids. (#3620) (#3658)

  • Workflow-backend waves (claude-orchestration, BETA) no longer strand executor commits on worktree-wf_* branches: the emitted Workflow script now returns each agent's worktree metadata, and the orchestrator records it into the wave manifest so the existing merge-and-cleanup step lands every plan's commits. Missing metadata now halts the wave loudly instead of reporting success with an empty worklist. (#3450)

  • Global Claude Code installs now load their referenced workflow context — gsd-core/workflows/*.md and other spec-tree files previously emitted @$HOME/.claude/... @-file-references, a form Claude Code's @-import resolver silently drops (only ~/ and absolute paths resolve). Every such reference now resolves on ~/, matching the already-working skill/command surface; double-quoted shell $HOME references are untouched. (#3544) (#3551)

  • A phase with more than one *-VERIFICATION.md no longer reports the wrong one — verification-report discovery took the alphabetically-first match, so an ad-hoc worksheet such as 03-CORRECTION-VERIFICATION.md beat the real 03-VERIFICATION.md sitting beside it and the phase could report missing while a passing report existed. Three further copies of the same lookup picked whichever file the filesystem happened to list first, making phase status and the reported verification_path vary between machines. All five now share one resolver that prefers the canonically-named report and is deterministic when it has to fall back. (#3357) (#3513)

  • gsd-review no longer creates empty gsd-review-context.md / gsd-review-research.md section files (or hangs waiting on input) when a phase has no CONTEXT/RESEARCH notes: the build_prompt guards now test the glob expansion itself instead of probing with ls, which the block's nullglob setting had made always-true. (#3454)

  • Milestone phase counts no longer drop every letter-named phase directory — getMilestonePhaseFilter now includes letter-named phase directories (Phase A:…Phase L:, GSD's own non-numeric phase convention per ADR-612) in milestone progress and plan counts. A greedy regex previously captured the whole hyphenated directory name (A-tool-output-contract was read as A-tool-output-contract instead of A), so every letter-named phase silently fell out of its milestone and the progress/plan totals were fabricated over whatever numeric directory happened to survive — a well-formed, plausible number that could even look correct at a phase boundary. Numeric and milestone-prefixed phases are unchanged. (#3213) (#3368)

  • phase-plan-index silently drops short-form depends_on references (e.g. ["01"]), collapsing every plan into wave 1. The planner template's two worked dependency examples taught exactly that broken short form; they now teach the full-form plan id (e.g. ["01-01"]) the file's own frontmatter comment and other examples already document, so newly authored plans keep resolvable dependency edges. Resolver-side short-form handling is tracked separately in #3473. (#3475)

  • Recorded why install materialization stays three loops, not one — an architecture decision for epic #2866 phase 6. Measuring the three sites showed they diverge in mechanism rather than duplicate each other, so unifying them would have broken a prune that structurally cannot delete user files. (#3574) (#3575)

  • Dev-dependency js-yaml bumped to the patched 4.3.1, resolving a high-severity quadratic-CPU advisory — the lockfile now pins the backported !!omap fix (GHSA-5p4m-2wfm-xmqj, CVSS 7.5), reachable via eslint. A non-breaking in-range bump (no overrides, no major bump, one package moved); production npm audit --omit=dev is unaffected (devDependency only). (#3238) (#3246)

  • Executor dispatch no longer blocks on a plugin-marketplace install — the compiled runtime library is a build artifact produced at publish time and gitignored, so a plugin or git-clone install materializes a tree that never has it. Every hook that required one of those modules did so without the existing self-heal build seam that the CLI entrypoint already calls, so the agent-isolation guard's missing-module error landed in its fail-closed catch and was reported as could not read or resolve dispatch-isolation configuration — blocking every gsd-executor dispatch from the first dispatch of a session, while the statusline and update-check worker crashed at module load on the same tree. All seven affected hook files now self-heal first: the isolation guards surface the build seam's own actionable error instead of a misleading config message and stay fail-closed, and the cosmetic hooks degrade quietly rather than taking down the prompt. The guards also now emit a machine-readable reason_code alongside the human message. Installs from npm are unaffected — the seam's already-built fast path returns immediately. (#3582) (#3629)

  • npm-global installs can now actually fail the agents-installed gate — checkAgentsInstalled resolved the claude agents directory relative to its own install location, so an npm-global install validated the package's bundled agents/ against itself and agents_installed could never be false, silently disabling the halt/warn gates in new-project and new-milestone. When the install-relative path lies inside a node_modules tree the claude runtime now resolves getGlobalConfigDir('claude')/agents like every other runtime, honouring CLAUDE_CONFIG_DIR; repo runs and runtime-config-dir installs are unchanged, and the GSD_AGENTS_DIR override stays priority 1. (#3203) (#3229)

  • Plan files with Windows-style CRLF line endings now correctly enforce their must_haves contract — truths, artifacts, key_links, and prohibitions blocks previously parsed to an empty list on any CRLF-authored plan file, silently degrading goal-backward verification to LLM-derived truths instead of the authored contract, with no error surfaced for the most common failure shape. (#3360) (#3420)

  • Installing GSD for Claude at both global and local scope no longer silently hides your project's specs. Claude Code always resolves the personal skill over the project command, so a project with a local install previously ran the global workflow specs with no warning. The install now prints which scope wins and /gsd-health surfaces the same as diagnostic W028; at global scope, the winning skill's workflow reference now resolves your project's own specs first when present. (#2218) (#3537)

  • Installer no longer crashes when a source file disappears mid-copy. copyWithPathReplacement used to throw an unhandled ENOENT if a listed workflow/command file was deleted between its directory listing and the actual read — a rare filesystem race that could abort an entire install. It now skips the vanished file and continues installing everything else. (#3333) (#3341)

  • verify plan-structure now recognizes task child elements that carry attributes on their opening tag (e.g. ), so plans annotating verify mode (auto vs human) or other child-tag attributes no longer produce false "missing " / "missing " / etc. warnings. Bare tags continue to validate exactly as before. (#3433)

Security

  • MCP server configs are now explicitly flagged as unconfined in the capability consent prompt — a capability's MCP servers can legitimately point at commands, args, env, and working directories anywhere on the machine (unlike its hooks, which are confined to the installed bundle), and the consent disclosure now says so plainly for every spawned server instead of leaving the asymmetry unstated. (#3515) (#3517)
  • Hook security hardening — shared injection patterns + fail-closed force-add guard — the prompt-injection pattern list is now one shared module used by both the write-guard and the read-scanner hooks, so the two surfaces can no longer silently drift apart; and the opt-in workflow guard's force-add block on agent branches now fails closed on internal error instead of silently allowing. (#3504) (#3510)
  • Capability installs no longer fetch from internal hosts, and unpinned installs say so in the consent prompt — the URL importer refuses loopback/link-local/metadata hosts (including the cloud metadata addresses and localhost) before any bytes leave, and an http:// tarball URL fails with a clear https-only reason instead of a raw protocol error. Installs without an integrity pin now show a distinct 'NO PINNED HASH — staged unverified' line in the consent disclosure. (#3514) (#3516)
  • Path validation no longer accepts a symbolic link whose target is missing — validatePath canonicalizes a path with realpath, and for a path that does not exist yet it fell back to canonicalizing the parent directory instead. A symlink inside the project pointing at a non-existent location outside it took that fallback and was accepted, while a symlink pointing at an existing outside location was correctly refused — a difference an attacker could use to test whether arbitrary absolute paths exist. Such a link is now refused outright. The same fallback also compared an uncanonicalized path against a canonicalized base when several leading directories were missing, wrongly refusing legitimate not-yet-created paths on any non-canonical working directory (every macOS temp directory, for one); it now canonicalizes from the nearest existing ancestor. (#3493) (#3506)
  • verify key-links can no longer be hung by a plan's key_links pattern — the pattern was compiled straight from plan frontmatter with a backtracking engine and tested against whole file contents, so a nested-quantifier pattern such as (a+)+$ pinned a CPU core indefinitely and stalled any verify-phase run that reached it. Untrusted patterns now execute on the RE2 engine, whose match time is linear in the input length, and a pattern that cannot be compiled is refused outright rather than guessed at — a refused pattern can never report a match. (#3477) (#3496)
  • A reviewer lane's invoke fields are now disclosed at install and bound to the consent signature — an installed third-party capability could declare env on its reviewer lane and have those variables applied to the spawned reviewer process without that ever appearing in the consent prompt, which makes NODE_OPTIONS=--require ./evil.js an undisclosed code-execution path. Overlay reviewer lanes only became executable in #3062, and the disclosure did not move with them. The consent prompt now shows each env key and value (highlighting names that are execution primitives) and the manifest's own defaultHost, which the runtime uses whenever the configured host key resolves to nothing — previously such a lane displayed "(unresolved)" while still sending plan and review text to the address the manifest chose. Every other declared invoke field is covered by a residual, so a future field cannot repeat this. No already-installed capability is re-prompted by this change — consent is bound to the bundle's content hash, not to the disclosure signature. What changes is that an upgrade which edits any declared invoke field now counts as an executable-surface change and asks for consent again, where before it could alter what the lane runs in silence. (#2483) (#2493)
  • verify key-links no longer reads files outside the project — from: and to: were taken verbatim from plan frontmatter and resolved with path.join(cwd, …), which normalizes ../ rather than rejecting it, so a plan carried in an untrusted repository could name any file the process could read and learn from the reported result whether a supplied pattern matched its contents. Both paths now resolve through the project's realpath-based confinement seam; a path that escapes is refused without being read, reported as path_rejected, and never counts as verified. (#3493) (#3506)

[1.10.0] - 2026-08-08

Added

  • A blocking catastrophic-shrink guard now protects curated .planning/ artifacts from whole-file Write clobbers — the new PreToolUse hook gsd-write-guard.js compares the pending Write payload against the file on disk and hard-blocks (decision: 'block', exit 2) when the payload would collapse ROADMAP.md, a milestone roadmap (.planning/milestones/*-ROADMAP.md), or STATE.md below 40% of its current line count (files under 40 lines are exempt). The check is stateless per Write — each payload is compared against the file's current on-disk count, so the single-shot collapse is blocked while iterative erosion across individually-tolerated Writes is a disclosed non-goal. This is fix 3 of #973 — the only one enforced by code rather than by instructions to a model: fixes 1 and 2 (PR #989) are prose an agent may reason past and protect only audited agents, and #973 records an agent reading the existing advisory and reasoning past it while destroying three milestones of roadmap history. The guarantee is bounded, and the bound is worth stating precisely: this blocks accidental and single-shot collapse, and does not stop a determined agent — the sentinel below is a plain file, so an agent that would reason past an advisory can arm one with a single Bash call it is already permitted to make. What ships is the conversion of ignore a sentence into take one deliberate, path-bound, single-use, auditable action — a real improvement against the confused-agent threat #973 records, not a defense against an evader. Legitimate milestone resets bypass the guard mechanically: the workflow step writes the target's path into the single-use sentinel .planning/.gsd-allow-shrink, which the guard verifies (fresh, path-bound) and consumes — a per-step env var cannot reach a PreToolUse hook, so the sentinel is the transport code consults rather than prose an agent obeys; interactively, GSD_ALLOW_PLANNING_SHRINK=1 still bypasses once. Both are named in the block message. Registered on the Claude plugin surface, the settings-json runtimes, Kimi, and the OpenCode/Kilo plugin buses; on Kimi the guard normalizes the native payload shape (WriteFile, path) and writes its block reason to stderr, so it engages there from day one (the #2304 dormancy class). (#2255) (#2301)

  • New how-to: Take over a capability, reviewer lane, or EoS integration. The capability ecosystem documented a complete forward lifecycle — develop, publish, version, import, update, remove, turn off — but nothing covering a change of maintainer for an entry that already exists. There is no gsd capability transfer command and no rename tooling, and docs/registries/README.md specifies submission and the narrow removal policy but never transfer, so a would-be adopter had no documented path and a reviewing maintainer had no stated bar.

    The guide defines four takeover modes and the PR shape each one takes. T1 — consensual handoff keeps the id and the entry, changes only repo / author / install / uninstall, and requires a permalink to the outgoing author's public handoff comment in the entry's Discussion. That permalink is mandatory rather than advisory because entry-update authorship is not verified anywhere: scripts/registry-schema.cjs and npm run validate:registry check an entry's shape, not who is changing it, and the registry-entry PR template's "repo links to a repository I own" is a self-attestation — so a PR repointing repo and author at an unrelated account passes every automated gate, and the reviewing maintainer is the only control. T2 — adoption fork takes a new id, opens a new Discussion, and leaves the original entry untouched, because the narrow removal policy removes an entry only for illegal content, malware, spam, or a dead link and never for staleness or abandonment: an abandoned-but-working entry can never be reclaimed, so adoption is always additive and the original id stays taken. T3 — first-party absorption routes through approved-feature plus an ADR, lands under capabilities/<id>/capability.json per ADR-894, annotates rather than deletes the registry entry, and requires a migration note telling existing users to gsd capability remove <old-id> first — config keys are exclusive to one capability and skill/agent stems must be unique, so a first-party capability that collides with an installed overlay wins silently, leaving the user running code they did not think they were running. T4 — retirement is restricted to the four narrow grounds with evidence in the PR body.

    Around the modes the guide adds an evidence pack, license and reserved-prefix and consent gates, a per-surface snapshot of the inherited user-visible contract (loopExtensionPoints / hookKinds / configKeys / requires / runtimeCompat for Feature Capabilities; slug / flags / reviewsSection uniqueness across the merged first-party and overlay set for reviewer lanes; protocolVersion, interfacePoints, profile and the eight ADR-1239 axes for EoS integrations), and an install-continuity checklist covering the failure modes that break existing consumers — id continuity, since consent is stored per (realpath(projectRoot), capability id) and an id change re-prompts every installed project and orphans the update path; re-stating integrity and provenance after a rebuild under new ownership; holding the executable-surface set steady so the handoff is not itself a consent event; and not narrowing engines.gsd without a matching compatVersions row. Post-takeover obligations note that a Release must be cut under the new repo, since there is no re-registration and both the shields badge and the releases/latest permalink render live from repo. The two enforcement gaps — unverified entry-update authorship, and the absence of any id migration path — are stated explicitly in the guide so the process is not mistaken for something CI verifies.

    Fixed alongside: .github/PULL_REQUEST_TEMPLATE/registry-entry.md directed contributors to file their Discussion in a Registry category that does not exist. docs/registries/README.md names the category EoS Registry and explicitly notes the name is misleading because it carries threads for all three catalogs. Because discussion is a required field, the thread must exist before the entry's PR is opened — so a contributor following the template stalled at the first required step of the submission process. (#2999) (#3000)

  • MCP-capable hosts can now browse GSD's own workflows, references, and commands through the companion server — the workflow and reference tree is served as MCP resources and the /gsd-* commands as MCP prompts, so a host lists and fetches just the content it needs instead of relying on the copied file tree alone. Workflow resources arrive composed exactly as the installer writes them; the file-copy install is unchanged and stays the default on every runtime. (#3072) (#3083)

  • UAT checkpoint frames now cover 9 more languages — response_language values of Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, or Indonesian render a localized checkpoint banner/instruction instead of silently falling back to the English frame (#2530). (#2564)

  • Agent-dispatch isolation guard. An executor subagent dispatch that would run outside an isolated worktree is now hard-blocked when this dispatch's resolved isolation is harness-worktree, closing the #260-class main-checkout write path a prose-only instruction could silently skip — while correctly leaving legitimate sequential or orchestrator-managed dispatches (project opt-out, submodule intersection, diverged-base auto-degrade) untouched, since the guard reads the workflow's own resolved per-dispatch decision instead of a host's general capability. Covers a missing isolation="worktree" parameter on the Agent()/Task() dispatch, as well as a subagentStart dispatch whose session is not actually running in an isolated worktree, verified structurally since a session-level worktree flag carries no per-dispatch isolation parameter to check. (#3045) (#3069)

  • Unresolved deferred-items.md entries now reach the milestone-close audit. auditOpenArtifacts gains deferred_items as a ninth scanned category, so an out-of-scope discovery a phase agent correctly recorded rather than fixed surfaces in /gsd-complete-milestone's pre-close report alongside the other eight, and the existing [R] Resolve / [A] Acknowledge / [C] Cancel prompt applies to it. #2287 made the file readable at the phase boundary (audit-uat, /gsd-progress check 7); one boundary up it was still invisible, and phase directories archive to milestones/vX.Y-phases/ by default (#1871), so an unresolved entry left the live tree without ever being triaged. The resolved/unresolved predicate is not reimplemented — the scanner lazily requires uat.cjs's exported parseDeferredItems, so both boundaries agree by construction about what "open" means. Behaviour change worth noting before you upgrade: a project carrying unresolved deferred items will now see the [R]/[A]/[C] prompt at milestone close where close previously proceeded silently. That is the intended correction, but it surfaces pre-existing debt on the first run. (#2646) (#2983)

  • Reviewer lanes can now be listed for discovery — a new how-to walks lane authors through publishing to the Reviewer Lane Registry: which of the three catalogs applies, opening the required discussion thread first, the fields that reject entries most often, and why registering once means GitHub Releases become the update channel. (#2904) (#2917)

  • Workflow markdown can now fragmentize into per-runtime-composed sections. Authors can mark sections of a workflow file with in-file <!-- gsd:section id= when= --> markers; per-runtime emission strips the markers and composes the marked sections back byte-identical-or-smaller, piloted on execute-phase.md. (#2930) (#2972)

  • gsd_run query context-predicates — targeted lookups against the CONTEXT.md fact-store — search predicates live by class, id prefix, or substring instead of reading the whole file, with a CI-guarded docs/CONTEXT-INDEX.json index kept in sync automatically. (#2928) (#2938)

Changed

  • Agent definitions now share the workflow fragment pipeline — a <!-- gsd:section --> marker in an agents/*.md file is stripped at install time on every emission path instead of shipping verbatim into the runtime, and the largest agents move their reference material into gsd-core/references/ so they regain headroom under their size caps. (#2995) (#3058)
  • Windsurf command install no longer fails on an oversized description, and emitted artifacts are now checked against their host's byte limit — the Windsurf workflow converter truncates a long description instead of throwing, matching the bound its sibling skill converter already applied, and a new per-runtime cap gate measures what each runtime actually receives rather than what the source files weigh. (#2931) (#2984)
  • /gsd:debug now initializes in one round-trip instead of three — the workflow previously made three separate gsd-tools calls to assemble its context (state.load, resolve-model, and config-get workflow.tdd_mode); it now makes a single init.debug call carrying the same resolved values. (#3149) (#3154)
  • Documented the widened when= vocabulary and the per-workflow section manifest. docs/reference/workflow-fragments.md now lists all 14 closed when= atoms, the two admission gates a new atom must clear, the manifest artifact's per-workflow {workflows:{<name>:[...]}} shape (absent key = degraded, empty array = computed-empty), and that boolean-flag membership in InvocationFacts.flags is token-presence, not value-truthiness. Also added the missing --reset-phase-numbers flag to /gsd-new-milestone's argument-hint. (#2992) (#3013)
  • The extracted workflow fragment tree is now inventoried — the 47 step files and 13 mode files that live under gsd-core/workflows/<workflow>/ were invisible to docs/INVENTORY-MANIFEST.json, so a new one could ship with no row and no gate firing. They now have their own manifest families. (#2996) (#3061)
  • Workflow guidance now loads only the branch your invocation actually took — thirteen more large workflows moved onto the fragment model, so running /gsd-code-review without --fix no longer loads the fix-dispatch branch, /gsd-progress without --forensic no longer loads the forensic audit, and so on across every migrated workflow. (#2994) (#3030)
  • Flag-gated workflow guidance is now actually loaded on demand — /gsd-plan-phase reads its PRD-express, ADR-ingest, reviews-prerequisite, research-only and chunked-planning guidance only when the matching flag or config is active, instead of always inlining all six branches. This also repairs /gsd-execute-phase --wave, whose section gating never took effect because the workflow never forwarded the flag to the init bundle, so wave-filtering guidance was silently skipped on every run. (#2993) (#3019)
  • Budget-aware content composition is now a shared context-composer seam — the priority-ordered trimming that kept cross-AI review prompts inside a model's context window was locked inside that one pipeline. It is now a reusable seam with an injectable budget unit, so later work can right-size what ships to each runtime. Review-prompt output is unchanged, proven byte-for-byte against a 50-case corpus captured from the previous implementation. (#2929) (#2958)
  • /gsd-execute-phase now loads only the branch guidance your invocation actually uses. Running it without --wave no longer pulls the wave-filtering instructions into context, and a plain integer phase no longer loads the decimal-phase gap-closure branch. The init bundle reports which sections apply to each invocation and the workflow reads only those, so the orchestrator spends its context on the path it is actually taking. (#2932) (#2987)

Fixed

  • Installing or updating GSD no longer destroys a user-authored package.json at the runtime config root — the CommonJS marker ({"type":"commonjs"}) that pins GSD's staged .js scripts is now written into the directories GSD itself fills (hooks/, and plugins//extensions/ for the runtimes with a native plugin adapter) instead of over <configRoot>/package.json. Previously every install and every /gsd-update re-install overwrote that file unconditionally — no existence check, no merge, no backup — permanently destroying any name, type, dependencies, or scripts the user or host tool had put there. This hit 11 runtimes and was worst on OpenCode and Kilo, where the config-root package.json is the documented place to declare local-plugin npm dependencies. Install and uninstall now share one ownership predicate, so a package.json GSD did not write is never overwritten and never removed; uninstall still retires the marker left behind by earlier versions. (#2544) (#2593)

  • workstream progress / workstream status / workstream list no longer report a workstream's CURRENT milestone as "milestone complete" / 100% while phases in that milestone are unstarted, in progress, or failing verification. Three coupled defects in the shared inventory derivation are fixed. (1) The shipped signal was project-lifetime rather than milestone-scoped — workstreamMilestoneShipped() returned true if ANY *-ROADMAP.md snapshot existed or SHIPPED appeared anywhere in ROADMAP.md, and since every previously shipped milestone leaves a permanent collapsed <summary>✅ … SHIPPED</summary> block, any workstream that had ever shipped was pinned to "milestone complete" forever (an over-correction from #1913). It now requires the CURRENT version's archived milestones/<version>-ROADMAP.md snapshot, or the current milestone's own ROADMAP line marked shipped; REQUIREMENTS snapshots are deliberately not accepted because they can be written at milestone start. (2) The completion percentage silently excluded phases declared for the current milestone but never scaffolded, while completed PRIOR-milestone phase directories inflated the numerator — both numerator and denominator are now scoped to the current milestone, whose phase set is read from the ROADMAP ## Progress table (which lists phases with no directory) via the canonical findTableWithColumns parser, with the current version taken from the workstream STATE.md milestone: field rather than ROADMAP in-progress markers, which can be stale. (3) Phase completeness ignored the verification verdict — a phase with SUMMARY count ≥ PLAN count now counts as in_progress rather than complete when its verdict is an explicit failing one (gaps_found/human_needed); missing/unknown/stale are intentionally untouched so verifier-disabled projects do not regress to never-complete.

    Two further denominator gaps are closed. A phase declared as a ## Progress table row with no ### Phase N heading was dropped by the heading-only count even when other headings existed (the regex counts 1 for a "1 heading + 1 table-only" roadmap), and milestone scoping could not cover it because a flat Progress table carries no per-phase milestone attribution — so greenfield and single-milestone projects kept the faulty count. When scoping cannot engage, the denominator is now the union of the Progress table's declared phase numbers and the phase directories, so neither source can shrink it. Separately, a sub-phase directory inserted mid-milestone (30.1-… under a table-declared phase 30) has no table row of its own and previously had no milestone attribution at all; it now inherits its parent phase's milestone and joins BOTH sides of the calculation — numerator-only would let completed_phases exceed a denominator that never counted it and cap back to 100%, reintroducing the reported defect. Attribution is one-directional (a sub-phase counts only when its parent is in the current milestone), so a follow-up created in a later milestone under an older parent is excluded rather than misattributed.

    Membership and the denominator are derived from a single canonical phase-key surface, promoted to the phase-id owner module as phaseKeyFromToken / phaseKeyFromDir / phaseKeyFromProse / parentPhaseKey (previously a private pair in state.cts). Deriving one side of a comparison with a bespoke regex was itself a way to reproduce this issue: a padded | 01. … | table row never matched a 1-slug directory, and a project-code-prefixed PROJ-05-… directory matched nothing at all — each silently zeroing or pinning the rollup while phases[] reported the opposite. Directory membership additionally consults getMilestonePhaseFilter, the module that owns milestone-phase filtering, which now accepts a workstream name so its planningDir resolution can target .planning/workstreams/<ws>/ (a loop over workstreams cannot express that through GSD_WORKSTREAM) and exposes versionScoped so its phase count is never mistaken for a current-milestone denominator on an unversioned roadmap. "Milestone shipped" detection likewise moved to that module as isMilestoneShippedInRoadmap: heading and <summary> lines only — a bullet such as - [x] 03-01: ship the v2.0 login endpoint ✅ is prose about a phase, not a milestone verdict — with the version token boundary-matched so a shipped v2.0.1 heading cannot close v2.0. A ROADMAP row whose Milestone cell is blank or malformed now stays in the denominator instead of vanishing from both sides, and a stale directory colliding on phase number with a current one (Bug #2445's scenario) counts once; the Builder asserts completed_phases <= denominator and throws rather than letting Math.min round a contradiction up to 100%. getMilestonePhaseFilter still applies its own internal phase-id normaliser for directory matching rather than routing through phase-id.cts; the two signals are OR'd, so a divergence can only widen membership, never narrow it — but they remain two normalisers, not one.

    A declared-but-empty current milestone is scoped rather than treated as unscoped. STATE.md's milestone: field updates the moment /gsd-new-milestone writes the heading, while the ## Progress table and phase sections land later; in that window nothing attributes a phase to the current milestone, scoping switched off entirely, and the fallback counted the project's whole phase history as both numerator and denominator — reporting 100% for a milestone with no work done, the same symptom by a different route. Three witnesses now distinguish that state, each covering a ROADMAP shape the others miss: getMilestonePhaseFilter gained versionSectionFound (the milestone's section exists but declares no phases — versionScoped cannot answer this, because a located-but-empty section falls through to the zero-count pass-all degrade that resets it), the existing missingExplicitVersion (a versioned roadmap with no section for this version), and a Progress table attributing every row to another milestone. A ROADMAP that attributes no versions anywhere matches none of them — its rows parse unattributed and stay in the current milestone — so free-form legacy projects keep their whole-roadmap count instead of regressing to 0%. Within an empty milestone, membership inverts: a phase directory belongs unless another milestone's row claims it, so a phase scaffolded before the roadmap catches up is counted rather than dropped from both sides. Scoping is now stated by the caller (milestoneScoped) instead of inferred from currentMilestonePhaseCount > 0, which could not represent "scoped and legitimately zero-phase".

    status is cross-validated against the milestone's own artifacts, not asserted from the shipped marker alone. Scoping the marker to the current milestone stopped a PRIOR milestone pinning status to "milestone complete", but the marker was still echoed as fact for the current one — so a single payload could report status: "milestone complete" beside progress_percent: 67, which is this issue's own symptom reached through status. The two shipped signals are now distinguished and cross-checked at different strengths, because one check cannot serve both. A heading signal (an operator-typed ✅ SHIPPED in the LIVE roadmap) is refused when the milestone's completion ratio is short, which also catches phases declared but never scaffolded. A snapshot signal (milestones/<version>-ROADMAP.md) is not gated on that ratio alone: the milestone complete run that writes it also moves the milestone's phase directories into milestones/<version>-phases/ while copying — never truncating — the live ROADMAP, so a CLEAN archive reads 0/N by construction and a bare ratio gate would strip "milestone complete" from every archived milestone in every project. But a phase directory still present under phases/ means the archive is not clean — a phase was added or reopened after it, reachable because milestone complete does not advance STATE.md's milestone: field — and once that is true the ratio is meaningful again, so the snapshot check is the conjunction of the two. The legacy project-lifetime fallback is ungated by signal, as before. The cross-check as a whole engages only when milestone scoping is active, for ALL three signals and not just legacy: with scoping off the denominator is the whole-roadmap count and membership is everything, so there is no current-milestone artifact set to check a current-milestone claim against. When a marker is refused, the STATE.md field is not accepted as a fallback claim of completion either — in this window it commonly asserts the same thing — so against contradicting artifacts neither source can report the milestone complete.

    Greenfield roadmaps without a versioned Progress table, and projects whose current milestone version cannot be determined, keep the previous behaviour. (#2562)

    Behaviour change for consumers of the inventory JSON: roadmap_phase_count, completed_phases and progress_percent now describe the workstream's CURRENT milestone rather than its lifetime, and there is no schema signal marking the change. Anything reading those fields — including getOtherActiveWorkstreamInventories, which filters completed workstreams out of the active list — sees real movement: a post-v1.0 workstream that reported milestone complete / 100% will now report its actual in-flight progress. The inventory also gains milestone_shipped_unverified: true when a shipped marker fired for the current milestone but its artifacts contradicted it. It is distinct from status_conflict, which continues to report only the derived-vs-STATE.md-field disagreement. workstream list, workstream status and workstream progress all project the new field, so a refused marker is visible at the CLI rather than collapsing silently into a fallback status. (#2588)

  • Cursor now shows each GSD workflow once in the slash menu while keeping skills available for contextual model invocation — Upgrades safely retire manifest-managed legacy commands/gsd-*.md duplicates, back up modified managed copies, and preserve unknown user-authored commands. (#2812)

  • Deleting a phase's verification report can no longer inflate workstream completion once that report has been seen. Removing a *-VERIFICATION.md file after a failing gaps_found or human_needed verdict was recorded used to be indistinguishable from never having verified the phase at all, so completed_phases and progress_percent silently rose. workstream status/list/progress now remember the last real verdict observed per phase in a new .verification-ledger.json file alongside each workstream's STATE.md, so a failing verdict a prior read has already seen can't be erased by deleting its report. This adds a small write side effect to those previously read-only commands, and the file is a new tracked artifact under .planning/workstreams/<name>/ for projects that commit their planning docs.

    The ledger fails closed, not open: once a workstream has adopted it (the ledger file exists), a phase with no remembered entry — including one whose ledger entry can't be read because the file is corrupt or unreadable — is treated as not-yet-verified-and-blocking, not as safe-to-complete. A workstream that has never used the verifier is untouched (no ledger file is ever created for it), which is what keeps every existing project from dropping to in_progress the moment this ships.

    Three limitations, disclosed rather than silently left: this is prospective only — a phase verified and its report deleted before this fix ships has no ledger entry and can't be recovered retroactively. Deleting the ledger file itself, not just the report, still returns that phase to pre-adoption behavior; this is inherent to any design where a wholly-absent ledger must be safe (the alternative is gating every never-verified phase in every project on upgrade), and is not something ledger design alone can close. And the ledger is not tamper-proof: anyone with write access to .planning/workstreams/<name>/.verification-ledger.json can hand-edit an entry to "passed" and the remembered value is trusted indefinitely — this is a different and arguably worse way to inflate completion than deleting the ledger (which at least resets to a visibly pre-adoption, untracked state), since an edited entry looks like genuine durable history. Integrity-checking the ledger's own content is out of scope for this fix. (#2645) (#3016)

  • graphify query --budget <N> now reports whether the budget was met — the response carries budget_met and budget_estimate when a budget is requested. The estimate measures the response as emitted (the pretty-printed payload the caller is handed, wrapper keys included), so budget_met is a claim about the bytes you actually receive rather than about a smaller internal form. Seeds are retained unconditionally, so the seed set is a floor the edge-tier reduction cannot go below; previously a request for 500 tokens could return a ~119k-token payload with no signal that the budget was missed. The tier loop also now recomputes reachability and the estimate after each tier removal, so it stops as soon as the pruned result fits instead of dropping the next, higher-confidence tier unnecessarily. --budget 0, which the CLI accepts and forwards, is now honored as a (necessarily unmeetable, reported) budget instead of being silently treated as no budget. (#2738) (#2819)

  • /gsd-spec-phase now actually runs its edge-completeness and prohibition-completeness probes — every gate-passed path reaches Step 5.5, and Step 5.5 now falls through to Step 5.6 instead of jumping past it. Previously all four gate-passed transitions went straight to SPEC generation and Step 5.5's own "all edges resolved" gate skipped the prohibition probe, so a SPEC could ship with an empty Edge Coverage section, an empty Prohibitions section, or both — and a weaker model following the prose literally would never notice. Since the probes are what carry must-NOT constraints and data-shape edges into must_haves, the plan and the verifier inherited the gap too. (#2733) (#2779)

  • The api-coverage detector's negation-suppression check no longer takes superlinear time on long prose, which was hanging the verification gate (#2784, #3127). It also no longer fails to suppress a negated pair ("this phase integrates no external API") when the negation sits in any clause other than the first on a line — a latent offset bug made negation suppression a no-op for every clause after the first. (#3124)

  • execute-phase now warns when local commits are ahead of origin — forking the phase branch from origin/$DEFAULT_BRANCH silently missed unpushed local commits (e.g. plan/research docs). A loud WARNING now names the divergence before the fork. (#2639) (#2981)

  • broken-windows capability no longer claims ship blocking is unconditional — the description now states that /gsd-ship blocking applies only when workflow.windows_enforce is enabled (default false); ledger tracking is unaffected. (#2787) (#2814)

  • Worktree safety gates no longer report success when they could not check — a git command that timed out (a locked index, a stalled network mount) was treated the same as "this is not a git repository", so the base-divergence gate answered "safe to run parallel worktrees" without ever resolving the fork base, and worktree-context resolution silently fell back to the current directory. The base-divergence gate now degrades to sequential execution instead of assuming safety. Worktree-context resolution still falls back to the current directory (there is no safer default), but now surfaces a loud warning that planning artifacts may be written to the wrong tree instead of silently trusting it. Worktree creation also no longer skips its root-confinement check when the caller omits the root. (#3050) (#3054)

  • roadmap.update-plan-progress no longer deletes hand-written annotations — bumping the plan count used to swallow the rest of the Plans line, silently deleting any prose a human wrote after the count. The verb now replaces only the count token and leaves trailing text intact. (#2853) (#2916)

  • detectApiIntegration no longer triggers on negated prose — a clause pairing an integration verb with an API noun but also containing a negation qualifier (no, not, without, neither, nor, etc.) is now suppressed. "This phase integrates no external API" no longer fires a false positive that halts verification. (#2784) (#3127)

  • Bug-report template version guidance corrected — the template pointed reporters at npm list -g, which does not track what /gsd-update installs into the runtime home. It now points at the gsd-file-manifest.json version field that the installer writes. (#2998) (#3100)

  • windows append/waive/fixed no longer destroy prose below the JSON ledger — the writer reconstructed the file from the parsed JSON ledger only, silently dropping any human-authored prose sections below the closing fence. The writer now preserves trailing prose across all write operations. (#2893) (#2975)

  • Gap-closure plans generated by /gsd-plan-phase --gaps now deterministically carry gap_closure: true — the planner's frontmatter validator previously only checked plans against a schema that never required this field, so a gap-closure plan could silently omit it and /gsd-execute-phase --gaps-only would then match zero plans with no error. (#2847) (#3018)

  • phase complete no longer advances next_phase into 999.x backlog headings — the roadmap heading scan (stage 2 of the next-phase cascade) accepted any higher-numbered heading without checking the sentinel convention, so a Phase 999.1: Backlog Item heading was treated as the next real phase. Sentinel phase ids (999.x backlog, 0.x drafts) are now skipped. (#2786) (#3130)

  • A split-parent phase marked complete in the ROADMAP is no longer permanently reported as current_phase — a phase split into sub-phases (parent kept as shared context, zero plans by design) was stuck as researched because the roadmap-checkbox override required completion.phase_complete (always false for zero-plan phases). The override now fires for zero-plan phases when the roadmap checkbox is checked. (#3033) (#3114)

  • Dispatch flattening now honors the declared nesting depth budget, so runtimes that cannot host a backgrounded orchestrator plus a delegated leaf run inline instead of producing an unsupported depth-2 tree — shouldFlattenDispatch checked only the two background booleans, so a host advertising maxDepth:1 was told it may background, which under Codex MultiAgent V2 produced a depth-2 orchestration tree the declared contract forbids. The decision now also requires nested + a full subagent toolkit + a depth budget greater than 1 or unbounded, reusing the convention already in degradationFor and _normalizeDispatchCallSpan. Runtimes lacking any of those — codex at maxDepth:1, kimi with nested:false, kimi-code with a built-in-only toolkit — now correctly run inline, the safer path that keeps worktree isolation and verification in force; only cursor remains background-eligible. (#2939) (#3063)

  • pi installs no longer trigger pi's deprecated-directory startup warning, respect PI_CODING_AGENT_DIR, and never lose custom files during an update — the shared hook bundle now installs to gsd-hooks/ instead of hooks/ (which pi reserves for its own deprecated extension location and warns about on every startup), with an upgrade migration retiring the old directory; pi's own PI_CODING_AGENT_DIR override is now honored when resolving where GSD writes; and /gsd-update's custom-file detection now recognizes the renamed bundle, so user files placed under it are backed up before a clean install instead of being silently wiped. (#3023) (#3175)

  • current_phase no longer rewinds to an archived phase when STATE.md carries a historical Phase: line — a stale Phase: or **Phase:** line in an archive section of a long-lived STATE.md silently overwrote current_phase on every state write, and because current_phase drives gsd-progress and --next routing the rewind sent work to the wrong phase. Phase extraction is now scoped to the ## Current Position section (mirroring the existing ## Session scoping for Stopped At / Paused At). (#2956) (#2961)

  • /gsd-update --sync no longer fails with MODULE_NOT_FOUND — the sync-skills workflow shelled out to gsd-core/bin/install.js, which the installer never copies. Now uses gsd-tools query skills-root (which IS shipped) to resolve skills roots. (#3024) (#3195)

  • Pi no longer emits a typebox unavailable warning at every startup — the warning fired because the Pi adapter attempts to require('typebox') (not a gsd-core dependency) and falls back to a plain JSON-Schema object on every startup. The fallback is the normal path; the warning is now suppressed. (#3022) (#3111)

  • fish_add_path no longer skips a directory whose name starts with a dash — fish parses a leading-dash token as an option, so the suggested command silently added nothing; it now passes the end-of-options separator. Also fixes a config.toml written unparseable when a value carried a newline or NUL, an installer PATH hint that printed a header with nothing under it, and a reviewer lane that crashed instead of degrading when its conversation cache file held the literal null. (#3118) (#3124)

  • A halted plan no longer leaves its dependents on the runnable work list — when a plan reaches a designed stop and its SUMMARY records status: halted, plans that depend on it (directly or transitively) are now reported as blocked, with the halted plan(s) named, instead of being offered to the executor as ordinary incomplete work. (#2830) (#3038)

  • Plan-phase now auto-recovers from a stalled planner or plan-checker spawn instead of hanging indefinitely — when a planner/plan-checker subagent produces no completion marker and no fresh on-disk plan activity for a configurable threshold (planner.stall_threshold_minutes, default 10 minutes, checked every planner.stall_detect_interval_minutes, default 5), plan-phase now automatically surfaces the existing accept-plans/retry/stop recovery choice instead of waiting for a manual interrupt. Trade-off: a planner/plan-checker that finishes quickly is no longer detected instantly — completion is observed at most one stall_detect_interval_minutes (default 5 min) after it happens, in exchange for eliminating the previously-indefinite hang. (#2650)

    Hardened a repo-wide test-portability pattern (maintainer-authorized scope expansion): ten test files that extract a fenced bash block from a workflow .md file and execute it via spawnSync/execFileSync now normalize CRLF to LF at the point of reading the file, before any fence-slicing or regex runs. A raw readFileSync followed by a bare \n-based regex against markdown fences is fragile by construction — it silently assumes LF regardless of how the bytes actually arrived — and this normalization removes that assumption at a single shared readFileNormalized() helper in tests/helpers.cjs, used by all ten call sites, so the next .md-extraction test is correct by default instead of needing to rediscover the fix independently. (Correction: this was NOT the cause of this PR's own windows-latest CI failure — .gitattributes' blanket * text=auto eol=lf means a Windows checkout of this repo never receives CRLF in the first place. That failure was a separate bash -c argv-transport defect in the #2650 test file itself, fixed alongside this.) (#3015)

  • gsd-tools windows no longer crashes on CRLF ledgers — on repos with core.autocrlf=true (Windows default), the frontmatter parser threw on the last key of a CRLF WINDOWS.md, making the broken-windows status/waive/fixed subcommands unusable. (#3116) (#3137)

  • Installed third-party reviewer lanes can now be selected, planned, and invoked — an installed role:"reviewer" capability was roster-visible and disclosed at install but /gsd-review (gsd-tools review-lane sections|flags|plan|invoke) built its lane map from the static first-party set only, so every third-party lane failed with "no such declared lane". The invocation surface now merges installed overlay reviewer lanes (first-party wins on collision, ADR-2782 D8). (#2927) (#3062)

  • Completing a phase no longer checks the box for a requirement the traceability table records as deferred or blocked — the phase-completion write flipped the REQUIREMENTS.md checkbox unconditionally and kept the flip when the traceability row existed but rejected the same completion, so a requirement recorded as Deferred or Blocked read as shipped. The checkbox now rolls back when a row exists but rejects the write, matching the existing requirements mark-complete behavior so the two surfaces never silently disagree. (#3073)

  • Hotfix branches with auto cherry-pick no longer abort on already-applied commits — cutting a hotfix from a tag whose chore: sync next package version commit applied empty (already present by content) aborted the entire create run. The cherry-pick error handler now distinguishes empty picks (no unmerged paths → skip) from genuine conflicts (unmerged paths → abort), and the job summary lists skipped-as-empty commits separately. (#2913) (#2970)

  • Installer --help now documents every supported runtime — --pi and --gemini were accepted but omitted from the help output, making them invisible to users discovering runtime support via --help. A parity test now guards against future drift. (#3026) (#3112)

  • Several guards that could not verify something previously reported the same result as everything is fine: a duplicate external job could dispatch past a corrupt sibling manifest, state rebuild could report success while phase-table reconciliation never ran, an unreadable lock body was treated as freely stealable at the same short window as a genuinely empty one, a staleness check that itself failed reported not stale, and git base-branch returned main whether it verified that or every git query timed out. These now fail closed instead of silently succeeding. (#3057) (#3088)

  • Updating GSD on Codex no longer deletes user settings from config.toml — the config merge preserved content before the GSD marker block but discarded everything after it, so any model preference, MCP server, or profile added after a fresh install was wiped on every update. The merge now preserves genuine user TOML after the block by routing it through the existing section stripper, which removes only GSD-owned sections while keeping user tables, and #2406's leaked-section de-dup still holds. Re-merging is idempotent. (#3067)

  • --kimi-code reviewer lane is now selectable in /gsd:review — the lane was declared, documented, and its flag resolved, but the review workflow's CLI detection and flag list omitted it (hardcoded to 11 of 12 lanes). Both now include Kimi CLI detection and the --kimi-code flag. (#3035) (#3115)

  • query commit no longer silently switches to a phase/milestone branch — git checkout -b both created AND switched HEAD, resurrecting merged-and-deleted phase branches. Now uses git branch (create-only, no switch); the commit always lands on the current branch. Callers that want to be on the phase branch should use execute-phase's branching step. (#3079) (#3141)

  • Package-legitimacy docs now match the registry-API gate — security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md, FEATURES.md, and the planner's STRIDE template described the pre-ADR-0656 design (slopcheck as the install-or-degrade gate, unavailability degrading every package to [ASSUMED]). Docs now describe the actual registry-API verdict gate (npm/PyPI/crates.io), with slopcheck as an optional escalate-only adapter. The ja-JP mirror is fully aligned, and the mechanical portion of the same drift (command strings, table headers, and already-attested-term swaps) is corrected in the zh-CN, ko-KR, and pt-BR mirrors as well; the prose-composition remainder in those three locales is tracked separately in #3002. (#2775) (#3010)

  • milestone complete no longer silently disarms its unstarted-phase guard when STATE.md's milestone: field drifts — the guard now runs whenever the ROADMAP can be scoped for the requested version (independent of STATE), and a STATE mismatch emits a WARNING naming both values instead of skipping the scan. (#2946) (#3081)

  • Trae IDE is now detected as its own runtime — /gsd-new-project and /gsd-ingest-docs no longer fall through to the Claude default when run inside Trae, and a --trae install no longer writes a malformed .claude/.trae/rules/ or .trae/.trae/rules/ instruction-file path; it now resolves to the concrete .trae/rules/rules.md. (#2658) (#3006)

  • Project-local agents are detected across non-Claude runtimes — GSD status and workflows now use a manifest-backed local installation before the global fallback. (#2623)

  • Malformed predicate declarations are now reported instead of silently dropped. A doubled-dot id, a space in an id, a lowercase-leading class, and a value with an embedded CR/LF are each surfaced as a distinct malformed diagnostic reason instead of vanishing with no trace; the example parser (examples/dynamic-context-management/) was also brought back into parity with production and its own index is now drift-guarded by a new lint script. (#2944) (#2950)

  • Workflow shell blocks no longer abort under zsh when a glob matches nothing — an unmatched glob inside a for word list aborted the entire shell block under zsh (macOS default shell), silently bypassing every statement after it, including the verify-phase decision-coverage gate. Each affected bash block now enables nullglob portably (shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null) so an unmatched glob expands to nothing and the loop is skipped cleanly under both shells. (#2962) (#3087)

  • docs/json-errors.md now documents the ExitError plain-text carve-out — the page previously claimed every CLI error emits a structured JSON envelope on stderr, but usage errors (ExitError) intentionally emit plain text with their own exit code. The structured-envelope guidance is now scoped to non-usage failures, with the carve-out stated explicitly and a characterization test pinning both paths. (#2979) (#3093)

  • review-lane rejects an unknown subcommand instantly instead of after a dozen subprocess spawns — an unrecognized subcommand fell through to the usage-error branch only after loading the capability registry and building a per-lane plan, which spawns one child process per lane. The error now fires before any of that work starts (~119ms instead of ~1288ms). (#3148) (#3192)

  • Worktree timeout guards now fire on Windows — the checks that detect a timed-out git command required the process to report a SIGTERM signal, which Node does not guarantee on every platform, so on Windows they could silently never fire and the guard they protect would pass without having verified anything. The check is now a single shared predicate keyed on the timeout code alone. (#3050) (#3060)

  • roadmap.analyze now discovers non-numeric-leading phase ids — the phase-heading and checklist discovery regexes required a digit-first id (e.g. 07), so a project using letter-prefixed ids (e.g. B7) got phase_count: 0 even though get-phase/execute-phase resolved the same ids fine. The regexes now accept an optional leading letter prefix. (#3036) (#3117)

  • progress.completed_plans no longer stays pinned after a gap-closure cycle — when plan-phase re-planned a phase and added gap-closure plans, total_plans corrected upward but completed_plans was restored to its pre-growth value, so STATE.md showed completed_plans < total_plans permanently even after every plan (including the gap-closure ones) was summarized. completed_plans and completed_phases now ratchet up to the disk-derived count under the plan-phase progress opt-in (never deriving downward, preserving the curated-progress ratchet for unrelated edits). (#2969) (#3091)

  • /gsd-audit-uat now sees archived phases and table-shaped artifacts — three silent false negatives are fixed: (1) the audit scanned only .planning/phases/, so a project whose milestones had been archived to .planning/milestones/<version>-phases/ silently omitted those phases, and one with ALL phases archived hard-errored with "No phases directory found" instead of reporting its outstanding items; (2) a deferred-items.md recording entries as a GFM table yielded zero items; (3) a table-shaped ## Gaps section likewise yielded zero items. Results now carry archived_milestone so consumers can label provenance. Same false-negative family as #2286/#2287, one document shape further out. (#2766) (#3082)

  • A Codex surface re-stage no longer creates a duplicate skill tree — re-staging skills on a global Codex install wrote them to $CODEX_HOME/skills while the installer had correctly placed them in $HOME/.agents/skills, leaving two active GSD skill trees and no signal which one was live. The re-stage and the legacy dev-preferences migration now resolve the same destination the installer uses. (#2911) (#3049)

  • /gsd-verify-work diagnosis and interactive plan execution no longer halt on a stale worktree fork base — when worktrees are enabled and local HEAD has advanced past origin/HEAD (the GSD steady state of committing every step and pushing only on request), the spawned debug/executor agent used to fork from the stale ref and hit a base-mismatch fatal mid-investigation with no recovery. Both dispatch sites now run the same pre-dispatch worktree.base-check gate the executor and quick-task paths already run, auto-degrading to sequential main-tree dispatch with an explanatory message. (#2649) (#2955)

  • Phases no longer leak archived data from another workstream — resolving a phase in one workstream whose own directory doesn't exist yet no longer falls back to an unrelated workstream's (or a flat-mode project's) same-numbered archived phase; it correctly resolves as pending. (#2855) (#3008)

  • npm test no longer writes into the developer's live config directory — TEST_ENV_BASE scrubbed 14 session-identity vars but omitted CLAUDE_CONFIG_DIR, GSD_RUNTIME, and CODEX_HOME (config-location vars that decide WHERE a child writes). The config-home resolver consults these before HOME, so an ambient value won unconditionally over a sandboxed HOME. All three are now blanked. (#2665) (#3134)

  • Multi-paragraph changeset bodies no longer truncate and lose their PR trailer — serializeChangelog wrote bullet bodies verbatim, so an embedded newline became a column-0 line that parseChangelog treated as the end of the bullet, silently dropping the continuation and the (#NNNN) trailer. Continuation lines are now indented so the round-trip preserves content and attribution. (#3001) (#3101)

  • Cross-AI reviewer lanes no longer silently drop on Windows — deps.spawn used shell: false with a bare binary name, which fails with ENOENT on Windows .cmd shims (npm-installed CLIs). Now applies the #2667 cmd.exe /d /s /c shim gate. Spawn errors (ENOENT, ETIMEDOUT) are also surfaced in the reviewer err file instead of being silently dropped. (#3086) (#3142)

  • A phase stranded between its last plan and verification can now be recovered — if every plan carried a SUMMARY but the run never reached the verify step (most often because a checkpoint plan was retired yet still summarized), re-running execute-phase exited immediately and could never produce the missing VERIFICATION.md, so the recommended recovery command silently did nothing. It now resumes at the phase gates instead, with the code-review and regression gates still running. (#2868) (#3041)

  • Planning artifacts whose frontmatter is preceded by a UTF-8 byte-order mark no longer lose all their frontmatter fields — the frontmatter parser's fence check required the opening dashes at byte zero, so a BOM written by Windows PowerShell or several editors made every field silently disappear. A leading BOM is now stripped before the check, so the fields parse identically to the no-BOM case. The no-frontmatter and thematic-break cases stay silent and empty as before. (#3076)

  • GSD-2 import no longer duplicates frontmatter in the generated SUMMARY.md — importing a GSD-2 project whose task summaries were authored with CRLF line endings emitted the original GSD-2 frontmatter a second time, as body text, below the new one. Stripping now goes through the canonical line-ending-tolerant parser. (#2703) (#3027)

  • Codex skill adapter collaboration-tool vocabulary corrected — the generated adapter documented an obsolete wait(ids) call (the real tool is collaboration.wait_agent), unconditionally instructed close_agent without a tool-visibility gate, and omitted the required task_name field and the fork_turns parameter. The adapter now names the real wait tool, disambiguates it from the unrelated exec-cell functions.wait, gates close_agent on schema visibility, and covers task_name + fork_turns. (#3004) (#3104)

  • Documentation now shows the command form that actually works — reader-facing docs instructed users to type /gsd:<command>, a form no runtime registers, so copying it produced an unrecognized command. All 178 occurrences across 53 files, including the Japanese, Korean, Portuguese and Chinese mirrors, now use /gsd-<command>. A new lint keeps it from drifting back, while leaving the colon form intact where it is load-bearing — source artifacts, where install-time converters key on it — and preserving the genuine /gsd-core:<command> plugin namespace. (#2903) (#3047)

  • The composer's load-bearing-fragment guarantee is now enforced, not just documented — ADR-1671 promised a deterministic gate proving no load-bearing content is dropped or shrunk when context is trimmed to fit a budget; only synthetic unit tests existed. The gate now runs against real declared strategies and fails if it would ever assert over nothing. (#3065) (#3068)

  • phase.complete no longer closes a phase while its plans are silently unexecuted — a phase could previously close "complete" with an arbitrary number of plans missing a completion record (a confirmed incident closed a phase with 6/30 plans unexecuted, including its entire final scope). phase.complete now refuses, naming the unexecuted plans, unless they are explicitly retired via status: superseded frontmatter. (#2648) (#2953)

  • A worktree whose owner could not be probed is no longer deleted — an orphan lock holding a process id above 2147483647 made the liveness check throw a type error rather than an errno error, which read as "owner is dead" and removed the worktree. Only "no such process" now means dead; every unrecognized outcome leaves the worktree alone. An unreadable lock timestamp also reported "too fresh", advising a wait that could never help, and now reports its own reason. (#3103) (#3106)

  • Spec-phase edge resolution vocabulary realigned to the code's Status enum — the workflow prose in spec-phase.md, plan-phase.md, and ui-phase.md used the retired covered/backstop-as-status vocabulary that validateResolution rejects. Now uses resolved + verification: explicit|backstop. (#3132) (#3138)

  • Gate predicate artifact-frontmatter-equals is now implemented — declared gates that use it are evaluated instead of erroring on an unrecognized kind. (#2785) (#2816)

  • /gsd commands in Pi now display their output — the command handler returned output as a bare string, which Pi's ExtensionAPI silently dropped. It now returns Pi's structured { content: [{ type: 'text', text }] } display shape (matching the gsd_invoke tool's proven contract), so success output and error messages are visible. (#2991) (#3097)

  • gsd-code-fixer no longer creates its review-fix worktree outside the project tree on Windows — the worktree was hardcoded to a /tmp/sv-... mktemp path, which on Git Bash landed outside the repository (every file read inside it prompted for permission) and produced an un-removable short path. The worktree now lives repo-relative under .claude/worktrees/, the same location the executor worktrees use. (#2647) (#2942)

  • /gsd-code-review no longer picks a wrong diff base from unanchored commit-message grep — the diff-base fallback searched all commit messages for the bare phase number as a substring, matching version strings, dates, and issue refs, then took the oldest match. The grep is now anchored to the phase-mention convention (Phase N with a word boundary), so the fail-closed branch is reachable when no commit genuinely references the phase. (#2989) (#3096)

  • roadmap.analyze no longer silently drops phases when the phase-listing heading isn't version-bearing — if the phase list lives under a plain ## Phases heading (the shipped greenfield template's own shape) and a later version-bearing progress/notes heading exists, the milestone scope previously latched onto the later heading and stripped every ### Phase N: detail from the preamble, returning phase_count: 0 with exit 0 and empty stderr. Phase details in the preamble are now preserved when the selected milestone section has none of its own. (#2947) (#3084)

  • execGit now reports timedOut on every result, and its return type is no longer misdeclared — three modules hand-copied the shape of execGit's result because the canonical type was not exported, and two of those copies declared exitCode as nullable when it can never be null. The shape is now declared once and reused, so a consumer can no longer be written against a contract the function does not honor. (#3071) (#3077)

  • Heavy workflow skills no longer fail on Claude with thinking disabled — effort: max in plan-phase, execute-phase, and autonomous SKILL.md frontmatter was rejected by the Anthropic API (400: effort 'max' is not supported when thinking is disabled). The installer now clamps max/xhigh to high for Claude-runtime skills, the maximum value that works in both thinking states on all supported models. (#3039) (#3119)

  • state.* writes no longer flip the milestone or rewrite progress with whole-project counts — when the stored milestone had no matching non-shipped ROADMAP heading, buildStateFrontmatter auto-derived a confidently-wrong milestone and clobbered the stored value + progress on every write. The disk scan now scopes to the STORED milestone explicitly, so a state write that doesn't change progress leaves the milestone and progress block untouched. (#3017) (#3105)

  • Project configs no longer inherit runtime from the machine-wide ~/.gsd/defaults.json — on machines with 2+ runtimes installed (e.g. Codex + Claude Code), the last installer's runtime value poisoned every new project, resolving agents to wrong model IDs. The key is now excluded from the defaults spread. (#2840) (#2985)

  • Nine compiled .cjs runtime artifacts under gsd-core/bin/lib/ are no longer tracked in git — they are ADR-457 build outputs of src/*.cts sources and were missing from .gitignore, letting the committed bytes silently drift from source (as happened to api-coverage.cjs in #2653). They now build fresh from source like their ~160 already-gitignored siblings. (#2657) (#3011)

  • Secret-scan no longer reports a false positive on the zh-CN verification-patterns translation — the translated document carries the same illustrative placeholder examples as its English source, but the exclusion was never extended to the translation. The strict-mode scan now passes. (#3044) (#3122)

  • state planned-phase no longer overwrites authoritative last_activity_desc — when the frontmatter and body had the same activity date but different descriptions, the write path preserved the date but overwrote the frontmatter's description with stale body prose. Same-date frontmatter desc is now preserved. (#3052) (#3140)

  • Claude Code plugin installs no longer silently disable all hooks — the plugin manifest (.claude-plugin/plugin.json) explicitly declared hooks/hooks.json, which Claude Code also auto-loads by default, causing a duplicate-declaration rejection that silently disabled every hook (security guards, monitors, injection scanners). The redundant declaration is removed; Claude Code's auto-load path handles it. (#3029) (#3113)

  • scripts/lint-compiled-artifact-sync.cjs no longer fails on containerized checkouts owned by a different uid — its internal git calls now scope safe.directory to the repo root per-invocation, so the guard runs instead of erroring with "detected dubious ownership" in any CI lane where the checkout owner differs from the running user. (#2657) (#3011)

  • roadmap validate now performs real structural validation — it previously returned {"warnings":[]} (exit 0) for every input including empty files, garbage text, and missing files, providing false assurance. It now checks file existence/readability, emptiness, frontmatter well-formedness, and the presence of at least one phase entry, exiting non-zero on any warning (per its documented contract). The existing opt-in milestone-prefix consistency check is preserved. (#2978) (#3092)

  • Codex local capability metadata now matches project-scoped installs — Remove the inert user-home override from the local skills descriptor, document the global/local skill roots, and reject user-home overrides across all local artifact-layout entries. (#2777) (#2831)

  • Documentation now consistently warns about --dangerously-skip-permissions — the flag was presented without a caveat in the user guide, the onboarding tutorial, and all four translated locales (ja-JP, zh-CN, ko-KR, pt-BR), while the English first-project tutorial carried a proper caution. All occurrences now carry the same [!CAUTION] block. (#3043) (#3121)

  • phase remove now reports accurate state_updated and keeps STATE.md progress counters in sync — the command reported state_updated: true based on file existence (always true) rather than actual content change, and the frontmatter progress.total_phases/completed_phases/percent counters went stale when the STATE.md body lacked a Total Phases: field (the no-op write guard skipped the frontmatter resync). (#2640) (#2974)

  • Unusable last_activity now emits a diagnostic — a present-but-unparseable last_activity in STATE.md silently suppressed the idle-stranded recommendation. The fallback (stale_activity: false) stays for continuity, but a last_activity_unparseable warning is now emitted so the degradation is visible. (#3099) (#3139)

  • Statusline now shows GSD state in workstream mode — the GSD-state segment used to silently disappear in workstream-mode projects with no root STATE.md, even with an active workstream selected; it now resolves the active workstream (env var or stored pointer) and shows its milestone/phase/progress, or an explicit "no active workstream" message when nothing resolves. (#2850) (#3012)

  • Published installs no longer crash on a script that can't load — scripts/gen-emitted-baseline.cjs shipped in the npm tarball but required three modules from tests/ (which does not ship), producing MODULE_NOT_FOUND at load time. The script is repo-only CI tooling and is now excluded from the tarball. A class-extinction guard test ensures no shipped script can require outside the shipped tree going forward. (#2858) (#2968)

  • A --kimi-code install now configures hooks in Kimi Code, not Kimi CLI — installing GSD for Kimi Code wrote its lifecycle hooks, hook bundle and CommonJS marker into Kimi CLI's ~/.kimi/config.toml, so Kimi Code itself received no hooks at all and a machine with only Kimi Code got a config file no product reads. Each Kimi product now uses its own root and its own environment override (KIMI_SHARE_DIR for Kimi CLI, KIMI_CODE_HOME for Kimi Code), and uninstalling one no longer removes the other's hooks. (#2755) (#3032)

  • Contributor PRs stop conflicting on a file they never meaningfully changed — the emitted-drift acknowledgment moves from one shared tests/emitted-drift-ack.json every PR rewrote wholesale to per-PR fragments under tests/emitted-drift-acks/, so two PRs needing an acknowledgment can no longer collide with each other; the legacy file's 35 spent entries are migrated (not deleted) into a fragment so nothing is lost, and a next-only push guard now fails if the legacy shared file itself ever reappears, since every entry is scoped to the diff that introduced it and is spent the moment it merges. (#2914) (#2923)

  • Slug no longer ends with a trailing hyphen when truncated — long titles whose 60-character cut landed on a word separator produced a slug ending in -, which then leaked into phase directory and branch names. The trailing-hyphen strip now runs after truncation. (#2849) (#2967)

  • MemPalace capture no longer silently disables itself when capture_artifacts is unset — the skill gate used !== true (treating absent as disabled), but the capability schema defaults to enabled. Fixed to === false (disabled only on explicit false). (#2641) (#2982)

  • Completing the last phase of a milestone no longer advances into a 0.x backlog sentinel row — the phase-completion cascade's lowest-outstanding-phase override had no sentinel filter, so an unchecked backlog row like Phase 0.1 sorted below every real phase and was selected as the next phase, corrupting STATE.md and desyncing the current phase number from its name. The override now excludes sentinel-range phase ids via the existing isSentinelPhaseId predicate, so a real lower-numbered outstanding phase is still selected while backlog sentinels are skipped and the milestone completes cleanly. (#3070)

  • Research agents no longer call a context7 tool that doesn't exist — four shipped docs instructed agents to call mcp__context7__get-library-docs, a tool the context7 MCP server does not register (it exposes only resolve-library-id and query-docs). Every research workflow that loaded the canonical doc-lookup reference either errored, fell back to the ctx7 CLI, or fabricated a result. All sites now name query-docs with the registered libraryId/query params, the CLI-fallback rationale now describes the real project-scoped .mcp.json mechanism, and a parity guard fails the build if the banned name returns. (#2943) (#2963)

  • Workflow-backend worktree branches (worktree-wf_*) are now recognized by all worktree guards — the Claude-orchestration Workflow backend created worktrees on branches none of the four guards recognized, causing the path-containment hook to fail open and the cleanup/executor commands to reject or silently drop entries. All four sites now accept the worktree-wf_ namespace alongside agent-* / worktree-agent-*. (#3021) (#3109)

  • phase_id_convention set in .planning/config.json is no longer silently dropped — the config loader's resolved-config constructor omitted the key despite it being in the valid-keys manifest, so the milestone-prefix validation check could only be activated via the ROADMAP frontmatter fallback. The key now survives resolution. (#2997) (#3098)

  • Agent-skills warnings now suggest the global: prefix when a bare name matches a global skill — configuring a skill by bare name (e.g. patch-coverage-check) that exists as a global skill was silently skipped with no hint that the fix is global:patch-coverage-check. The skip warning now appends a hint when the bare name matches an existing global skill. (#2941) (#2973)

  • /gsd-spike no longer blends unrelated ideas' requirements together — .planning/spikes/MANIFEST.md now scopes each idea's paragraph and Requirements under its own idea key, and /gsd-spike --wrap-up only emits a feature area's owning idea's requirements instead of the whole file. (#1700) (#3014)

  • worktree cleanup-wave no longer aborts the rest of a wave when one entry is blocked — a blocked entry (mismatched branch/base, a deletion, a dirty worktree, or a failed merge/removal) now stays blocked with its existing reason code, while every other independently-clean entry in the wave still merges and is removed instead of being stranded unattempted. (#2852) (#3009)

  • Local lint:changeset and lint:docs-required now diff against next instead of main — the local fallback was the release branch (main), which lags far behind the integration branch (next), so the lint always passed by finding fragments from other already-merged PRs in the oversized diff range. The local invocation now matches the base CI uses. (#2988) (#3095)

  • Non-Latin phase and milestone titles no longer produce empty slugs — a Cyrillic title used to reduce to an empty slug, creating unnamed phase directories (bare numeric prefix like 01-) and empty milestone_slug fields. Titles are now transliterated to ASCII before the slug filter, so a non-Latin title yields a usable slug. Latin-script output is unchanged. (#2848) (#2934)

  • graphify version detection now verifies tool identity — a foreign binary named graphify on PATH that printed a plausible version string would silently report compatible: true with no warning. The check now confirms the graphifyy Python package via importlib.metadata before trusting the version, emitting a clear warning naming the mismatch when identity cannot be confirmed. (#3020) (#3107)

Security

  • Prompt-injection scan no longer misses single-quoted eval()/exec() payloads on macOS, and no longer flags ordinary prose — the patterns used a GNU-grep-only \\x27 escape that BSD/macOS grep read as four literal characters, so single-quoted code-execution payloads went undetected there while passing on CI; separately, several patterns lacked a left word boundary and matched inside ordinary words (fact as a, retrieval(, Jordan mode). (#3023) (#3175)
  • Production dependency tree is clear of known advisories — three transitive packages reached by @anthropic-ai/claude-agent-sdk carried published advisories: fast-uri (host confusion via a backslash authority introducer), ip-address (three SSRF / trust-boundary bypasses via leading-zero octets, CIDR-suffix suppression, and IPv4-mapped address misclassification), and hono. All three are lockfile-only, semver-in-range updates. (#2755) (#3032)
  • A directory name containing $(…) or a backtick no longer becomes a live command in your shell startup file — the PATH-persistence suggestion escaped its export PATH="…" line for the echo that carries it, not for the rc file it lands in, so a substitution in the target directory survived into ~/.bashrc and ran on every new shell. (#3118) (#3124)

[1.9.1] - 2026-07-31

Added

  • Reviewer lanes are now documented as an authorable capability surface — a new how-to walks capability authors through declaring a reviewer body so /gsd-review discovers, invokes, and renders their external review CLI or model endpoint, and the manifest reference's invoke row now lists the full accepted vocabulary for both transports. (#2782) (#2906)
  • Third-party reviewer lanes can now be listed in a discoverability catalog. ADR-2782 made a reviewer lane installable by a third party, but the two existing registries could not hold one — the Community Capability Registry requires a non-empty loopExtensionPoints, which a lane registers on none of, and the EoS Registry is for host integrations. A new Reviewer Lane Registry (docs/registries/reviewers.json → docs/registries/reviewer-registry.md) gives lanes a home, with an entry schema describing the lane itself: slug, flags, transport, evidence class, and REVIEWS.md section. (#2904) (#2912)

Fixed

  • Fallow structural pre-pass no longer silently no-ops on Windows — run-with-timeout now mediates .cmd/.bat/.exe spawns via an explicit cmd.exe /c argv array (Node's CVE-2024-27980 hardening requires a shell for these on Windows), and the fallow pre-pass names the failure kind so a Windows spawn failure is not mistaken for an absent binary. The existing bash -c callers and POSIX behavior are unchanged. (#2667) (#2897)
  • A clean Codex install now applies balanced model settings to agent TOMLs on the first run — ~/.gsd/defaults.json (resolve_model_ids + runtime) is now written before agent TOML generation, so the runtime-aware model resolver knows the target runtime during the first pass. Previously a second install was required. (#2834) (#2900)
  • verify-summary no longer reports a valid SUMMARY as failed because of a path mentioned in prose — file-claim extraction is now bound to a creation/modification claim (a Created:/Modified:/key-files line), so a prose mention of a future deliverable is not checked for existence; and verify-summary now resolves the project root, so invoking it from a subdirectory no longer manufactures missing files. (#2910)
  • findProjectRoot no longer silently resolves to a parent project across a git-repo boundary — when invoked from a nested git repository that has no .planning/ of its own, resolution stays within the caller's repo (or falls back to the start directory) instead of crossing into an ancestor GSD project. The existing plain-descendant and co-located .git+.planning cases are unchanged. (#2909)
  • A requirement row stranded at Gaps Found can now be completed again, and requirements mark-complete no longer reports false success on a row it could not move — the completion guards now accept Gaps Found (so revert-phase's stranded rows are recoverable instead of permanently blocking the milestone), and when a traceability table has a row for an ID, mark-complete counts it as updated only if the row actually moved (not merely because the checkbox flipped). (#2788) (#2902)
  • /gsd-code-review --fix now honors workflow.use_worktrees — when the setting is false, the fixer edits and commits in the main checkout instead of creating a git worktree (matching the other writer workflows), and the spec forbids rm -rf on a possible Windows reparse point so an improvised worktree teardown can no longer delete the real node_modules. The REVIEW-FIX report also records where verification ran. (#2905)

[1.9.0] - 2026-07-31

Added

  • New kimi-code runtime (Node Kimi Code CLI) registered as a distinct EoS capability — Kimi Code users running --kimi --global were silently installing the Python kimi-cli agent YAMLs (which Kimi Code ignores) and getting an empty gsd-tools query agent-skills response. The split adds a kimi-code descriptor with runtime: "node", dispatch.namedDispatch: false, builtInSubagents: [coder, explore, plan], and registers it across every drift-guarded surface (allRuntimes, runtimeMap, FALLBACK_ALIASES, RUNTIME_LABELS, model-catalog, runtime-aliases manifest, capability-registry, capability-matrix, CONTEXT.md glossary). runtimeFlags('kimi-code').isKimiCode === true; --kimi-code selects kimi-code without interactive prompt; existing kimi (Python kimi-cli) users see no behavior change beyond the corrected localConfigDir: ".kimi". (#2511) (#2519)
  • --kimi-code --global now installs a working Agent Skills surface at ~/.kimi-code/skills/gsd-*/SKILL.md — previously the kimi-code descriptor (Phase 1) carried an empty artifactLayout, so the install produced zero skills and Kimi Code's merge_all_available_skills = true auto-discovery found nothing. Phase 2 adds the convertClaudeCommandToKimiCodeSkill converter, fills the descriptor's artifactLayout.global with the skills kind entry, and removes the Phase 1 SKIP_INSTALL_CONTRACT skip by setting the install contract surface to flat-skills (NOT kimi-skills-agents — Kimi Code has no custom agents). Kimi Code auto-discovers the skills on next launch; no agents/gsd.yaml or subagents/*.yaml installed. (#2509) (#2520)
  • gsd-tools query agent-skills <name> returns the installed agent's prompt content on non-Claude runtimes — previously, when a non-Claude runtime (kimi, kimi-code, opencode, kilo, etc.) had no explicit agent_skills config entry, buildAgentSkillsBlock returned empty and the ${AGENT_SKILLS_*} workflow injection carried no persona. Phase 3 adds a fallback in cmdAgentSkills: on non-Claude runtimes, when the configured block is empty, resolve the runtime's agents directory via checkAgentsInstalled(runtime) and read <agentsDir>/<agentType>.md as the block. Gated to runtime !== 'claude' (Claude supports named dispatch and its ${AGENT_SKILLS_*} contract is a skills-injection path, not a persona fallback). (#2510) (#2521)
  • Runtime-aware subagent dispatch for built-in-only runtimes (kimi-code) — workflows calling Agent(subagent_type="gsd-*") now resolve the type for the current runtime via gsd_run query resolve-dispatch-type --requested <name> --raw before dispatch. On named-dispatch runtimes (Claude/OpenCode/…) the gsd-* name is returned unchanged; on built-in-only runtimes (kimi-code — three built-in subagents coder/explore/plan, no custom registration) it maps to the closest built-in by role-suffix heuristic (-planner→plan, -researcher/-checker/-auditor→explore, everything else→coder). The persona rides ${AGENT_SKILLS_<ROLE>} (Phase 3) regardless of the resolved type. Adds the resolveDispatchType function to host-integration, the query to gsd-tools, a reference doc, and the resolution preamble to 26 workflow files. Pivot from the epic's original Option B (PreToolUse hook remap) after research confirmed Kimi Code's hook API supports only allow/deny, not tool_input rewriting. (#2508) (#2525)
  • The installer now distinguishes Kimi CLI (Python) from Kimi Code (Node) at install time — running --kimi or --kimi-code prints a one-line description of each product, and if the selected variant doesn't match the detected ~/.kimi/config.toml vs ~/.kimi-code/config.toml, the installer warns with the correct --kimi-code / --kimi re-run command. Catches the "ran --kimi --global but actually on Kimi Code" mistake that produced inert YAMLs and empty agent-skills before the Phase 1 descriptor split. (#2513) (#2535)
  • New docs/migration/kimi-to-kimi-code.md migration guide + built-in-only subagent-toolkit enum value — users who installed via --kimi but are actually on Kimi Code (Node CLI) now have a step-by-step migration path (re-install with --kimi-code, remove inert YAMLs, verify skills, verify agent-skills query). The built-in-only enum value replaces the undocumented sentinel on the kimi-code descriptor's subagentToolkit axis, making the descriptor self-documenting: Kimi Code's three built-in subagents (coder/explore/plan) are now a first-class negotiated value rather than an escape hatch. (#2512) (#2538)
  • npm run regen:derived regenerates every derived artifact in one command — replacing several separate invocations (build, gen:registry, gen-adr-index, gen-capability-matrix, gen-inventory-manifest, sync-manifest-versions, gen:install-tree) with one dependency-ordered command. (#2721) (#2730)
  • /gsd:review --kimi-code reviews your plans with Kimi Code CLI — the new lane joins the cross-AI reviewer roster and is included by --all when detected. Detection distinguishes Kimi Code from the legacy Python kimi-cli, which ships a binary of the same name, so a host with only the legacy tool reports the lane unavailable instead of registering a reviewer that cannot serve it. (#2718) (#2861)
  • /gsd:update now offers to restore the user-added files it backs up — files you added inside GSD-managed directories were copied to gsd-user-files-backup/ before the clean install and then left there forever; only --reapply (a different bucket, gsd-local-patches/) had a restore path. The update now lists what it backed up, runs a compatibility pass against the newly installed release, and offers to put the files back. Declining leaves the backup untouched, and the backup is never deleted. (#1854) (#2679)
  • Phase effort estimation against a calibrated smart-zone budget — plans can now be sized against a configurable token budget (workflow.smart_zone_tokens, default 100000) instead of a static heuristic, and the estimate self-corrects against measured reality. Adds the estimate-check and estimate-calibration query verbs. (#2630) (#2661)
  • Bracket phase-ID core grammar lands behind an opt-in flag — parsePhaseId/renderPhaseId/toDir add one pure round-trippable PhaseId model inside the ADR-2121 canonical owner (src/phase-id.cts), gated on phase_id_convention: 'bracket', with generative round-trip properties; legacy null/milestone-prefixed paths stay byte-untouched (epic #612 PR-1). (#2249) (#2258)
  • Reviewer CLIs now honor GSD's configured reasoning effort instead of silently inheriting your global CLI default — cross-AI review runs previously picked up whatever effort sat in your own ~/.codex/Claude/OpenCode config, so the same project produced 1-3 minute review cycles on one machine and 12-15+ minute cycles on another with no in-project way to influence it. GSD now resolves one effort value from the effort.* cascade and passes it to each reviewer in that CLI's own syntax; a host with no documented reasoning setting is left untouched rather than given a guessed flag. (#2481) (#2490)
  • List the gsd-cursor EoS host integration in the registry — six phase-aware Cursor profiles (max / hybrid / value / budget / frontier / openweight), added to docs/registries/eos.json with a versioned v1.1.0 install command. (#2581)
  • Plans now carry a calibrated effort estimate — every generated PLAN.md includes an estimate block, and /gsd-plan-phase flags a phase projected to exceed the smart-zone budget with a concrete split recommendation. Advisory only; it never blocks planning. (#2631) (#2670)
  • Reviewer lanes can be declared as capability manifest data — a capability may now carry a reviewer body describing a cross-AI review lane (slug, flags, transport, probe, invocation shape, timeout floor, output policy), and a new role: "reviewer" declares a lane that is not an install target. The registry validates the body against closed vocabularies and enforces slug, flag, and section uniqueness across first-party and installed capabilities, so two lanes can no longer silently share a REVIEWS.md heading. A capability with no reviewer body is unaffected. (#2795) (#2823)
  • Reviewer lanes now ship as capability declarations — the eleven cross-AI reviewer lanes are declared as manifest data instead of a half-derived, half-hardcoded roster. Five reviewers GSD never installs into (Gemini, CodeRabbit, Ollama, LM Studio, llama.cpp) become lane-only capabilities with no install surface, and the six hosts that are also reviewers gain a reviewer body alongside their runtime descriptor. gsd capability list shows the five new lanes. The roster itself is unchanged — the same eleven reviewers, derived rather than hardcoded — and runtime.hostBehaviors.reviewerCli keeps working for one release. (#2798) (#2837)
  • Parallel execute-phase waves now run on Codex, OpenCode, Kimi and Kimi Code — previously only Claude Code could execute a wave's independent plans concurrently, because worktree isolation relied on its harness-native isolation="worktree" primitive and every other runtime failed closed to sequential. Executor isolation is now a negotiated capability: runtimes whose harness isolates executors (Claude Code, Cursor) use their own flag, and runtimes exposing a headless exec with a working directory (Codex, OpenCode, Kimi, Kimi Code) get worktrees that GSD creates, validates and merges itself. Runtimes with no isolation primitive still run sequentially, and an unknown declaration always degrades to sequential rather than to an unisolated parallel run. (#2627) (#2635)
  • Codex host-plugin binding + negotiated executor-worktree isolation — ADR-1239 gains a Codex worked-binding amendment and a new dispatch.isolation capability (harness- vs orchestrator-managed git worktrees) enabling parallel execute-phase waves on non-Claude runtimes. (#2600) (#2600)
  • Estimates now calibrate against reality — the executor records what a phase actually cost into SUMMARY.md, and /gsd:extract-learnings computes the estimate-vs-actual correction so future plan estimates improve for your project. (#2632) (#2672)

Changed

  • The phase researcher must now read and cite in-repo values before calling them verified — an enum, schema or type union, error code, status constant, or filesystem path earns a [VERIFIED: path:line-range] tag only if the researcher opened the source-of-truth file with Read during the run and quoted the values verbatim in the <interfaces> block; every value used in a code skeleton must appear in that quote, and anything else stays [ASSUMED]. Previously the tag could be earned from training memory or a web search alone, so a plausible-but-drifted enum could pass into RESEARCH.md, get copied into PLAN.md, and fail only at the executor's parse()/typecheck — a mid-execution deviation, the most expensive place to discover it. (#1699) (#2768)
  • Completing a phase now warns when its SUMMARY claims files that never landed — phase complete runs the artifact check that verify-summary has always applied to the research SUMMARY against the completing phase's own SUMMARY.md files, and reports any referenced path that is not on disk through its existing warnings[] channel. Previously the check was wired to exactly two call sites, both pointed at .planning/research/SUMMARY.md, so the summaries that actually assert "I created these files" were never verified and an interrupted phase counted toward 100% silently. Advisory only: it never blocks completion. Paths are recovered heuristically from the SUMMARY body, so globs, URLs, bare hostnames, and paths resolving outside the project are skipped rather than reported; the key-files: frontmatter block and commit hashes are deliberately not read. (#2572) (#2685)
  • The UI consideration probe now asks about loading and error states for interactive controls — a UI surface classified only as an interactive control (a button, toggle, switch, or slider, with no accompanying form or list) previously had only its long-text state probed, so a spec could omit what the control shows while its action is in flight or when it fails and still pass. Control-only surfaces are now probed for their in-flight and failure states too. (#2151) (#2575)
  • Reviewer lane flags and section titles are now gated across every documentation surface — /gsd:review reviewer flags were hand-enumerated in five docs and three workflow files that had silently drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md entirely. The lane roster is now the single declared source: workflows derive their flag lists from a new review-lane flags query, and a parity gate fails the build when any documented flag or reviewer section title diverges from it. The capability manifest reference also gains the previously undocumented reviewer body and hostBehaviors field. (#2800) (#2882)
  • Cross-AI reviewer lanes are now declared data rather than hand-written per-CLI blocks — every lane's binary, prompt and output channel, timeout, probe and empty-output policy comes from its capability manifest, so a reviewer can be shipped as a plugin instead of a core patch. Two user-visible consequences: a reviewer that returns only whitespace is now reported as a failed lane on every reviewer (previously only on LM Studio and llama.cpp, so elsewhere a blank reply was rendered as a clean review), and an OpenAI-compatible lane whose configured host has changed since you consented to it is blocked with an explanation rather than silently sending your plans to the new destination. jq, curl and GNU timeout are no longer required on PATH for any lane. (#2782) (#2861)
  • Reviewer config keys are now owned by their reviewer-lane capabilities — review.models.<lane>, review.<lane>_host and review.max_prompt_tokens_per_reviewer.<lane> moved from the central config schema to federated slices on the lanes that use them. Key names and existing .planning/config.json files are unchanged and no migration is needed. Two consequences are user-visible: a review.models.<x> or review.max_prompt_tokens_per_reviewer.<x> key naming something that is not a declared lane is now rejected by config-set where it was previously accepted and silently ignored; and clearing one of these keys now reads back as its declared default rather than reporting key-not-found, because a federated key always resolves — an empty string for the model and host keys, and -1 for a per-lane token budget (a deliberate sentinel, since 0 already means "do not trim this lane" and must stay distinguishable from unset). review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy across lanes and deliberately remain central. (#2797) (#2841)
  • Reviewer lanes are disclosed and consent-gated before install — a capability that declares a reviewer lane now discloses what it will run and what it will be sent, and blocks on consent before any file is promoted. A spawned lane discloses its binary and its full arguments; an OpenAI-compatible lane discloses its destination host and the config key naming it, including a localhost destination. Both name the egress payload classes — plan text, requirements, research findings, and CONTEXT.md decisions. Changing a lane's binary, arguments, destination, prompt channel, or handler forces re-consent on update; a capability with no reviewer lane is unaffected and its consent record is unchanged. (#2796) (#2826)
  • Raw and calibrated phase-estimate token counts are now distinct types — the two states of an estimate (the planner's uncorrected projection and the same figure with the project's calibration factor applied) could previously be swapped at any seam without complaint, because both are plain positive integers. That produced two shipped defects in epic #1952: a doubly-applied correction (factor squared) and a calibration loop that measured against its own output and never converged. Both are now compile errors. No behavior, output, or schema change. (#2671) (#2676)
  • The emitted-attribution size ratchet now tells you how to clear it — a PR that only grew a workflow or agent file used to fail with a byte delta and the word "acknowledgment", without naming tests/emitted-drift-ack.json, saying it does not exist yet, giving its schema, or stating that the key is the bare filename. All three failing branches now print a minimal valid document and repeat that nothing is regenerated. (#2778) (#2780)

Removed

  • npm run gen:golden, UPDATE_GOLDEN, npm run size:baseline, and npm run setup:merge-driver are removed — the committed golden-install-parity fixtures and the two per-file size baselines they regenerated are deleted. The differential attribution check (tests/emitted-attribution.test.cjs) is now the sole gate for both emitted-content propagation and workflow/agent size growth; editing shipped content requires zero manual fixture regeneration. npm run regen:derived and npm run gen:install-tree are unaffected. (#2724) (#2767)

Fixed

  • Permission errors on phase and milestone directories now surface instead of looking empty — an unreadable phase directory used to be silently reported as "no CONTEXT.md" (so the discuss/plan gates wrongly skipped context) and an unreadable milestones/ directory as "no archives" (so active-milestone resolution and archived-phase filtering misbehaved), because both scans treated a permission or I-O failure the same as a genuinely empty directory. (#1883) (#2802)
  • Worktree branch guards now accept Claude Code's agent-<id> namespace — the worktree record-agent command, the spawn-time branch check, the cleanup-wave manifest reader, and the force-add/path/workflow guards all accept both the current agent-<id> and the legacy worktree-agent-<id> branch naming. Previously, Claude Code's rename from worktree-agent-<id> to agent-<id> caused every executor sub-agent to fail its branch check (false-positive FATAL / exit 42) and silently dropped valid cleanup-manifest entries (empty_manifest), blocking merge-back. (#1995) (#2548)
  • secure-phase, validate-phase, and next workflows now scope their query commit calls — all three pass --files with the specific artifact path, preventing the blanket git add .planning/ default branch from sweeping unrelated staged or unstaged files into a commit whose message describes a single artifact. Previously, these three call sites (out of 65 total) were the only ones omitting --files, causing #2112's commit-scoping fix to never reach them. (#2269) (#2549)
  • /gsd-map-codebase Update mode now refreshes all date stamps — the **Analysis Date:** line, the *... analysis: ...* footer, and the <!-- refreshed: ... --> header are set to the current date on every run, overwriting any prior date. Previously, Update runs only replaced [YYYY-MM-DD] placeholder tokens, which don't exist in already-generated files (they contain concrete dates from the prior run), so stamps silently retained the original mapping date. (#2279) (#2550)
  • All seven guard hooks now normalize Kimi's payload shape — the five JS guards (gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard, gsd-read-injection-scanner, gsd-workflow-guard) and the two shell hooks (gsd-graphify-update.sh, gsd-phase-boundary.sh) normalize Kimi's native payload shape before their checks: the tool name (WriteFile → Write, StrReplaceFile → Edit, ReadFile → Read, Shell → Bash, bare or module-qualified), the tool-input fields (path → file_path, edit.old/edit.new — single or list — → old_string/new_string), and the PostToolUse tool_output field → tool_response, matching kimi-cli's actual tool and hook-event schemas. The two blocking guards (worktree path and workflow) also write their block reason to stderr, which is what Kimi feeds back to the model on exit 2. Previously the Kimi [[hooks]] matcher was translated to Kimi's vocabulary but the scripts' payload checks were not, leaving every guard — including the prompt-injection read scanner — dormant on Kimi while appearing registered. (#2304) (#2518)
  • parseCoverageMatrix now scopes table parsing to recognized coverage matrices — pipe-tables outside the matrix (e.g., summary tables) are ignored instead of being silently parsed as data rows, multi-section matrices with repeated headers are supported, and inline markdown emphasis (**OPT-OUT**) on decision cells is stripped before validation. Previously, the parser scanned every |-prefixed line file-wide with a latching header flag, causing silent phantom-capability corruption from unrelated tables, false rejection of multi-section matrices, and rejection of bold-emphasized decisions. (#2366) (#2551)
  • state.planned-phase now warns on no-op transitions and syncs progress.total_plans — when STATE.md's Current Position has no recognized labels (narrative prose), the command emits a warning field so the workflow can detect the no-op instead of continuing with stale state. When a plan count is provided, progress.total_plans in the YAML frontmatter is updated alongside the body Total Plans in Phase field, preventing contradictory state between the two representations. Previously, the command silently returned success with an empty updated array and zero bytes written, and left progress.total_plans at 0 while the body reported the actual count. (#2400) (#2552)
  • Codex --local installation no longer writes skills to $HOME/.agents/skills — the skills-kind home override (which redirects skills to the user-global .agents directory) is now only applied for --global scope. When --local is specified, skills are installed under the project-local config directory, matching the scope the user selected. Previously, a --local Codex install created a split installation: project-local config but user-global skills. (#2429) (#2553)
  • use_worktrees: false is now honored at the worktree dispatch gate — the per-plan dispatch condition checks BOTH the project-level USE_WORKTREES flag AND the per-plan USE_WORKTREES_FOR_PLAN variable. Previously, the dispatch gate checked only the per-plan variable (derived from submodule intersection), so plans that didn't touch submodules would still fork isolation="worktree" agents even when the project-level setting disabled worktrees entirely. The fix is net-negative in file size (prose compression offsets the added shell condition). (#2474) (#2561)
  • The Gemini and Claude reviewer legs now fail loudly instead of silently dropping out of the cross-AI review — both blocks capture stderr to a .err sidecar instead of discarding it to /dev/null, and write a diagnostic stub with the captured error when the lane produces no output. Previously they were the only two of the ten prompt-fed reviewer legs with neither guard, so any failure that wrote no stdout (CLI missing, unauthenticated, rate-limited, crashed) left a zero-byte review file that write_reviews rendered as a reviewer that had run cleanly with nothing to report — quietly degrading an N-reviewer consensus to N-1 while present_results reported success. The guard matches the shape the Codex and Cursor legs already use. (#2494) (#2592)
  • gsd-ui-auditor no longer documents an uncallable Playwright-MCP capture path — the agent's tools: allowlist grants no MCP namespace, so the <playwright_mcp_approach> block it presented as "preferred" could never dispatch: the availability check had a fixed answer, the three mcp__playwright__* calls were unreachable, and the CLI fallback was the only branch that ever ran. The dead block is removed, leaving the CLI screenshot path as the sole documented approach, and a new consistency test fails any agents/*.md that documents an mcp__<server>__* namespace its own tools: line withholds. Session-level Playwright-MCP capture in /gsd-ui-review is unaffected — that path is genuinely runtime-detected. The same documented-vs-granted drift is corrected one layer out in docs/AGENTS.md, where 26 of 34 per-agent Tools rows disagreed with the agent's frontmatter — 22 omitting Skill, 7 omitting Edit, 8 omitting MCP grants entirely (7 of them abbreviating up to eight distinct servers as "mcp (context7)"), and one still naming Task, a tool that no longer exists — with a parity guard added so the role cards and the frontmatter cannot drift apart again. (#2526) (#2594)
  • query commit --files no longer silently checks out the wrong phase branch mid-commit — the phase-token extraction is now anchored to the directory segment under .planning/phases/ and reuses the project-code-aware extractPhaseToken helper instead of an unanchored regex, so a project_code ending in a digit (e.g. PROJECT_V2) no longer makes …/PROJECT_V2-07-name/… match the 2- inside V2- and resolve to the wrong phase. The commit-path branch auto-switch also no longer silently force-switches an already-checked-out working branch onto a different existing phase branch (it creates-if-absent only, per the original #1278 intent); the only prior trace of the silent switch was a git reflog entry. (#2539) (#2669)
  • Reviewer/workflow config lookups no longer silently drop the configured value on machines without jq — review.md, plan-phase.md, ship.md, debug.md, autonomous.md, ai-integration-phase.md, and eval-review.md now resolve config-get scalars with the native --raw flag and resolve-model / resolve-execution / verification.status object fields with --pick, instead of piping through jq. Previously, on a stock Windows/Git-Bash box with no jq on PATH, the … | jq … stage failed (exit 127), the failure was swallowed by 2>/dev/null || <default>, and the configured per-lane model/host/budget came back empty — so the lane fell back to CLI defaults (e.g. ~/.codex/config.toml instead of the configured review.models.codex) with no diagnostic, and the autonomous.md verify gate could misroute on an empty status. The legitimate structured-JSON jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile) are untouched — only the jq-replaceable config/model/verify lookups moved to the native flags. Because those sites remain, /gsd-review now probes for jq up front and reports the ollama, lm_studio, llama_cpp, opencode, and antigravity lanes as unavailable with an install hint when it is missing, instead of running them into empty output; the gemini, claude, codex, coderabbit, qwen, and cursor lanes stay selectable with no jq installed. (#2589) (#2673)
  • Upgrading a Claude-global GSD install now uses the new version's skill content instead of the previous version's — the installer read a .gsd-source marker that still pointed at the prior install's source location before rewriting it, so on an upgrade every converted skill was generated from the old version's command definitions (while the file manifest faithfully recorded the stale content's hash as correct). The marker is now written before anything reads it. (#2624) (#2811)
  • phase complete and state begin-phase no longer rewrite current_phase_name to the name's own parenthetical — transitions that already hold the exact display name now pass it to syncStateFrontmatter as an authoritative override, so the lossy body-prose re-derivation never runs the final word on a field the transition just resolved. Previously, completing into a phase named Closer-ruling measurement (D1a) wrote current_phase_name: D1a (the prose parser's paren-over-dash preference harvested the name's own parenthetical), and every downstream consumer of the scalar inherited the mangled name. parsePhaseFromProse also gains status-keyword-aware precedence (the #1695 AC #3 residual) for genuinely unknown prose: the em-dash name wins when it is not a status keyword or Milestone: tail, so 48 — Closer-ruling measurement (D1a) now parses to Closer-ruling measurement instead of D1a. (#2736) (#2821)
  • Seven dangling references in the ADR corpus and contributor docs now resolve — (1) docs/adr/1239-gsd-embeddable-orchestration-engine.md linked the host-integration capability matrix as reference/… from inside docs/adr/, resolving to the nonexistent docs/adr/reference/; all three occurrences now use ../reference/…, and the two whose link text promises §codex now carry the matching #codex fragment. (2) src/plan-drift-guard.cts cited docs/adr/0022-source-grounding-drift-guard.md, a path that has never existed — corrected to the real docs/adr/22-plan-drift-guard.md; because the file is compiled into the shipped payload, the bad citation was shipping to users. (3) CONTRIBUTING.md and docs/contributor-standards.md illustrated the ADR naming convention with issue #3485, a pre-rename number from get-shit-done-redux that does not resolve in open-gsd/gsd-core — the worked example now uses #2264, which does, and the one genuinely historical #3485 reference is annotated rather than rewritten. (4) docs/adr/857-capability-system.md's H1 still carried a [Proposed] status bracket contradicting its Accepted — ratified 2026-07-17 Status field; the ADR index generator strips the bracket for display, so the contradiction was invisible to the gate. (5) scripts/gen-adr-index.cjs's back-link comment still described ADR-857 as Proposed and its claim over ADR-0011/ADR-58 as a supersession — both restated at the 2026-07-17 ratification, when the claim became Subsumes and the reciprocal back-links were added. (6) docs/how-to/install-on-your-runtime.md linked that same capability matrix as a bare host-integration-capability-matrix.md from inside docs/how-to/ in its ZCode and pi sections — the identical defect as (1), so both now use ../reference/…. (7) docs/CONFIGURATION.md cited ADR-1244 as adr/1244-runtime-capability-registry-overlay.md; the file is adr/1244-capability-ecosystem.md. (#2691) (#2692)
  • roadmap get-phase no longer drops success criteria that wrap onto a second line — the parser broke the criteria run at any indented continuation line, truncating the wrapped criterion (losing its trailing [REQ-ID] tag) and silently dropping every criterion below it. verify-work and plan-phase consumed the shortened list, so a phase could be planned and certified complete against a strict subset of its own success criteria with nothing reporting the gap. Continuation lines now fold into their criterion; blank-line-separated criteria still parse. (#2522) (#2637)
  • The host-integration capability matrix now documents the kimi-code runtime — kimi-code shipped as a distinct runtime but its section was never added, so its hostIntegration axes had no cited source. Sourcing each axis against Kimi Code CLI's own docs also corrected three values that had been inherited from the unrelated Python kimi CLI: embeddingMode is declarative (plugins are a manifest plus markdown Skills, with no in-process API), dispatch.nested is true (the coder built-in dispatches nested sub-agents), and dispatch.maxDepth is undocumented (no depth bound is published). (#2603) (#2687)
  • /gsd-profile-user now writes the runtime-native instruction file on Codex and other AGENTS-native runtimes — generate-claude-profile hardcoded .claude/CLAUDE.md for both project and global scope, ignoring the runtime-aware resolution that #3163 wired into the sibling generate-claude-md handler. The #3163 fix diverged when it didn't propagate here, so running $gsd-profile-user --refresh on a Codex install created/modified Claude configuration instead of producing a Codex AGENTS.md profile. The command now resolves its target through the shared runtime policy: project scope uses getProjectInstructionFile(runtime) and global scope derives ~/.<config-home>/<instruction-basename>, so codex lands at ~/.codex/AGENTS.md. Claude behaviour is preserved. A parity test guards against future re-divergence between the two handlers. (#2659) (#2659)
  • The plan-phase decision-coverage gate can no longer silently pass when its context-path argument is missing — the handler now fails closed on an empty/missing argument (a caller error), and the plan-phase workflow recomputes the CONTEXT.md path in the same Bash block that runs the gate (the variable set in the init block did not survive into the gate block). A genuinely-absent CONTEXT.md still produces the legitimate green skip. Previously the gate reported passed without ever checking coverage. (#2770) (#2881)
  • /gsd-code-review no longer silently drops CRLF-saved artifacts — the Tier-2 file-scope extractor (and every REVIEW/REVIEW-FIX frontmatter reader in the code-review and code-review-fix workflows) used a literal \n to find the YAML block, so any SUMMARY.md/REVIEW.md saved with CRLF line endings (default on Windows) contributed zero files with no warning. The boundary now normalizes CRLF first, so a mixed CRLF/LF phase reviews the union of its files instead of an incomplete set. (#2694) (#2839)
  • Dev-dependency brace-expansion bumped to patched versions (1.1.18 / 5.0.9), resolving the high-severity DoS/OOM advisories — the lockfile now pins the 2026-07-30 patch backports reachable via eslint and stryker. A non-breaking in-range bump (no overrides, no major bumps); production npm audit --omit=dev is unaffected (devDependency only). (#2765) (#2888)
  • The markdown-parsing lint rule now catches the stricter cell-regex spelling it previously missed — a hand-rolled table scan written as [^|\n] (excluding both the pipe and the newline, which is the more correct form) slipped past the guard entirely, so STATE.md field replacement kept parsing tables with a local regex and rewriting the whole document. The rule now flags any pipe-excluding character class, and the STATE.md field writer edits a bounded byte range instead. (#2880) (#2889)
  • Workstream-scoped config reads now inherit from the project root config — config-get under an active workstream (GSD_WORKSTREAM) now resolves a key absent from the workstream's own config to the project-root value before falling back to schema defaults, instead of reporting 'Key not found'. A workstream config still overrides root for any key it sets; root only fills gaps. Previously a key set only at root was silently lost under a workstream, causing shipped workflow boolean guards (e.g. use_worktrees, plan_review_convergence) to apply their hardcoded fallback and silently invert the user's setting. (#2833)
  • The Claude-orchestration Workflow backend can now actually dispatch a wave — every script emitWorkflowScript generated was rejected by the Workflow tool. It omitted the required export const meta = {…} first statement (fatal on its own), called resumeFromRunId() and budget() which are a tool input parameter and a read-only object rather than script functions, and passed parallel(agent(…), agent(…)) where an array of thunks is required. Two further defects meant the script was never even reached: nothing resolved the Agent SDK version, so the gate ladder returned agent_sdk_version_unknown on every automated run while capability state still reported the capability active; and the runtime fallback diverged from the canonical GSD_RUNTIME > config.runtime > 'claude' chain, so any invocation without --runtime reported runtime_not_claude. The router now resolves the installed SDK version itself and defers to the canonical runtime resolver, and the emitted script is valid ES module syntax with phase() titles matching meta.phases. (#2590) (#2681)
  • Releases no longer fail their own emitted-parity gate — cutting any release ran the differential attribution check against a baseline built at a different version, so the install-time hook version stamp made all 364 emitted hook paths look like unexplained drift and every finalize/rc run hard-failed before tagging or publishing. (#2891) (#2894)
  • Merging an emitted-drift acknowledgment no longer turns the mainline red. An acknowledgment is now scoped to the diff that introduced it, so once its ripple is absorbed into the base it goes inert instead of reporting as stale — which had reddened next for five consecutive commits and every pull request branching off it. (#2789) (#2803)
  • Discuss-phase advisor mode now spawns the registered gsd-advisor-researcher subagent instead of general-purpose — resolving a contradiction with the universal-anti-patterns rule (injected into the same context) that forbids non-GSD agent types. The manual "read the agent def" prompt line is dropped (spawning by type auto-loads it). (#2771; the sibling assumptions-site needs a design decision — filed as #2883) (#2886)
  • Subagent spawns no longer fail on non-Claude runtimes when no model resolves — 15 workflows told the orchestrator to pass a model parameter without saying to drop it when nothing resolved, so 43 dispatch sites sent an empty model and the spawn 404'd. That was the default state on Codex, OpenCode, Gemini CLI, Kilo, Qwen and Hermes, where GSD sets resolve_model_ids: "omit" on install. Every dispatching workflow now carries the rule. (#2711) (#2713)
  • The statusline now renders GSD state correctly on Windows-authored (CRLF) STATE.md — parseStateMd no longer drops the entire frontmatter block on CRLF input. The fence regex and downstream splits now accept CRLF line endings, matching the canonical extractFrontmatter parser. Previously a CRLF STATE.md silently produced an empty GSD-state segment (no status, phase, or milestone) with no error. (#2754) (#2865)
  • api-coverage now ships the #2366 coverage-matrix fix — the tracked gsd-core/bin/lib/api-coverage.cjs build artifact had drifted four days behind src/api-coverage.cts, so the module that actually ships still parsed non-coverage tables as data, mishandled multi-section matrices with repeated headers, and failed to parse **OPT-OUT**. Regenerated, plus a new lint:generated-sync check that fails when any tracked compiled artifact no longer matches its source. Also prunes two stale entries from the no-phantom-issue-refs guard: GitHub numbers issues and PRs from one shared counter, so both had since become real merged PRs, and the guard was rejecting accurate citations of them. (#2653) (#2656)
  • OpenCode/Kilo no longer spawn the context-monitor subprocess on every tool call when context warnings are disabled — the adapter now reads the existing hooks.context_warnings toggle in-process and skips the child-process spawn entirely when it is set to false, instead of paying a Node boot per tool call only to read the flag and exit inside the child. Behavior is unchanged when the toggle is absent or enabled (the default). (#2824)
  • Editing src/ no longer trips an undocumented changeset-lint failure — CONTRIBUTING.md listed the Changeset Required triggers without src/, the path that compiles into every gsd-core/bin/lib/*.cjs, so contributors touching it hit a CI failure the docs said could not happen — and a local run of the lint reported success regardless, because it silently requires GITHUB_BASE_REF to see the branch at all. Both are now documented, and the config-loader test-helper that reset only one of its two warning-dedup sets now resets both. (#2674) (#2678)
  • The .planning/ write reminder can no longer be suppressed or fabricated by a model-supplied file_path — the phase-boundary hook now treats tool_input.path (the field kimi-cli actually executes on) as authoritative and file_path as the fallback, reaching the same "path authoritative" outcome the JS guards establish via upstream normalization (#2595). Previously a model-controlled decoy file_path could silence the reminder for a genuine .planning/ write or raise one naming a file never touched. (#2752) (#2860)
  • test:/chore:/ci:/docs:/refactor:/perf:/revert: PRs no longer publish under the user-facing Enhancement heading in release notes — the release-notes classifier now routes recognized non-user-facing conventional-commit types to an Internal bucket and omits them from the published GitHub release notes (and the Discord announcement's user-facing sections). Previously these internal-work PRs rendered as Enhancements alongside genuinely user-facing changes. feat:/fix: classification is unchanged, and untyped or anchor-defeated titles still fall back to Enhancement. (#2838)
  • /gsd-execute-phase and /gsd-quick branches no longer auto-track origin/master — the branch-creation git checkout -b <branch> origin/$DEFAULT_BRANCH omitted --no-track, so with the default branch.autoSetupMerge=true git wired the new branch's upstream to refs/heads/$DEFAULT_BRANCH. A subsequent GUI sync (GitHub Desktop, VS Code) then pushed the branch's commits straight onto origin/$DEFAULT_BRANCH, bypassing PR review — in one project every commit of a 7-plan phase landed on origin/master. --no-track is now passed; the first git push -u origin <branch> sets up correct same-name tracking. (#2498) (#2628)
  • Cursor CLI sessions now detect .planning/ — the sessionStart and stop hooks resolved the project from process.cwd(), which under the cursor-agent CLI is the Cursor config dir (~/.cursor), not the workspace. Every CLI session therefore reported "no .planning/ workflow found" even with .planning/STATE.md present, and the stop hook's verify-work reminder could never fire. Both hooks now read workspace_roots from the hook payload they already buffered but never parsed, preferring the root that actually carries .planning/STATE.md (multi-root workspaces) and falling back to the first root, then cwd so IDE invocations are unchanged. (#2587) (#2680)
  • Plan, summary, verification, and state validators now reject NUL-corrupted files — frontmatter validate, verify plan-structure, and state validate now fail loud (valid:false) when a file contains embedded NUL bytes, with an error naming the encoding problem and its downstream consequence. Previously such a file passed as valid:true but was silently skipped by recursive/binary-skipping search tools (rg, grep -I), reading downstream as 'file absent' rather than 'file corrupt.' (#2829)
  • OpenCode no longer declares background subagent dispatch it does not have — capabilities/opencode/capability.json advertised dispatch.background and dispatch.backgroundDispatch as true, but OpenCode's native subagent dispatch is synchronous: the Task tool's background parameter is hidden from the model behind the opt-in OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS flag, which defaults to false, and the session loop still handles one subtask at a time. Since negotiateHostCapabilities and every degradationFor consumer trusts these per-field values, declaring an absent capability overstated it — the opposite of the fail-closed posture the negotiation exists to enforce. Both fields are now false, and the host-integration capability matrix carries the corrected values with current upstream citations. (#2598) (#2682)
  • Debug sessions now commit their session docs — with commit_docs: true, finishing a /gsd:debug session left the session doc (and sometimes the fix's own code changes) sitting untracked in the working tree. The session manager, which owns the end of a debug session, never had a commit step — only the single-spawn debugger path did. Terminal sessions now commit the doc and any uncommitted in-session fix code, still respecting commit_docs; sessions that pause mid-investigation deliberately do not. (#2568) (#2731)
  • Codex installs now ship the complete update-check hook set — the --codex installer (both --profile=core and --profile=full) now installs and refreshes all four hook files the update-check/context-monitor feature needs (gsd-check-update.js, gsd-check-update-worker.js, managed-hooks-registry.cjs, gsd-context-monitor.js) together, instead of only the two parent scripts. Previously a registered parent hook pointed at a worker and registry the same installer never delivered. (#2695) (#2822)
  • Discuss-phase no longer carries four internal text contradictions — auto-mode removed a dead max_discuss_passes config read that contradicted its single-pass rule; the gate-prompts reference now matches the actual context-handling options and drops the 'Let Claude decide' cop-out that conflicted with the workflow's no-skip rule; the auto_advance fallback no longer routes back to the already-run confirm_creation step; and the assumptions workflow's answer_validation is re-synced to the canonical parent block. (#2886)
  • Corrected the legacy ADR range documentation — the legacy zero-padded ADR range is now stated once (in docs/adr/README.md, as 0001–0012) and referenced rather than restated by docs/contributor-standards.md, so the two can no longer drift. The two zero-padded files that look legacy but are not (0174, 0656) are now identified as modern, mis-padded issue-numbered ADRs. Previously the two documents disagreed and neither matched disk. (#2836)
  • Agents and workflows no longer instruct a bare gsd-tools that fails on a shim-only install — command-position gsd-tools invocations in the shipped agent/workflow source are now the portable gsd_run resolver (already defined in those files), so they resolve the runtime-local shim on installs with no gsd-tools binary on PATH. Previously only the Codex install-conversion pipeline rewrote these; the Claude-facing source shipped them verbatim and failed with command not found. (#2751) (#2851)
  • query commit --files now accepts absolute paths — cmdCommit used path.join(cwd, file), which concatenates instead of resetting on an absolute path, so absolute --files entries (e.g. the absolute phase_dir emitted by init phase-op since #2428) were joined to cwd+absPath (non-existent) and silently dropped as nothing_to_commit — and a mixed relative/absolute list committed the relative entries while reporting committed:true. Absolute paths are now normalized to repo-relative before staging/branch-detection, so they commit correctly and the phase-branch detection no longer matches digit-hyphen runs in the absolute prefix. (#2523) (#2638)
  • /gsd:review no longer silently drops a reviewer you asked for — naming a reviewer with an explicit flag (--gemini --qwen) on a host where that lane could not run reported an info note and reviewed with a thinner set, while the run reported success; a cross-AI review that quietly loses a lane is blind in one eye. An explicitly-named lane that cannot run — CLI absent, jq missing, or local server unreachable — is now an error. --all and review.default_reviewers are unchanged and still skip undetected lanes with an info note. The Qwen lane also now captures stderr to a sidecar and includes it in its failure stub, matching every other lane, so a missing binary and an auth prompt are no longer indistinguishable from an empty review. (#2794) (#2820)
  • EoS Registry entries carrying the documented effortSurface axis are no longer rejected — the registry validator required an exact eight-key axes object, so an entry that faithfully mirrored its upstream descriptor's optional ninth effortSurface key (argv or none, added by ADR-1239 amendment #2481) failed validation outright. (#2810) (#2813)
  • The claude-orchestration Workflow backend now honors your model settings — with that BETA capability enabled, every plan was dispatched with no model at all, so model_overrides, model_policy and model_profile were silently ignored and each agent ran on whatever the session happened to be using. Plans now run on the same model the normal dispatch path would have used, and the generated script states which model was applied. Two consequences to expect: agents that were inheriting the session model will now run on the model your profile selects, and the first run after upgrading re-executes any in-flight resumable run, because the dispatch options changed. (#2686) (#2715)
  • A truncated or half-written frontmatter file is no longer silently read as "no metadata" — a document whose --- fence was opened and never closed used to return exactly the same empty result as a file that legitimately has no frontmatter, so a crash mid-write left every phase/state reader proceeding with empty contracts and no signal. GSD now names the offending file on stderr while returning the same value as before, so nothing that consumed the old result changes. A Markdown horizontal rule at the top of a document — including one above a labelled line such as Note: or Author: — is not mistaken for a truncated fence. (#1882) (#2712)
  • Refusing to run a phase from an executor worktree now tells you how to recover your work — when GSD stopped because the session had drifted into an executor worktree, it only said to re-run from the orchestrator's worktree. If that worktree held commits or uncommitted changes, following that advice silently abandoned them. The refusal now lists the commits and files that exist only there, and gives the exact steps to integrate them before continuing. (#1856) (#2727)
  • An unreadable ROADMAP.md is no longer reported as a brand-new project — a permission or I/O error reading .planning/ROADMAP.md used to return the same "phase not found" and v1.0 / milestone values as a project that simply has no roadmap yet, so workflows synthesized a blank phase or skipped requirement extraction with no signal. GSD now names the unreadable file on stderr while returning exactly what it returned before. A project that genuinely has no ROADMAP.md stays silent. (#1881) (#2729)
  • A corrupt .planning/config.json no longer silently discards your entire configuration — a single trailing comma used to fall back to built-in defaults with no signal, indistinguishable from having no config file at all, so a project could run for weeks on defaults while its model profile, workflow toggles and branching strategy sat unread on disk. GSD now tells you the file could not be used and that its settings were not applied, and reports the cause (config_unparseable / config_unreadable) distinctly from genuine absence. The same applies to an unreadable file and to the global ~/.gsd/defaults.json. (#1880) (#2688)
  • --validate is no longer documented for /gsd-plan-phase and /gsd-execute-phase — both commands silently ignored the flag (only /gsd-quick implements it), so the docs promised a state-validation step that never ran. The false flag-table rows, CLI examples, and the manager.flags.execute: "--validate" config example are removed across the English docs and the ja-JP/zh-CN/ko-KR/pt-BR mirrors; the config example now shows --cross-ai (a flag execute-phase actually parses). /gsd-quick's --validate docs are unchanged. (#2197) (#2574)
  • /gsd-plan-phase no longer 404s on non-Claude runtimes with model_profile:"inherit" + resolve_model_ids:"omit" — the workflow passed model="{planner_model}" (and researcher_model/checker_model) verbatim into Agent() calls, so when the resolved model was empty it sent model="" and the runtime fell back to an unavailable Claude model → 404. plan-phase now mirrors execute-phase: when a *_model is "inherit" or empty, the model= param is omitted so the subagent inherits the orchestrator model. (#2517) (#2634)
  • Stale todos/done references in workflows and docs now read todos/completed — the todos/done → todos/completed rename (commit 447d17a9) under-swept 14 descriptive lines across check-todos.md, the /gsd-help tree, ARCHITECTURE.md, and USER-GUIDE.md (en + 4 locales). Those stale references steered agents and users to archive closed todos into done/ — a directory nothing in gsd-core reads — so closed todos became invisible to ID sequencing and to anything that inventories closed work. All 14 sites now read completed/, matching the canonical code path (cmdTodoComplete). A CI guard now blocks future under-sweeps. (#2491) (#2626)
  • Verification-status next-step commands now use the command surface each runtime actually installs — on a Codex project, a phase blocked on verification suggested /gsd:execute-phase, which Codex does not install; the correct form is $gsd-execute-phase. The routing table stored hard-coded, deprecated colon-form strings with no runtime context, so phase complete and query verification.status relayed them verbatim to every runtime. All four routed states (missing, unknown, gaps_found, stale) now project through the shared runtime formatter. (#2617) (#2700)
  • A failed LM Studio or llama.cpp reviewer leg is now visible instead of silently dropped — when a local OpenAI-compatible endpoint was unreachable or returned empty content, /gsd-review wrote no review file at all, so the reviewer's section was omitted from the final review and the result was indistinguishable from that reviewer never having been selected. Both legs now emit a diagnosable stub carrying curl's stderr and the raw response body, matching the guard the claude/gemini/codex legs already had. (#2605) (#2689)
  • /gsd-execute-phase now auto-closes pending todos for single-digit phases — the close_phase_todos step normalizes both the phase number and each todo's resolves_phase value before comparing, so a todo tagged resolves_phase: 5 is recognized when phase 05 completes. Previously the step compared the zero-padded PHASE_NUMBER (e.g. "05") against the unpadded value new-milestone wrote (e.g. "5") as literal strings, so every single-digit phase (1-9) silently failed to auto-close its todos — they stayed stuck in pending/ forever despite their resolving phase completing. Decimal sub-phases (4.1 vs 04.1), letter suffixes, and quoted YAML values are now handled too. (#2576) (#2597)
  • The host-integration capability matrix now documents the effortSurface axis for every runtime — the axis shipped in #2481 with real values in 19 runtime descriptors, but the matrix that ADR-1239 designates its cited source of truth had no legend entry and not one per-runtime row, so every committed value was undocumented in the one place meant to explain it. (#2615) (#2698)
  • STATE.md frontmatter is no longer silently overwritten by stale field lines in archive sections — buildStateFrontmatter extracted Last Activity, Paused At, and the other current-state fields from the entire STATE.md body via stateExtractField, which matches the first Field: line anywhere. A historical line in an archive section further down the file silently overwrote the correct frontmatter value on every sync, and because the poisoning line stayed in the body it regressed again on the next write — so each repair looked successful and then silently reverted, with the offending line hundreds of lines away from the frontmatter. Field extraction is now scoped: current-state fields read from the body preamble before the first ## heading, and session fields read from ## Session. This generalizes the #2444 fix, which scoped Stopped At to ## Session but did not propagate to the sibling fields. (#2660) (#2660)
  • A commit whose git add fails now says so, instead of partially committing or reporting "nothing to commit" — when staging failed (an unwritable index in a linked worktree, permissions, or a timeout), GSD discarded the error: a multi-file request silently committed only the paths that happened to stage, and a total failure surfaced as nothing_to_commit or a downstream pathspec error naming an innocent file. Staging failures are now collected and reported as staging_failed (or staging_timeout) with the offending file and git's original stderr, before any commit is attempted, and the index is rolled back to its prior state. Applies to scoped (--files) commits, default .planning/ commits, and sub-repo commits alike. (#2608) (#2693)
  • Cursor, Windsurf, and Codex hooks no longer fail with require is not defined under an ESM config root — GSD now writes the {"type":"commonjs"} marker into the hooks directory alongside the staged .js scripts for these three runtimes (it already did for every other runtime), so Node loads them as CommonJS regardless of the runtime config's "type". (#2717) (#2846)
  • The portability linter now catches Windows-path failures in membership and substring assertions — no-path-literal-in-assert flags .includes/.indexOf/.startsWith/.endsWith/.match over a path-returning receiver (including through a .map() hop), not just equality assertions. Previously these passed lint and failed on Windows CI; the rule now surfaces them at lint time. (#2764) (#2879)
  • /gsd-review's codex lane no longer passes the hook-trust bypass flag or runs its capability probe — host-harness safety classifiers denied invocations carrying them, and flagless invocations work in steady state. A genuine untrusted-hook failure still surfaces as a dropped lane with diagnosable stderr. (#2479) (#2536)
  • /gsd-plan-phase --reviews now actually replans in chunked mode instead of silently skipping every plan — the per-plan resume-check skips existing plans for crash-resume, but now exempts --reviews (whose purpose is to replan with review feedback). Also fixed the outline resume-check, which looked for a marker the agent only returned (never wrote to the file), so the outline always re-ran. (#2762) (#2887)
  • pi no longer silently hijacks non-Anthropic providers' model choices — pi/gsd.cjs's before_provider_request handler unconditionally rewrote payload.model to the built-in pi/sonnet tier default (claude-sonnet-5) via the model-catalog fallback, breaking every outgoing request for pi users on non-Anthropic providers (kimi-coding, zai, openrouter, openai-codex, minimax). The handler now inspects model_profile_overrides.pi[tier] explicitly before calling resolveTierEntry (whose catalog fallback previously masked the "user did not opt in" signal) and fail-opens (return undefined) when the user has not set an override — including explicit null and '' (clearing a previously-set value). An explicit opt-in via model_profile_overrides.pi[tier] still steers, preserving the legitimate use case. (#2460) (#2499)
  • GSD_AUDIT=1 now actually produces an audit trail — the reference dispatch logger is wired onto the live command seam, so opting in yields the documented structured stderr line and the .planning/.gsd-trace.jsonl audit trail. Previously the seam built its dispatch hub without a logger, so it fell back to a no-op and the opt-in signal was inert with no indication why. With observability off, dispatch output is byte-for-byte unchanged. (#2620) (#2621)
  • Codebase scan and ship-time capability hooks now honor your model settings — /gsd:scan dispatched its mapper agent with a model placeholder nothing resolved, and ship-time capability hooks did the same, so model_overrides and model_policy were silently ignored at both and the agent ran on whatever the session happened to be using. Both now resolve a real model, and omit the model parameter entirely when it resolves to "inherit" or empty rather than passing an empty value that fails on non-Claude runtimes. Note: the scan mapper now runs on the model your profile selects rather than inheriting the session's. (#2684) (#2710)
  • State sync now reports the correct total phase count on a flat unmilestoned roadmap — progress.total_phases no longer falls back to the on-disk phase-directory count when the roadmap has no versioned milestone heading; it uses the authoritative roadmap count, matching the write-path and resolving the contradiction between smart-entry's total_phases and roadmap_total_phases. (#2828) (#2892)
  • Worktree cleanup-wave now rescues uncommitted SUMMARY.md — the rescue step's git cat-file -e HEAD:<path> check assumed an absent path returns exit 1, but git returns 128, so rescue never fired: the executor's uncommitted <id>-SUMMARY.md blocked cleanup as worktree_dirty and risked silent loss on worktree remove --force. Rescue now fires on any non-zero exit (only exit 0 = committed → skip), so uncommitted SUMMARYs are copied into the main tree before the dirty check. (#2556) (#2611)
  • Code-review now scopes repository-root and extensionless build files (Dockerfile, Makefile, .gitlab-ci.yml, renovate.json, AGENTS.md) — the SUMMARY.md file extractor no longer silently drops every root-level path and every extensionless build file, and a partial SUMMARY scope is now cross-checked against git diff with a warning naming any changed files it missed. (#2666) (#2895)
  • execute-phase.md now has ~3.3 KB of byte-budget headroom — the offer_next step body (terminal reporting + next-phase routing prose) was extracted to gsd-core/references/offer-next.md and eagerly @-referenced, restoring the headroom the frozen size ceiling exists to provide. Previously the ceiling had only ~32-137 bytes of margin, so any bugfix touching execute-phase.md had to extract unrelated content or raise the ceiling. Runtime behavior is unchanged (the @-reference loads eagerly). (#2537) (#2642)

Security

  • Malformed and shadowing Kimi payloads no longer disarm the guards that block — normalizeKimiPayload (inlined in all five PreToolUse/PostToolUse guard hooks) rebuilt old_string/new_string with String(e.old ?? ''). Two inputs crashed it, and because normalization runs before any tool dispatch, both crashes landed in each guard's outer catch { process.exit(0) } — which emits the same exit code as "nothing to report", turning a should-block call into a silent allow. First, ?? guards the value and not the dereference, so a nullish entry (edit: [null]) threw on the property read. Second, coercion itself can throw: {"toString": null} is valid JSON that raises Cannot convert object to primitive value, so even a well-formed edit object could crash normalization. Two hard blocks were bypassable through either route: gsd-worktree-path-guard's cross-git-root write block (the same write is correctly blocked with a well-formed edit list), and gsd-workflow-guard's force-add block on agent-* branches (via a Shell payload carrying a spurious edit field the Bash path never even reads). Fixed with e?.old / e?.new plus a guarded coercion, landed identically across all five copies; the coercion is wrapped rather than type-tested so that stringification is unchanged for every value that can coerce. Three model-supplied fields are now authoritative rather than merely defaulted. Normalization used to fill file_path, old_string and new_string only when the key was === undefined, so any value the model chose to include won — while kimi-cli executes on path and edit. Its StrReplaceFile schema is path + edit only (src/kimi_cli/tools/file/replace.py @ 4a550ef) and carries none of those three keys, so each one appearing in a Kimi payload is always model-supplied. A cross-root path paired with a spurious file_path: "" left gsd-worktree-path-guard reading an empty string and exiting 0 while the identical write without the extra key blocked; likewise a new_string: "" — or any benign non-empty decoy, which a type test would not have caught — left gsd-prompt-guard's injection scan reading empty content and returning at its if (!content) guard before it ever saw the real edit[].new. All three are now reconstructed unconditionally, which can only ever narrow what a guard inspects to what will actually be written. Reachability is not speculative: kimi-cli's soul/toolset.py json-parses the model's raw tool arguments and passes the dict verbatim as tool_input to PreToolUse, doing typed validation only later inside tool.call() — so the model controls extra keys at the moment the hook decides. Separately, the guards now read payload path fields typed. A non-string file_path ([], {}) is truthy, so it survived each guard's if (!filePath) early-out and then threw inside path.isAbsolute() / .includes() / .replace(), reaching the same fail-open catch — crash-to-allow through the guard's own read rather than through normalization, and live on native Claude Code payloads too, since normalization returns early for non-Kimi tool names and so never masked the bad value there. Previously this was closed only as a side effect of a valid string path overwriting file_path; it is now closed unconditionally at all six read sites (the five normalized guards plus gsd-windsurf-pre-write, which already read typed), and a source-level invariant (tests/kimi-guard-typed-payload-reads.test.cjs) fails if any hook regresses to an untyped read. The native Claude Code contract (file_path governs) is unchanged. Scope on Kimi: normalization makes each guard's checks run; it does not make every guard enforceable. What can actually block on Kimi is what runs at PreToolUse — the worktree cross-root write block and the workflow force-add block. gsd-read-injection-scanner is a PostToolUse hook, and kimi-cli's dispatch never inspects PostToolUse hook results (soul/toolset.py fires them as a detached task and returns the tool result without awaiting it), so no output shape the scanner emits can block or flag a Kimi tool call; its prompt-injection block is not enforceable on Kimi under Kimi's current hook architecture. Regression coverage is negative-controlled against the pre-fix guards, and a property test (tests/kimi-normalize-payload.property.test.cjs) backs the totality claim generatively. next-only — released versions carry no Kimi normalization at all. (#2547) (#2595)
  • Dev-tooling js-yaml bumped past the merge-key DoS advisory — js-yaml was pinned ^4.2.0, inside the vulnerable 4.0.0 - 4.2.0 range of GHSA-52cp-r559-cp3m (quadratic CPU on YAML merge-key chains). It is a devDependency with no shipped-runtime reachability, but scripts/workflow-policy.cjs parses workflow frontmatter in CI, which is attacker-controlled on a fork PR. Now ^4.2.1. (#2654) (#2655)

[1.8.0] - 2026-07-22

Added

  • A default-off, BETA, claude-only "Claude orchestration" capability — adopts Claude Code's Workflow tool (/effort ultracode, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, restoring the wave parallelism + plan-checker + verifier that the #853 backgrounded-agent nesting limitation forces inline on Claude Code, and folding the existing gsd-ultraplan-phase plan-offload under the same runtime gate. When claude_orchestration.enabled is on AND the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets the floor (claude_orchestration.min_agent_sdk_version, default 0.3.149), execute-phase emits a generated Workflow script (waves → parallel() barriers, plans → agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap → separate sequential stages, resumeFromRunId wired to the phase run id, shared budget pool) that composes the SAME executor agent + worktree isolation the inline path uses, so artifacts/commits are produced identically. Detection is pure and fail-closed (any miss → inline), so on any runtime lacking the Workflow tool behaviour is byte-identical to today. Adds a pure module gsd-core/bin/lib/claude-orchestration.cjs (detectWorkflowBackend, emitWorkflowScript), the capabilities/claude-orchestration/ declaration with two gated loop contributions (execute:wave:post, plan:post) and a claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow), federated config keys, and an ADR-1143 implementation amendment. (#1143) (#2044)
  • Phases that integrate an external API/SDK/service can no longer seal without a decided coverage matrix — a new api-coverage gate on the ai-integration capability blocks /gsd:verify-work until the phase produces a COVERAGE.md enumerating the API's full capability surface, with every non-integrated capability an explicit, reasoned opt-out. Full coverage is the default; the matrix is the subtraction record, so "we integrated the API" can no longer silently mean "we integrated whatever the first use case exercised." Toggleable via workflow.api_coverage_gate (on by default). (#1562) (#2065)
  • OpenCode installs now auto-register the GSD companion MCP server (mcp.gsd) — --opencode install writes a mcp.gsd entry (local stdio → gsd-mcp-server) into opencode.json, so OpenCode drives GSD's command + planning-state surface over MCP with no bespoke plugin (ADR-1239 Phase D / #1682). Idempotent and non-clobbering; a user-defined mcp.gsd is preserved. (#1682) (#1929)
  • OpenCode plugin handles session.idle + the opencode-subset hook dialect is implemented — the GSD OpenCode plugin now recognizes session.idle (↔ Claude Stop lifecycle point), completing the compaction/idle pair (#1914 shipped compaction). The reserved opencode-subset dialect gains a consumer — hookEventSurfaceFor() in host-integration.cts — describing OpenCode's session/tool/file event subset (no workflow-phase events; the engine owns phase sequencing, ADR-1239 §OpenCode binding). Adds a Claude-parity test asserting the plugin covers the full declared subset. (#1682) (#1930)
  • GSD now warns when model config changed without re-running the installer on static-frontmatter runtimes — on codex and opencode, editing model_overrides or model_profile_overrides or model_policy.runtime_tiers in .planning/config.json or ~/.gsd/defaults.json previously had no effect until the user re-ran gsd install <runtime>, and the failure was silent: the sub-agent kept using the base model. Workflow entry points like gsd-tools init * now emit a one-line stderr warning naming the changed config file and the exact remediation command when they detect the config is newer than the baked agent files. The guard is read-only and warning-only by default, dedup'd per session, and skipped entirely on Claude Code because Claude Code resolves models at spawn time. Resolves #1688 as the structural follow-up to #1650. (#1692)
  • gsd-tools state rebuild — new subcommand that re-derives STATE.md body structure from canonical sources (frontmatter + .planning/phases/ disk scan), reconciling drifted ## Current Position prose, dropping orphaned rows from the **By Phase:** table, clearing template-placeholder field values, and de-duplicating ## Session Continuity Archive blocks. Every mutation is recorded in a ## Rebuild Log audit section. Idempotent (running twice on a clean file is a no-op). Supports --dry-run (preview) and --verbose (tee log to stderr). Heavier, manual counterpart to the lightweight auto-triggered state sync. (#1830)
  • graphify.graph_path makes the knowledge-graph location configurable so one umbrella graph can serve multiple projects — a new .planning/config.json key (path relative to project root, or absolute) overrides where /gsd-graphify query|status|diff read the graph, letting a single curated cross-repo umbrella graph serve every sibling sub-project without N drifting ~5 MB mirror copies. Previously the graph location was hardcoded to <cwd>/.planning/graphs/ with no override; the only workaround was copying the umbrella graph.json into each project (which drifted, wasted disk, and could be silently overwritten by an in-project build). The diff snapshot travels with the configured graph; build stays project-scoped; unset → byte-identical default; a configured-but-missing file yields an actionable error naming the path. (#1825) (#2013)
  • Claude Sonnet 5 is now the standard (sonnet) tier model. The model catalog and provider presets resolve the sonnet/standard tier to claude-sonnet-5 (GA 2026-06-30) across the Anthropic-backed runtimes (claude, copilot, and the anthropic/anthropic-fable presets), plus the OpenRouter-style anthropic/claude-sonnet-5 for opencode/hermes, replacing the superseded claude-sonnet-4-6. Opus and Haiku tier defaults are unchanged (the haiku high-effort preset's escalation slot tracks the current sonnet model). Shipped in 1.6.1. (#1847) (#1848)
  • gsd-debugger now guards fix acceptance with a multi-signal anti-overfitting gate — a fix that greens the target test can no longer be silently accepted. The debugger now runs a five-signal guardrail before accepting a fix (target test, mutation check via Stryker, no-op/behavior-deleting diff detector, adjacent/held-out tests, and revert-and-reconfirm), degrades gracefully when Stryker or a test suite is absent (each skip is logged, never a silent pass), records every signal's result under Resolution.verification in the debug file, and returns a FIX REJECTED BY GUARDRAIL outcome that gsd-debug-session-manager surfaces for revise / accept-as-documented-debt / abandon. Full rules live in gsd-core/references/debugger-fix-acceptance.md. (#1958) (#2396)
  • gsd-debugger now ranks suspect code by Ochiai suspiciousness before forming hypotheses — when a runnable test suite with per-test coverage exists (≥1 failing and ≥1 passing test), the debugger computes a spectrum-based fault-localization (Ochiai) ranking over the coverage and seeds the top-N suspicious locations into the Evidence section as first-class hypothesis candidates, narrowing the search space deterministically before any LLM reasoning. Tarantula is documented as a fallback formula. The step degrades cleanly (logged, never a silent pass) when there is no test suite, no failing tests, or no per-test coverage, and it is explicitly not trusted on flaky/Heisenbug spectra (pairs with the Phase 2B bug-taxonomy routing). Full rules live in gsd-core/references/debugger-sbfl.md. (#1959) (#2403)
  • gsd-debugger now branches root-cause analysis instead of chaining, guarding against 5-Whys single-cause bias — before committing root_cause, the debugger enumerates candidate causes across ≥2 Ishikawa categories (code / config / environment / data) rather than a single linear "why" chain, and explicitly answers an AND-gate question ("could this failure require more than one contributing condition simultaneously?"). When the AND-gate fires, every contributing cause is recorded — so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may now hold one OR a small set of contributing causes (additive; a single-cause session still records exactly one root_cause while the reasoning_checkpoint gains two RCA fields populated in every session). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields, and debugger-philosophy.md adds the single-cause-bias trap to its cognitive-bias table. Full rules live in gsd-core/references/debugger-rca-branching.md. (#1960) (#2405)
  • gsd-debugger now classifies each failure by bug class and routes the investigation technique accordingly, replacing the flat 11-technique menu with selection-by-class — at a new Phase 1.75 the debugger assigns a bug_class (Bohrbug / Heisenbug-Mandelbug / Concurrency) and consults an explicit, inspectable routing table: Bohrbugs route to deterministic reproduction + SBFL (Phase 1.25) + git bisect; Heisenbugs/Mandelbugs route to record-replay (rr) + stability-stress + statistical sampling and explicitly skip SBFL (a flaky spectrum poisons the ranking); Concurrency bugs surface the atomicity/order/deadlock checklist before general techniques. The 11 techniques remain as routed targets, not an undifferentiated list (supersede, not append). bug_class + chosen strategy are written to the debug file; the common-bug-patterns catalog is cross-referenced to the taxonomy. Full rules live in gsd-core/references/debugger-bug-taxonomy.md. (#1961) (#2407)
  • gsd-debugger now hardens regression tests via PBT shrinking, explicit oracle classification, and boundary neighbors — extending Minimal Reproduction and Test-First Debugging. When a bug triggers on a class of inputs, the debugger wraps the failing input in a property (fast-check for JS/TS, Hypothesis for Python) and lets the shrinker auto-minimize the counterexample, storing the minimized input as the regression seed; before writing the assertion it classifies the oracle as specified / derived (contract/model) / metamorphic / implicit (crash — weakest, never the silent default) and records it under Resolution.oracle_type; and it generates boundary neighbors (off-by-one, min/max, empty/singleton) around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — which is what the Phase 1A mutation guardrail needs to bite. Degrades gracefully to manual minimization when no PBT framework is present. Full rules live in gsd-core/references/debugger-repro-hardening.md. (#1962) (#2409)
  • gsd-debugger now emits a blameless-postmortem Prevention block at resolution, closing the loop on bug-class prevention — at archive_session the debugger produces three blame-free components: a branching 5-Whys causal chain (branching per the Phase 2A RCA discipline, not a single linear chain; "agent error" prompts "why was that error possible?", never blame), a "why wasn't this caught?" answer naming the existing gate (test/typecheck/lint/review/verify) that missed it, and a concrete recurrence guard (a regression test / assertion / lint rule / knowledge-base pattern). The knowledge-base entry gains two structured fields — why_not_caught and recurrence_guard — so a future Phase-0 recall surfaces not just the prior fix but the prior prevention (additive; old entries without the fields still load). The session-manager's compact summary surfaces a one-line prevention summary. Full rules live in gsd-core/references/debugger-prevention.md; kept minimal — a block, not an incident-management subsystem. (#1963) (#2410)
  • Third-party capability gates now actually fire via a generic command-exit-zero predicate. — a capability's declared check.predicate gate was rendered for display but never evaluated (only built-in check.query gates were enforced, and the security capability's gate worked solely via a hard-coded ship.md branch). A new generic evaluator (gsd_run check predicate) now evaluates check.predicate blocks by kind; the first built-in kind command-exit-zero runs a bounded sh -c command at the project root and blocks the loop on non-zero exit (timeout → block, fail-closed). The execute:wave:post, execute:post, and plan:post gate-dispatch sites route predicate gates to the new evaluator automatically. (#2008) (#2011)
  • GSD's lifecycle hooks now run under Kimi CLI — installing GSD into Kimi wires its session-state, phase-boundary, graphify, and guard hooks into Kimi's own native config.toml [[hooks]] bus (Beta on Kimi's side) instead of silently no-op'ing, and GSD's Kimi subagents can now run in the background. Kimi's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2095) (#2159)
  • GSD is now installable on pi — npx @opengsd/gsd-core --pi installs the GSD extension to ~/.pi/agent/extensions/gsd.cjs, and /gsd <family> <subcommand> now dispatches real commands through the embedded engine (the reference binding previously could only run query help). Drives pi through the negotiated imperative Host-Integration adapter, with active-model steering and the full pi lifecycle-event surface. (#2102) (#2205)
  • The EoS Registry now lists GSD for Oh My Pi — discover the independently maintained tchivs/gsd-omp protocol-v1 host integration, including exact install and uninstall commands, supported interface points, and negotiated host axes. (#2448)
  • Broken-windows ledger — /gsd:ship now blocks (when workflow.windows_enforce=true, opt-in) while .planning/WINDOWS.md has any open entry, and the executor auto-populates the ledger with stubs, skipped tests, and unrun verifies as it works. Each window can be waived only with a recorded reason (auditable) or fixed (removed from the blocking set); /gsd:progress surfaces the open + waived counts. Backward-compatible: projects with no ledger ship cleanly (open_count starts at 0), and enforcement is off by default so tracking can precede the gate. Enable with gsd config-set workflow.windows_enforce true. (#1950) (#2441)
  • GSD now ships a pi extension — a real, jiti-loadable ExtensionAPI module (pi/gsd.cjs) that registers /gsd (dispatches through the GSD command-routing hub) + gsd_invoke tool + tool_call event, installable at ~/.pi/agent/extensions/. A reachability test proves the /gsd handler dispatches through the engine (keystone wired, not just registered on a mock). (#1965) (#1965)
  • plan-phase now authors edge and prohibition predicates into PLAN.md must_haves when a phase SPEC omits ## Edge Coverage / ## Prohibitions, so goal-backward verification still has predicates to check on a spec-less phase (ADR-857 Phase 6). Gated by the new default-on workflow.specless_probe_fallback toggle — disable it to skip the fallback (the skip is recorded visibly in the plan). Spec-less prohibitions are authored descriptor-less and disposed flagged/unverified (honest verifier #1154), never a silent pass. (#1835)
  • Discover third-party GSD Capabilities in a new Community Capability Registry. — A non-endorsing discoverability catalog where authors register a Capability via a documentation PR; each entry carries a live latest-release badge and a per-entry GitHub Discussion for community ranking and comments. (#2188) (#2188)
  • GSD now warns when a stale global CLI (e.g. a retired @gsd-build/sdk canary) shadows your project-local install — the gsd-tools CLI startup detects when the running binary is outside the project root while a project-local install exists, and prints a remediation warning to stderr (non-blocking). (#1754) (#1755)
  • gsd-mcp-server — companion MCP server (interface points 1 + 5) — a new bin command (npx @opengsd/gsd-core gsd-mcp-server) runs a stdio JSON-RPC 2.0 MCP server exposing gsd_invoke_command (→ the GSD command-routing hub) + gsd_read_state / gsd_write_state (→ .planning/ state), so any MCP-consuming host (Claude Code, Codex, OpenCode, VS Code, Gemini CLI, Cursor, Cline, Hermes) can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681). Dependency-free (hand-rolled JSON-RPC). How-to: docs/how-to/connect-gsd-mcp-server.md. (#1810)
  • Opt-in absolute token count on the statusline context meter — new statusline.show_context_tokens config (default false). When enabled, the meter shows the absolute context total after the percentage, e.g. "████░░░░░░ 46% (156k)", summing input, cache-creation, cache-read, and output tokens from the hook payload (a broader basis than the meter's percentage, which is derived from used_percentage and excludes output tokens — the two figures can diverge slightly). Default meter output is unchanged. (#2161) (#2174)
  • Long-running compute can now be externalized as async external jobs instead of blocking the agent turn — a default-off external-job capability lets executors submit SLURM jobs, commit a .planning/async-jobs manifest, defer SUMMARY.md, and return external_job_waiting; the core loop already reconciles these manifests (#1165), so this adds the producer half (SLURM adapter, pure manifest module, planner/executor fragments, operation policy). (#1105) (#1998)
  • GSD now ships a repo-local VS Code extension — a buildable extension (vscode/extension.js + vscode/package.json) that registers gsd.invoke (dispatches through the GSD command-routing hub) in the VS Code command palette. A reachability test proves the handler dispatches through the engine (keystone wired). Not Marketplace-published; mirrors the OpenCode plugin's bar. (#1966) (#1966)
  • Discover third-party GSD Embeddable Orchestration System (EoS) integrations in a new EoS Registry. — A non-endorsing discoverability catalog where host-integration authors register via a documentation PR; each entry declares its Host-Integration interface points, negotiated axes, and protocol version, with a live release badge and a per-entry GitHub Discussion for ranking and comments. (#2193) (#2193)
  • GSD Core ships a .claude-plugin/marketplace.json marketplace manifest so Claude-plugin-compatible runtimes (ZCODE et al.) can discover and install gsd-core from a custom marketplace source. Additive — the existing .claude-plugin/plugin.json and the Claude Code install path are unchanged. The catalog version (plugins[0].version) tracks package.json via the release version-sync. (#1861)
  • GSD now drives VS Code through the Embeddable Orchestration System — the VS Code extension is rewired through the negotiated imperative Host-Integration adapter (active vscode.lm model, engine hook bus, sandboxed storage), gains native Language Model Tools (GSD skills as #gsd-* tools) and #runSubagent dispatch, and runs as a Web Extension (no Node APIs). (#2103) (#2210)
  • /gsd:next smart-entry workflow — adds a state-aware entry point that classifies the current project situation (no-project, blocked, verify-failed, planning, executing, verify-pending, complete, and more) and recommends the right next GSD command. The gsd-tools smart-entry [--json] classifier handles phase ordering including decimal phase IDs; the /gsd:next skill surfaces the workflow with tiered fallback behavior. (#1798)
  • OpenCode now runs GSD's lifecycle safety hooks (prompt-injection guard, read-before-edit guard, injection scanner, worktree/workflow guards, context monitor) via a native plugin installed to ~/.config/opencode/plugins/gsd-core.js. OpenCode declares hooksSurface: 'none', so these hooks were previously inert; the plugin bridges OpenCode's event bus onto GSD's existing hook scripts. Installed automatically by npx @opengsd/gsd-core --opencode and removed on uninstall. (#1923)
  • Opt-in compact GSD-state statusline format — new statusline.state_format config, enum full|compact (default full, the existing rendering). compact renders " · P/ · " (e.g. "v1.12 · P7/12 · executing"), dropping the milestone name and progress bar and collapsing narrative statuses to the canonical vocabulary from normalizeStateStatus() — the canonical stuck state paused renders uppercase as PAUSED. Solves the unbounded-width problem where free-text status sentences push the context meter off the line. (#2162) (#2175)
  • <precondition> task element (Design by Contract) — plans may now declare a runnable/checkable fact a task assumes (env var set, prior-phase artifact present, external-setup done) that plan ordering does not guarantee; the executor asserts it before running the task and halts with a checkpoint on unmet instead of building on a broken assumption. Plans that omit <precondition> behave exactly as today. (#1949) (#2422)
  • Config-gated provider escalation when a run hits a quota or rate limit — an executor killed by a provider throttle stopped the phase and waited for a manual restart; escalating a tier did not help because the same throttled provider was still in play. Set dynamic_routing.provider_escalation to an ordered list of fallback model IDs and GSD now switches provider on a quota-exceeded failure, logs the swap (sonnet → gpt-5), honors the provider's Retry-After, caps the walk at max_escalations, and names every model tried once the list is spent. Opt-in — unset, quota failures keep today's manual recovery prompt. (#2296) (#2458)
  • Host-integration descriptors now carry an extensionEvents vocabulary — the extension-system event surface (OpenCode, pi) is a separate descriptor field from managed hookEvents, so OpenCode declares extensionEvents:opencode without conflicting with the hooksSurface:none invariant. (#1946) (#1946)
  • /gsd-review now supports custom reviewer instances — run one model-capable adapter (e.g. OpenCode) as several independent reviewer identities via a bounded review.reviewer_instances config, so two different models can review in a single pass without manually swapping config or hand-merging REVIEWS.md. (#1517) (#1766)
  • Opt-in git branch and working-state segment in the statusline — the shell prompt's branch/dirty-state signal is hidden for the whole session under the Claude Code TUI, so wrong-branch commits and ship-time push rejections surface only after the fact. New statusline.show_git config (default false) renders the branch name plus staged/unstaged/untracked/ahead/behind markers (or ✓ when clean and in sync) after the directory segment. When disabled, no git subprocess is spawned and output is unchanged. (#2163) (#2183)
  • /gsd:onboard guides brownfield setup — existing repos now have a top-level onboarding command that routes through codebase mapping, docs ingest, project initialization, and an onboarding summary without silently overwriting planning files. (#1994)
  • Plural/optional/chosen assumption-delta checkpoint during planning — when a phase makes something plural, optional, or chosen that used to be singular, required, or derived, the planner is now prompted to re-ask whether the primary key / identity model still names the right thing, preventing silent architectural drift from accumulating into a later user-facing bug. Advisory (non-blocking); fires only on a detected signal. Toggle with workflow.assumption_delta. (#1561) (#1767)
  • /gsd-ui-phase now probes UI state coverage — a new ui-consideration-probe (the third probe-core adapter) enumerates the shape-rooted UI states a UI-SPEC must resolve (empty/loading/error/populated/partial/overflow/zero-one-many/long-text). After the UI checker approves, the probe surfaces applicable considerations for each element, records a ## UI Considerations section in the UI-SPEC, and plan-phase lifts each resolved consideration into must_haves — so a purely-visual state with no wired test routes to insufficient_spec → human_needed at verify rather than a silent pass. (#1979)
  • Host-Integration Interface (ADR-1239 Phase A) — a versioned, negotiated capability contract (runtime.hostIntegration) over the six host-integration points (command, dispatch, model, hooks, state, artifact). Adds an in-process negotiateHostCapabilities handshake that fail-closes on undeclared/unknown/undocumented values (effective ⊆ host-declared ∩ engine-known), a typed degradation ladder, host-capability profiles, and a documentation-sourced per-CLI capability matrix for all 16 runtimes. Interface-definition only — no change to install behaviour. (#1690)
  • ZCode (Z.ai) is now an installable runtime — a desktop Agentic Development Environment for the GLM-5.2 model can now be targeted with --zcode, landing GSD skills at ~/.zcode/skills/<name>/SKILL.md plus slash commands and subagents. ZCode ships as a pure declarative capability descriptor (capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode' branches, reusing the Claude skill converter — the de-hardcoded, data-driven runtime path that 1.7.0 (ADR-1016 / ADR-1239) enables. (#1925) (#2039)
  • Reversibility tagging for planning decisions — decisions can now be rated reversible, costly, or one-way by how expensive they are to undo. A one-way decision (one whose undo needs a data migration, breaks a published contract, or is impossible) earns a checkpoint:decision before the task that implements it, so an unattended run pauses for your sign-off instead of walking through the door. costly decisions are flagged in the plan without blocking; reversible ones flow as before. Pass --no-reversibility-gates to /gsd:plan-phase to suppress the checkpoint on runs you mean to leave unattended — ratings are still recorded either way. (#1951) (#2471)

Changed

  • gsd-debugger now recalls prior resolved sessions semantically via MemPalace instead of keyword overlap — at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions as candidate hypotheses, catching the same-root-cause / different-wording cases keyword overlap missed (a prior "requests hang under load" now surfaces for "API times out when many users connect"). Resolved sessions are indexed into MemPalace at archive (symptoms + root cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching against it (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Full rules live in gsd-core/references/debugger-semantic-recall.md. (#1964) (#2416)
  • The GSD CLI now self-heals a missing runtime build. The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts (ADR-457) that ship prebuilt in the npm tarball but are absent on a Claude Code plugin-marketplace / git-clone install, which never runs npm run build:lib. Previously every command died at load with Cannot find module './lib/cli-exit.cjs'. The gsd-tools entrypoint now detects the missing output and compiles it once, on demand (lock-guarded so parallel invocations don't race), then proceeds — a single no-op check on the already-built npm path. When TypeScript is genuinely unavailable it prints an actionable npm install && npm run build:lib message instead of crashing. (#2036)
  • Internal: Claude Code's installer is now driven through the public Host-Integration Interface (ADR-1239 / EoS). bin/install.js routes claude install/uninstall through the imperative adapter (createImperativeAdapter) instead of calling the engine directly, and its 13 hardcoded runtime === 'claude' / runtime !== 'claude' branches are folded into descriptor-driven runtime.hostBehaviors on capabilities/claude/capability.json (permission schema, settings.local.json scope routing, .gsd-source marker, effort frontmatter, canonical-workflow authorship, and more). Install/uninstall output is byte-identical for both the global skills layout and the local legacy layout (golden-parity asserted for both scopes); no other runtime changes. Removes the "add-a-host tax" of scattered string-equality checks for the tier-1 reference host. No user-facing change. (#2086) (#2106)
  • OpenCode is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). OpenCode and its Kilo sibling previously installed via a bespoke runtime === 'opencode'/isOpencode branch in bin/install.js; its commands+skills+plugin install now runs through the imperative adapter → the engine's combined-family install path (installRuntimeArtifacts), and every hardcoded runtime === 'opencode' branch is folded into descriptor-driven runtime.hostBehaviors. Install/uninstall output is byte-identical (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) background dispatch — OpenCode shipped experimental background subagents in v1.15 and made them default-on in v1.17, so dispatch.background/backgroundDispatch flip to true; GSD no longer force-flattens OpenCode-hosted wave dispatch (shouldFlattenDispatch now returns false), letting agents run concurrently where the host supports it. (2) expanded event surface — the OpenCode plugin now subscribes to permission.asked, permission.replied, and session.error (added to EXTENSION_EVENT_SURFACES.opencode), wiring the declared surface for future permission/error-aware bindings. (#2087) (#2108)
  • Codex is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS). Codex previously installed via hardcoded runtime === 'codex'/isCodex projection in bin/install.js; its config.toml / agent-.toml / hooks.json install now runs through the declarative embedding adapter and descriptor-driven runtime.hostBehaviors, with zero positive isCodex gates and zero runtime === 'codex' branches remaining (source-guarded). Install/uninstall output stays byte-parity-gated (tests/fixtures/golden-install-parity/codex.json). Three Context7-verified upgrades land, each with a test driving the user-reachable surface: (1) skill root — GSD skills now install to Codex's canonical $HOME/.agents/skills (via a skills-kind home override) instead of the deprecated $CODEX_HOME/skills fallback, and pre-move installs are migrated (stale ~/.codex/skills/gsd-* cleaned on both install and uninstall, user-owned content preserved); (2) hook events — GSD registers the six documented Codex lifecycle events it previously skipped (PreToolUse, PermissionRequest, PreCompact, PostCompact, SubagentStop, UserPromptSubmit, in addition to the existing SessionStart/SubagentStart/Stop/PostToolUse) in hooks.json, so gsd-context-monitor fires at the same points as in Claude Code, and the descriptor extendedHookEvents is reconciled from [] to the schema-valid wired subset; (3) dispatch tuning — [agents] max_depth = 1 is written explicitly into the managed config.toml block to pin the negotiated dispatch.maxDepth: 1 axis (degradationFor flattens GSD-hosted waves to single-level), and validateCodexConfigSchema now permits a known-scalar-only [agents] AgentsToml table (coexisting with the flattened [agents.gsd-*] role sub-tables) while still rejecting the [[agents]] and unknown-key break-forms from #2760. (#2088) (#2110)
  • Cursor is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). Cursor previously installed via hardcoded runtime === 'cursor'/isCursor branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded cursor branch is folded into descriptor-driven runtime.hostBehaviors (reapplyCommand, frontmatterDialect, hooksJsonSurface, skipSharedHooksInstall, reportCommandsDir, managedHookEvents). Install/uninstall output is byte-identical (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) expanded hook-bus coverage — GSD registers all 6 managed lifecycle events in Cursor's hooks.json (preToolUse, stop, subagentStart, subagentStop in addition to the original sessionStart/postToolUse), driven by a new descriptor-driven adapter module (src/host-integration-adapters/imperative-hook-bus.cts) that reads hostBehaviors.managedHookEvents instead of a hardcoded event pair; cite https://cursor.com/docs/hooks. (2) named/background nested subagent dispatch — Cursor's dispatch.background/backgroundDispatch/nested are all true with maxDepth: 2, so shouldFlattenDispatch(cursor) returns false and GSD's wave-based execution drives Cursor's native background + depth-2 nested subagent invocation instead of flattening to inline sequential calls; cite https://cursor.com/docs/subagents + https://cursor.com/docs/sdk/typescript. (#2089) (#2120)
  • Cline is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). Cline previously installed via hardcoded runtime === 'cline'/isCline branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded cline branch is folded into descriptor-driven runtime.hostBehaviors (reapplyCommand, frontmatterDialect, skipSharedHooksInstall, localTargetIsProjectRoot, clineRulesSurface, localCommandsViaRules). Install/uninstall output is byte-identical (golden parity asserted for cline + claude/cursor/codex/opencode). Two Context7-verified upgrades land: (1) AgentPlugin.hooks.beforeTool planning guard — the .clinerules/hooks/PreToolUse file-convention hook (#787) is re-implemented as a real Cline SDK AgentPlugin that cancels write-class calls targeting .planning/ (same fail-open semantics), driven by a new descriptor-driven adapter module (src/host-integration-adapters/cline-sdk-binding.cts); cite https://github.com/cline/cline/blob/main/docs/sdk/plugins.mdx. (2) createAgentModel model overrides — DefaultGateway.createAgentModel({providerId, modelId}) is wired so GSD's per-subagent model_overrides/model_profile_overrides resolution applies to Cline subagents (modelMode: active); cite https://github.com/cline/cline/blob/main/docs/sdk/reference/gateway.mdx. Cline's dispatch deliberately stays degraded/flat (maxDepth: 1, read-only, no nested spawning) per the documented host restriction — never silently upgraded to full nested/background. (#2090) (#2132)
  • Hermes Agent is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS). Hermes previously installed via hardcoded runtime === 'hermes'/isHermes branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded hermes branch is folded into descriptor-driven runtime.hostBehaviors. Three upgrades land: (1) real plugin hook vocabulary — GSD registers a new extensionEvents: "hermes" dialect carrying the 13 documented Hermes plugin events (pre_tool_call, post_tool_call, pre_llm_call, post_llm_call, on_session_start, on_session_end, on_session_finalize, on_session_reset, subagent_start, subagent_stop, pre_gateway_dispatch, pre_approval_request, transform_tool_result), replacing the borrowed hookEvents: "claude" 6-event surface that silently never fired; cite https://github.com/nousresearch/hermes-agent/blob/main/website/docs/user-guide/features/hooks.md. (2) dispatch posture — Hermes' dispatch.nested: true with maxDepth: 1 is correctly negotiated (not silently flattened). (3) branding/category metadata — DESCRIPTION.md category descriptions, version: frontmatter, and branding rewrites are now descriptor-driven rather than hardcoded. Install/uninstall output is byte-identical (golden parity asserted for all runtimes). (#2091) (#2134)
  • Qwen Code now projects GSD's specialist agents as native subagents — installing GSD into Qwen Code writes ~/.qwen/agents/gsd-*.md files you can invoke directly (planner, executor, code-reviewer, …) instead of reaching them only through skill prose, and a SubagentStart hook now fires alongside SubagentStop. Qwen's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2092) (#2153)
  • Kilo Code now supports native hooks, active-model routing, and named subagent dispatch — installing GSD into Kilo wires a lifecycle-hook plugin, keeps each agent's requested model instead of dropping it, projects GSD's specialist agents as invokable subagents, and documents the GSD MCP companion. Kilo's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2093) (#2156)
  • GSD skills installed for Trae now carry SOLO stage metadata — Trae's SOLO Agent can recognize GSD skills as workflow-stage skills for auto-invocation instead of requiring manual triggering. Several of Trae's install branches (shared-hooks gating, path rewrites) also move onto its capability descriptor. Note: the stage-metadata field is a best-effort/inferred shape — Trae publishes no formal schema. (#2094) (#2157)
  • Installing GSD into Antigravity now writes the permissions.allow rules its CLI documents — so GSD's own reads and hooks aren't stuck on interactive prompts — and registers GSD's companion MCP server via a standalone mcp_config.json (best-effort: Antigravity's raw config schema isn't published, so this uses the Gemini-CLI-successor format). Antigravity's install is now driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2096) (#2165)
  • Augment Code now installs through its capability descriptor, with a native MCP companion — installing GSD into Augment registers the GSD companion server in Augment's settings.json mcpServers and drives command/skill/agent conversion from Augment's negotiated descriptor instead of hardcoded runtime special-cases. (#2097) (#2166)
  • CodeBuddy now wires GSD's full extended lifecycle hook set and is driven by its capability descriptor — installing GSD into CodeBuddy now registers SubagentStart, SubagentStop, Stop, and PreCompact hooks in its settings.json (it previously had none of these), matching the coverage Qwen/Kimi already ship, and CodeBuddy's install is fully descriptor-driven instead of via residual hardcoded runtime branches. (#2098) (#2169)
  • GitHub Copilot now wires GSD's full lifecycle hook bus and is driven by its capability descriptor — installing GSD into Copilot registers preToolUse, postToolUse, userPromptSubmitted, and sessionEnd handlers in its hooks/gsd-session.json (beyond today's sessionStart-only advisory), and Copilot's residual hardcoded runtime branches are folded onto descriptor-driven hostBehaviors. (#2099) (#2172)
  • Windsurf now enforces GSD's write/command safety guards through Cascade's native hook bus — installing GSD into Windsurf registers blocking pre_write_code/pre_run_command hooks in .windsurf/hooks.json (exit-code-2 blocking) and drives Windsurf's install from its capability descriptor instead of hardcoded runtime branches. (#2100) (#2190)
  • ZCode's install is now driven and regression-tested through its capability descriptor — ZCode joins the dogfooded declarative-adapter reference hosts with a byte-identical install, and its shared-hooks exclusion is folded onto hostBehaviors instead of a hardcoded runtime branch. (Hook-automation and MCP upgrades remain blocked on ZCode publishing its on-disk config formats.) (#2101) (#2195)
  • Codex/OpenAI default models advance to the GPT-5.6 family (Sol/Terra/Luna) — the Codex runtime tier defaults and the openai provider preset now resolve to current-generation model IDs instead of the superseded GPT-5.4/5.5 line, so Codex users on default profiles get improved agentic coding (Sol) and lower costs (Terra/Luna) without changing any config. (#2122) (#2146)
  • Internal: the installer's program (display-name) + command (slash-invocation) chains are now single-source lookups — the 14-line program chain (an exact duplicate of runtimeLabel) → getRuntimeLabel, and the 14-line command chain (the per-runtime /gsd-new-project syntax: gemini /gsd:, codex $, cursor skill-mention, kimi /skill:, default /gsd-new-project) → new getRuntimeNewProjectCommand(runtime) helper (ADR-1239 Phase B / #1679 AC2 slice 4). runtime === count in bin/install.js: 53 → 25 (cumulative this session: 129 → 25). Stdout strings preserved byte-for-byte; no install-output change (golden-parity 16/16). No user-facing change. (#1813)
  • Internal: the installer's per-function is<Runtime> flag-declaration blocks are now a single runtimeFlags lookup — the four duplicated const isX = runtime === 'x' blocks in bin/install.js (uninstall / writeManager / install / a fourth helper — 48 branches) are collapsed into one runtimeFlags(runtime) helper in runtime-name-policy.cts (ADR-1239 Phase B / #1679 AC2 slice 3). The add-a-host tax for flags is removed (one RUNTIME_FLAG_IDS entry, not four declaration blocks). Install output is byte-identical for all 16 runtimes (golden-parity asserted); runtime === count in bin/install.js: 101 → 53. No user-facing change. (#1811)
  • Internal: third-party descriptor loader enforces configHome write-confinement at load time — loadRegistry({includeInstalled:true, configHome}) now rejects (skip + warn, fail-closed) any installed third-party host-plugin descriptor whose declared destSubpath resolves outside the supplied configHome, before it is composed into the registry (ADR-1239 Phase C-2 / #1681 slice 2). The configHome option is optional and backward-compatible (omitted → no load-time check; install-time gate still bounds writes). No user-facing change for existing flows. (#1808)
  • Internal: agent install for cursor/windsurf/augment/trae/codebuddy now flows through the descriptor path — ADR-1235 step 1 routes the trivial-converter runtime group's agents off the inline install() loop onto the descriptor-driven installRuntimeArtifacts path, applying the cross-cutting steps uniformly (pre-converter, no workflow-stamp). Agent output is byte-identical for all 16 runtimes (golden-parity asserted, global + local verified); no user-facing change. (#1764)
  • gsd-ui-checker gains an adversarial FORCE stance (LLM-playbook principle 16) — the only verdict-producing critic that lacked one now resists rubber-stamping UI-SPEC contracts, with BLOCK/FLAG/PASS classification. Based on arXiv 2505.23840 (third-person objective persona), 2506.04975 (objective-not-hostile persona). (#1584)
  • Internal: the declarative embedding adapter is now named + bound behind a minimal HostIntegrationInterface — createDeclarativeAdapter({runtime}) (new src/adapter-declarative.cts) delegates in-process to install-engine's installRuntimeArtifacts/uninstallRuntimeArtifacts, formalizing today's projection path as one of the two embedding adapters behind a common contract (ADR-1239 Phase C-1 / #1680 AC1). Output is byte-identical to today's install (gated by golden-install-parity). The full 6-point interface binding surface is deferred until the imperative adapter (AC2) fixes the shape (ADR-1239 open wire-shape question). No user-facing change — the adapter is not yet wired to any runtime path. (#1802)
  • Internal: getDirName is now derived from a documented runtime.localConfigDir descriptor field — each runtime's local content-rewrite directory (e.g. cursor→.cursor, copilot→.github) moved from a hand-maintained if-chain into its capability descriptor (ADR-1239 Phase B), so it can no longer drift from the registry. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1757)
  • Internal: copyWithPathReplacement converter selection is now data-driven — the installer's back-compat content-copy path replaced its 13 hardcoded runtime === 'x' flag chains with a single per-runtime dispatch table (ADR-1239 Phase B). Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1759)
  • Phase-completion now writes Status: All phases complete instead of the overloaded bare Milestone complete — the phase-level completion verb (completePhaseCore) was writing the same bare 'Milestone complete' string that the milestone-close verb uses for terminal state, causing a phase-level verb to own a milestone-level field. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts); milestone termination (' milestone complete' / 'Awaiting next milestone') remains solely with the milestone-close verb. (#2204) (#2259)
  • #853 dispatch-flatten is now data-driven (ADR-1239 Phase B) — whether GSD backgrounds the plan/execute orchestrator is decided from a documentation-sourced backgroundDispatch capability per host (via gsd_run query dispatch-should-flatten) instead of a hardcoded runtime === 'codex' check. Cursor now backgrounds the orchestrator (its docs document backgrounded subagent nesting); codex unchanged; all other hosts run inline. Fail-closed to inline on any uncertainty. (#1719)
  • Internal: companion MCP server module (interface points 1 + 5) — handleMessage/runServer (new src/mcp-server.cts) is a minimal, dependency-free stdio JSON-RPC 2.0 server exposing gsd_invoke_command (→ the command-routing hub) + gsd_read_state/gsd_write_state (→ the Phase 3 stateIO seam), so any MCP-consuming host can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681 slice 3a). Bin entry / packaging deferred to slice 3b. No user-facing change — the server is not yet wired to a bin entry. (#1809)
  • requirements mark-complete reports a per-surface write-set — the command now returns a per-requirement write_set (checkbox + traceability surfaces) and a write_set_complete that is true only when every surface of every requirement applied, so a partial (checkbox-only) reconcile can no longer masquerade as full success even inside a multi-ID batch. Introduces the reusable ADR-2143 §5/§6 Result / WriteSet contract. (#2251) (#2251)
  • Internal: the imperative embedding adapter now composes the capability registry behind the same HostIntegrationInterface — createImperativeAdapter({runtime}) (new src/adapter-imperative.cts) calls loadRegistry({includeInstalled:true}) (first-party-wins + consent + fail-closed — identical trust semantics to the CLI) and binds the engine surface behind the same contract the declarative adapter (AC1) satisfies, plus a registry accessor for an in-process host to bind its primitives to (ADR-1239 Phase C-1 / #1680 AC2). Concrete host binding is deferred to Phase 5. No user-facing change — the adapter is not yet wired to any runtime path. (#1803)
  • Internal: the model adapter seam exposes passive + active adapters selected by modelMode — createModelAdapter({modelMode}) (new src/model-adapter.cts): passive formalizes today's tier routing (delegates to model-resolver.resolveModelForTier), active is a host-supplied sendRequest seam (VS Code vscode.lm / pi providers), fail-closed until Phase 5 binds a concrete provider (ADR-1239 Phase C-1 / #1680 AC3). No user-facing change — the seam is not yet wired to any runtime path. (#1804)
  • Internal: derive the non-Claude runtime list from the capability registry — NON_CLAUDE_RUNTIMES is now computed from the capability registry instead of a hand-maintained literal, so it can no longer drift from the per-runtime descriptors. No user-visible behavior change (the list is identical). (#1728)
  • Honest verifier — verify-phase now abstains on non-inferable backstop truths instead of confidently false-passing them (#1154). When the spec's edge-probe marks a truth non-inferable (verification: backstop) and the verifier cannot confirm it with explicit evidence (a passing wired held-out/property test, or a directly-observed behavior), it now reports human_needed with reason insufficient_spec ("unverified — held-out test recommended") rather than a silent passed. Autonomous runs complete with "N unverified non-inferable checks"; interactive runs route to the end-of-phase human checkpoint. Inferable truths are never abstained (over-abstention guard); abstention is exogenous (driven by the tag, not self-judgment). Truth-axis mirror of the prohibition judgment-tier (ADR-550 D4). (#1738)
  • Document Claude Code's advisor-tool inheritance in the model-profiles reference: the session-level advisor is inherited by all GSD subagents and composes with per-agent tiering, with candidate executor/advisor pairings, when it is worth enabling, and the session-level (no per-agent control) constraint. (#1922)
  • Extraction discipline for strict-format agents (LLM-playbook principle 8) — gsd-doc-classifier and gsd-doc-synthesizer apply taxonomy/precedence rules directly without inventing content, reducing reasoning-induced format drift. Based on arXiv 2504.05081 (few-shot beats CoT for pattern tasks), 2506.00069 (terminal instruction placement), 2505.14810, 2505.11423. (#1584)
  • Internal: extracted the runtime-artifact install engine from bin/install.js — installRuntimeArtifacts/uninstallRuntimeArtifacts/installOpencodeFamilySkills and their helpers now live in a dedicated gsd-core/bin/lib/install-engine.cjs module (ADR-1239 Phase B), so adapters can import the install pipeline instead of reaching into the 12k-line installer. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing behaviour change. (#1735)
  • MemPalace memory_mode kg_backend and replace are now functional — selecting either mode now routes recall through the palace instead of silently behaving like augment: kg_backend treats the palace temporal KG as the primary knowledge-graph source (native .planning/graphs/ as fallback), and replace resolves recall through the palace as the source of truth. Every mode stays default-resilient — an unreachable palace falls back to native memory and no memory is lost. (#2010) (#2010)
  • /gsd:surface and --materialize now produce byte-identical agent output to a fresh install — surface-path agents for descriptor-driven runtimes (cursor, windsurf, augment, trae, codebuddy, copilot, antigravity) now receive the same path-prefix rewrite, Co-Authored-By attribution, runtime-specific conversion, and body normalization as the install path. Copilot and Antigravity agents are now installed via the descriptor-driven path (copilot agents get the .agent.md filename rename). Cline remains on the inline loop (rules-only local branch). (#1575) (#2040)
  • Internal: hook-bus + stateIO adapter seams — createHookBus({bus}) (new src/hook-bus.cts, host/engine/none — engine is in-process pub/sub, host fail-closed, none silent) + createStateIO({io}) (new src/state-io.cts, filesystem/sandboxed-storage/session-log-append — filesystem delegates to fs, the rest are fail-closed seams) (ADR-1239 Phase C-1 / #1680 AC4). Completes the Phase 3 adapter seam layer; concrete host binding is Phase 5. No user-facing change. (#1805)
  • Long-context model names render compactly in the statusline — the verbose " (1M context)" suffix Claude Code appends to the model display name now collapses to a compact " (1M)" badge (tolerant of future window sizes and the abbreviated "ctx" variant: "(500K context)" → "(500K)", "(1M ctx)" → "(1M)"). Lossless — the long-context signal stays, the 12 characters of width don't. (#2160) (#2173)
  • Lazy-split plan-phase.md into a steps/ directory — ~4.7 KB lighter eager context per /gsd-plan-phase call via byte-invariant progressive disclosure (ADR-1610). (#1852) (#1934)
  • GSD subagents now self-load configured agent_skills regardless of orchestrator bash — projects that map skills via .planning/config.json agent_skills.<agent-type> no longer silently lose them on /gsd-autonomous or Cursor, where Skill()-delegated workflow bash init did not reliably run. Each of the 22 consumer agents queries its own type at init and reads the listed skills, with a dedup guard so runtimes that also inject orchestrator-side (Claude Code) never carry two copies. (#1866) (#1868)
  • Internal: install/uninstall runtime labels are now sourced from a single getRuntimeLabel lookup — the two duplicated runtimeLabel assignment chains in bin/install.js (uninstall + install) are collapsed into one curated label table in runtime-name-policy.cts, sibling to the registry-derived getDirName (ADR-1239 Phase B, #1679). Install output is byte-identical for all 16 runtimes (golden-parity asserted). Two console-label inconsistencies are normalized as a side effect: kimi shows 'Kimi CLI' in both sites, and cline uninstall no longer falls through to 'Claude Code'. (#1800)
  • Phase plans now lead with a verified end-to-end "tracer" slice by default — every plan starts with one thin, production-quality slice wired through every layer, which the executor verifies before building out the remaining tasks, so an architectural dead-end surfaces after one commit instead of after ten. Pass --no-tracer to restore the previous horizontal-layer default; --mvp now layers user-story framing and the Walking Skeleton on top of the tracer-first ordering. (#1945) (#2294)
  • Internal: external-descriptor trust gate — load-time configHome confinement — assertDescriptorConfined(descriptor, configHome) (new src/external-descriptor-trust.cts) fail-closed rejects any installed third-party host-plugin descriptor whose declared destSubpath resolves outside the user-approved configHome, before its install plan runs (ADR-1239 Phase C-2 / #1681 slice 1). Defense-in-depth load-time twin of Phase 2's install-time assertDestWithinConfigHome. Not yet wired into the loader (slice 2). No user-facing change. (#1806)
  • Internal: the installer's runtime → global-config-home hook-pathogen fragment is now a single getGlobalConfigHomeFragment lookup — the 14-branch if (runtime === 'x') return "'...'" chain in getConfigDirFromHome (bin/install.js, the hook path.join() codegen mapping) is collapsed into one table in runtime-name-policy.cts, sibling to getRuntimeLabel (ADR-1239 Phase B, #1679 AC2 slice 2). Generated hook output is byte-identical for all 16 runtimes (golden-parity asserted); antigravity's dynamic env-overridable resolution is preserved in the caller. No user-facing change. (#1801)

Removed

  • Removed the sunset Gemini CLI runtime — use Antigravity CLI instead — Google discontinued Gemini CLI on 2026-06-18, so npx gsd-core --gemini now prints a deprecation notice and points you to Antigravity CLI (the official successor), which GSD already ships as a first-class runtime. (#1928) (#1996)

Fixed

  • The verify-work security-blocked presentation no longer offers next-phase planning. When security enforcement blocks phase advancement (no SECURITY.md produced), the workflow now routes only to the current-phase fix instead of competing /gsd:plan-phase {next} and /gsd:execute-phase {next} options. (#1687)
  • milestone complete and roadmap analyze now exclude the Phase 0 / Phase 999 backlog sentinels. A milestone whose only directory-less ROADMAP heading is a backlog sentinel can be completed without --force, and roadmap analyze no longer counts the sentinel in phase_count or routes next_phase into it. Completes the ^999 exclusion #1445 added to the progress denominators. (#1691)
  • config-set no longer silently coerces values into something the disk never sees — Number.isFinite replaced !isNaN in the value parser so Infinity/-Infinity are no longer coerced to non-finite numbers that JSON.stringify then renders as null on disk while the CLI echoes Infinity (output ≠ disk). context_window now has a per-key validator requiring a finite positive integer (rejects Infinity, 0, negatives, non-integers with a non-zero exit), and project_code is always persisted as a string so a leading-zero code like 007 survives verbatim instead of collapsing to 7. Numeric coercion for genuine numeric keys (e.g. granularity 42) is unchanged. (#1581) (#2023)
  • phase.complete no longer reports a false is_last_phase on a <details>-wrapped checkbox checklist (#1591, #1752) — when the active milestone's phase checklist was written as - [ ] Phase N: checkbox items inside a <details> block and the next phase had no directory on disk yet (still in planning), phase.complete's isLastPhase roadmap-enumeration fallback used a heading-only pattern (/#{2,4}\s*Phase…/) that never matched checkbox items. It returned is_last_phase: true, next_phase: null on a mid-milestone phase and — via the milestone-complete cascade — wrongly flipped STATE.md to Milestone complete and decremented progress.total_phases (e.g. 8 → 7). The pattern now matches heading-style (### Phase N:), plain checkbox-list phases (- [ ] Phase N: / - [x] Phase N:), and the canonical bold checklist form the roadmap template emits (- [ ] **Phase N: Name**); extractCurrentMilestone already surfaces the <details>-wrapped checklist correctly, so no parser change was needed. Only the reproduced phase.complete fallback is changed; the heading-only sibling patterns elsewhere in phase.cts are untouched. (#1819)
  • The <agent_skills> block emitted by gsd init no longer leaks backslash paths into @-reference skill paths on Windows. The global skill directory (a native path.join result) was interpolated into the generated markdown without POSIX normalization, producing references like @C:\…\skills\name/SKILL.md; the reference is now normalized at the emit site so skill references use forward slashes on every platform. (#1736)
  • /gsd-settings no longer warns about four search-provider keys on fresh projects (#1747) — buildNewProjectConfig emits seven search-provider availability flags and research-provider.cts providerAvailability() consumes all seven, but only three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json). Running /gsd-settings on a freshly generated .planning/config.json printed unknown config key(s) … tavily_search, ref_search, perplexity, jina — these will be ignored even though the user never hand-edited the config. The four missing keys are now registered alongside brave_search/firecrawl/exa_search and documented in docs/CONFIGURATION.md; a drift guard in tests/bug-2530-valid-config-keys.test.cjs now requires every config-driven research-provider flag to be in the schema, so a future provider addition cannot reintroduce the drift. (#1814)
  • gsd-tools state json no longer reports conflated progress for an unversioned milestone (#1761) — the ADR-1769 Phase 7 fix (#1794) taught state sync to leave Progress untouched when a milestone version is asserted but the ROADMAP has no versioned heading for it, but the state json read path still rebuilt progress via buildStateFrontmatter, whose phase-heading count fell back to the whole document and summed sibling milestones. state json therefore reported a conflated total_phases (e.g. 8 = 4+4 across two milestones) plus a derived percent, contradicting the sync guard on the very same project. The read path now mirrors the sync guard: when the asserted milestone cannot be bounded to a versioned ROADMAP heading, total_phases falls back to the on-disk phase-dir count and percent is omitted. Bounded milestones (versioned ROADMAP, or no milestone asserted) are unchanged; the signal rides on the existing _diskScanCache so extractCurrentMilestone's return contract and its other callers are untouched. (#1818)
  • gsd-graphify-update.sh now reads the full multi-line command in Gate 2 (#1772) — the PostToolUse auto-update hook joined tool_name + \n + tool_input.command and extracted the command with sed -n '2p' (line 2 only). Agent runtimes (Claude Code's Bash tool among them) routinely emit HEAD-advancing commits as multi-line scripts (cd /path, then git add, then git commit …), so line 2 was the cd, Gate 2's *"git commit"* match failed, and the rebuild silently no-op'd on real commits even with graphify.auto_update: true. The failure was invisible in manual probes because a single-line git commit -m x passes line 2 verbatim. The hook now captures line 2 through EOF (sed -n '2,$p') so the case glob sees the full command string; single-line behavior is unchanged and multi-line commands without a HEAD-advancing op still no-op cleanly. (#1815)
  • /gsd-thread close|resume now writes the thread status/updated frontmatter (#1778) — the thread workflow's CLOSE and RESUME branches invoked frontmatter.set with the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>), but since 1.6 the dispatcher parses the file positionally and reads field/value from the named flags --field/--value via parseNamedArgs. The positional form left field/value undefined, cmdFrontmatterSet errored file, field, and value required, and the writes were silently skipped — so closing a thread never marked it status: resolved and resuming never marked it status: in_progress, with the error scrolling past on every thread command. All four sites (CLOSE status+updated, RESUME status+updated) now use the 1.6 hybrid form that verify-work.md already uses (frontmatter.set <file> --field <field> --value <value>). (#1816)
  • The installer no longer copies dead lifecycle hook scripts for ZCode — it declares hooksSurface: 'none' and has no plugin surface, so the staged hooks/*.js, hooks/*.sh, hooks/lib/ and the CommonJS package.json marker were dead weight in ~/.zcode/. The hook-copy guards in install.js now exclude ZCode alongside the other no-hook runtimes. OpenCode, which also declares hooksSurface: 'none', is deliberately kept: its native plugin adapter (#1914) spawns those staged hooks via OpenCode's event bus and needs both them and the marker. (This fix originally excluded Kilo too, on the premise that it had no plugin surface; that premise was wrong — Kilo's native plugin spawns the staged guard hooks, exactly like OpenCode's — and #2327 reverses the Kilo half.) (#2057)
  • Test gates can no longer hang forever on a watch-mode test runner. vitest defaults to watch mode in an interactive terminal (exactly where gsd-execute-phase runs), so a resolved npm test / pnpm test that maps to vitest never exited and the orchestrator waited indefinitely until the user manually intervened. Every GSD test-command gate — the regression gate, the post-merge gate, the audit-fix gate, and the verify-phase gate — now routes the resolved command through a shared normalize-test-command helper that rewrites it to a one-shot form (direct vitest → vitest run; jest --watch → --watchAll=false; a package-manager test script backed by watch-vitest → CI=true prefix; already-one-shot commands are left unchanged). The three gates that previously hung or silently continued — the regression, post-merge, and audit-fix gates — additionally bound execution with a configurable workflow.test_gate_timeout (default 600s), aborting or surfacing the cause on timeout instead of hanging; the verify-phase gate was already bounded (a fixed 5-minute limit) and keeps it, now naming watch mode on timeout. The normalizer only rewrites a runner named as a standalone command token (so paths/targets like run-vitest.js are never mangled), is length-capped and linear-time on adversarial input, and only reads a regular-file package.json. (#2060)
  • settings-advanced.md no longer has an orphan </step> around §8 Model Policy — the §8 Model Policy block ended with a closing </step> but had no matching opening tag (5 opens / 6 closes), leaving its content as loose inter-step prose that could fail to execute reliably. Added the missing <step name="model_policy"> opener so the section is a proper step. A new workflow <step>-tag-balance regression guard (fenced-code-stripped) now blocks any future orphan tag across all top-level workflows. (#1864) (#2014)
  • The runtime launcher now honors CLAUDE_CONFIG_DIR — the gsd_run preamble embedded in every workflow/agent resolved the Claude global install only at $HOME/.claude/gsd-core/bin/, while the installer honored CLAUDE_CONFIG_DIR, so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every gsd_run call (every GSD command failed with gsd-tools.cjs not found). The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude} — matching the installer and the other runtimes' ${VAR:-default} pattern — so a custom CLAUDE_CONFIG_DIR is found and the default $HOME/.claude path is unchanged. Re-synced into all 95 workflows/agents; two capped workflows trimmed to stay under their byte budgets. (#1865) (#2024)
  • Node-test prohibition proofs now require a clean-fixture causation control — a node-test prohibition's fail-first proof no longer accepts a deceptive content-independent negative test (one that reds merely because GSD_PROHIB_SUBJECT is set, ignoring the subject's content). The check_clean_fixture control is now mandatory for the node-test kind: a descriptor that omits it is un-provable and hard-gates, rather than greening on the violation alone. Breaking (Hyrum): a previously-green node-test prohibition with no clean fixture now hard-gates — blast radius is zero in-tree (no node-test prohibition ships today). The lint-rule kind is unchanged (its subject IS the linted file, no GSD_PROHIB_SUBJECT indirection). (#1906) (#2001)
  • Third-party capabilities now work on installed layouts. capability install no longer rejects capabilities with a real engines.gsd range as "incompatible with GSD 0.0.0" — the host version is now read from the authoritative gsd-core/VERSION file across every runtime and the capability install CLI. The installer also now ships the registry generator scripts (gen-capability-registry.cjs, gen-loop-host-contract.cjs), so installed third-party capabilities actually compose into the loop instead of being silently discarded. (#1938)
  • /gsd:verify-work preserves verification state across gap-closure execution and no longer auto-promotes deferred follow-ups into blocking gaps — resuming after /gsd:execute-phase --gaps-only used to lose the verification state: the UAT ## Gaps still read status: failed even after their fix plans executed, so verify-work re-diagnosed them as fresh blockers, spawned a new gap plan, and reported only the new plan as verified. A state contract now links each gap to its fix plan: every UAT gap carries a stable gap_id (G-{phase}-{N}), gap-closure plans tag the ids they address in their frontmatter (gap_ids: […]), and a new reconcile_gaps step on resume marks a gap status: resolved when its plan has a matching *-SUMMARY.md — so fixed gaps aren't re-diagnosed and the phase can close. Separately, a deferred-follow-up branch captures future-work ideas (signals like "later", "next version", "out of scope") into a ## Deferred Follow-Ups section instead of creating a blocking gap/plan. (#1921) (#2025)
  • roadmap update-plan-progress no longer counts stray non-plan *-SUMMARY.md files against phase completion — remediation/gap-closure summaries (e.g. 30-FIX-CR02-SUMMARY.md, 30-GAPCLOSURE-SUMMARY.md) inflated summary_count, and once summary_count >= plan_count the phase silently flipped to Complete (checkbox checked, date stamped) even though several plans had no summary. A new countMatchedSummaries helper (core-utils) pairs summaries to plans via the PLAN→SUMMARY marker swap + the <stem>-SUMMARY.md form (layout-agnostic across root, bare, and nested layouts), so only a summary that corresponds to a real plan counts. Wired into scanPhasePlans (fixing roadmap listing, state sync, verification, workstream inventory at once) and cmdRoadmapUpdatePlanProgress. (#1988) (#2016)
  • milestone complete --ws requirements archive header now points at the workstream REQUIREMENTS.md — the archive header string hardcoded the root path (`…see .planning/REQUIREMENTS.md`), so a workstream archive directed readers at the wrong file even though #1917 had already fixed the archive locations to land inside the workstream. The display path is now derived from the same workstream-aware reqPath the writer uses (path.relative(cwd, reqPath)), so root behavior is byte-identical and the workstream case correctly reads .planning/workstreams/<ws>/REQUIREMENTS.md. (#1993) (#2015)
  • Load-failed capability gates now fail open with a loud warning instead of blocking the whole project — when an installed overlay (third-party) capability failed to load (e.g. an incompatible engines.gsd range) but had declared a gate-kind loop hook, the loop resolver injected a blocking synthetic gate (blocking:true, onError:halt) at every point where that capability declared a gate. A single incompatible capability therefore halted every ship:pre and verify:post in the project — unrelated to what the gate would have checked, and with no remediation surfaced. The resolver now injects no gate and instead emits a loud warning — to stderr and in the loop render-hooks envelope's warnings array — naming the load-failure reason and the exact gsd capability remove <id> remediation, and the loop proceeds (fail open). The capability id embedded in that remediation is validated against the canonical id shape first, so a malformed overlay directory name cannot inject shell metacharacters into the surfaced command. The loader still records _overlay.blockedGates; only the consequence changes from block to warn. step/contribution overlays were already skip-open. (#2009) (#2075)
  • phase.complete now updates the ## Progress rollup row even when an earlier phase-numbered table precedes it — the Progress-row writer used a non-global regex that matched any table row starting with the phase number, so it bound to the first such row (e.g. a | Phase | Requirements | Count | coverage table), no-op'd on the wrong 3-column row, and never reached the real Progress row. The regex is now scoped to the ## Progress section so it binds to the correct table. The command still returned roadmap_updated: true (that field is fs.existsSync(ROADMAP.md)), masking the silent failure. (#2012) (#2032)
  • context7 now works for plugin-marketplace installs (8 agents regained doc lookup) — the agents granted only mcp__context7__*, which matches a standalone context7 MCP server but not the official Claude Code plugin-marketplace install (context7@claude-plugins-official), whose tools are named mcp__plugin_context7_context7__*. The grant never matched, so advisor/ai/domain/phase/project/ui-researcher + planner + executor silently lost documentation lookup and fell back to WebSearch. All 8 agents now grant both forms, the researcher profile table is updated, and a parity guard asserts no agent grants the standalone form without the plugin form. (#2017) (#2029)
  • applySurface no longer deletes every gsd-* agent when the skills manifest resolves empty — the agent-prune loop in _syncGsdDir deleted any gsd-*.md not in the staged set, and when the manifest was empty/unresolvable (null manifest, no array entries, no files key, or an unresolvable install source root), the staged set was empty → every agent was pruned. Skills were guarded by pruneSkillDirs's manifest-membership check (conservative preservation on empty manifest); agents had no equivalent. The agent-prune loop is now skipped when the manifest is empty/absent, so agents are preserved while copy (adding genuinely new agents) still runs. (#2018) (#2031)
  • planning-config.md global-learnings path corrected to ~/.gsd/knowledge/ — the features.global_learnings row directed users to ~/.gsd/learnings/, but the implementation (src/learnings.cts, execute-phase.md) stores and reads global learnings from ~/.gsd/knowledge/. Anyone following the docs to inspect, back up, or seed their global learnings looked in a directory the code never touches. (#2019) (#2026)
  • Removed dead SDK file references from runtime-loaded markdown that triggered an infinite find.exe storm on Windows — agents/gsd-executor.md pointed at sdk/src/query/QUERY-HANDLERS.md and gsd-core/workflows/reapply-patches.md at sdk/dist/cli.js, both retired with the SDK package (ADR-0174). AI runtimes that resolve doc references by filesystem search ran find / -iname …; on Git Bash for Windows / maps to the drive root, so find.exe traversed the whole disk (14h+, orphaned processes, 4M+ open handles each, unkillable). The references now resolve to live paths, and a new regression guard asserts no sdk/src|sdk/dist|sdk/handlers file references remain in agents/workflows/references markdown. (#2020) (#2027)
  • roadmap update-plan-progress no longer checks the phase checkbox without verification — the command stamped the phase-level ROADMAP checkbox and completion date the moment the last plan summary landed (called routinely after every wave and every plan), with no verification gate — unlike phase.complete which correctly requires readVerificationStatus(...).status === 'passed'. Now isComplete requires both all plan summaries AND a passed verification, matching the cmdPhaseComplete contract, so the checkbox only fires after gsd-verifier has confirmed the phase. (#2022) (#2030)
  • phase complete no longer marks a milestone done out of order, nor silently writes root state in workstream mode. Completing the numerically-highest phase while an earlier phase was still outstanding wrongly flipped STATE.md to Status: Milestone complete (the milestone-end check only looked for higher-numbered phases, so an out-of-order completion — e.g. Phase 10 before Phase 9 — read as the end). It now reports milestone-end only when every lower-numbered phase in the milestone is checked complete. Separately, in workstream mode with no active workstream, phase complete previously fell back to root .planning and wrote STATE.md/ROADMAP.md (and the mislabel) into the shared root other workstreams read; it now fails safe — asking for --ws <name> or an active workstream — mirroring the existing init progress guard. (#2066) (#2066)
  • Phase directories whose slug begins with a single digit now resolve correctly. A phase like 46-6-rs-pipeline-orchestrator (roadmap name "6 Rs Pipeline Orchestrator") had its phase token over-collected as 46-6 instead of 46, so gsd-tools phase-by-number lookups resolved phase_dir=null / has_context=false (breaking init.plan-phase, init.phase-op, and downstream execute/verify/ship). Numeric phase-token components must now be zero-padded (≥2 digits), so a single-digit slug word is no longer absorbed into the token. Fixed consistently across every same-class implementation — extractPhaseToken, PHASE_TOKEN_FROM_DIR_RE and canonicalPlanStem (health checks / plan pairing), isDirInMilestone's numeric matcher (milestone filtering), and extractCanonicalPlanId — so the health-check and milestone-filter subsystems are fixed alongside phase resolution. (#2059)
  • gsd-tools config-set <key> null now clears (removes) the key instead of persisting the literal string "null". The documented "Clear" action previously fell through the value parser and stored "null" — a truthy value — so "cleared" keys stayed set and config-get returned "null"; for secret keys (brave_search/firecrawl/exa_search) a masked success line hid a truthy value on disk that integrations could pass along as a real credential. config-set <key> null now deletes the key (short-circuiting the typed per-key validators so clearing an enum/boolean/number key removes it rather than being rejected), making the "Clear" flows in settings-integrations.md / settings-advanced.md actually clear. (#2058)
  • init plan-phase no longer collapses foreign-prefixed task/workstream IDs into numeric phases — a query like MEM-01 (where MEM is not the configured project_code) used to have its prefix stripped and resolve to the unrelated numeric Phase 01; it now reports phase_found: false unless a phase directory or roadmap entry literally carries that prefix. The configured project_code's own prefixed phases (e.g. LKML-01 under project_code: LKML) continue to resolve as before. (#2056) (#2105)
  • phase complete no longer ticks the wrong phase's ROADMAP checkbox — completing a phase whose number also appears in a later phase's description (e.g. an idempotent re-run of an already-complete phase) used to mark the wrong phase done, because the checkbox-matching regex greedily spanned from ] to any later "Phase N" mention instead of only the immediately-following phase title. (#2067) (#2079)
  • gsd-tools effort sync no longer crashes in an installed runtime. In any global install (e.g. ~/.claude/gsd-core/), effort sync threw Cannot find module '../../../bin/install.js' — the command reached into the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home, so that file is never present there. As a result, effort config changes (routing_tier_defaults / agent_overrides) silently never reached installed agents without a full reinstall. The two resolvers (readGsdEffectiveEffortConfig + resolveInstallTimeEffort, with their helpers) are now extracted into a shipped gsd-core/bin/lib/install-effort-resolver.cjs that both effort sync and the installer import — a single source of truth that is always present in the installed tree. (#2076) (#2076)
  • model_overrides and per-phase-type models now actually apply to the assumptions-analyzer, code-reviewer, and code-fixer agents on Claude Code. Previously model_overrides["gsd-code-reviewer"] / ["gsd-assumptions-analyzer"] / ["gsd-code-fixer"] (and models.verification / models.discuss / models.execution) were accepted and resolved but silently dropped — the workflows spawned these agents with no model, so they inherited the session model and the configured routing never took effect (no warning). Every spawn now threads its resolved model: discuss-phase-assumptions, code-review, and code-review-fix (both the re-review and the two fixer spawns) resolve it inline, and quick's review step uses the code-reviewer's own resolved model instead of the executor's. The stale "discuss — reserved, no subagent" model-profile docs are corrected to list gsd-assumptions-analyzer, and the verification row now includes gsd-code-reviewer. (#2074) (#2074)
  • /gsd-review's Antigravity CLI reviewer no longer fails silently on large prompts, unavailable pinned models, or pre-session stalls — the agy invocation now uses a file-reference prompt to avoid exec arg-list overflow, is wrapped in an external wall-clock timeout paired with --print-timeout because --print-timeout cannot fire before agy creates a session, passes --model from review.models.agy when set as an escape hatch for a 404'd pinned model, and its empty-output stub now surfaces an agy cli.log diagnostic instead of a bare generic message. Supersedes the #687 "no external killer / inline $(cat)" contract, which predated agy gaining --model and predated its own guidance to pair --print-timeout with a terminal timeout. (#2073) (#2109)
  • init execute-phase, init verify-work, and init phase-op no longer collapse foreign-prefixed task IDs to numeric phases — MEM-01 under project_code: LKML was silently stripped to 01 and resolved to the unrelated numeric Phase 01, because the #2056 guard was applied only to init plan-phase. The guard is now extracted into shared helpers (guardedFindPhase / guardedGetRoadmapPhase) that delegate to the canonical isForeignPrefixedPhaseQuery from phase-id.cts, and all four init commands route through them. (#2104) (#2149)
  • commit --files now commits only the declared paths — gsd-tools commit --files A B previously ran a bare git commit that absorbed the entire staged index, silently sweeping in unrelated files the caller never named. The commit now appends a pathspec (-- <paths>) so only the staged subset of --files lands in the commit; the no---files default path is unchanged. Missing tracked files are still skipped (not committed as deletions, #2014), and when all declared files are missing the function short-circuits to nothing_to_commit instead of absorbing the index. (#2112) (#2148)
  • Fixed unresolvable bare require('gsd-core/...') in gsd-surface command doc — the four require() examples now derive the engine path from runtimeConfigDir (resolvable at runtime), and the reinstall hint corrects npm i -g gsd-core to npm i -g @opengsd/gsd-core. (#2116) (#2213)
  • milestone complete --dry-run now prints a preview plan instead of silently mutating — gsd-tools milestone complete --dry-run was neither parsed nor rejected, so a caller expecting a preview triggered the full destructive mutation (archive phases, move audit artifacts, rewrite STATE.md) with no way to back out. The --dry-run flag is now honored: it returns a JSON plan listing would_archive (roadmap, requirements, audit, phase dirs) and would_update (MILESTONES.md, STATE.md) targets with zero filesystem mutations. (#2118) (#2155)
  • /gsd-secure-phase now has a single SECURITY.md writer — the gsd-security-auditor subagent previously held Write/Edit tools and was instructed to "write SECURITY.md" with no padded <N>- prefix and no template frontmatter, while the orchestrator's Step 6 also wrote the phase-scoped <N>-SECURITY.md from templates/SECURITY.md. The auditor is now return-only (drops Write/Edit, returns a structured verdict with threats_open); the orchestrator is the sole file writer. The workflow's Step 5 spawn constraints explicitly forbid the auditor from writing SECURITY.md. (#2119) (#2154)
  • Dead security scan exports removed; injection-scan docs corrected to match reality — scanEntropyAnomalies and shannonEntropy were dead code with zero production callers (live hooks inline their own patterns for independence). REQ-SCAN-INJ-02/-03 now accurately describe what runs live (injection patterns, invisible Unicode) vs CI-only (base64-decode, codebase scan). (#2198) (#2211)
  • Post-merge, regression, and other GSD test/build gates no longer fail with a spurious "command not found" on stock macOS. These gates hardcoded GNU coreutils' timeout, which stock macOS ships neither as timeout nor gtimeout; a passing build or test run now completes under a portable, coreutils-independent run-with-timeout wrapper instead of exiting 127 and being misreported as a failure. (#2351) (#2426)
  • Installed third-party capability skills now materialize on OpenCode and Kilo — capability install + capability set --runtime opencode (or kilo) could report a capability as installed: true, surfaced: true, active: true while its skill was never written to skills/gsd-<stem>/SKILL.md: the OpenCode/Kilo combined-family install path never called the seam #2322 fixed for other runtimes. Installed capability skills now materialize the same way there too, bound to their declaring capability, with first-party skills always winning a name collision. (#2362) (#2434)
  • Shared requirement IDs across multiple plans no longer read Complete before every declaring plan (and phase verification) has finished — execute-plan.md now gates completion on sibling plans' SUMMARY.md files via a new read-only requirements ready-ids check, and a gaps_found phase verification reverts any requirement ID this phase owns back out of Complete before the gap report renders. Single-plan requirement IDs are unaffected — no added latency. (#2388) (#2424)
  • phase.add no longer silently mistakes a goal-shaped description for a phase title — a long or multi-sentence description used to land verbatim in the ### Phase N: header with no signal anything was off; phase.add now returns a warning field when the description looks goal-shaped, and the phase-number auto-detect docs now correctly point callers at the orchestrating workflow instead of implying gsd-tools.cjs resolves it itself. (#2390) (#2425)
  • response_language now reaches orchestrator-owned prompts across most workflows and the UAT verification checkpoint frame — previously only subagent prompts honored a configured response_language; the orchestrator's own questions (verify-work, new-project, new-milestone, quick, manager, and others) and the hardcoded English UAT checkpoint banner stayed in English regardless of configuration. Both now render in the configured language, with output byte-identical to before when unset. (#2402) (#2457)
  • Codex installer no longer double-registers each agent role in config.toml, eliminating one duplicate-role startup warning per agent — generateCodexConfigBlock stopped emitting [agents.gsd-*] tables whose config_file pointed back at the same standalone TOMLs Codex already auto-discovers under $CODEX_HOME/agents/; reinstalling over an existing config also drops any legacy managed role tables left by a prior install while preserving unrelated user config and the user's own AgentsToml scalars. (#2406) (#2432)
  • Production dependency tree carries no known advisories — five advisories disclosed against the transitive tree under @anthropic-ai/claude-agent-sdk → @modelcontextprotocol/sdk were cleared: fast-uri (GHSA-4c8g-83qw-93j6, high) and hono (GHSA-xgm2-5f3f-mvvc, GHSA-hvrm-45r6-mjfj, GHSA-w62v-xxxg-mg59) re-resolved to patched releases inside their already-declared ranges with no package.json change, and @hono/node-server (GHSA-frvp-7c67-39w9) pinned to >=2.0.5 via overrides because @modelcontextprotocol/sdk@1.29.0 — already the latest published version — still declares the vulnerable ^1.19.9 range. npm audit --omit=dev reports zero advisories. (#2496) (#2497)
  • Custom STATE.md frontmatter keys are no longer dropped on every mutating verb — syncStateFrontmatter rebuilt the frontmatter from a fixed schema, silently dropping any custom key. It now carries forward existing keys the schema does not own. (#2202) (#2233)
  • Non-frontend phases with UI hint: no are no longer blocked by the UI-SPEC gate — the UI safety gate's token list included the bare token UI, which matched GSD's own **UI hint**: no metadata line and false-detected a UI, blocking backend/infra phases at /gsd-plan-phase. An explicit UI hint: yes|no is now authoritative and the hint line is no longer token-sniffed. (#2150) (#2222)
  • OpenCode reviewer no longer silently yields an empty review on large prompts — /gsd-review --opencode now invokes opencode run --format json and reconstructs the review from the assistant text parts, so a large-prompt run where the default build agent ends its turn with zero output tokens no longer produces an empty stub. When the agent genuinely emits no text, the stub now reports the stop reason, output-token count, and captured stderr instead of a generic message. (#1936) (#1992)
  • OpenCode's first-time install baseline now protects pre-existing files under the commands/ directory, not just the legacy command/ alias — after #2329 moved OpenCode command materialization to commands/, the baseline scan that guards a machine's very first GSD-tracked install still only knew about the legacy command/ directory, so a pre-existing, unrelated commands/gsd-*.md file was silently deleted by ordinary command materialization instead of blocking the install for an explicit keep/remove choice — the same protection command/ already had. The scan now covers both directories. Kilo is unaffected and keeps using command/. (#2354)
  • api-coverage detector no longer false-positives non-API phases (and no longer fails open) — the external-API-integration detector behind the blocking verify:pre seal gate required only same-line co-occurrence of an integration verb and an API noun, treated / as a word boundary (so first-party Next.js src/app/api/… route paths matched), and read any capitalized word before API/SDK/REST/GraphQL as a service name (so threat-model prose like "Resolver-only API" fired). It is now fail-closed: the compound rule requires the integration verb and API noun to share one clause (the clause boundary is the whole relationship test — no fragile word-gap cap that a genuine long integration clause would trip); fenced code, inline code spans, and path-shaped tokens are excluded before matching while external hosts like api.stripe.com/v1 still count; and the <Service> API surface rule rejects stopwords, locality/protocol descriptors ("Internal API", "REST API"), compound modifiers, and first-party-qualified services, so a real vendor name (Stripe API) fires from any clause position. A phase that integrates no external API can declare it first-class in COVERAGE.md — No external API integration: <reason> — instead of fabricating a matrix row; when the detector still finds signals, the declaration overrides but the gate surfaces the overridden signals so the contradiction is visible. Because a false positive is cheaply dismissed by that declaration while a false negative silently slips a real API phase past the gate, the detector deliberately leans toward detecting. (#2365) (#2397)
  • stale-bake-guard hermeticity fix (test-isolation) — the readGsdEffectiveModelOverrides subtest no longer reads the developer's real ~/.gsd/defaults.json; the resolver now accepts a homedir seam so the test sandboxes HOME. (#2152) (#2223)
  • /gsd-surface (list/status) works on Claude Code global installs — the installer now writes a .gsd-source marker pointing at its commands/gsd source, so findInstallSourceRoot resolves on the global skills layout (which ships no commands/gsd tree) instead of throwing could not locate commands/gsd. (#1487) (#1487)
  • Cursor no longer shows every /gsd-* command twice — a --cursor install wrote both a skill and a slash command for each action, so every GSD entry appeared twice in Cursor's / menu. GSD now installs Cursor skills as user-invocable: false (matching the existing CodeBuddy behavior), so the slash command is the single / entry point while skills remain model-invocable. (#2341) (#2386)
  • phase complete --phase N now works alongside the positional form — the phase verb family treated the first positional as the phase number, so --phase 12 was passed as the literal phase name and failed with 'Phase --phase not found'. The phase family now accepts the --phase flag consistently with the state family, and unrecognized flags yield a usage error. (#2201) (#2231)
  • Third-party capability skills now surface correctly after install — a skills-only role: feature capability installed active but its skills never reached the runtime surface, capability enable/set rejected it as unknown capability, and capability list disagreed with capability state. resolveSurface now unions the composed registry's capabilityClusters into the surfaced skill set (no on-disk linking), the writer validates against the composed overlay-aware registry, and capability list carries a surfaced field matching capability state. (#2054)
  • /gsd-ship no longer emits a 100%-missing TDD Audit noise table — the TDD Audit PR-body section was always emitted, but the execute pipeline only writes gate_status: git trailers when TDD mode is active. Without TDD mode (the default), every commit was counted missing and the table was pure noise with no way to disable it. The section is now gated behind workflow.tdd_mode: when TDD mode is off, both the TDD Audit section and the aggregate gate_status: trailer are skipped entirely; when on, the existing behavior is preserved. (#2467)
  • phases.clear now archives phase history under the outgoing milestone version, not the newly-switched one — because new-milestone advances the milestone before clearing leftover phases, the phase-history archive was silently misfiled under the new milestone's <version>-phases/ directory. A new --archive-version override on phases.clear (threaded from the new-milestone workflow) files the archive under the previous milestone's version; without it, behavior is unchanged. (#2288) (#2323)
  • Fixed: probe-core's runProbeCli now fails closed on per-item adapter garbage inside a well-shaped report envelope, matching its documented 'fails closed on adapter garbage' contract. (#1910)
  • Deferred out-of-scope findings logged to deferred-items.md are now surfaced — the executor's SCOPE BOUNDARY convention writes discoveries to a phase directory's deferred-items.md, but nothing read it back, so those items were permanently invisible. /gsd-progress's forensic audit and audit-uat now glob .planning/phases/*/deferred-items.md and surface unresolved entries. (#2287) (#2318)
  • /gsd:verify-work no longer silently terminates when all remaining UAT tests are blocked — sessions with blocked_count > 0 and pending_count == 0 now route to complete_session as expected, enabling the zero-issues auto-transition path. (#1722)
  • state record-metric no longer appends per-plan rows into the By-Phase velocity table — it now maintains its own Per-Plan Metrics table (self-created on first use), and its auto-create scaffold header is corrected. (#2253) (#2253)
  • Dynamic routing now escalates the model, not just effort — with dynamic_routing.enabled, retry attempts advanced the reasoning effort but the model stayed pinned to the default tier because resolve-execution resolved the model without consulting dynamic_routing. resolve-execution now resolves the model per-attempt through the tier ladder (e.g. standard→heavy on attempt 1, capped at max_escalations); resolution is unchanged when dynamic routing is disabled. (#2068) (#2334)
  • /gsd-next no longer reports a project as complete while phases are still unchecked — smart-entry's completion check now grounds in ROADMAP.md's actual Progress table (global, authoritative) instead of STATE.md's stale milestone-scoped total_phases, and its status regex requires milestone-level language (milestone complete / all phases complete / complete) instead of matching any per-phase shipped or done substring. Together these fix the false-complete misclassification that could route /gsd-next toward /gsd-new-milestone — which archives still-pending phase directories. (#2466)
  • Codex reviewer now captures the review via codex's --output-last-message flag instead of redirecting stdout, so Windows process-teardown output no longer pollutes the review file and slips past the empty-output guard. (#1709)
  • last_activity now shows your local calendar day — the clock seam derived the date by slicing a UTC instant, so in negative-UTC-offset zones during UTC's early evening the date-only last_activity field jumped a day ahead of the operator's actual date (and of last_updated's local date). Operator-facing date fields now use a host-local calendar day while internal/cosmetic stamps stay UTC. (#2136) (#2216)
  • A phase with a deliberately-unexecuted (superseded) plan no longer stays stuck below 100% — a plan reassigned or dropped mid-phase can never gain a matching SUMMARY, yet plan-scan counted it forever, so the phase read In Progress and the milestone sat below 100% permanently — the plan-level analogue of the retired-phase bug (#1514). Mark such a plan status: superseded in its PLAN.md frontmatter and it is now excluded from both the plan and summary counts, so the phase completes honestly (a 13-plan phase with 2 superseded reads 11/11). Plans without the marker are unchanged. (#2349) (#2404)
  • milestone_name is no longer clobbered with a delimiter-led fragment — getMilestoneInfo's ## heading regex was unanchored, so it matched a heading quoted inside backticks in the Milestones bullet and wrote garbage like — Active Milestone over the curated milestone name on every phase transition. Now consults the 🚧 marker first, anchors the regex to line start, strips the leading delimiter, and widens the preserve guard so a bad derive keeps the existing name. (#2135) (#2215)
  • init milestone-op now counts project_code-prefixed phase directories correctly — fully shipped milestones using the standard prefixed directory layout no longer report completed_phases: 0 or stay falsely incomplete. (#1844) (#1844)
  • /gsd-mempalace-capture no longer crashes on first invocation — the skill's own documented rooms: example wrote a flat list of bare strings, but mempalace's miner expects each entry as a dict with a name key, so following the example verbatim and running mempalace mine crashed with TypeError: string indices must be integers, not 'str'. Both skills/gsd-mempalace-capture/SKILL.md and commands/gsd/mempalace-capture.md now ship the corrected - name: <room> shape, so the documented example runs successfully end-to-end. (#2464)
  • /gsd-quick no longer halts with a stale-base worktree mismatch — the worktree executor now degrades to sequential execution when its fork base has diverged from origin/HEAD, instead of spawning a worktree guaranteed to fail the base-mismatch guard. (#1991)
  • GSD_ALLOW_SYMLINKED_DEST=1 lets users with intentional symlinked configHome layouts install/update again — v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like skills/ or hooks/) was a pre-existing symlink, with no opt-out. Three legitimate user-owned layouts were blocked: multi-account configs with symlinked shared skills/hooks (POSIX symlinks), Windows Junctions to shared skills dirs, and dotfiles-managed configHome (e.g. nix-darwin symlinking ~/.claude itself to a version-controlled dir). The new env var follows user-owned symlinks instead of refusing them, while preserving the two load-bearing refusals from the original threat model: path-traversal in the destSubpath string itself (../../etc-style), and a symlink resolving to the install root itself (would let the prune pass wipe it). (#2393) (#2445)
  • state record-session no longer silently drops inserted fields on a CRLF STATE.md — the section-rewrite regexes in cmdStateRecordSession used literal \n which couldn't match a CRLF STATE.md (---\r\n), so when a canonical session field (Resume file / Stopped at / Last session) was missing and had to be inserted via the section-rewrite path, the CRLF-tolerant detector entered the branch, the writer regex silently no-op'd, but updated.push(...) ran unconditionally. The command returned {"recorded": true, "updated": ["Resume File"]} while the field was never written to disk. With core.autocrlf=input, the CRLF working-tree file produced no git diff/git status change, so the bug was invisible. Both regexes now use the CRLF-tolerant \r?\n form (same canonical pattern already in use elsewhere), and a new defensive invariant gates updated.push(...) on the replace callback actually firing — so a future detector/writer drift will surface as missing updated entries rather than re-arming this silent-success class. (#2482)
  • /code-review no longer skips a phase whose SUMMARY.md records ~/-prefixed file paths — such a path was silently dropped as "deleted" (bash never tilde-expands a ~ that arrives as a variable's value), emptying the review scope and reporting "no source files changed" as a false success. Tilde paths are now expanded to $HOME/… before the deleted-file filter runs. (#2419)
  • Setting external_job.submit_timeout_ms / poll_timeout_ms / artifact_dir in .planning/config.json now actually configures the SLURM adapter — the keys were declared by the external-job capability but the adapter only read env vars, so config edits silently had no effect. The adapter now resolves them through the canonical capability-config seam (env override > config > registry default), surfaces the resolved artifact_dir in submit output, documents why the contribution registers at execute:wave:post (#1164 asks for wave:pre, which execute-phase.md does not dispatch today; wiring it is a core-loop change #1164 explicitly defers), and gains unit coverage for the CLI surface (parseFlags, findPlanningDir, resolveExternalJobSettings, formatShowReport). (#1164) (#2006)
  • The Antigravity reviewer in /gsd-review no longer reviews blind — agy -p never granted the agent the repo under review, so it frequently anchored on its own scratch directory and returned plan-text-only verdicts counted at full consensus weight. The reviewer is now granted the repo (capability-probed --add-dir) and anchored to the absolute repo root; a review that still runs without repo access is stamped [reviewed-without-repo-access] and down-weighted in the Consensus Summary. The cursor-agent prompt gains the same absolute-root anchor. (#2176) (#2184)
  • Non-Claude installs no longer brand all GSD output as Claude — the installer never persisted runtime: <id> into ~/.gsd/defaults.json for non-Claude runtimes, so resolveRuntime() (precedence: GSD_RUNTIME env > config.runtime > 'claude') fell through to the hard-coded 'claude' default. A non-Claude install showed agent_runtime: "claude" and Claude-formatted /gsd-* slash hints with no env or config hand-set. The installer now persists runtime: <runtime> into ~/.gsd/defaults.json for non-Claude runtimes, mirroring the existing resolve_model_ids: "omit" write at the same call site. Claude is the fallback so it needs no write; an explicit pre-existing runtime value is always preserved. (#2395) (#2446)
  • Autonomous reruns now skip phases with deferred verification until you resume them explicitly — if a prior /gsd-autonomous run recorded verification_deferred_human or verification_deferred_gaps, later reruns no longer drop back into the same prompt loop and instead point you at the saved resume command. (#1846) (#1846)
  • requirements mark-complete no longer reports silent success when the traceability row is missing — it OR-ed its checkbox and table-row writes into one flag, so a checkbox-only reconcile returned a payload byte-identical to a full reconcile while the traceability row stayed Pending (and re-run masked it as already-complete). It now surfaces table_unmatched for IDs whose checkbox reconciled but whose table row is absent, and treats a checked box with no table row as partial rather than done. (#2140) (#2219)
  • state prune now resolves the current phase from the canonical location — frontmatter current_phase, the Current Phase field, or the prose Phase: line scoped to the ## Current Position section — instead of extracting Phase over the whole document, where stateExtractField's pipe-table fallback could latch onto an unrelated | Phase | N | row (e.g. a historical verification table) and compute a wrong prune cutoff. (#1832)
  • model_overrides Claude model IDs now resolve to Agent-tool aliases on the claude runtime — a full Claude model ID (e.g. claude-sonnet-5) in model_overrides was returned verbatim and silently dropped by the Claude Agent tool (whose model parameter documents only tier aliases), causing the spawned subagent to inherit the parent session model instead of the configured one. It now maps to the tier alias (sonnet/opus/haiku/fable), consistent with the model_policy path (#1144). Bare aliases, non-Claude values, and non-Claude runtimes are unchanged; a Claude ID with no alias warns once and falls through to tier resolution. (#2041) (#2048)
  • validate health no longer false-flags the adaptive model profile, and now warns when a models.<phase_type> tier is invalid — health reported W004 invalid model_profile "adaptive" for a profile that has been valid since v1.40, and a typo like "planning": "opuss" was accepted in silence while the resolver quietly ignored it. Health now sources its profile list from the model catalog and emits W022 for unknown phase types and invalid tier values. (#2336)
  • Phase dirs whose slug leads with a multi-digit number (e.g. a year) resolve again — a phase like 14-2026-photos-performance (roadmap name "2026 Photos & Performance") had its phase token over-collected as 14-2026, so init.plan-phase, init.execute-phase, phase-plan-index, state.planned-phase, and roadmap.annotate-dependencies reported phase_dir=null / plan_count=0 while the directory existed. Continuation segments of a phase token are now capped at the exactly-2-digit zero-padded form the write side emits, via a single shared grammar source consumed by all five parsing sites (the residual case from #2043). (#2232) (#2254)
  • Phase headers that place a parenthetical tag before the colon (### Phase 26 (Cluster B): Title) now resolve and enumerate the same as untagged headers. Previously the resolver returned not-found and roadmap analyze/listing silently dropped the phase (wrong phase_count, progress, and next_phase). Tag tolerance is applied at every phase-header read site; untagged and all existing header formats parse unchanged. (#1765)
  • /gsd-stats no longer misreports a phase as Not Started when two directories collide on the same phase key — cmdStats now folds colliding statuses by precedence (Complete > Needs Review > Executed > In Progress > Planned > Not Started) instead of overwriting last-write-wins, so the furthest-along status wins regardless of fs.readdirSync order. Separately, /gsd-health now emits a new W023 warning whenever two or more real phase directories collide on the same normalized phase key, naming both directories and their independently-computed statuses (neutral wording — never guesses which is the real one). (#2461)
  • Executor and milestone-summary/forensics workflows now call state.* commands with named flags so the named-only router records metrics, decisions, blockers, and session continuity instead of silently dropping positional args. (#1873)
  • bug-1367 install test no longer fails on Windows CI when hooks/dist isn't pre-built — the test ran install.js without building its hooks/dist precondition (a gitignored build artifact the unit lane doesn't build), so on a lane without pre-built hooks the installer hit "Failed to install hooks: directory is empty" and the before-hook threw. The test now builds hooks in its own before() (mirroring golden-install-parity). (#1926) (#1927)
  • /gsd-fast now appends Quick Task rows to STATE.md again — the log_to_state column-count guard used an off-by-one awk formula (NF-1) that was always one too high, so the schema gate rejected the very table quick.md creates and silently skipped the STATE.md update. Also now supports the 6-column validate-mode table. (#2133) (#2214)
  • Build the gitignored hooks/dist/ artifact once upfront in scripts/run-tests.cjs (the same chokepoint as ensureBuiltArtifacts), before any concurrent install test spawns install.js. Closes the scoped-CI first-build empty-dir race that intermittently failed install tests with Failed to install hooks: directory is empty (e.g. bug-3683-workflow-colon-namespace-leak). (#1967) (#1968)
  • workstream progress no longer reports shipped milestones as executing — gsd-tools workstream progress now derives each workstream's status from authoritative shipped signals (an archived milestone snapshot under milestones/, or a SHIPPED marker in the workstream ROADMAP) instead of trusting the mutable STATE.md Status field, so a stale field can never hide a shipped/archived milestone. The output adds status_source (field | derived) and status_conflict (true when the derived value disagrees with the stale field). (#1913) (#1916)
  • Windows install/upgrade/state-write operations no longer fail on transient antivirus/indexer file locks — the fs.renameSync atomic-publish sites (install state, hooks config, capability ledger/lifecycle, phase/workstream/milestone dirs, roadmap, planning/state locks) now retry EPERM/EBUSY/EACCES via retryRenameSync instead of propagating the transient lock; enforced by the new local/require-fs-op-fallback lint rule (ADR-1703 Phase 6). (#1740) (#1742)
  • reconstructFrontmatter now emits valid YAML for scalars and block-array items that were previously serialized unescaped. Values carrying a YAML indicator plus a literal quote/backslash, embedded control characters, the empty string, a leading YAML indicator, or leading/trailing whitespace are now routed through a properly escaped double-quoted form, so frontmatter round-trips through strict parsers (js-yaml, PyYAML) instead of corrupting the block on the next state sync. (#1807)
  • phase remove no longer destroys the Progress table when removing the last phase — deleting a phase used a whole-document regex whose scan, on the final phase, ran past the section and swept away the ## Progress heading and its entire tracking table; the deletion is now structurally bounded to the phase’s own section. (#2253) (#2253)
  • Subagent prompts embedding orchestrator-relative planning paths now resolve correctly when the spawned subagent's own working directory differs from the orchestrator's (e.g. a git worktree) — init.* (and state.load) command handlers now emit state_path, roadmap_path, phase_dir, project_path, research_dir, codebase_dir, intel_dir, conflicts_path, debug_dir, and similar fields as absolute paths anchored on the project root, and the planner/checker/verifier/synthesizer/roadmapper/debugger/mapper/classifier subagent-prompt blocks that previously hardcoded bare .planning/... literals now reference those fields instead; a subagent spawned into a different cwd would previously report real, already-committed files as missing. (#2376) (#2428)
  • phases clear archives phase directories instead of destroying them — at a milestone switch, committed phase directories were hard-deleted (rmSync) with no archive, silently losing browsable phase history (the #1447 dirty-tree guard was a no-op for the common committed case). Phase directories are now moved to milestones/<version>-phases/ (collision-safe; timestamp fallback when no version resolves), so history survives the switch. The #1447 uncommitted-changes guard is retained as a secondary backstop. (#1871) (#1919)
  • /gsd-review and /gsd:ship temp files are now scoped to a single per-run directory — both workflows previously wrote prompt, section, and reviewer-output files to /tmp/gsd-review-*-{phase}.* keyed only on the bare phase number, so two projects sharing a phase number (or a crashed run's leftover file) could collide and silently feed a reviewer another project's stale content with no error; every temp path now lives under one mktemp-created run directory that's removed after the review completes. (#2358) (#2433)
  • Cross-AI review no longer silently drops the Codex/Claude/Gemini lanes on large plan sets — the prompt-fed reviewer blocks in review.md invoked each CLI with no explicit timeout, so a slow source-grounded review was killed at the host default (~2 min) and the lane was silently lost. The workflow now directs a high Bash timeout and frames an empty output as a timeout (not the crash it was misdiagnosed as). (#2194) (#2226)
  • Runtime brand-swap no longer mislabels <runtime_compatibility> comparison tables — every runtime installer that rebrands "Claude Code" to its own name (Cursor, Windsurf, Trae, Cline, CodeBuddy, Qwen, Hermes) also swapped it inside the runtime-comparison tables in shipped workflows, where "Claude Code" is a compared-runtime label, not a host self-reference — corrupting the comparison. Branding now protects <runtime_compatibility> regions while still rebranding genuine self-references. (#2284) (#2309)
  • check tdd.review-checkpoint no longer silently skips TDD plans with CRLF line endings — the frontmatter regex at src/check-command-router.cts:751 used literal \n which couldn't match a CRLF PLAN.md delimiter (---\r\n), so a Windows-authored type: tdd plan was silently classified as "no type:tdd plans found" and the advisory gate short-circuited to a confident pass with no violations table. The regex now uses the same CRLF-tolerant form (/^---\r?\n([\s\S]*?)\r?\n---/) already in use elsewhere in the same file (line 205, extractPlanDesignatedSections). With core.autocrlf=input, the triggering CRLF was invisible to git diff/git status, so the contributor had no way to tell their plan was being misclassified. (#2477)
  • Phase verification no longer reads stale from filesystem timestamps alone — staleness is now derived from git commit times instead of file mtimes, so a phase whose report declares status: passed stays passed across a fresh git clone, cp -R, or an unrelated touch/reformat, instead of being silently downgraded to stale by a checkout-order mtime skew. (#2348) (#2394)
  • /gsd-progress no longer reports a stale root milestone in workstream mode — in a multi-workstream project with no active workstream set, gsd-tools query init.progress silently fell back to root .planning/STATE.md (often stale) and reported it confidently. It now fails safe with an actionable error naming the available workstreams and the --ws/workstream set fix, so a stale root value is never reported. Flat mode and --ws <name> are unchanged. (#1912) (#1918)
  • Kilo installs now stage the shared PreToolUse guard hooks the native plugin spawns — Kilo's capability descriptor declared both a nativePlugin (which spawns gsd-prompt-guard, gsd-read-guard, and gsd-worktree-path-guard as subprocesses) and skipSharedHooksInstall: true (which suppressed staging those scripts into the Kilo config dir), so every guard silently no-opped on every Kilo install. The skip flag is removed (Kilo now stages the same hooks bundle as OpenCode, whose byte-identical plugin was unaffected), and the plugin's runHook now warns loudly — once per hook file — when a guard script is missing instead of treating the absence as a silent allow. Resolves #2305. (#2327)
  • The decision-coverage gate no longer fails open on unrecognized decision-ID prefixes — check.decision-coverage-plan classified a populated <decisions> block as "no trackable decisions" (a clean pass) whenever its IDs used a prefix the parser couldn't read (e.g. D5-01 instead of D-01), silently skipping the gate on real decisions. The gate now recognizes any bold-lead-in decision bullet as evidence and fails loud (could-not-parse) when it can't read a populated block, instead of passing. (#2347) (#2389)
  • Windows: stop double-quoting $CLAUDE_PROJECT_DIR-anchored managed node hook paths during the #2979 legacy rewrite, which produced ""$CLAUDE_PROJECT_DIR"/..." and broke every node managed hook with MODULE_NOT_FOUND (PreToolUse-guard deadlock). (#1746)
  • /gsd-stats and STATE.md progress no longer freeze stale total_plans — the progress ratchet was applied to the whole progress record, so any single counter decreasing (e.g. completed_plans) froze every field including total_plans. Now total_plans always takes the freshly derived value (joining total_phases from #1446), so it corrects in both directions — upward when a new phase adds plans, downward when a milestone reorganization removes phases. The write-path applyStatePreservation also switched from wholesale block restore to per-field merge, so state planned-phase writes a consistent total_plans instead of the pre-transform stale value. (#2468)
  • phase complete now updates STATE progress on milestone-grouped roadmaps — deriveProgressFromRoadmap parses the ## Progress table by header (column-by-name) instead of a fixed 4-column layout, so the 5-column milestone-grouped shape is no longer silently unparsed. (#2168)
  • Windows Claude Code hooks now work under PowerShell — when Claude Code's hook runner resolves to PowerShell (not Git Bash), every GSD-installed hook failed with Unexpected token because the installer emitted bare quoted paths with no PowerShell call operator. The fix adds a hookShell parameter to the hook-command projection chain; when hookShell='powershell', the & call operator is prepended. Default behavior (Git Bash, no prefix) is unchanged. (#2236) (#2261)
  • /gsd-debug now auto-resumes instead of stopping mid-investigation — when the debug session-manager's own turn ended before the investigation was complete, the orchestrator treated the intermediate progress summary as completion and returned control to the user. It now recognizes a non-terminal CONTINUE_REQUIRED return, auto-resumes from the on-disk checkpoint, and only stops for genuine terminal conditions (with a no-progress anti-loop guard). (#2257) (#2300)
  • Installing a non-Claude runtime no longer breaks Claude's model resolution in no-project sessions — the installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json, which any runtime read back, so install order silently flipped Claude's adaptive tier aliases (executor→sonnet, planner→opus) to an empty model string. Resolution is now scoped to the runtime actually resolving, via a per-install .gsd-runtime marker: Claude ignores a global-defaults omit and keeps its tier aliases, non-alias runtimes still omit, and an explicit project-level omit/true is always honored. (#2297) (#2332)
  • check.decision-coverage-plan no longer false-blocks on decisions cited in <read_first>/<behavior>/<verify>/<acceptance_criteria>/<done> — the gate scanned only <objective>/<tasks>/<task>/<action> tag bodies while its remediation message claimed "(or body)". A decision faithfully cited in any of the five other planner-canonical tags (the natural place for "read this CONTEXT decision before editing" pointers, verification steps, acceptance criteria, etc.) was reported as uncovered with a misleading fix-hint that sent the fixer to "the body" — where a re-citation still failed. The scan now covers all nine planner-canonical tag bodies AND the message names the surfaces it actually scans, so message and behavior cannot drift apart again. (#2372) (#2443)
  • capability state and loop render-hooks now accept --runtime to override the auto-detected runtime — previously both commands parsed only --config-dir, so the runtime config dir was derived from the persisted .planning/config.json runtime (precedence GSD_RUNTIME → config.runtime → claude). A repo that persisted runtime:"codex" resolved the config dir to ~/.codex, where the Claude skill isn't installed, so every skill-bearing capability reported surfaced:false and execute:post/verify:post hooks silently no-op'd when the operator drove GSD from Claude Code. --runtime <r> (canonicalized, so aliases like codex-app work) now bypasses that fallback so the config dir resolves to the explicitly-named runtime's home. Behavior without the flag is unchanged. (#2003) (#2051)
  • /gsd now registers on pi — installing GSD for pi wrote its extension as gsd.cjs, a suffix pi's extension auto-discovery skips silently, so /gsd never appeared and nothing reported an error. The extension now installs as gsd.js, and upgrading removes the stale gsd.cjs. (#2470) (#2478)
  • phase complete no longer false-reports REQ-IDs as missing when the traceability table leads with a status column — the parser required the REQ-ID in the first column, so a table shaped | ☐ | REQ-01 | … matched zero rows and every body REQ-ID was reported missing. It now matches REQ-IDs in any column. (#2203) (#2234)
  • init milestone-op now ignores backlog 999.x headings when counting milestone phases — parked backlog items no longer inflate phase_count or pin all_phases_complete false for an otherwise finished milestone. (#1843) (#1843)
  • Phase archival is now wired end-to-end across the milestone lifecycle — finishes the #1871 follow-up: phases archive is now a real command (the half-wired alias is routed, no longer errors Unknown), milestone complete archives phase dirs by default (--no-archive-phases opts out), and new-milestone §6 stages the archive move + source removal in the same commit so history is preserved atomically rather than left as orphaned uncommitted deletions. (#1871) (#1924)
  • state update-progress no longer mangles the frontmatter and discards the progress suffix — its Progress: regex matched the raw STATE.md including frontmatter, so the YAML progress: key was hit first (corrupting the frontmatter) while the body line stayed stale and was silently reverted on the next write, and any descriptive suffix after the progress bar was destroyed. It now targets the body line only and preserves the suffix. (#2177) (#2224)
  • /gsd-plan-review-convergence no longer silently overrides configured reviewers with Codex — a bare invocation (no reviewer flags) now respects review.default_reviewers (and, transitively, review.reviewer_instances) per ADR-0011/ADR-0015, instead of always injecting --codex and bypassing the configured default. Users without review.default_reviewers configured still get --codex as before. The startup banner now shows what will actually run. (#2451)
  • /gsd-ship no longer silently drops the ship-status note from STATE on merge — the track_shipping step committed the STATE ship-note after creating the PR but never pushed it, so on a fast merge the note stayed local-only and never reached the default branch. The ship-note is now pushed onto the PR branch with a [ci skip] trailer so it lands on merge without a redundant pipeline. (#2138) (#2217)
  • /gsd-debug no longer stalls on a phantom background handoff — the orchestrator treated the foreground session-manager spawn as a background task and queried its agent ID via TaskOutput (which needs a task ID), then waited on a handoff that was never queryable. The workflow now states the spawn is foreground/blocking, forbids passing an agent ID to TaskOutput, and gives a lost-handoff recovery path. (#2196) (#2227)
  • Roadmap phase lookup now ignores fenced examples and the backlog sentinel lane — roadmap get-phase and init plan-phase no longer return fenced sample headings as real phases or treat 999.x backlog items as active milestone work. (#1845) (#1845)
  • phase complete no longer checks the wrong ROADMAP checkbox or writes the plan count into a shipped milestone — the roadmap mutators ran unanchored and un-milestone-scoped, so they could flip a bullet inside a backticked prose literal or a Backlog entry instead of the closing phase's, and write the plan count into a same-numbered phase in a shipped milestone. The checkbox flip is now line-anchored and both writers are scoped to the current milestone. (#2200) (#2229)
  • audit-uat no longer reports a false-clean total_items: 0 when real items exist — the parsers ignored two artifact shapes: a ## Gaps section recording open findings, and verification items declared in frontmatter (human_verification: array) or as ### N.+bold-paragraph blocks. audit-uat now surfaces unresolved ## Gaps entries and reads the frontmatter array / heading shape, so a phase with outstanding UAT/verification work is no longer waved through as clean. (#2286) (#2317)
  • claude_orchestration.enabled: true now actually routes execute-phase waves through the Workflow backend — the capability shipped registered-but-inert: nothing in /gsd-execute-phase ever called its backend detection, and the execute:wave:pre hook it needed was declared but never rendered, so enabling it had zero effect. execute-phase now renders execute:wave:pre before each wave and, when the capability is enabled and all gates pass, dispatches independent plans via the generated Workflow script; any gate miss or disabled config falls back to byte-identical inline dispatch. (#2285) (#2314)
  • roadmap get-phase resolves project-code-prefixed headings by bare number — a bare-number query (e.g. 29) now resolves a drifted ### Phase AB-29: heading, matching the internal resolver used by init.phase-op; previously the CLI returned empty. A bare sibling (### Phase 29:) still takes precedence. A project-code-prefixed heading present only as a summary/checklist line (no matching detail section) now reports a malformed_roadmap diagnostic — for both prefixed and bare-number queries — instead of a silent empty result. (#2114) (#2139)
  • query config-get now returns capability-registry defaults for absent keys — keys declared with a default in the capability registry (e.g. workflow.security_enforcement, which defaults to true) previously reported "Key not found" (exit 1) when missing from config.json, diverging from the runtime's own resolver and letting ... || echo false guards silently read the security gate as disabled. config-get now resolves these through the same registry defaults the runtime uses. (#2256) (#2299)
  • milestone complete --ws now archives into the workstream instead of root — the archive paths (MILESTONES.md, the milestones/ archive dir, and the per-version MILESTONE-AUDIT.md) were hardcoded to root .planning/, so a workstream milestone close scattered its artifacts into root and never produced a workstream-local archive. They now derive from the workstream-aware planning base (planningPaths(cwd).planning); flat-mode (no --ws) is unchanged. (#1911) (#1917)
  • /gsd:new-milestone --ws <name> no longer overwrites the shared PROJECT.md milestone heading — in workstream mode the shared .planning/PROJECT.md had its ## Current Milestone heading rewritten with one workstream's milestone, so with parallel workstreams whichever ran last silently won the shared heading. The milestone-state write in Step 4 is now skipped when a workstream is active, and the commit no longer stages PROJECT.md. The --ws flag is also now parsed into ${GSD_WS}, which previously expanded to empty and silently dropped workstream scope from the suggested next-step routing hints. (#2338)
  • The context-monitor hook no longer fails Codex's Stop hook — GSD wires gsd-context-monitor to Codex lifecycle events including Stop, but the hook emitted a hookSpecificOutput.additionalContext envelope that Codex's Stop schema rejects ("hook returned invalid stop hook JSON output") exactly when context was low. The hook now emits that envelope only for context-injection events (PostToolUse / AfterTool) and exits silently for Stop and every other lifecycle event, while its debounce and critical-session bookkeeping still run. (#2289) (#2324)
  • phase complete now reads milestone-grouped ROADMAP progress tables — progress reported 0% on projects whose Progress table carries a Milestone column, because the reader assumed a fixed column position; it now resolves progress columns by name so both flat and milestone-grouped tables work (#2137). Quick Tasks logging via /gsd:fast also appends schema-correct, lock-safe rows instead of guessing the column count in shell (#2133). (#2248) (#2248)
  • Managed hooks no longer break after a volta node upgrade or prune — on machines using volta to manage Node, the installer baked a version-pinned node path into every managed hook command. Once volta pruned that node version, every hook failed to spawn with No such file or directory at the start of each session, until the installer was re-run. Hook commands now resolve through volta's stable shim, which survives version changes. (#2335) (#2375)
  • Todo severity is now captured and surfaced end-to-end — /gsd-capture (add-todo) now confirms a severity (blocker/major/minor/cosmetic) before writing a todo instead of silently omitting it, and gsd-tools list-todos / init todos now include the severity field in their JSON output (omitted for older todos that have none), so a backlog can be triaged by severity instead of by re-reading every file. (#2337) (#2381)
  • Skill-bearing capabilities now surface correctly on flat command-layout installs — on an install using the flat commands/gsd-<stem>.md source layout (e.g. a Claude Code local project install with no commands/gsd/ subdir), every skill-bearing capability (nyquist, code-review, security, ui, mempalace, ai-integration, profile-pipeline) was silently reported surfaced:false/enabled:false/active:false, so their loop hooks (verify:post, execute:post, etc.) never fired even with the corresponding workflow.* toggle on. The skill-manifest resolver now detects the flat layout and produces the same stems the nested commands/gsd/*.md loader does. (#1858) (#2049)
  • Claude Code installs now pre-approve .planning/ and STATE.md writes — the installer wrote Write(.planning/*)/Write(STATE.md) permission rules, but Claude Code has no standalone Write gate (file edits are gated via Edit(pattern)), so those rules never matched and every fresh install still hit first-run approval prompts (and a session-start warning). The installer now writes Edit(...) rules and migrates the stale Write(...) entries away on the next run. (#2278) (#2302)
  • Roadmap, requirements, and state table edits are confined to the right table — the last ad-hoc table writers (phase completion updating roadmap progress, requirements mark-complete, and state record-metric/velocity) now route through the shared markdown-table seam, so a stray decoy table elsewhere in a document can no longer swallow a phase-progress update, a single ragged neighbouring row no longer silently aborts the whole edit, and per-plan metric recording no longer drops trailing section content or duplicates the section. (#2253) (#2253)
  • Installed third-party capability skills now materialize as real slash commands — a capability could pass every check (installed: true, surfaced: true, active: true) and still never exist on disk: the registry layer counted the capability's skill as surfaced, but the file-copy step only ever scanned gsd-core's own bundled commands, so nothing was ever written to the runtime's skills/ directory and the command was never invocable. Installed capability skills are now staged from where they live, bound to the capability that actually declared and registered them (never inferred from directory listing order), and are subject to the same runtime-targeted body rewrites as first-party skills — first-party skills still win any name collision. (#2340)
  • /gsd:plan-review-convergence can now use the Antigravity CLI reviewer — its reviewer-flag whitelist predated the 1.7.0 Antigravity adapter and silently dropped --agy/--antigravity, so convergence fell back to --codex only and the working adapter was unreachable (especially after Gemini CLI's upstream shutdown). Both flags are now recognized and passed through to /gsd-review unchanged. (#2293) (#2325)
  • npm run lint:ci (and every npm script banner) on next and feature branches cut from next no longer reports a stale pre-release version after a final release — the release pipeline's finalize job shipped X.Y.0 to npm latest but never bumped next to match, so next carried the last rc.N placeholder indefinitely (observed: 1.7.0-rc.6 lingering after 1.7.0 shipped). The finalize job now runs scripts/sync-next-version.cjs — the same step the rc job already ran — keeping next at the last published release for every release type as scripts/sync-next-version.cjs:6-9 always promised. (#2423) (#2437)
  • verify plan-structure no longer false-flags checkpoint tasks for missing <action>/<verify>/<done> — every <task type="checkpoint:*"> was reported as a structural error because the verifier unconditionally required the auto-task fields. It now branches on the task's type attribute: checkpoint:human-verify requires its canonical triple (<what-built>/<how-to-verify>/<resume-signal>), checkpoint:decision requires <decision>/<options>/<resume-signal>, checkpoint:human-action requires <action>/<instructions>/<verification>/<resume-signal> (per gsd-core/references/checkpoints.md), and unknown checkpoint:* subtypes require only the universal <resume-signal>. Non-checkpoint tasks keep the historical <action>/<verify>/<done>/<files> requirements unchanged. (#2473)
  • Hermes installs now project named-agent dispatch onto delegate_task instead of asserting a nonexistent Agent tool — installed Hermes workflows brand-swapped "Claude Code"→"Hermes Agent" but kept literal Agent(...) calls and falsely claimed "The Agent tool IS available", which Hermes doesn't expose. A Hermes .md converter now rewrites named dispatch onto Hermes's delegate_task contract (embedding the resolved role prompt since Hermes has no named-agent lookup, mapping background dispatch, dropping unsupported per-call model), driven by the runtime's documented dispatch facts, and fails closed if a referenced role prompt is missing. (#2284) (#2309)
  • phase complete no longer silently drops requirement IDs the roadmap cites but REQUIREMENTS.md never defined — completing a phase whose **Requirements**: line named an unregistered REQ-ID reported requirements_updated: true with zero warnings while the file was left byte-for-byte unchanged, indistinguishable from a run that wrote everything. Ghost IDs now raise a warning, requirements_updated reflects whether a write actually landed, an active heading like ## v1 Requirements is no longer mistaken for a deferred section, and a phase whose every cited ID is unregistered still reports its missing-requirement rows instead of "No requirements or decisions to check." (#2339)
  • ~/.gsd/defaults.json no longer silently drops model_policy, model_profile_overrides, and runtime — the global-defaults path of config load now forwards these three keys identically to a project's .planning/config.json, so a machine-wide model policy / runtime / overrides specified globally is honored even outside a project. (#2069) (#2442)
  • ROADMAP phase edits can no longer escape their section — completing a phase updated its plan count and per-plan checkboxes with whole-document regexes that could bleed into a neighbouring phase; those per-phase writes are now structurally bounded to the phase own section via a new withSection / withPhaseSection seam (#2130, #2067, #2080). (#2250) (#2250)
  • close_phase_todos no longer leaves moved todos as phantom unstaged deletions in git status — the workflow step moved resolved todos from .planning/todos/pending/ to .planning/todos/completed/ with a plain mv, then committed by listing only the destination directory in --files. Git's index still tracked the moved file at its old pending/ path, so the deletion was never staged and the moved-away file lingered as an unstaged deletion in git status until some later broad git add -A happened to catch it. The step's commit --files list now includes BOTH directories so git add .planning/todos/pending/ stages the deletion atomically with the new completed/ copy in the same commit. (#2415) (#2447)
  • STATE.md ## Session fields now resolve on Windows — the session-section reader used a \n-only heading regex that silently failed on a CRLF ## Session heading, nulling all session state on Windows checkouts; it now reads through the CRLF-safe section seam. (#2253) (#2253)
  • Bullet/em-dash ROADMAP phases no longer resolve to Phase null — the roadmap phase lookup matched only ATX headings with a colon, so a bullet entry like - [ ] **Phase N — Name** (which the roadmapper emits) failed to resolve and Phase null landed in STATE.md; a bullet-only ROADMAP also broke the milestone phase count. Phase lookup and the milestone filter now accept bullet/checkbox entries with an em-dash/en-dash/hyphen/colon separator. (#2199) (#2228)
  • Linuxbrew users no longer lose all GSD-managed hooks after brew upgrade node — normalizeNodePath only recognized macOS Homebrew Cellar paths, so on Linux the version-pinned node path stayed baked into hook commands and 404'd after a node bump (and reinstall couldn't repair it). It now rewrites any Homebrew Cellar path — Intel, Apple Silicon, Linuxbrew, custom HOMEBREW_PREFIX — to the stable <prefix>/bin/node symlink. (#2185) (#2225)
  • milestone complete no longer corrupts the recorded phase — closing a milestone (e.g. v0.5) previously overwrote current_phase in STATE.md with the version's minor digit, and a follow-up state complete-phase mined a bogus 0.5 token and rewrote the file; phase resolution is now anchored so the real phase is preserved and a milestone-closure line is rejected. (#2111) (#2131)
  • Headless MemPalace capture no longer fails silently — the headless invocation mempalace mine <path> --wing <wing> --room <room> used a --room flag that does not exist on the mine subcommand (only search accepts --room), causing every headless/no-MCP capture run to fail with unrecognized arguments: --room and silently skip (onError: skip). The fix replaces the flag with MemPalace's documented room-assignment mechanism: stage the artifact under a room-named subfolder with a mempalace.yaml taxonomy so detect_room() assigns it via folder-path match. (#2220) (#2260)
  • Codex agents no longer fail to launch with an unsupported-model error — GSD was writing an Anthropic tier name (opus/sonnet/haiku/fable) or a claude-* id into each Codex agent's .toml model field, which Codex rejects — fatally on a ChatGPT account (The 'sonnet' model is not supported when using Codex with a ChatGPT account). GSD now never writes an Anthropic-flavored model to a Codex agent: an explicit real-Codex model pin is kept, anything else is omitted so the agent inherits the working session model. (#2310) (#2312)
  • Fixed: a hand-authored non-inferable backstop truth with a stray trailing space or surrounding quotes no longer silently grades green — it correctly abstains (insufficient_spec), restoring the #1154 honest-verifier guarantee. (#1909)
  • commit_docs no longer silently disables on CRLF .gitignore repos — git check-ignore falsely reports a trailing-slash path (e.g. .planning/) as ignored when the .gitignore has CRLF line endings with blank lines. isGitIgnored now strips trailing slashes before querying, so the false positive cannot occur. (#2206) (#2235)
  • Phase-directory resolution fails loud on cross-project collisions — when two unrelated GSD projects share a .planning/phases/ tree, a bare phase number silently resolved to the first 0N-* directory found. The fix detects multiple matches and surfaces an ambiguous_matches result. (#2237) (#2262)
  • Build/test gates no longer report a false failure on repos with no detectable build/test tooling — the post-merge, regression, verify-phase, and audit-fix gates read config-get workflow.build_command/workflow.test_command without --raw, so an unset key returned the literal 2-byte string "" rather than empty output. The [ -z "$CMD" ] guard then saw a non-empty value, skipped the Makefile/Cargo/go.mod/package.json auto-detection cascade, and executed the literal "" as a command → exit 127, misread as a build/test failure (docs-only or planning-only repos, or any repo before its first build file). All of these reads now pass --raw, restoring the intended "no command detected — skip" no-op. (#2350) (#2399)
  • scanPhasePlans no longer counts PLAN-REVIEW artifacts as executable plans — *-PLAN-REVIEW.md files were counted by the loose /PLAN/i fallback. The fix adds a PLAN_REVIEW_RE exclusion before the fallback. (#2252) (#2263)
  • Dependency tree no longer carries a known body-parser advisory — GHSA-v422-hmwv-36x6 (low-severity DoS via invalid limit value, published 2026-07-20) in body-parser@2.2.2 was pulled transitively via @anthropic-ai/claude-agent-sdk → @modelcontextprotocol/sdk → express and surfaced by npm audit --omit=dev. Re-resolved body-parser to 2.3.0 in package-lock.json within express's already-declared ^2.2.1 range; no overrides block needed, package.json is unchanged. (#2473)
  • Milestone audit no longer flags a not-yet-validated phase as a Nyquist failure — a phase that was planned but never run through validate-phase now reports as NOT-VALIDATED (a "run validate-phase" TODO) instead of collapsing into PARTIAL alongside phases whose validation genuinely failed. (#2117) (#2209)
  • CI gates no longer fail with no merge base on branches behind the base. The mutation, changeset-required, and docs-required workflows shallow-fetched the base ref, truncating the ancestry their three-dot origin/<base>...HEAD diffs depend on — so the mutation gate reported failure and silently skipped its Stryker shards, leaving the 80% threshold unverified on any PR not already level with next. (#2452) (#2485)
  • OpenCode slash commands now install to the supported commands/ directory instead of OpenCode's legacy command/ alias — GSD wrote all ~71 /gsd-* commands to command/ (singular), which OpenCode's docs list only as a backwards-compatibility alias for the documented commands/ (plural) convention. Commands now land in ~/.config/opencode/commands/ (global) and .opencode/commands/ (local), and upgrading migrates the legacy directory, preserving any files you put there yourself. OpenCode currently resolves both names, so this is an alignment rather than a rescue — it takes GSD off a path the vendor may withdraw. Kilo is unaffected. (#2354)

Security

  • gate="blocking-human" checkpoints are no longer auto-approved by the execute-phase orchestrator — the package-legitimacy gate (#2827) spans two layers: gsd-executor refuses to auto-approve a gate="blocking-human" checkpoint and escalates it via checkpoint_return_format so a human can vet the package, and execute-phase's checkpoint_handling step decides what happens next. That step dispatched purely on checkpoint type and never read gate, so under --auto / --chain it immediately auto-approved the very checkpoint the executor had just refused to auto-approve (human-verify → {user_response} = "approved"). The slopsquatting defence was therefore inert in exactly the unattended mode where nobody is watching: an [ASSUMED]/[SUS] package reached install with no human ever seeing the verification prompt. checkpoint_handling now carves out gate="blocking-human" (and the package-legitimacy what-built markers) ahead of every auto-mode branch, routing those checkpoints to the standard present-to-user flow regardless of type. references/checkpoints.md documents the gate attribute and its two values for the first time — previously blocking-human appeared nowhere outside agents/gsd-executor.md, so no planner had a documented way to author a checkpoint that auto-mode could not bypass. The existing regression test asserted the executor half only; it now asserts the orchestrator half too, which is why it stayed green while the gate was open. (#2107) (#2113)
  • Patched a transitive denial-of-service advisory in the production dependency tree — body-parser reached GSD via the Claude Agent SDK's MCP dependency and, on versions through 2.2.2, silently stopped enforcing request size limits when given an invalid limit value (GHSA-v422-hmwv-36x6). Pinned to >=2.3.0. (#2470) (#2478)
  • phases.clear --archive-version and milestone complete <version> now reject version labels containing path separators or .. — the milestone version becomes a filesystem directory name that phase directories are moved into, so an unvalidated value could relocate phase history outside .planning/milestones/. Both now validate against a strict version-token pattern and fail loudly. (#2288) (#2323)
  • query config-get no longer leaks secret values or walks the prototype chain — the --default fallback path printed secret-named keys (e.g. brave_search) in plaintext instead of masking them, and dotted-key traversal used raw property access so config-get __proto__/constructor resolved to JavaScript internals at exit 0 instead of erroring. Both absent-key resolution and traversal are now masked and own-property-gated. (#2256) (#2299)
  • Hardened phase/roadmap/plan markdown parsing against quadratic-time (ReDoS) CPU exhaustion — a crafted ROADMAP.md, STATE.md, or PLAN.md with large runs of unclosed (, [, <tag>, <!--, or <details> could drive the phase-header, Plans-count, files_modified, and <tag>-block parsers into O(n²) scans (tens of seconds on a ~1.5 MB file). Every affected regex is now linear: header tag/bracket clauses are length-bounded, the Plans-count scan is section-local, and all <tag>…</tag> extraction routes through a single ReDoS-safe seam. (#2128) (#2141)
  • Installer writes are now confined to the declared config home — the workflow/skill emit path (copyWithPathReplacement) and the Codex config writer (installCodexConfig) now reject any destination that escapes the install root: crafted or absolute paths, path-separator agent names, and pre-existing symlinks are refused before any delete or write. Fail-closed: an install write with no declared root is rejected rather than written unconfined. (#1725)
  • Install write-confinement (ADR-1239 Phase B) — the installer now rejects any runtime-descriptor destSubpath that would write or delete outside the user's config home (path traversal, the config root itself, NUL bytes) and refuses to follow a pre-existing symlink that escapes it. Hardening only; no change to legitimate installs. (#1706)

[1.7.0] - 2026-07-15

Added

  • A default-off, BETA, claude-only "Claude orchestration" capability — adopts Claude Code's Workflow tool (/effort ultracode, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, restoring the wave parallelism + plan-checker + verifier that the #853 backgrounded-agent nesting limitation forces inline on Claude Code, and folding the existing gsd-ultraplan-phase plan-offload under the same runtime gate. When claude_orchestration.enabled is on AND the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets the floor (claude_orchestration.min_agent_sdk_version, default 0.3.149), execute-phase emits a generated Workflow script (waves → parallel() barriers, plans → agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap → separate sequential stages, resumeFromRunId wired to the phase run id, shared budget pool) that composes the SAME executor agent + worktree isolation the inline path uses, so artifacts/commits are produced identically. Detection is pure and fail-closed (any miss → inline), so on any runtime lacking the Workflow tool behaviour is byte-identical to today. Adds a pure module gsd-core/bin/lib/claude-orchestration.cjs (detectWorkflowBackend, emitWorkflowScript), the capabilities/claude-orchestration/ declaration with two gated loop contributions (execute:wave:post, plan:post) and a claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow), federated config keys, and an ADR-1143 implementation amendment. (#1143) (#2044)
  • Phases that integrate an external API/SDK/service can no longer seal without a decided coverage matrix — a new api-coverage gate on the ai-integration capability blocks /gsd:verify-work until the phase produces a COVERAGE.md enumerating the API's full capability surface, with every non-integrated capability an explicit, reasoned opt-out. Full coverage is the default; the matrix is the subtraction record, so "we integrated the API" can no longer silently mean "we integrated whatever the first use case exercised." Toggleable via workflow.api_coverage_gate (on by default). (#1562) (#2065)

OpenCode installs now auto-register the GSD companion MCP server (mcp.gsd) — --opencode install writes a mcp.gsd entry (local stdio → gsd-mcp-server) into opencode.json, so OpenCode drives GSD's command + planning-state surface over MCP with no bespoke plugin (ADR-1239 Phase D / #1682). Idempotent and non-clobbering; a user-defined mcp.gsd is preserved. (#1682) (#1929)

OpenCode plugin handles session.idle + the opencode-subset hook dialect is implemented — the GSD OpenCode plugin now recognizes session.idle (↔ Claude Stop lifecycle point), completing the compaction/idle pair (#1914 shipped compaction). The reserved opencode-subset dialect gains a consumer — hookEventSurfaceFor() in host-integration.cts — describing OpenCode's session/tool/file event subset (no workflow-phase events; the engine owns phase sequencing, ADR-1239 §OpenCode binding). Adds a Claude-parity test asserting the plugin covers the full declared subset. (#1682) (#1930)

  • GSD now warns when model config changed without re-running the installer on static-frontmatter runtimes — on codex and opencode, editing model_overrides or model_profile_overrides or model_policy.runtime_tiers in .planning/config.json or ~/.gsd/defaults.json previously had no effect until the user re-ran gsd install <runtime>, and the failure was silent: the sub-agent kept using the base model. Workflow entry points like gsd-tools init * now emit a one-line stderr warning naming the changed config file and the exact remediation command when they detect the config is newer than the baked agent files. The guard is read-only and warning-only by default, dedup'd per session, and skipped entirely on Claude Code because Claude Code resolves models at spawn time. Resolves #1688 as the structural follow-up to #1650. (#1692)
  • gsd-tools state rebuild — new subcommand that re-derives STATE.md body structure from canonical sources (frontmatter + .planning/phases/ disk scan), reconciling drifted ## Current Position prose, dropping orphaned rows from the **By Phase:** table, clearing template-placeholder field values, and de-duplicating ## Session Continuity Archive blocks. Every mutation is recorded in a ## Rebuild Log audit section. Idempotent (running twice on a clean file is a no-op). Supports --dry-run (preview) and --verbose (tee log to stderr). Heavier, manual counterpart to the lightweight auto-triggered state sync. (#1830)
  • graphify.graph_path makes the knowledge-graph location configurable so one umbrella graph can serve multiple projects — a new .planning/config.json key (path relative to project root, or absolute) overrides where /gsd-graphify query|status|diff read the graph, letting a single curated cross-repo umbrella graph serve every sibling sub-project without N drifting ~5 MB mirror copies. Previously the graph location was hardcoded to <cwd>/.planning/graphs/ with no override; the only workaround was copying the umbrella graph.json into each project (which drifted, wasted disk, and could be silently overwritten by an in-project build). The diff snapshot travels with the configured graph; build stays project-scoped; unset → byte-identical default; a configured-but-missing file yields an actionable error naming the path. (#1825) (#2013)
  • Claude Sonnet 5 is now the standard (sonnet) tier model. The model catalog and provider presets resolve the sonnet/standard tier to claude-sonnet-5 (GA 2026-06-30) across the Anthropic-backed runtimes (claude, copilot, and the anthropic/anthropic-fable presets), plus the OpenRouter-style anthropic/claude-sonnet-5 for opencode/hermes, replacing the superseded claude-sonnet-4-6. Opus and Haiku tier defaults are unchanged (the haiku high-effort preset's escalation slot tracks the current sonnet model). Shipped in 1.6.1. (#1847) (#1848)
  • Third-party capability gates now actually fire via a generic command-exit-zero predicate. — a capability's declared check.predicate gate was rendered for display but never evaluated (only built-in check.query gates were enforced, and the security capability's gate worked solely via a hard-coded ship.md branch). A new generic evaluator (gsd_run check predicate) now evaluates check.predicate blocks by kind; the first built-in kind command-exit-zero runs a bounded sh -c command at the project root and blocks the loop on non-zero exit (timeout → block, fail-closed). The execute:wave:post, execute:post, and plan:post gate-dispatch sites route predicate gates to the new evaluator automatically. (#2008) (#2011)
  • GSD's lifecycle hooks now run under Kimi CLI — installing GSD into Kimi wires its session-state, phase-boundary, graphify, and guard hooks into Kimi's own native config.toml [[hooks]] bus (Beta on Kimi's side) instead of silently no-op'ing, and GSD's Kimi subagents can now run in the background. Kimi's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2095) (#2159)
  • GSD is now installable on pi — npx @opengsd/gsd-core --pi installs the GSD extension to ~/.pi/agent/extensions/gsd.cjs, and /gsd <family> <subcommand> now dispatches real commands through the embedded engine (the reference binding previously could only run query help). Drives pi through the negotiated imperative Host-Integration adapter, with active-model steering and the full pi lifecycle-event surface. (#2102) (#2205)

GSD now ships a pi extension — a real, jiti-loadable ExtensionAPI module (pi/gsd.cjs) that registers /gsd (dispatches through the GSD command-routing hub) + gsd_invoke tool + tool_call event, installable at ~/.pi/agent/extensions/. A reachability test proves the /gsd handler dispatches through the engine (keystone wired, not just registered on a mock). (#1965) (#1965)

  • plan-phase now authors edge and prohibition predicates into PLAN.md must_haves when a phase SPEC omits ## Edge Coverage / ## Prohibitions, so goal-backward verification still has predicates to check on a spec-less phase (ADR-857 Phase 6). Gated by the new default-on workflow.specless_probe_fallback toggle — disable it to skip the fallback (the skip is recorded visibly in the plan). Spec-less prohibitions are authored descriptor-less and disposed flagged/unverified (honest verifier #1154), never a silent pass. (#1835)
  • Discover third-party GSD Capabilities in a new Community Capability Registry. — A non-endorsing discoverability catalog where authors register a Capability via a documentation PR; each entry carries a live latest-release badge and a per-entry GitHub Discussion for community ranking and comments. (#2188) (#2188)
  • GSD now warns when a stale global CLI (e.g. a retired @gsd-build/sdk canary) shadows your project-local install — the gsd-tools CLI startup detects when the running binary is outside the project root while a project-local install exists, and prints a remediation warning to stderr (non-blocking). (#1754) (#1755)
  • gsd-mcp-server — companion MCP server (interface points 1 + 5) — a new bin command (npx @opengsd/gsd-core gsd-mcp-server) runs a stdio JSON-RPC 2.0 MCP server exposing gsd_invoke_command (→ the GSD command-routing hub) + gsd_read_state / gsd_write_state (→ .planning/ state), so any MCP-consuming host (Claude Code, Codex, OpenCode, VS Code, Gemini CLI, Cursor, Cline, Hermes) can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681). Dependency-free (hand-rolled JSON-RPC). How-to: docs/how-to/connect-gsd-mcp-server.md. (#1810)
  • Opt-in absolute token count on the statusline context meter — new statusline.show_context_tokens config (default false). When enabled, the meter shows the absolute context total after the percentage, e.g. "████░░░░░░ 46% (156k)", summing input, cache-creation, cache-read, and output tokens from the hook payload (a broader basis than the meter's percentage, which is derived from used_percentage and excludes output tokens — the two figures can diverge slightly). Default meter output is unchanged. (#2161) (#2174)
  • Long-running compute can now be externalized as async external jobs instead of blocking the agent turn — a default-off external-job capability lets executors submit SLURM jobs, commit a .planning/async-jobs manifest, defer SUMMARY.md, and return external_job_waiting; the core loop already reconciles these manifests (#1165), so this adds the producer half (SLURM adapter, pure manifest module, planner/executor fragments, operation policy). (#1105) (#1998)

GSD now ships a repo-local VS Code extension — a buildable extension (vscode/extension.js + vscode/package.json) that registers gsd.invoke (dispatches through the GSD command-routing hub) in the VS Code command palette. A reachability test proves the handler dispatches through the engine (keystone wired). Not Marketplace-published; mirrors the OpenCode plugin's bar. (#1966) (#1966)

  • Discover third-party GSD Embeddable Orchestration System (EoS) integrations in a new EoS Registry. — A non-endorsing discoverability catalog where host-integration authors register via a documentation PR; each entry declares its Host-Integration interface points, negotiated axes, and protocol version, with a live release badge and a per-entry GitHub Discussion for ranking and comments. (#2193) (#2193)
  • GSD Core ships a .claude-plugin/marketplace.json marketplace manifest so Claude-plugin-compatible runtimes (ZCODE et al.) can discover and install gsd-core from a custom marketplace source. Additive — the existing .claude-plugin/plugin.json and the Claude Code install path are unchanged. The catalog version (plugins[0].version) tracks package.json via the release version-sync. (#1861)
  • GSD now drives VS Code through the Embeddable Orchestration System — the VS Code extension is rewired through the negotiated imperative Host-Integration adapter (active vscode.lm model, engine hook bus, sandboxed storage), gains native Language Model Tools (GSD skills as #gsd-* tools) and #runSubagent dispatch, and runs as a Web Extension (no Node APIs). (#2103) (#2210)
  • /gsd:next smart-entry workflow — adds a state-aware entry point that classifies the current project situation (no-project, blocked, verify-failed, planning, executing, verify-pending, complete, and more) and recommends the right next GSD command. The gsd-tools smart-entry [--json] classifier handles phase ordering including decimal phase IDs; the /gsd:next skill surfaces the workflow with tiered fallback behavior. (#1798)
  • OpenCode now runs GSD's lifecycle safety hooks (prompt-injection guard, read-before-edit guard, injection scanner, worktree/workflow guards, context monitor) via a native plugin installed to ~/.config/opencode/plugins/gsd-core.js. OpenCode declares hooksSurface: 'none', so these hooks were previously inert; the plugin bridges OpenCode's event bus onto GSD's existing hook scripts. Installed automatically by npx @opengsd/gsd-core --opencode and removed on uninstall. (#1923)
  • Host-integration descriptors now carry an extensionEvents vocabulary — the extension-system event surface (OpenCode, pi) is a separate descriptor field from managed hookEvents, so OpenCode declares extensionEvents:opencode without conflicting with the hooksSurface:none invariant. (#1946) (#1946)
  • /gsd-review now supports custom reviewer instances — run one model-capable adapter (e.g. OpenCode) as several independent reviewer identities via a bounded review.reviewer_instances config, so two different models can review in a single pass without manually swapping config or hand-merging REVIEWS.md. (#1517) (#1766)
  • Opt-in git branch and working-state segment in the statusline — the shell prompt's branch/dirty-state signal is hidden for the whole session under the Claude Code TUI, so wrong-branch commits and ship-time push rejections surface only after the fact. New statusline.show_git config (default false) renders the branch name plus staged/unstaged/untracked/ahead/behind markers (or ✓ when clean and in sync) after the directory segment. When disabled, no git subprocess is spawned and output is unchanged. (#2163) (#2183)
  • /gsd:onboard guides brownfield setup — existing repos now have a top-level onboarding command that routes through codebase mapping, docs ingest, project initialization, and an onboarding summary without silently overwriting planning files. (#1994)
  • Plural/optional/chosen assumption-delta checkpoint during planning — when a phase makes something plural, optional, or chosen that used to be singular, required, or derived, the planner is now prompted to re-ask whether the primary key / identity model still names the right thing, preventing silent architectural drift from accumulating into a later user-facing bug. Advisory (non-blocking); fires only on a detected signal. Toggle with workflow.assumption_delta. (#1561) (#1767)
  • /gsd-ui-phase now probes UI state coverage — a new ui-consideration-probe (the third probe-core adapter) enumerates the shape-rooted UI states a UI-SPEC must resolve (empty/loading/error/populated/partial/overflow/zero-one-many/long-text). After the UI checker approves, the probe surfaces applicable considerations for each element, records a ## UI Considerations section in the UI-SPEC, and plan-phase lifts each resolved consideration into must_haves — so a purely-visual state with no wired test routes to insufficient_spec → human_needed at verify rather than a silent pass. (#1979)
  • Host-Integration Interface (ADR-1239 Phase A) — a versioned, negotiated capability contract (runtime.hostIntegration) over the six host-integration points (command, dispatch, model, hooks, state, artifact). Adds an in-process negotiateHostCapabilities handshake that fail-closes on undeclared/unknown/undocumented values (effective ⊆ host-declared ∩ engine-known), a typed degradation ladder, host-capability profiles, and a documentation-sourced per-CLI capability matrix for all 16 runtimes. Interface-definition only — no change to install behaviour. (#1690)
  • ZCode (Z.ai) is now an installable runtime — a desktop Agentic Development Environment for the GLM-5.2 model can now be targeted with --zcode, landing GSD skills at ~/.zcode/skills/<name>/SKILL.md plus slash commands and subagents. ZCode ships as a pure declarative capability descriptor (capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode' branches, reusing the Claude skill converter — the de-hardcoded, data-driven runtime path that 1.7.0 (ADR-1016 / ADR-1239) enables. (#1925) (#2039)

Changed

The GSD CLI now self-heals a missing runtime build. The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts (ADR-457) that ship prebuilt in the npm tarball but are absent on a Claude Code plugin-marketplace / git-clone install, which never runs npm run build:lib. Previously every command died at load with Cannot find module './lib/cli-exit.cjs'. The gsd-tools entrypoint now detects the missing output and compiles it once, on demand (lock-guarded so parallel invocations don't race), then proceeds — a single no-op check on the already-built npm path. When TypeScript is genuinely unavailable it prints an actionable npm install && npm run build:lib message instead of crashing. (#2036)

  • Internal: Claude Code's installer is now driven through the public Host-Integration Interface (ADR-1239 / EoS). bin/install.js routes claude install/uninstall through the imperative adapter (createImperativeAdapter) instead of calling the engine directly, and its 13 hardcoded runtime === 'claude' / runtime !== 'claude' branches are folded into descriptor-driven runtime.hostBehaviors on capabilities/claude/capability.json (permission schema, settings.local.json scope routing, .gsd-source marker, effort frontmatter, canonical-workflow authorship, and more). Install/uninstall output is byte-identical for both the global skills layout and the local legacy layout (golden-parity asserted for both scopes); no other runtime changes. Removes the "add-a-host tax" of scattered string-equality checks for the tier-1 reference host. No user-facing change. (#2086) (#2106)
  • OpenCode is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). OpenCode and its Kilo sibling previously installed via a bespoke runtime === 'opencode'/isOpencode branch in bin/install.js; its commands+skills+plugin install now runs through the imperative adapter → the engine's combined-family install path (installRuntimeArtifacts), and every hardcoded runtime === 'opencode' branch is folded into descriptor-driven runtime.hostBehaviors. Install/uninstall output is byte-identical (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) background dispatch — OpenCode shipped experimental background subagents in v1.15 and made them default-on in v1.17, so dispatch.background/backgroundDispatch flip to true; GSD no longer force-flattens OpenCode-hosted wave dispatch (shouldFlattenDispatch now returns false), letting agents run concurrently where the host supports it. (2) expanded event surface — the OpenCode plugin now subscribes to permission.asked, permission.replied, and session.error (added to EXTENSION_EVENT_SURFACES.opencode), wiring the declared surface for future permission/error-aware bindings. (#2087) (#2108)
  • Codex is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS). Codex previously installed via hardcoded runtime === 'codex'/isCodex projection in bin/install.js; its config.toml / agent-.toml / hooks.json install now runs through the declarative embedding adapter and descriptor-driven runtime.hostBehaviors, with zero positive isCodex gates and zero runtime === 'codex' branches remaining (source-guarded). Install/uninstall output stays byte-parity-gated (tests/fixtures/golden-install-parity/codex.json). Three Context7-verified upgrades land, each with a test driving the user-reachable surface: (1) skill root — GSD skills now install to Codex's canonical $HOME/.agents/skills (via a skills-kind home override) instead of the deprecated $CODEX_HOME/skills fallback, and pre-move installs are migrated (stale ~/.codex/skills/gsd-* cleaned on both install and uninstall, user-owned content preserved); (2) hook events — GSD registers the six documented Codex lifecycle events it previously skipped (PreToolUse, PermissionRequest, PreCompact, PostCompact, SubagentStop, UserPromptSubmit, in addition to the existing SessionStart/SubagentStart/Stop/PostToolUse) in hooks.json, so gsd-context-monitor fires at the same points as in Claude Code, and the descriptor extendedHookEvents is reconciled from [] to the schema-valid wired subset; (3) dispatch tuning — [agents] max_depth = 1 is written explicitly into the managed config.toml block to pin the negotiated dispatch.maxDepth: 1 axis (degradationFor flattens GSD-hosted waves to single-level), and validateCodexConfigSchema now permits a known-scalar-only [agents] AgentsToml table (coexisting with the flattened [agents.gsd-*] role sub-tables) while still rejecting the [[agents]] and unknown-key break-forms from #2760. (#2088) (#2110)
  • Cursor is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). Cursor previously installed via hardcoded runtime === 'cursor'/isCursor branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded cursor branch is folded into descriptor-driven runtime.hostBehaviors (reapplyCommand, frontmatterDialect, hooksJsonSurface, skipSharedHooksInstall, reportCommandsDir, managedHookEvents). Install/uninstall output is byte-identical (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) expanded hook-bus coverage — GSD registers all 6 managed lifecycle events in Cursor's hooks.json (preToolUse, stop, subagentStart, subagentStop in addition to the original sessionStart/postToolUse), driven by a new descriptor-driven adapter module (src/host-integration-adapters/imperative-hook-bus.cts) that reads hostBehaviors.managedHookEvents instead of a hardcoded event pair; cite https://cursor.com/docs/hooks. (2) named/background nested subagent dispatch — Cursor's dispatch.background/backgroundDispatch/nested are all true with maxDepth: 2, so shouldFlattenDispatch(cursor) returns false and GSD's wave-based execution drives Cursor's native background + depth-2 nested subagent invocation instead of flattening to inline sequential calls; cite https://cursor.com/docs/subagents + https://cursor.com/docs/sdk/typescript. (#2089) (#2120)
  • Cline is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS). Cline previously installed via hardcoded runtime === 'cline'/isCline branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded cline branch is folded into descriptor-driven runtime.hostBehaviors (reapplyCommand, frontmatterDialect, skipSharedHooksInstall, localTargetIsProjectRoot, clineRulesSurface, localCommandsViaRules). Install/uninstall output is byte-identical (golden parity asserted for cline + claude/cursor/codex/opencode). Two Context7-verified upgrades land: (1) AgentPlugin.hooks.beforeTool planning guard — the .clinerules/hooks/PreToolUse file-convention hook (#787) is re-implemented as a real Cline SDK AgentPlugin that cancels write-class calls targeting .planning/ (same fail-open semantics), driven by a new descriptor-driven adapter module (src/host-integration-adapters/cline-sdk-binding.cts); cite https://github.com/cline/cline/blob/main/docs/sdk/plugins.mdx. (2) createAgentModel model overrides — DefaultGateway.createAgentModel({providerId, modelId}) is wired so GSD's per-subagent model_overrides/model_profile_overrides resolution applies to Cline subagents (modelMode: active); cite https://github.com/cline/cline/blob/main/docs/sdk/reference/gateway.mdx. Cline's dispatch deliberately stays degraded/flat (maxDepth: 1, read-only, no nested spawning) per the documented host restriction — never silently upgraded to full nested/background. (#2090) (#2132)
  • Hermes Agent is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS). Hermes previously installed via hardcoded runtime === 'hermes'/isHermes branches in bin/install.js; its install/uninstall now runs through the imperative adapter, and every hardcoded hermes branch is folded into descriptor-driven runtime.hostBehaviors. Three upgrades land: (1) real plugin hook vocabulary — GSD registers a new extensionEvents: "hermes" dialect carrying the 13 documented Hermes plugin events (pre_tool_call, post_tool_call, pre_llm_call, post_llm_call, on_session_start, on_session_end, on_session_finalize, on_session_reset, subagent_start, subagent_stop, pre_gateway_dispatch, pre_approval_request, transform_tool_result), replacing the borrowed hookEvents: "claude" 6-event surface that silently never fired; cite https://github.com/nousresearch/hermes-agent/blob/main/website/docs/user-guide/features/hooks.md. (2) dispatch posture — Hermes' dispatch.nested: true with maxDepth: 1 is correctly negotiated (not silently flattened). (3) branding/category metadata — DESCRIPTION.md category descriptions, version: frontmatter, and branding rewrites are now descriptor-driven rather than hardcoded. Install/uninstall output is byte-identical (golden parity asserted for all runtimes). (#2091) (#2134)
  • Qwen Code now projects GSD's specialist agents as native subagents — installing GSD into Qwen Code writes ~/.qwen/agents/gsd-*.md files you can invoke directly (planner, executor, code-reviewer, …) instead of reaching them only through skill prose, and a SubagentStart hook now fires alongside SubagentStop. Qwen's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2092) (#2153)
  • Kilo Code now supports native hooks, active-model routing, and named subagent dispatch — installing GSD into Kilo wires a lifecycle-hook plugin, keeps each agent's requested model instead of dropping it, projects GSD's specialist agents as invokable subagents, and documents the GSD MCP companion. Kilo's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2093) (#2156)
  • GSD skills installed for Trae now carry SOLO stage metadata — Trae's SOLO Agent can recognize GSD skills as workflow-stage skills for auto-invocation instead of requiring manual triggering. Several of Trae's install branches (shared-hooks gating, path rewrites) also move onto its capability descriptor. Note: the stage-metadata field is a best-effort/inferred shape — Trae publishes no formal schema. (#2094) (#2157)
  • Installing GSD into Antigravity now writes the permissions.allow rules its CLI documents — so GSD's own reads and hooks aren't stuck on interactive prompts — and registers GSD's companion MCP server via a standalone mcp_config.json (best-effort: Antigravity's raw config schema isn't published, so this uses the Gemini-CLI-successor format). Antigravity's install is now driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2096) (#2165)

Augment Code now installs through its capability descriptor, with a native MCP companion — installing GSD into Augment registers the GSD companion server in Augment's settings.json mcpServers and drives command/skill/agent conversion from Augment's negotiated descriptor instead of hardcoded runtime special-cases. (#2097) (#2166)

CodeBuddy now wires GSD's full extended lifecycle hook set and is driven by its capability descriptor — installing GSD into CodeBuddy now registers SubagentStart, SubagentStop, Stop, and PreCompact hooks in its settings.json (it previously had none of these), matching the coverage Qwen/Kimi already ship, and CodeBuddy's install is fully descriptor-driven instead of via residual hardcoded runtime branches. (#2098) (#2169)

GitHub Copilot now wires GSD's full lifecycle hook bus and is driven by its capability descriptor — installing GSD into Copilot registers preToolUse, postToolUse, userPromptSubmitted, and sessionEnd handlers in its hooks/gsd-session.json (beyond today's sessionStart-only advisory), and Copilot's residual hardcoded runtime branches are folded onto descriptor-driven hostBehaviors. (#2099) (#2172)

  • Windsurf now enforces GSD's write/command safety guards through Cascade's native hook bus — installing GSD into Windsurf registers blocking pre_write_code/pre_run_command hooks in .windsurf/hooks.json (exit-code-2 blocking) and drives Windsurf's install from its capability descriptor instead of hardcoded runtime branches. (#2100) (#2190)
  • ZCode's install is now driven and regression-tested through its capability descriptor — ZCode joins the dogfooded declarative-adapter reference hosts with a byte-identical install, and its shared-hooks exclusion is folded onto hostBehaviors instead of a hardcoded runtime branch. (Hook-automation and MCP upgrades remain blocked on ZCode publishing its on-disk config formats.) (#2101) (#2195)
  • Codex/OpenAI default models advance to the GPT-5.6 family (Sol/Terra/Luna) — the Codex runtime tier defaults and the openai provider preset now resolve to current-generation model IDs instead of the superseded GPT-5.4/5.5 line, so Codex users on default profiles get improved agentic coding (Sol) and lower costs (Terra/Luna) without changing any config. (#2122) (#2146)
  • Internal: the installer's program (display-name) + command (slash-invocation) chains are now single-source lookups — the 14-line program chain (an exact duplicate of runtimeLabel) → getRuntimeLabel, and the 14-line command chain (the per-runtime /gsd-new-project syntax: gemini /gsd:, codex $, cursor skill-mention, kimi /skill:, default /gsd-new-project) → new getRuntimeNewProjectCommand(runtime) helper (ADR-1239 Phase B / #1679 AC2 slice 4). runtime === count in bin/install.js: 53 → 25 (cumulative this session: 129 → 25). Stdout strings preserved byte-for-byte; no install-output change (golden-parity 16/16). No user-facing change. (#1813)
  • Internal: the installer's per-function is<Runtime> flag-declaration blocks are now a single runtimeFlags lookup — the four duplicated const isX = runtime === 'x' blocks in bin/install.js (uninstall / writeManager / install / a fourth helper — 48 branches) are collapsed into one runtimeFlags(runtime) helper in runtime-name-policy.cts (ADR-1239 Phase B / #1679 AC2 slice 3). The add-a-host tax for flags is removed (one RUNTIME_FLAG_IDS entry, not four declaration blocks). Install output is byte-identical for all 16 runtimes (golden-parity asserted); runtime === count in bin/install.js: 101 → 53. No user-facing change. (#1811)
  • Internal: third-party descriptor loader enforces configHome write-confinement at load time — loadRegistry({includeInstalled:true, configHome}) now rejects (skip + warn, fail-closed) any installed third-party host-plugin descriptor whose declared destSubpath resolves outside the supplied configHome, before it is composed into the registry (ADR-1239 Phase C-2 / #1681 slice 2). The configHome option is optional and backward-compatible (omitted → no load-time check; install-time gate still bounds writes). No user-facing change for existing flows. (#1808)
  • Internal: agent install for cursor/windsurf/augment/trae/codebuddy now flows through the descriptor path — ADR-1235 step 1 routes the trivial-converter runtime group's agents off the inline install() loop onto the descriptor-driven installRuntimeArtifacts path, applying the cross-cutting steps uniformly (pre-converter, no workflow-stamp). Agent output is byte-identical for all 16 runtimes (golden-parity asserted, global + local verified); no user-facing change. (#1764)
  • gsd-ui-checker gains an adversarial FORCE stance (LLM-playbook principle 16) — the only verdict-producing critic that lacked one now resists rubber-stamping UI-SPEC contracts, with BLOCK/FLAG/PASS classification. Based on arXiv 2505.23840 (third-person objective persona), 2506.04975 (objective-not-hostile persona). (#1584)
  • Internal: the declarative embedding adapter is now named + bound behind a minimal HostIntegrationInterface — createDeclarativeAdapter({runtime}) (new src/adapter-declarative.cts) delegates in-process to install-engine's installRuntimeArtifacts/uninstallRuntimeArtifacts, formalizing today's projection path as one of the two embedding adapters behind a common contract (ADR-1239 Phase C-1 / #1680 AC1). Output is byte-identical to today's install (gated by golden-install-parity). The full 6-point interface binding surface is deferred until the imperative adapter (AC2) fixes the shape (ADR-1239 open wire-shape question). No user-facing change — the adapter is not yet wired to any runtime path. (#1802)
  • Internal: getDirName is now derived from a documented runtime.localConfigDir descriptor field — each runtime's local content-rewrite directory (e.g. cursor→.cursor, copilot→.github) moved from a hand-maintained if-chain into its capability descriptor (ADR-1239 Phase B), so it can no longer drift from the registry. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1757)
  • Internal: copyWithPathReplacement converter selection is now data-driven — the installer's back-compat content-copy path replaced its 13 hardcoded runtime === 'x' flag chains with a single per-runtime dispatch table (ADR-1239 Phase B). Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1759)
  • Phase-completion now writes Status: All phases complete instead of the overloaded bare Milestone complete — the phase-level completion verb (completePhaseCore) was writing the same bare 'Milestone complete' string that the milestone-close verb uses for terminal state, causing a phase-level verb to own a milestone-level field. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts); milestone termination (' milestone complete' / 'Awaiting next milestone') remains solely with the milestone-close verb. (#2204) (#2259)
  • #853 dispatch-flatten is now data-driven (ADR-1239 Phase B) — whether GSD backgrounds the plan/execute orchestrator is decided from a documentation-sourced backgroundDispatch capability per host (via gsd_run query dispatch-should-flatten) instead of a hardcoded runtime === 'codex' check. Cursor now backgrounds the orchestrator (its docs document backgrounded subagent nesting); codex unchanged; all other hosts run inline. Fail-closed to inline on any uncertainty. (#1719)
  • Internal: companion MCP server module (interface points 1 + 5) — handleMessage/runServer (new src/mcp-server.cts) is a minimal, dependency-free stdio JSON-RPC 2.0 server exposing gsd_invoke_command (→ the command-routing hub) + gsd_read_state/gsd_write_state (→ the Phase 3 stateIO seam), so any MCP-consuming host can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681 slice 3a). Bin entry / packaging deferred to slice 3b. No user-facing change — the server is not yet wired to a bin entry. (#1809)
  • requirements mark-complete reports a per-surface write-set — the command now returns a per-requirement write_set (checkbox + traceability surfaces) and a write_set_complete that is true only when every surface of every requirement applied, so a partial (checkbox-only) reconcile can no longer masquerade as full success even inside a multi-ID batch. Introduces the reusable ADR-2143 §5/§6 Result / WriteSet contract. (#2251) (#2251)
  • Internal: the imperative embedding adapter now composes the capability registry behind the same HostIntegrationInterface — createImperativeAdapter({runtime}) (new src/adapter-imperative.cts) calls loadRegistry({includeInstalled:true}) (first-party-wins + consent + fail-closed — identical trust semantics to the CLI) and binds the engine surface behind the same contract the declarative adapter (AC1) satisfies, plus a registry accessor for an in-process host to bind its primitives to (ADR-1239 Phase C-1 / #1680 AC2). Concrete host binding is deferred to Phase 5. No user-facing change — the adapter is not yet wired to any runtime path. (#1803)
  • Internal: the model adapter seam exposes passive + active adapters selected by modelMode — createModelAdapter({modelMode}) (new src/model-adapter.cts): passive formalizes today's tier routing (delegates to model-resolver.resolveModelForTier), active is a host-supplied sendRequest seam (VS Code vscode.lm / pi providers), fail-closed until Phase 5 binds a concrete provider (ADR-1239 Phase C-1 / #1680 AC3). No user-facing change — the seam is not yet wired to any runtime path. (#1804)
  • Internal: derive the non-Claude runtime list from the capability registry — NON_CLAUDE_RUNTIMES is now computed from the capability registry instead of a hand-maintained literal, so it can no longer drift from the per-runtime descriptors. No user-visible behavior change (the list is identical). (#1728)
  • Honest verifier — verify-phase now abstains on non-inferable backstop truths instead of confidently false-passing them (#1154). When the spec's edge-probe marks a truth non-inferable (verification: backstop) and the verifier cannot confirm it with explicit evidence (a passing wired held-out/property test, or a directly-observed behavior), it now reports human_needed with reason insufficient_spec ("unverified — held-out test recommended") rather than a silent passed. Autonomous runs complete with "N unverified non-inferable checks"; interactive runs route to the end-of-phase human checkpoint. Inferable truths are never abstained (over-abstention guard); abstention is exogenous (driven by the tag, not self-judgment). Truth-axis mirror of the prohibition judgment-tier (ADR-550 D4). (#1738)
  • Document Claude Code's advisor-tool inheritance in the model-profiles reference: the session-level advisor is inherited by all GSD subagents and composes with per-agent tiering, with candidate executor/advisor pairings, when it is worth enabling, and the session-level (no per-agent control) constraint. (#1922)
  • Extraction discipline for strict-format agents (LLM-playbook principle 8) — gsd-doc-classifier and gsd-doc-synthesizer apply taxonomy/precedence rules directly without inventing content, reducing reasoning-induced format drift. Based on arXiv 2504.05081 (few-shot beats CoT for pattern tasks), 2506.00069 (terminal instruction placement), 2505.14810, 2505.11423. (#1584)
  • Internal: extracted the runtime-artifact install engine from bin/install.js — installRuntimeArtifacts/uninstallRuntimeArtifacts/installOpencodeFamilySkills and their helpers now live in a dedicated gsd-core/bin/lib/install-engine.cjs module (ADR-1239 Phase B), so adapters can import the install pipeline instead of reaching into the 12k-line installer. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing behaviour change. (#1735)
  • MemPalace memory_mode kg_backend and replace are now functional — selecting either mode now routes recall through the palace instead of silently behaving like augment: kg_backend treats the palace temporal KG as the primary knowledge-graph source (native .planning/graphs/ as fallback), and replace resolves recall through the palace as the source of truth. Every mode stays default-resilient — an unreachable palace falls back to native memory and no memory is lost. (#2010) (#2010)
  • /gsd:surface and --materialize now produce byte-identical agent output to a fresh install — surface-path agents for descriptor-driven runtimes (cursor, windsurf, augment, trae, codebuddy, copilot, antigravity) now receive the same path-prefix rewrite, Co-Authored-By attribution, runtime-specific conversion, and body normalization as the install path. Copilot and Antigravity agents are now installed via the descriptor-driven path (copilot agents get the .agent.md filename rename). Cline remains on the inline loop (rules-only local branch). (#1575) (#2040)
  • Internal: hook-bus + stateIO adapter seams — createHookBus({bus}) (new src/hook-bus.cts, host/engine/none — engine is in-process pub/sub, host fail-closed, none silent) + createStateIO({io}) (new src/state-io.cts, filesystem/sandboxed-storage/session-log-append — filesystem delegates to fs, the rest are fail-closed seams) (ADR-1239 Phase C-1 / #1680 AC4). Completes the Phase 3 adapter seam layer; concrete host binding is Phase 5. No user-facing change. (#1805)
  • Long-context model names render compactly in the statusline — the verbose " (1M context)" suffix Claude Code appends to the model display name now collapses to a compact " (1M)" badge (tolerant of future window sizes and the abbreviated "ctx" variant: "(500K context)" → "(500K)", "(1M ctx)" → "(1M)"). Lossless — the long-context signal stays, the 12 characters of width don't. (#2160) (#2173)
  • Lazy-split plan-phase.md into a steps/ directory — ~4.7 KB lighter eager context per /gsd-plan-phase call via byte-invariant progressive disclosure (ADR-1610). (#1852) (#1934)
  • GSD subagents now self-load configured agent_skills regardless of orchestrator bash — projects that map skills via .planning/config.json agent_skills.<agent-type> no longer silently lose them on /gsd-autonomous or Cursor, where Skill()-delegated workflow bash init did not reliably run. Each of the 22 consumer agents queries its own type at init and reads the listed skills, with a dedup guard so runtimes that also inject orchestrator-side (Claude Code) never carry two copies. (#1866) (#1868)
  • Internal: install/uninstall runtime labels are now sourced from a single getRuntimeLabel lookup — the two duplicated runtimeLabel assignment chains in bin/install.js (uninstall + install) are collapsed into one curated label table in runtime-name-policy.cts, sibling to the registry-derived getDirName (ADR-1239 Phase B, #1679). Install output is byte-identical for all 16 runtimes (golden-parity asserted). Two console-label inconsistencies are normalized as a side effect: kimi shows 'Kimi CLI' in both sites, and cline uninstall no longer falls through to 'Claude Code'. (#1800)
  • Internal: external-descriptor trust gate — load-time configHome confinement — assertDescriptorConfined(descriptor, configHome) (new src/external-descriptor-trust.cts) fail-closed rejects any installed third-party host-plugin descriptor whose declared destSubpath resolves outside the user-approved configHome, before its install plan runs (ADR-1239 Phase C-2 / #1681 slice 1). Defense-in-depth load-time twin of Phase 2's install-time assertDestWithinConfigHome. Not yet wired into the loader (slice 2). No user-facing change. (#1806)
  • Internal: the installer's runtime → global-config-home hook-pathogen fragment is now a single getGlobalConfigHomeFragment lookup — the 14-branch if (runtime === 'x') return "'...'" chain in getConfigDirFromHome (bin/install.js, the hook path.join() codegen mapping) is collapsed into one table in runtime-name-policy.cts, sibling to getRuntimeLabel (ADR-1239 Phase B, #1679 AC2 slice 2). Generated hook output is byte-identical for all 16 runtimes (golden-parity asserted); antigravity's dynamic env-overridable resolution is preserved in the caller. No user-facing change. (#1801)

Removed

  • Removed the sunset Gemini CLI runtime — use Antigravity CLI instead — Google discontinued Gemini CLI on 2026-06-18, so npx gsd-core --gemini now prints a deprecation notice and points you to Antigravity CLI (the official successor), which GSD already ships as a first-class runtime. (#1928) (#1996)

Fixed

  • The verify-work security-blocked presentation no longer offers next-phase planning. When security enforcement blocks phase advancement (no SECURITY.md produced), the workflow now routes only to the current-phase fix instead of competing /gsd:plan-phase {next} and /gsd:execute-phase {next} options. (#1687)
  • milestone complete and roadmap analyze now exclude the Phase 0 / Phase 999 backlog sentinels. A milestone whose only directory-less ROADMAP heading is a backlog sentinel can be completed without --force, and roadmap analyze no longer counts the sentinel in phase_count or routes next_phase into it. Completes the ^999 exclusion #1445 added to the progress denominators. (#1691)
  • config-set no longer silently coerces values into something the disk never sees — Number.isFinite replaced !isNaN in the value parser so Infinity/-Infinity are no longer coerced to non-finite numbers that JSON.stringify then renders as null on disk while the CLI echoes Infinity (output ≠ disk). context_window now has a per-key validator requiring a finite positive integer (rejects Infinity, 0, negatives, non-integers with a non-zero exit), and project_code is always persisted as a string so a leading-zero code like 007 survives verbatim instead of collapsing to 7. Numeric coercion for genuine numeric keys (e.g. granularity 42) is unchanged. (#1581) (#2023)
  • phase.complete no longer reports a false is_last_phase on a <details>-wrapped checkbox checklist (#1591, #1752) — when the active milestone's phase checklist was written as - [ ] Phase N: checkbox items inside a <details> block and the next phase had no directory on disk yet (still in planning), phase.complete's isLastPhase roadmap-enumeration fallback used a heading-only pattern (/#{2,4}\s*Phase…/) that never matched checkbox items. It returned is_last_phase: true, next_phase: null on a mid-milestone phase and — via the milestone-complete cascade — wrongly flipped STATE.md to Milestone complete and decremented progress.total_phases (e.g. 8 → 7). The pattern now matches heading-style (### Phase N:), plain checkbox-list phases (- [ ] Phase N: / - [x] Phase N:), and the canonical bold checklist form the roadmap template emits (- [ ] **Phase N: Name**); extractCurrentMilestone already surfaces the <details>-wrapped checklist correctly, so no parser change was needed. Only the reproduced phase.complete fallback is changed; the heading-only sibling patterns elsewhere in phase.cts are untouched. (#1819)
  • The <agent_skills> block emitted by gsd init no longer leaks backslash paths into @-reference skill paths on Windows. The global skill directory (a native path.join result) was interpolated into the generated markdown without POSIX normalization, producing references like @C:\…\skills\name/SKILL.md; the reference is now normalized at the emit site so skill references use forward slashes on every platform. (#1736)
  • /gsd-settings no longer warns about four search-provider keys on fresh projects (#1747) — buildNewProjectConfig emits seven search-provider availability flags and research-provider.cts providerAvailability() consumes all seven, but only three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json). Running /gsd-settings on a freshly generated .planning/config.json printed unknown config key(s) … tavily_search, ref_search, perplexity, jina — these will be ignored even though the user never hand-edited the config. The four missing keys are now registered alongside brave_search/firecrawl/exa_search and documented in docs/CONFIGURATION.md; a drift guard in tests/bug-2530-valid-config-keys.test.cjs now requires every config-driven research-provider flag to be in the schema, so a future provider addition cannot reintroduce the drift. (#1814)
  • gsd-tools state json no longer reports conflated progress for an unversioned milestone (#1761) — the ADR-1769 Phase 7 fix (#1794) taught state sync to leave Progress untouched when a milestone version is asserted but the ROADMAP has no versioned heading for it, but the state json read path still rebuilt progress via buildStateFrontmatter, whose phase-heading count fell back to the whole document and summed sibling milestones. state json therefore reported a conflated total_phases (e.g. 8 = 4+4 across two milestones) plus a derived percent, contradicting the sync guard on the very same project. The read path now mirrors the sync guard: when the asserted milestone cannot be bounded to a versioned ROADMAP heading, total_phases falls back to the on-disk phase-dir count and percent is omitted. Bounded milestones (versioned ROADMAP, or no milestone asserted) are unchanged; the signal rides on the existing _diskScanCache so extractCurrentMilestone's return contract and its other callers are untouched. (#1818)
  • gsd-graphify-update.sh now reads the full multi-line command in Gate 2 (#1772) — the PostToolUse auto-update hook joined tool_name + \n + tool_input.command and extracted the command with sed -n '2p' (line 2 only). Agent runtimes (Claude Code's Bash tool among them) routinely emit HEAD-advancing commits as multi-line scripts (cd /path, then git add, then git commit …), so line 2 was the cd, Gate 2's *"git commit"* match failed, and the rebuild silently no-op'd on real commits even with graphify.auto_update: true. The failure was invisible in manual probes because a single-line git commit -m x passes line 2 verbatim. The hook now captures line 2 through EOF (sed -n '2,$p') so the case glob sees the full command string; single-line behavior is unchanged and multi-line commands without a HEAD-advancing op still no-op cleanly. (#1815)
  • /gsd-thread close|resume now writes the thread status/updated frontmatter (#1778) — the thread workflow's CLOSE and RESUME branches invoked frontmatter.set with the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>), but since 1.6 the dispatcher parses the file positionally and reads field/value from the named flags --field/--value via parseNamedArgs. The positional form left field/value undefined, cmdFrontmatterSet errored file, field, and value required, and the writes were silently skipped — so closing a thread never marked it status: resolved and resuming never marked it status: in_progress, with the error scrolling past on every thread command. All four sites (CLOSE status+updated, RESUME status+updated) now use the 1.6 hybrid form that verify-work.md already uses (frontmatter.set <file> --field <field> --value <value>). (#1816)
  • The installer no longer copies dead lifecycle hook scripts for Kilo and ZCode — both declare hooksSurface: 'none' and have no plugin surface, so the staged hooks/*.js, hooks/*.sh, hooks/lib/ and the CommonJS package.json marker were dead weight in ~/.kilo/ and ~/.zcode/. The two hook-copy guards in install.js now exclude Kilo and ZCode alongside the other no-hook runtimes. OpenCode, which also declares hooksSurface: 'none', is deliberately kept: its native plugin adapter (#1914) spawns those staged hooks via OpenCode's event bus and needs both them and the marker. (#2057)
  • Test gates can no longer hang forever on a watch-mode test runner. vitest defaults to watch mode in an interactive terminal (exactly where gsd-execute-phase runs), so a resolved npm test / pnpm test that maps to vitest never exited and the orchestrator waited indefinitely until the user manually intervened. Every GSD test-command gate — the regression gate, the post-merge gate, the audit-fix gate, and the verify-phase gate — now routes the resolved command through a shared normalize-test-command helper that rewrites it to a one-shot form (direct vitest → vitest run; jest --watch → --watchAll=false; a package-manager test script backed by watch-vitest → CI=true prefix; already-one-shot commands are left unchanged). The three gates that previously hung or silently continued — the regression, post-merge, and audit-fix gates — additionally bound execution with a configurable workflow.test_gate_timeout (default 600s), aborting or surfacing the cause on timeout instead of hanging; the verify-phase gate was already bounded (a fixed 5-minute limit) and keeps it, now naming watch mode on timeout. The normalizer only rewrites a runner named as a standalone command token (so paths/targets like run-vitest.js are never mangled), is length-capped and linear-time on adversarial input, and only reads a regular-file package.json. (#2060)
  • settings-advanced.md no longer has an orphan </step> around §8 Model Policy — the §8 Model Policy block ended with a closing </step> but had no matching opening tag (5 opens / 6 closes), leaving its content as loose inter-step prose that could fail to execute reliably. Added the missing <step name="model_policy"> opener so the section is a proper step. A new workflow <step>-tag-balance regression guard (fenced-code-stripped) now blocks any future orphan tag across all top-level workflows. (#1864) (#2014)
  • The runtime launcher now honors CLAUDE_CONFIG_DIR — the gsd_run preamble embedded in every workflow/agent resolved the Claude global install only at $HOME/.claude/gsd-core/bin/, while the installer honored CLAUDE_CONFIG_DIR, so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every gsd_run call (every GSD command failed with gsd-tools.cjs not found). The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude} — matching the installer and the other runtimes' ${VAR:-default} pattern — so a custom CLAUDE_CONFIG_DIR is found and the default $HOME/.claude path is unchanged. Re-synced into all 95 workflows/agents; two capped workflows trimmed to stay under their byte budgets. (#1865) (#2024)
  • Node-test prohibition proofs now require a clean-fixture causation control — a node-test prohibition's fail-first proof no longer accepts a deceptive content-independent negative test (one that reds merely because GSD_PROHIB_SUBJECT is set, ignoring the subject's content). The check_clean_fixture control is now mandatory for the node-test kind: a descriptor that omits it is un-provable and hard-gates, rather than greening on the violation alone. Breaking (Hyrum): a previously-green node-test prohibition with no clean fixture now hard-gates — blast radius is zero in-tree (no node-test prohibition ships today). The lint-rule kind is unchanged (its subject IS the linted file, no GSD_PROHIB_SUBJECT indirection). (#1906) (#2001)
  • Third-party capabilities now work on installed layouts. capability install no longer rejects capabilities with a real engines.gsd range as "incompatible with GSD 0.0.0" — the host version is now read from the authoritative gsd-core/VERSION file across every runtime and the capability install CLI. The installer also now ships the registry generator scripts (gen-capability-registry.cjs, gen-loop-host-contract.cjs), so installed third-party capabilities actually compose into the loop instead of being silently discarded. (#1938)
  • /gsd:verify-work preserves verification state across gap-closure execution and no longer auto-promotes deferred follow-ups into blocking gaps — resuming after /gsd:execute-phase --gaps-only used to lose the verification state: the UAT ## Gaps still read status: failed even after their fix plans executed, so verify-work re-diagnosed them as fresh blockers, spawned a new gap plan, and reported only the new plan as verified. A state contract now links each gap to its fix plan: every UAT gap carries a stable gap_id (G-{phase}-{N}), gap-closure plans tag the ids they address in their frontmatter (gap_ids: […]), and a new reconcile_gaps step on resume marks a gap status: resolved when its plan has a matching *-SUMMARY.md — so fixed gaps aren't re-diagnosed and the phase can close. Separately, a deferred-follow-up branch captures future-work ideas (signals like "later", "next version", "out of scope") into a ## Deferred Follow-Ups section instead of creating a blocking gap/plan. (#1921) (#2025)
  • roadmap update-plan-progress no longer counts stray non-plan *-SUMMARY.md files against phase completion — remediation/gap-closure summaries (e.g. 30-FIX-CR02-SUMMARY.md, 30-GAPCLOSURE-SUMMARY.md) inflated summary_count, and once summary_count >= plan_count the phase silently flipped to Complete (checkbox checked, date stamped) even though several plans had no summary. A new countMatchedSummaries helper (core-utils) pairs summaries to plans via the PLAN→SUMMARY marker swap + the <stem>-SUMMARY.md form (layout-agnostic across root, bare, and nested layouts), so only a summary that corresponds to a real plan counts. Wired into scanPhasePlans (fixing roadmap listing, state sync, verification, workstream inventory at once) and cmdRoadmapUpdatePlanProgress. (#1988) (#2016)
  • milestone complete --ws requirements archive header now points at the workstream REQUIREMENTS.md — the archive header string hardcoded the root path (`…see .planning/REQUIREMENTS.md`), so a workstream archive directed readers at the wrong file even though #1917 had already fixed the archive locations to land inside the workstream. The display path is now derived from the same workstream-aware reqPath the writer uses (path.relative(cwd, reqPath)), so root behavior is byte-identical and the workstream case correctly reads .planning/workstreams/<ws>/REQUIREMENTS.md. (#1993) (#2015)
  • Load-failed capability gates now fail open with a loud warning instead of blocking the whole project — when an installed overlay (third-party) capability failed to load (e.g. an incompatible engines.gsd range) but had declared a gate-kind loop hook, the loop resolver injected a blocking synthetic gate (blocking:true, onError:halt) at every point where that capability declared a gate. A single incompatible capability therefore halted every ship:pre and verify:post in the project — unrelated to what the gate would have checked, and with no remediation surfaced. The resolver now injects no gate and instead emits a loud warning — to stderr and in the loop render-hooks envelope's warnings array — naming the load-failure reason and the exact gsd capability remove <id> remediation, and the loop proceeds (fail open). The capability id embedded in that remediation is validated against the canonical id shape first, so a malformed overlay directory name cannot inject shell metacharacters into the surfaced command. The loader still records _overlay.blockedGates; only the consequence changes from block to warn. step/contribution overlays were already skip-open. (#2009) (#2075)
  • phase.complete now updates the ## Progress rollup row even when an earlier phase-numbered table precedes it — the Progress-row writer used a non-global regex that matched any table row starting with the phase number, so it bound to the first such row (e.g. a | Phase | Requirements | Count | coverage table), no-op'd on the wrong 3-column row, and never reached the real Progress row. The regex is now scoped to the ## Progress section so it binds to the correct table. The command still returned roadmap_updated: true (that field is fs.existsSync(ROADMAP.md)), masking the silent failure. (#2012) (#2032)
  • context7 now works for plugin-marketplace installs (8 agents regained doc lookup) — the agents granted only mcp__context7__*, which matches a standalone context7 MCP server but not the official Claude Code plugin-marketplace install (context7@claude-plugins-official), whose tools are named mcp__plugin_context7_context7__*. The grant never matched, so advisor/ai/domain/phase/project/ui-researcher + planner + executor silently lost documentation lookup and fell back to WebSearch. All 8 agents now grant both forms, the researcher profile table is updated, and a parity guard asserts no agent grants the standalone form without the plugin form. (#2017) (#2029)
  • applySurface no longer deletes every gsd-* agent when the skills manifest resolves empty — the agent-prune loop in _syncGsdDir deleted any gsd-*.md not in the staged set, and when the manifest was empty/unresolvable (null manifest, no array entries, no files key, or an unresolvable install source root), the staged set was empty → every agent was pruned. Skills were guarded by pruneSkillDirs's manifest-membership check (conservative preservation on empty manifest); agents had no equivalent. The agent-prune loop is now skipped when the manifest is empty/absent, so agents are preserved while copy (adding genuinely new agents) still runs. (#2018) (#2031)
  • planning-config.md global-learnings path corrected to ~/.gsd/knowledge/ — the features.global_learnings row directed users to ~/.gsd/learnings/, but the implementation (src/learnings.cts, execute-phase.md) stores and reads global learnings from ~/.gsd/knowledge/. Anyone following the docs to inspect, back up, or seed their global learnings looked in a directory the code never touches. (#2019) (#2026)
  • Removed dead SDK file references from runtime-loaded markdown that triggered an infinite find.exe storm on Windows — agents/gsd-executor.md pointed at sdk/src/query/QUERY-HANDLERS.md and gsd-core/workflows/reapply-patches.md at sdk/dist/cli.js, both retired with the SDK package (ADR-0174). AI runtimes that resolve doc references by filesystem search ran find / -iname …; on Git Bash for Windows / maps to the drive root, so find.exe traversed the whole disk (14h+, orphaned processes, 4M+ open handles each, unkillable). The references now resolve to live paths, and a new regression guard asserts no sdk/src|sdk/dist|sdk/handlers file references remain in agents/workflows/references markdown. (#2020) (#2027)
  • roadmap update-plan-progress no longer checks the phase checkbox without verification — the command stamped the phase-level ROADMAP checkbox and completion date the moment the last plan summary landed (called routinely after every wave and every plan), with no verification gate — unlike phase.complete which correctly requires readVerificationStatus(...).status === 'passed'. Now isComplete requires both all plan summaries AND a passed verification, matching the cmdPhaseComplete contract, so the checkbox only fires after gsd-verifier has confirmed the phase. (#2022) (#2030)
  • phase complete no longer marks a milestone done out of order, nor silently writes root state in workstream mode. Completing the numerically-highest phase while an earlier phase was still outstanding wrongly flipped STATE.md to Status: Milestone complete (the milestone-end check only looked for higher-numbered phases, so an out-of-order completion — e.g. Phase 10 before Phase 9 — read as the end). It now reports milestone-end only when every lower-numbered phase in the milestone is checked complete. Separately, in workstream mode with no active workstream, phase complete previously fell back to root .planning and wrote STATE.md/ROADMAP.md (and the mislabel) into the shared root other workstreams read; it now fails safe — asking for --ws <name> or an active workstream — mirroring the existing init progress guard. (#2066) (#2066)
  • Phase directories whose slug begins with a single digit now resolve correctly. A phase like 46-6-rs-pipeline-orchestrator (roadmap name "6 Rs Pipeline Orchestrator") had its phase token over-collected as 46-6 instead of 46, so gsd-tools phase-by-number lookups resolved phase_dir=null / has_context=false (breaking init.plan-phase, init.phase-op, and downstream execute/verify/ship). Numeric phase-token components must now be zero-padded (≥2 digits), so a single-digit slug word is no longer absorbed into the token. Fixed consistently across every same-class implementation — extractPhaseToken, PHASE_TOKEN_FROM_DIR_RE and canonicalPlanStem (health checks / plan pairing), isDirInMilestone's numeric matcher (milestone filtering), and extractCanonicalPlanId — so the health-check and milestone-filter subsystems are fixed alongside phase resolution. (#2059)
  • gsd-tools config-set <key> null now clears (removes) the key instead of persisting the literal string "null". The documented "Clear" action previously fell through the value parser and stored "null" — a truthy value — so "cleared" keys stayed set and config-get returned "null"; for secret keys (brave_search/firecrawl/exa_search) a masked success line hid a truthy value on disk that integrations could pass along as a real credential. config-set <key> null now deletes the key (short-circuiting the typed per-key validators so clearing an enum/boolean/number key removes it rather than being rejected), making the "Clear" flows in settings-integrations.md / settings-advanced.md actually clear. (#2058)
  • init plan-phase no longer collapses foreign-prefixed task/workstream IDs into numeric phases — a query like MEM-01 (where MEM is not the configured project_code) used to have its prefix stripped and resolve to the unrelated numeric Phase 01; it now reports phase_found: false unless a phase directory or roadmap entry literally carries that prefix. The configured project_code's own prefixed phases (e.g. LKML-01 under project_code: LKML) continue to resolve as before. (#2056) (#2105)
  • phase complete no longer ticks the wrong phase's ROADMAP checkbox — completing a phase whose number also appears in a later phase's description (e.g. an idempotent re-run of an already-complete phase) used to mark the wrong phase done, because the checkbox-matching regex greedily spanned from ] to any later "Phase N" mention instead of only the immediately-following phase title. (#2067) (#2079)
  • gsd-tools effort sync no longer crashes in an installed runtime. In any global install (e.g. ~/.claude/gsd-core/), effort sync threw Cannot find module '../../../bin/install.js' — the command reached into the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home, so that file is never present there. As a result, effort config changes (routing_tier_defaults / agent_overrides) silently never reached installed agents without a full reinstall. The two resolvers (readGsdEffectiveEffortConfig + resolveInstallTimeEffort, with their helpers) are now extracted into a shipped gsd-core/bin/lib/install-effort-resolver.cjs that both effort sync and the installer import — a single source of truth that is always present in the installed tree. (#2076) (#2076)
  • model_overrides and per-phase-type models now actually apply to the assumptions-analyzer, code-reviewer, and code-fixer agents on Claude Code. Previously model_overrides["gsd-code-reviewer"] / ["gsd-assumptions-analyzer"] / ["gsd-code-fixer"] (and models.verification / models.discuss / models.execution) were accepted and resolved but silently dropped — the workflows spawned these agents with no model, so they inherited the session model and the configured routing never took effect (no warning). Every spawn now threads its resolved model: discuss-phase-assumptions, code-review, and code-review-fix (both the re-review and the two fixer spawns) resolve it inline, and quick's review step uses the code-reviewer's own resolved model instead of the executor's. The stale "discuss — reserved, no subagent" model-profile docs are corrected to list gsd-assumptions-analyzer, and the verification row now includes gsd-code-reviewer. (#2074) (#2074)
  • /gsd-review's Antigravity CLI reviewer no longer fails silently on large prompts, unavailable pinned models, or pre-session stalls — the agy invocation now uses a file-reference prompt to avoid exec arg-list overflow, is wrapped in an external wall-clock timeout paired with --print-timeout because --print-timeout cannot fire before agy creates a session, passes --model from review.models.agy when set as an escape hatch for a 404'd pinned model, and its empty-output stub now surfaces an agy cli.log diagnostic instead of a bare generic message. Supersedes the #687 "no external killer / inline $(cat)" contract, which predated agy gaining --model and predated its own guidance to pair --print-timeout with a terminal timeout. (#2073) (#2109)
  • init execute-phase, init verify-work, and init phase-op no longer collapse foreign-prefixed task IDs to numeric phases — MEM-01 under project_code: LKML was silently stripped to 01 and resolved to the unrelated numeric Phase 01, because the #2056 guard was applied only to init plan-phase. The guard is now extracted into shared helpers (guardedFindPhase / guardedGetRoadmapPhase) that delegate to the canonical isForeignPrefixedPhaseQuery from phase-id.cts, and all four init commands route through them. (#2104) (#2149)
  • commit --files now commits only the declared paths — gsd-tools commit --files A B previously ran a bare git commit that absorbed the entire staged index, silently sweeping in unrelated files the caller never named. The commit now appends a pathspec (-- <paths>) so only the staged subset of --files lands in the commit; the no---files default path is unchanged. Missing tracked files are still skipped (not committed as deletions, #2014), and when all declared files are missing the function short-circuits to nothing_to_commit instead of absorbing the index. (#2112) (#2148)
  • Fixed unresolvable bare require('gsd-core/...') in gsd-surface command doc — the four require() examples now derive the engine path from runtimeConfigDir (resolvable at runtime), and the reinstall hint corrects npm i -g gsd-core to npm i -g @opengsd/gsd-core. (#2116) (#2213)
  • milestone complete --dry-run now prints a preview plan instead of silently mutating — gsd-tools milestone complete --dry-run was neither parsed nor rejected, so a caller expecting a preview triggered the full destructive mutation (archive phases, move audit artifacts, rewrite STATE.md) with no way to back out. The --dry-run flag is now honored: it returns a JSON plan listing would_archive (roadmap, requirements, audit, phase dirs) and would_update (MILESTONES.md, STATE.md) targets with zero filesystem mutations. (#2118) (#2155)
  • /gsd-secure-phase now has a single SECURITY.md writer — the gsd-security-auditor subagent previously held Write/Edit tools and was instructed to "write SECURITY.md" with no padded <N>- prefix and no template frontmatter, while the orchestrator's Step 6 also wrote the phase-scoped <N>-SECURITY.md from templates/SECURITY.md. The auditor is now return-only (drops Write/Edit, returns a structured verdict with threats_open); the orchestrator is the sole file writer. The workflow's Step 5 spawn constraints explicitly forbid the auditor from writing SECURITY.md. (#2119) (#2154)
  • Dead security scan exports removed; injection-scan docs corrected to match reality — scanEntropyAnomalies and shannonEntropy were dead code with zero production callers (live hooks inline their own patterns for independence). REQ-SCAN-INJ-02/-03 now accurately describe what runs live (injection patterns, invisible Unicode) vs CI-only (base64-decode, codebase scan). (#2198) (#2211)
  • Custom STATE.md frontmatter keys are no longer dropped on every mutating verb — syncStateFrontmatter rebuilt the frontmatter from a fixed schema, silently dropping any custom key. It now carries forward existing keys the schema does not own. (#2202) (#2233)
  • Non-frontend phases with UI hint: no are no longer blocked by the UI-SPEC gate — the UI safety gate's token list included the bare token UI, which matched GSD's own **UI hint**: no metadata line and false-detected a UI, blocking backend/infra phases at /gsd-plan-phase. An explicit UI hint: yes|no is now authoritative and the hint line is no longer token-sniffed. (#2150) (#2222)
  • OpenCode reviewer no longer silently yields an empty review on large prompts — /gsd-review --opencode now invokes opencode run --format json and reconstructs the review from the assistant text parts, so a large-prompt run where the default build agent ends its turn with zero output tokens no longer produces an empty stub. When the agent genuinely emits no text, the stub now reports the stop reason, output-token count, and captured stderr instead of a generic message. (#1936) (#1992)
  • stale-bake-guard hermeticity fix (test-isolation) — the readGsdEffectiveModelOverrides subtest no longer reads the developer's real ~/.gsd/defaults.json; the resolver now accepts a homedir seam so the test sandboxes HOME. (#2152) (#2223)
  • /gsd-surface (list/status) works on Claude Code global installs — the installer now writes a .gsd-source marker pointing at its commands/gsd source, so findInstallSourceRoot resolves on the global skills layout (which ships no commands/gsd tree) instead of throwing could not locate commands/gsd. (#1487) (#1487)
  • phase complete --phase N now works alongside the positional form — the phase verb family treated the first positional as the phase number, so --phase 12 was passed as the literal phase name and failed with 'Phase --phase not found'. The phase family now accepts the --phase flag consistently with the state family, and unrecognized flags yield a usage error. (#2201) (#2231)
  • Third-party capability skills now surface correctly after install — a skills-only role: feature capability installed active but its skills never reached the runtime surface, capability enable/set rejected it as unknown capability, and capability list disagreed with capability state. resolveSurface now unions the composed registry's capabilityClusters into the surfaced skill set (no on-disk linking), the writer validates against the composed overlay-aware registry, and capability list carries a surfaced field matching capability state. (#2054)
  • Fixed: probe-core's runProbeCli now fails closed on per-item adapter garbage inside a well-shaped report envelope, matching its documented 'fails closed on adapter garbage' contract. (#1910)
  • /gsd:verify-work no longer silently terminates when all remaining UAT tests are blocked — sessions with blocked_count > 0 and pending_count == 0 now route to complete_session as expected, enabling the zero-issues auto-transition path. (#1722)
  • state record-metric no longer appends per-plan rows into the By-Phase velocity table — it now maintains its own Per-Plan Metrics table (self-created on first use), and its auto-create scaffold header is corrected. (#2253) (#2253)
  • Codex reviewer now captures the review via codex's --output-last-message flag instead of redirecting stdout, so Windows process-teardown output no longer pollutes the review file and slips past the empty-output guard. (#1709)
  • last_activity now shows your local calendar day — the clock seam derived the date by slicing a UTC instant, so in negative-UTC-offset zones during UTC's early evening the date-only last_activity field jumped a day ahead of the operator's actual date (and of last_updated's local date). Operator-facing date fields now use a host-local calendar day while internal/cosmetic stamps stay UTC. (#2136) (#2216)
  • milestone_name is no longer clobbered with a delimiter-led fragment — getMilestoneInfo's ## heading regex was unanchored, so it matched a heading quoted inside backticks in the Milestones bullet and wrote garbage like — Active Milestone over the curated milestone name on every phase transition. Now consults the 🚧 marker first, anchors the regex to line start, strips the leading delimiter, and widens the preserve guard so a bad derive keeps the existing name. (#2135) (#2215)
  • init milestone-op now counts project_code-prefixed phase directories correctly — fully shipped milestones using the standard prefixed directory layout no longer report completed_phases: 0 or stay falsely incomplete. (#1844) (#1844)
  • /gsd-quick no longer halts with a stale-base worktree mismatch — the worktree executor now degrades to sequential execution when its fork base has diverged from origin/HEAD, instead of spawning a worktree guaranteed to fail the base-mismatch guard. (#1991)
  • Setting external_job.submit_timeout_ms / poll_timeout_ms / artifact_dir in .planning/config.json now actually configures the SLURM adapter — the keys were declared by the external-job capability but the adapter only read env vars, so config edits silently had no effect. The adapter now resolves them through the canonical capability-config seam (env override > config > registry default), surfaces the resolved artifact_dir in submit output, documents why the contribution registers at execute:wave:post (#1164 asks for wave:pre, which execute-phase.md does not dispatch today; wiring it is a core-loop change #1164 explicitly defers), and gains unit coverage for the CLI surface (parseFlags, findPlanningDir, resolveExternalJobSettings, formatShowReport). (#1164) (#2006)
  • The Antigravity reviewer in /gsd-review no longer reviews blind — agy -p never granted the agent the repo under review, so it frequently anchored on its own scratch directory and returned plan-text-only verdicts counted at full consensus weight. The reviewer is now granted the repo (capability-probed --add-dir) and anchored to the absolute repo root; a review that still runs without repo access is stamped [reviewed-without-repo-access] and down-weighted in the Consensus Summary. The cursor-agent prompt gains the same absolute-root anchor. (#2176) (#2184)
  • Autonomous reruns now skip phases with deferred verification until you resume them explicitly — if a prior /gsd-autonomous run recorded verification_deferred_human or verification_deferred_gaps, later reruns no longer drop back into the same prompt loop and instead point you at the saved resume command. (#1846) (#1846)
  • requirements mark-complete no longer reports silent success when the traceability row is missing — it OR-ed its checkbox and table-row writes into one flag, so a checkbox-only reconcile returned a payload byte-identical to a full reconcile while the traceability row stayed Pending (and re-run masked it as already-complete). It now surfaces table_unmatched for IDs whose checkbox reconciled but whose table row is absent, and treats a checked box with no table row as partial rather than done. (#2140) (#2219)
  • state prune now resolves the current phase from the canonical location — frontmatter current_phase, the Current Phase field, or the prose Phase: line scoped to the ## Current Position section — instead of extracting Phase over the whole document, where stateExtractField's pipe-table fallback could latch onto an unrelated | Phase | N | row (e.g. a historical verification table) and compute a wrong prune cutoff. (#1832)
  • model_overrides Claude model IDs now resolve to Agent-tool aliases on the claude runtime — a full Claude model ID (e.g. claude-sonnet-5) in model_overrides was returned verbatim and silently dropped by the Claude Agent tool (whose model parameter documents only tier aliases), causing the spawned subagent to inherit the parent session model instead of the configured one. It now maps to the tier alias (sonnet/opus/haiku/fable), consistent with the model_policy path (#1144). Bare aliases, non-Claude values, and non-Claude runtimes are unchanged; a Claude ID with no alias warns once and falls through to tier resolution. (#2041) (#2048)
  • Phase headers that place a parenthetical tag before the colon (### Phase 26 (Cluster B): Title) now resolve and enumerate the same as untagged headers. Previously the resolver returned not-found and roadmap analyze/listing silently dropped the phase (wrong phase_count, progress, and next_phase). Tag tolerance is applied at every phase-header read site; untagged and all existing header formats parse unchanged. (#1765)
  • Executor and milestone-summary/forensics workflows now call state.* commands with named flags so the named-only router records metrics, decisions, blockers, and session continuity instead of silently dropping positional args. (#1873)
  • bug-1367 install test no longer fails on Windows CI when hooks/dist isn't pre-built — the test ran install.js without building its hooks/dist precondition (a gitignored build artifact the unit lane doesn't build), so on a lane without pre-built hooks the installer hit "Failed to install hooks: directory is empty" and the before-hook threw. The test now builds hooks in its own before() (mirroring golden-install-parity). (#1926) (#1927)
  • /gsd-fast now appends Quick Task rows to STATE.md again — the log_to_state column-count guard used an off-by-one awk formula (NF-1) that was always one too high, so the schema gate rejected the very table quick.md creates and silently skipped the STATE.md update. Also now supports the 6-column validate-mode table. (#2133) (#2214)
  • Build the gitignored hooks/dist/ artifact once upfront in scripts/run-tests.cjs (the same chokepoint as ensureBuiltArtifacts), before any concurrent install test spawns install.js. Closes the scoped-CI first-build empty-dir race that intermittently failed install tests with Failed to install hooks: directory is empty (e.g. bug-3683-workflow-colon-namespace-leak). (#1967) (#1968)
  • workstream progress no longer reports shipped milestones as executing — gsd-tools workstream progress now derives each workstream's status from authoritative shipped signals (an archived milestone snapshot under milestones/, or a SHIPPED marker in the workstream ROADMAP) instead of trusting the mutable STATE.md Status field, so a stale field can never hide a shipped/archived milestone. The output adds status_source (field | derived) and status_conflict (true when the derived value disagrees with the stale field). (#1913) (#1916)
  • Windows install/upgrade/state-write operations no longer fail on transient antivirus/indexer file locks — the fs.renameSync atomic-publish sites (install state, hooks config, capability ledger/lifecycle, phase/workstream/milestone dirs, roadmap, planning/state locks) now retry EPERM/EBUSY/EACCES via retryRenameSync instead of propagating the transient lock; enforced by the new local/require-fs-op-fallback lint rule (ADR-1703 Phase 6). (#1740) (#1742)
  • reconstructFrontmatter now emits valid YAML for scalars and block-array items that were previously serialized unescaped. Values carrying a YAML indicator plus a literal quote/backslash, embedded control characters, the empty string, a leading YAML indicator, or leading/trailing whitespace are now routed through a properly escaped double-quoted form, so frontmatter round-trips through strict parsers (js-yaml, PyYAML) instead of corrupting the block on the next state sync. (#1807)
  • phase remove no longer destroys the Progress table when removing the last phase — deleting a phase used a whole-document regex whose scan, on the final phase, ran past the section and swept away the ## Progress heading and its entire tracking table; the deletion is now structurally bounded to the phase’s own section. (#2253) (#2253)
  • phases clear archives phase directories instead of destroying them — at a milestone switch, committed phase directories were hard-deleted (rmSync) with no archive, silently losing browsable phase history (the #1447 dirty-tree guard was a no-op for the common committed case). Phase directories are now moved to milestones/<version>-phases/ (collision-safe; timestamp fallback when no version resolves), so history survives the switch. The #1447 uncommitted-changes guard is retained as a secondary backstop. (#1871) (#1919)
  • Cross-AI review no longer silently drops the Codex/Claude/Gemini lanes on large plan sets — the prompt-fed reviewer blocks in review.md invoked each CLI with no explicit timeout, so a slow source-grounded review was killed at the host default (~2 min) and the lane was silently lost. The workflow now directs a high Bash timeout and frames an empty output as a timeout (not the crash it was misdiagnosed as). (#2194) (#2226)
  • /gsd-progress no longer reports a stale root milestone in workstream mode — in a multi-workstream project with no active workstream set, gsd-tools query init.progress silently fell back to root .planning/STATE.md (often stale) and reported it confidently. It now fails safe with an actionable error naming the available workstreams and the --ws/workstream set fix, so a stale root value is never reported. Flat mode and --ws <name> are unchanged. (#1912) (#1918)
  • Windows: stop double-quoting $CLAUDE_PROJECT_DIR-anchored managed node hook paths during the #2979 legacy rewrite, which produced ""$CLAUDE_PROJECT_DIR"/..." and broke every node managed hook with MODULE_NOT_FOUND (PreToolUse-guard deadlock). (#1746)
  • phase complete now updates STATE progress on milestone-grouped roadmaps — deriveProgressFromRoadmap parses the ## Progress table by header (column-by-name) instead of a fixed 4-column layout, so the 5-column milestone-grouped shape is no longer silently unparsed. (#2168)
  • Windows Claude Code hooks now work under PowerShell — when Claude Code's hook runner resolves to PowerShell (not Git Bash), every GSD-installed hook failed with Unexpected token because the installer emitted bare quoted paths with no PowerShell call operator. The fix adds a hookShell parameter to the hook-command projection chain; when hookShell='powershell', the & call operator is prepended. Default behavior (Git Bash, no prefix) is unchanged. (#2236) (#2261)
  • capability state and loop render-hooks now accept --runtime to override the auto-detected runtime — previously both commands parsed only --config-dir, so the runtime config dir was derived from the persisted .planning/config.json runtime (precedence GSD_RUNTIME → config.runtime → claude). A repo that persisted runtime:"codex" resolved the config dir to ~/.codex, where the Claude skill isn't installed, so every skill-bearing capability reported surfaced:false and execute:post/verify:post hooks silently no-op'd when the operator drove GSD from Claude Code. --runtime <r> (canonicalized, so aliases like codex-app work) now bypasses that fallback so the config dir resolves to the explicitly-named runtime's home. Behavior without the flag is unchanged. (#2003) (#2051)
  • phase complete no longer false-reports REQ-IDs as missing when the traceability table leads with a status column — the parser required the REQ-ID in the first column, so a table shaped | ☐ | REQ-01 | … matched zero rows and every body REQ-ID was reported missing. It now matches REQ-IDs in any column. (#2203) (#2234)
  • init milestone-op now ignores backlog 999.x headings when counting milestone phases — parked backlog items no longer inflate phase_count or pin all_phases_complete false for an otherwise finished milestone. (#1843) (#1843)
  • Phase archival is now wired end-to-end across the milestone lifecycle — finishes the #1871 follow-up: phases archive is now a real command (the half-wired alias is routed, no longer errors Unknown), milestone complete archives phase dirs by default (--no-archive-phases opts out), and new-milestone §6 stages the archive move + source removal in the same commit so history is preserved atomically rather than left as orphaned uncommitted deletions. (#1871) (#1924)
  • state update-progress no longer mangles the frontmatter and discards the progress suffix — its Progress: regex matched the raw STATE.md including frontmatter, so the YAML progress: key was hit first (corrupting the frontmatter) while the body line stayed stale and was silently reverted on the next write, and any descriptive suffix after the progress bar was destroyed. It now targets the body line only and preserves the suffix. (#2177) (#2224)
  • /gsd-ship no longer silently drops the ship-status note from STATE on merge — the track_shipping step committed the STATE ship-note after creating the PR but never pushed it, so on a fast merge the note stayed local-only and never reached the default branch. The ship-note is now pushed onto the PR branch with a [ci skip] trailer so it lands on merge without a redundant pipeline. (#2138) (#2217)
  • /gsd-debug no longer stalls on a phantom background handoff — the orchestrator treated the foreground session-manager spawn as a background task and queried its agent ID via TaskOutput (which needs a task ID), then waited on a handoff that was never queryable. The workflow now states the spawn is foreground/blocking, forbids passing an agent ID to TaskOutput, and gives a lost-handoff recovery path. (#2196) (#2227)
  • Roadmap phase lookup now ignores fenced examples and the backlog sentinel lane — roadmap get-phase and init plan-phase no longer return fenced sample headings as real phases or treat 999.x backlog items as active milestone work. (#1845) (#1845)
  • phase complete no longer checks the wrong ROADMAP checkbox or writes the plan count into a shipped milestone — the roadmap mutators ran unanchored and un-milestone-scoped, so they could flip a bullet inside a backticked prose literal or a Backlog entry instead of the closing phase's, and write the plan count into a same-numbered phase in a shipped milestone. The checkbox flip is now line-anchored and both writers are scoped to the current milestone. (#2200) (#2229)
  • roadmap get-phase resolves project-code-prefixed headings by bare number — a bare-number query (e.g. 29) now resolves a drifted ### Phase AB-29: heading, matching the internal resolver used by init.phase-op; previously the CLI returned empty. A bare sibling (### Phase 29:) still takes precedence. A project-code-prefixed heading present only as a summary/checklist line (no matching detail section) now reports a malformed_roadmap diagnostic — for both prefixed and bare-number queries — instead of a silent empty result. (#2114) (#2139)
  • milestone complete --ws now archives into the workstream instead of root — the archive paths (MILESTONES.md, the milestones/ archive dir, and the per-version MILESTONE-AUDIT.md) were hardcoded to root .planning/, so a workstream milestone close scattered its artifacts into root and never produced a workstream-local archive. They now derive from the workstream-aware planning base (planningPaths(cwd).planning); flat-mode (no --ws) is unchanged. (#1911) (#1917)
  • phase complete now reads milestone-grouped ROADMAP progress tables — progress reported 0% on projects whose Progress table carries a Milestone column, because the reader assumed a fixed column position; it now resolves progress columns by name so both flat and milestone-grouped tables work (#2137). Quick Tasks logging via /gsd:fast also appends schema-correct, lock-safe rows instead of guessing the column count in shell (#2133). (#2248) (#2248)
  • Skill-bearing capabilities now surface correctly on flat command-layout installs — on an install using the flat commands/gsd-<stem>.md source layout (e.g. a Claude Code local project install with no commands/gsd/ subdir), every skill-bearing capability (nyquist, code-review, security, ui, mempalace, ai-integration, profile-pipeline) was silently reported surfaced:false/enabled:false/active:false, so their loop hooks (verify:post, execute:post, etc.) never fired even with the corresponding workflow.* toggle on. The skill-manifest resolver now detects the flat layout and produces the same stems the nested commands/gsd/*.md loader does. (#1858) (#2049)
  • Roadmap, requirements, and state table edits are confined to the right table — the last ad-hoc table writers (phase completion updating roadmap progress, requirements mark-complete, and state record-metric/velocity) now route through the shared markdown-table seam, so a stray decoy table elsewhere in a document can no longer swallow a phase-progress update, a single ragged neighbouring row no longer silently aborts the whole edit, and per-plan metric recording no longer drops trailing section content or duplicates the section. (#2253) (#2253)
  • ROADMAP phase edits can no longer escape their section — completing a phase updated its plan count and per-plan checkboxes with whole-document regexes that could bleed into a neighbouring phase; those per-phase writes are now structurally bounded to the phase own section via a new withSection / withPhaseSection seam (#2130, #2067, #2080). (#2250) (#2250)
  • STATE.md ## Session fields now resolve on Windows — the session-section reader used a \n-only heading regex that silently failed on a CRLF ## Session heading, nulling all session state on Windows checkouts; it now reads through the CRLF-safe section seam. (#2253) (#2253)
  • Bullet/em-dash ROADMAP phases no longer resolve to Phase null — the roadmap phase lookup matched only ATX headings with a colon, so a bullet entry like - [ ] **Phase N — Name** (which the roadmapper emits) failed to resolve and Phase null landed in STATE.md; a bullet-only ROADMAP also broke the milestone phase count. Phase lookup and the milestone filter now accept bullet/checkbox entries with an em-dash/en-dash/hyphen/colon separator. (#2199) (#2228)
  • Linuxbrew users no longer lose all GSD-managed hooks after brew upgrade node — normalizeNodePath only recognized macOS Homebrew Cellar paths, so on Linux the version-pinned node path stayed baked into hook commands and 404'd after a node bump (and reinstall couldn't repair it). It now rewrites any Homebrew Cellar path — Intel, Apple Silicon, Linuxbrew, custom HOMEBREW_PREFIX — to the stable <prefix>/bin/node symlink. (#2185) (#2225)
  • milestone complete no longer corrupts the recorded phase — closing a milestone (e.g. v0.5) previously overwrote current_phase in STATE.md with the version's minor digit, and a follow-up state complete-phase mined a bogus 0.5 token and rewrote the file; phase resolution is now anchored so the real phase is preserved and a milestone-closure line is rejected. (#2111) (#2131)
  • Headless MemPalace capture no longer fails silently — the headless invocation mempalace mine <path> --wing <wing> --room <room> used a --room flag that does not exist on the mine subcommand (only search accepts --room), causing every headless/no-MCP capture run to fail with unrecognized arguments: --room and silently skip (onError: skip). The fix replaces the flag with MemPalace's documented room-assignment mechanism: stage the artifact under a room-named subfolder with a mempalace.yaml taxonomy so detect_room() assigns it via folder-path match. (#2220) (#2260)
  • Fixed: a hand-authored non-inferable backstop truth with a stray trailing space or surrounding quotes no longer silently grades green — it correctly abstains (insufficient_spec), restoring the #1154 honest-verifier guarantee. (#1909)
  • commit_docs no longer silently disables on CRLF .gitignore repos — git check-ignore falsely reports a trailing-slash path (e.g. .planning/) as ignored when the .gitignore has CRLF line endings with blank lines. isGitIgnored now strips trailing slashes before querying, so the false positive cannot occur. (#2206) (#2235)
  • Phase-directory resolution fails loud on cross-project collisions — when two unrelated GSD projects share a .planning/phases/ tree, a bare phase number silently resolved to the first 0N-* directory found. The fix detects multiple matches and surfaces an ambiguous_matches result. (#2237) (#2262)
  • scanPhasePlans no longer counts PLAN-REVIEW artifacts as executable plans — *-PLAN-REVIEW.md files were counted by the loose /PLAN/i fallback. The fix adds a PLAN_REVIEW_RE exclusion before the fallback. (#2252) (#2263)
  • Milestone audit no longer flags a not-yet-validated phase as a Nyquist failure — a phase that was planned but never run through validate-phase now reports as NOT-VALIDATED (a "run validate-phase" TODO) instead of collapsing into PARTIAL alongside phases whose validation genuinely failed. (#2117) (#2209)

Security

  • gate="blocking-human" checkpoints are no longer auto-approved by the execute-phase orchestrator — the package-legitimacy gate (#2827) spans two layers: gsd-executor refuses to auto-approve a gate="blocking-human" checkpoint and escalates it via checkpoint_return_format so a human can vet the package, and execute-phase's checkpoint_handling step decides what happens next. That step dispatched purely on checkpoint type and never read gate, so under --auto / --chain it immediately auto-approved the very checkpoint the executor had just refused to auto-approve (human-verify → {user_response} = "approved"). The slopsquatting defence was therefore inert in exactly the unattended mode where nobody is watching: an [ASSUMED]/[SUS] package reached install with no human ever seeing the verification prompt. checkpoint_handling now carves out gate="blocking-human" (and the package-legitimacy what-built markers) ahead of every auto-mode branch, routing those checkpoints to the standard present-to-user flow regardless of type. references/checkpoints.md documents the gate attribute and its two values for the first time — previously blocking-human appeared nowhere outside agents/gsd-executor.md, so no planner had a documented way to author a checkpoint that auto-mode could not bypass. The existing regression test asserted the executor half only; it now asserts the orchestrator half too, which is why it stayed green while the gate was open. (#2107) (#2113)
  • Hardened phase/roadmap/plan markdown parsing against quadratic-time (ReDoS) CPU exhaustion — a crafted ROADMAP.md, STATE.md, or PLAN.md with large runs of unclosed (, [, <tag>, <!--, or <details> could drive the phase-header, Plans-count, files_modified, and <tag>-block parsers into O(n²) scans (tens of seconds on a ~1.5 MB file). Every affected regex is now linear: header tag/bracket clauses are length-bounded, the Plans-count scan is section-local, and all <tag>…</tag> extraction routes through a single ReDoS-safe seam. (#2128) (#2141)
  • Installer writes are now confined to the declared config home — the workflow/skill emit path (copyWithPathReplacement) and the Codex config writer (installCodexConfig) now reject any destination that escapes the install root: crafted or absolute paths, path-separator agent names, and pre-existing symlinks are refused before any delete or write. Fail-closed: an install write with no declared root is rejected rather than written unconfined. (#1725)
  • Install write-confinement (ADR-1239 Phase B) — the installer now rejects any runtime-descriptor destSubpath that would write or delete outside the user's config home (path traversal, the config root itself, NUL bytes) and refuses to follow a pre-existing symlink that escapes it. Hardening only; no change to legitimate installs. (#1706)

[1.6.1] - 2026-07-01

Added

  • Claude Sonnet 5 is now the standard (sonnet) tier model. The model catalog and provider presets resolve the sonnet/standard tier to claude-sonnet-5 (GA 2026-06-30) across the Anthropic-backed runtimes (claude, copilot, and the anthropic/anthropic-fable presets), plus the OpenRouter-style anthropic/claude-sonnet-5 for opencode/hermes, replacing the superseded claude-sonnet-4-6. Opus and Haiku tiers are unchanged. (#1847) (#1848)

Fixed

  • milestone complete and roadmap analyze now exclude the Phase 0 / Phase 999 backlog sentinels. A milestone whose only directory-less ROADMAP heading is a backlog sentinel can be completed without --force, and roadmap analyze no longer counts the sentinel in phase_count or routes next_phase into it. Completes the ^999 exclusion #1445 added to the progress denominators. (#1691)
  • phase.complete no longer reports a false is_last_phase on a <details>-wrapped checkbox checklist (#1591, #1752) — when the active milestone's phase checklist was written as - [ ] Phase N: checkbox items inside a <details> block and the next phase had no directory on disk yet (still in planning), phase.complete's isLastPhase roadmap-enumeration fallback used a heading-only pattern (/#{2,4}\s*Phase…/) that never matched checkbox items. It returned is_last_phase: true, next_phase: null on a mid-milestone phase and — via the milestone-complete cascade — wrongly flipped STATE.md to Milestone complete and decremented progress.total_phases (e.g. 8 → 7). The pattern now matches both heading-style (### Phase N:) and checkbox-list phases (- [ ] Phase N: / - [x] Phase N:); extractCurrentMilestone already surfaces the <details>-wrapped checklist correctly, so no parser change was needed. Only the reproduced phase.complete fallback is changed; the heading-only sibling patterns elsewhere in phase.cts are untouched. (#1819)
  • Windows: stop double-quoting $CLAUDE_PROJECT_DIR-anchored managed node hook paths during the #2979 legacy rewrite, which produced ""$CLAUDE_PROJECT_DIR"/..." and broke every node managed hook with MODULE_NOT_FOUND (PreToolUse-guard deadlock). (#1746)

[1.6.0] - 2026-06-24

Added

  • workflow.context_guard_mode config key — proactive context-exhaustion guard for execute-phase. Before each wave, the orchestrator self-assesses context pressure using the degradation signals defined in context-budget.md. Values: warn (default — emit warning and recommend /gsd:pause-work when POOR tier detected), auto (automatically invoke /gsd:pause-work before next wave), off (disable). Set via gsd config-set workflow.context_guard_mode auto for fully autonomous checkpoint behaviour. (#1452) (#1452)
  • agent-skills --json IR gains an additive value: { block, skills_count } field formalizing the Resolution<T> convention for config-interpreting read verbs; no breaking change. The new src/resolution.cts module exports Resolution<T> { value, configured, reason, warnings } (the canonical envelope) and makeResolution<T>() (the builder); AgentSkillsValue { block, skills_count } is the first adopter. All existing flat fields (agent_type, block, skills_count, warnings, configured, reason, source, degraded) are retained for back-compat. Capability-state and capability-writer keep their existing JSON shapes unchanged; only doc comments are added naming them the canonical read-verb and mutation-verb envelopes respectively. The shared seam across all shapes is warnings: string[]; a single generic across read+write verbs was rejected by the deletion test (ADR-1411 P3 amendment). (Part of #1411, P3 / #1416.) (#1425)
  • Added gsd capability outdated — a new subcommand that light-peeks each installed overlay capability's recorded source for the latest version that re-resolving that source would install and reports which have an update available (ADR-1244 D6 per-source matrix: git ls-remote --tags, npm view … version, local re-read; tarball → manual, registry → unknown). A capability is reported outdated only if re-resolving its recorded source would fetch a newer version: an npm range (@^1) resolves to the highest version matching the range (read from each npm view line's canonical version field, so a version-like substring in the package name never poisons the result), and a source pinned to an immutable ref (git #sha:/#tag:) or an exact npm version is reported pinned — never outdated, since update will not move it. A bare git ref (#<ref>) is classified at the remote with a bounded git ls-remote: a ref that resolves to a tag is pinned, while a mutable branch ref is never pinned (it degrades to unknown, since the installed commit is not recorded to compare against). Each capability is classified outdated / current / pinned / manual / unknown; subprocesses are bounded (git ≤30s, npm ≤60s) and a failing or unsupported peek degrades that row to unknown instead of crashing the command. --json emits the records array; the default prints a table. (#1463) (#1488)
  • gsd capability management command — install, update, remove, list, disable, and enable GSD capabilities (first-party and third-party overlays) from a registry / git / npm / tarball / local source, wiring the ADR-1244 lifecycle (source resolver, ledger, consent + integrity trust gate) to a user-facing CLI. (#1457) (#1457)
  • Runtime capability registry overlay — installed third-party capabilities (under ~/.gsd/capabilities/ or a project's .gsd/capabilities/) are now composed into the registry at runtime via loadRegistry({ includeInstalled }): validated against the same conformance invariants as first-party, first-party-wins on any collision, skipped-with-a-warning when incompatible with the running GSD version (engines.gsd), with gate-kind capabilities failing closed. Installed overlays are toggable via surface and federate their config keys (cwd-aware) exactly like first-party. Foundation (ADR-1244 Phase 2) for capability install/upgrade/remove. (#1440)
  • Capability manifests are now versioned — every capability.json carries a required semver version, plus optional engines.gsd, compatVersions, integrity and provenance fields, enforced by the capability conformance validator. First-party capabilities are version-stamped in lockstep with the GSD release. Foundation (ADR-1244 Phase 1) for installing, upgrading, and removing capabilities in later releases. (#1436)
  • /gsd-capture --list-seeds audits parked seeds — a new read-only listing of .planning/seeds/ showing each seed's ID, status, scope, and trigger, with an optional status filter (e.g. --list-seeds dormant). Backed by the gsd-tools list-seeds command. Previously seeds could only be created or auto-surfaced at /gsd-new-milestone, with no way to browse them on demand (#441). (#722)
  • Capability source resolver + install ledger — resolveCapabilitySource(spec) fetches a capability from a local path, git repo, npm package, or tarball URL, verifies it (sha512 integrity before staging, engines.gsd compatibility, full conformance validation) and stages a bundle without executing any capability code (copy/extract only — npm pack --ignore-scripts, never npm install; symlink/tar-slip/shell-metacharacter/unsafe-transport inputs rejected). A per-runtime ledger records what each install wrote for atomic, reversible upgrade/remove. Foundation (ADR-1244 Phase 3) for the upcoming gsd capability install command. (#1443)
  • Capability matrix reference — a generated catalogue (docs/reference/capability-matrix.md) of every first-party capability's role, tier, extension points, hook kinds, and engines.gsd, generated from the committed registry and kept honest by a CI drift guard so it can never fall out of sync with the actual capability set. (#1458) (#1458)
  • Third-party capabilities can ship dispatchable CLI commands (ADR-1244 Phase 5) — a capability that declares a commands family is now dispatched by gsd-tools <family> via the registry, the same seam the first-party graphify/intel/audit commands already use. Third-party command dispatch runs only for an installed, consented capability (a committed ledger entry) and loads the router module strictly from that capability's own install root (basename + realpath confinement, rejecting .. traversal and symlink escape); a bundle merely present on disk with no install record keeps its declarative surfaces but is never command-dispatchable. (#1450) (#1450)
  • Plugin installs now expose GSD skills — when GSD is installed as a Claude Code plugin (claude plugin install), its skills are available via gsd-core:<skill> the native way. Previously, plugin-only installs lacked the skill surface because bin/install.js never ran; agents that preload global:gsd-core:<skill> (PR #1261) now resolve against plugin-provided skills. (#1596) (#1597)
  • Added a validated gsd-tools worktree record-agent writer verb that appends a per-agent entry to the wave cleanup manifest, validating every field at write time with the same rules the cleanup-wave reader enforces (write-strict --agent-id) and failing loudly with a recovery hint instead of silently appending an under-populated entry. The execute-phase orchestrator now records each spawned worktree through this verb. (#1448) (#1448)
  • gap-analysis --phase-req-ids now expands numeric ID ranges — a same-prefix ascending equal-width range like SEL-01..SEL-03 expands to SEL-01, SEL-02, SEL-03 (zero-pad preserved) instead of being treated as one literal ID that gap-analysis then reports as missing. Ambiguous tokens (mismatched prefix, descending, differing width, non-numeric, >1000 span) stay literal. (#1269) (#1419)
  • /gsd-plan-phase now flags a stale codebase map before planning — the drift capability runs its codebase-drift check at plan:pre (non-blocking, warn-only), so a stale STRUCTURE.md is surfaced before the planner is spawned instead of being discovered mid-execution by the existing execute:wave:post gate. Gated on a new workflow.plan_drift_precheck toggle (default on), independent of workflow.schema_drift_gate, so autonomous/CI runs can silence the plan-time advisory without disabling the execute-time gates. (#1595)

Changed

  • Capability commands now emit dispatch audit records — graphify, intel, audit-uat, and audit-open now route through the Command Routing Hub per ADR-959 §III(B), so GSD_AUDIT=1 traces, the structured stderr JSON error envelope, and the typed Result contract cover them uniformly with all other command families. JSON-error reason values (usage, sdk_unknown_command) are preserved byte-identical. (#1646) (#1647)
  • /gsd-verify-work now routes UAT deterministically from a structured coverage: block on SUMMARY.md — deliverables proven by passing tests (human_judgment: false with a non-empty all-pass verification list) are auto-passed (source: automated, no prompt), and only judgment-dependent or unverified deliverables are presented for human sign-off. SUMMARYs without a coverage: block fall back to the previous prose-based extraction, byte-identical. Authored by execute-plan and validated by the new gsd-tools uat classify-coverage verb. (#1611)
  • Thread isGlobal install scope through the descriptor-driven convertedAgentsKind / stageAgentsForRuntimeWithConverter plumbing — a prerequisite for the ADR-1235 agent-conversion cutover. No runtime declares a converted agents kind yet; the capability.json wiring is deferred to a follow-up that first ships the ADR-1235 §0 byte-for-byte parity harness (so the /gsd:surface / --materialize consumer can mirror the legacy agent pipeline before the kind goes live). The legacy bin/install.js agent loop remains authoritative, so installed agent output is unchanged. (#1173) (#1438)
  • /gsd-review now asks external reviewers to verify plan claims against the source — the reviewer prompt requires opening the referenced files, citing file:line evidence + mechanism, and tracing asserted behavior, with a graceful-degradation clause for reviewers that have no file access. This turns every capable agentic reviewer into a real second source instead of a plan-text paraphraser. (#1318) (#1421)
  • eval-auditor scoring moved into a deterministic eval.score query verb (LLM-playbook principle 10) — coverage/infra/overall arithmetic and verdict banding are computed in code (gsd-tools query eval.score) instead of by the model. Based on arXiv 2601.15130 (Plausibility Trap / DPDM), 2508.15754 (Tool-Integrated Reasoning), 2507.10281 (Table Agent); 2504.00406 / 2510.15955 supporting. (#1583)
  • gsd-tools now resolves the project root from a descendant subdirectory — findProjectRoot walks up to the nearest ancestor directory containing .planning/ so config loads correctly when invoked outside the project root; previously it fell through to defaults for plain descendant paths (cwd-drift gap #1366). Sub_repos, multiRepo, and .git-based heuristics retain priority. (Part of #1411, P1 / #1414) (#1423)
  • verify-phase test-tier prohibition fail-first can now prove the RED is caused by the violation's content — the node-test machine-proof (#1279) confirmed a known-bad subject drives the negative test RED, but could not tell a genuine content-violation from a deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. An optional fifth flat scalar check_clean_fixture (→ CheckDescriptor.cleanFixture) threads a KNOWN-CLEAN control subject through projectProhibitions + descriptorFromProjection; when present the prover also runs the check against it and requires GREEN, so fail-first is proven only when the check is RED on the violation and GREEN on the clean subject (content-dependent). It is opt-in and additive: absent a clean fixture the prover behaves exactly as it did post-#1314 (no control, documented residual), preserving the zero-authoring compose path; the lint-rule kind needs no analog. (#1346) (#1518)
  • fish-shell support in the post-install PATH suggestion. When a directory is not on your PATH, the installer now prints a fish-native fish_add_path '<dir>' line alongside the zsh/bash suggestions (the previous export PATH=… commands are inert in fish). It also stops the false-positive "not on your PATH" warning for fish users whose fish_user_paths/config.fish already covers the directory, detected via a read-only probe of fish's config (no fish subprocess, no writes). No change for bash/zsh/PowerShell/cmd/Git-Bash users. (#727)

Fixed

  • Project-local Claude Code install now produces /gsd-<cmd> (hyphen) slash commands — the installer was writing command files to .claude/commands/gsd/<cmd>.md (subdirectory with bare names), causing Claude Code to namespace them as /gsd:<cmd> (colon form). The fix writes flat gsd-<cmd>.md files at .claude/commands/ level so Claude Code registers /gsd-<cmd> (hyphen form), matching hooks, statusline, and all cross-command references. Legacy commands/gsd/ directories from prior installs are cleaned up on reinstall and uninstall, with dev-preferences.md preserved. (#1367) (#1367)
  • execute-phase now re-checks the worktree fork base at the start of every wave and resets the wave manifest between waves (#1369) — two compounding issues caused wave N+1 worktrees to be created from the stale pre-wave-N commit. First, the worktree.base-check auto-degrade only ran once at initialize time; after Wave N merged and tracking commits advanced orchestrator HEAD past origin/HEAD, Wave N+1 worktrees were still forked from origin/HEAD (Claude Code's "fresh" base), causing both agents to immediately halt with FATAL: worktree base mismatch from the worktree_branch_check guard. Second, WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1 would reuse the consumed wave-N manifest file, causing the step 5.5 manifest guard (#3384) to block on subsequent waves. Two safeguards fix this: step 0.5 in the execute_waves "For each wave" loop re-runs worktree.base-check before every wave's dispatch (when divergence is detected, USE_WORKTREES is overridden to false for that wave); step 7c between waves unsets WAVE_WORKTREE_MANIFEST so wave N+1 creates a fresh per-wave manifest, and re-asserts worktree.baseRef:"head" (idempotent) so the Claude Code harness re-reads the live HEAD on the next dispatch. The permanent fix remains setting worktree.baseRef:"head" in .claude/settings.local.json (see #683). (#1369)
  • Workflow temp files now randomize correctly on BSD/macOS — several workflows called mktemp with templates where XXXXXX was followed by a .json/.md suffix (e.g. gsd-worktree-wave-XXXXXX.json, gsd-pr-body.XXXXXX.md). BSD/macOS mktemp only substitutes XXXXXX when it is the final path component, so those templates returned a literal, non-randomized path, letting concurrent workflow runs collide on the same temp manifest/body file (one run overwriting or consuming another's). The fix creates a suffixless temp then renames to add the extension — portable across BSD + GNU. Affected: execute-phase, quick, spec-phase, ship, profile-user. (#1520) (#1550)
  • Core-path file locks now verify the holder process is alive before stealing a stale lock (#1532) — the STATE.md write lock (acquireStateLock) and the .planning/ workspace lock (withPlanningLock) previously stole locks on a bare mtime timer with no liveness check, so a live-but-slow holder (e.g. a deep .planning/ scan on slow NFS) could have its lock stolen mid-write, corrupting STATE.md or losing an update. Both locks now gate stealing on process.kill(pid,0) liveness with a deadman ceiling above the wait budget (pid-reuse backstop), withPlanningLock no longer force-steals a live holder on timeout (and can no longer leak an uncaught EEXIST), writeStateMd computes its disk scan inside the lock, and acquireStateLock no longer leaks a file descriptor or strands an empty lock on a recoverable write error. The steal itself is now race-safe: a lock is never stolen while its body is still being written (the create→pid-write window), and stealing uses an atomic rename with an identity re-confirm so two waiters can no longer both reclaim the same lock and end up holding it concurrently. The uncontended path is unchanged. (#1532)
  • normalizeNodePath now maps pruned mise node paths to the stable shim (#1619) — resolveNodeRunner() bakes process.execPath into managed .js hook commands. Node realpaths execPath, so under mise it resolves to <data>/installs/node/<ver>/bin/node — a concrete version mise prunes on mise up, after which every managed hook fails to spawn (No such file or directory on every SessionStart and tool event), the same ephemeral-path failure #977 fixed for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise versioned install path to the stable sibling shim <data>/shims/node (.exe preserved on Windows) when that shim exists, deriving <data> from execPath so a custom MISE_DATA_DIR works, and falling back to the raw execPath unchanged otherwise. (#1619) (#1621)

fix(#1472): validate health is now workstream-aware — PROJECT.md and config.json are resolved from .planning/ root, while ROADMAP.md, STATE.md, and phases/ follow the workstream-scoped path; previously both sets were routed through planningDir() causing false E002/E003/E004/W003 when GSD_WORKSTREAM is set.

fix(#1454): validate health W017 no longer suggests removing the active session's worktree — stale-worktree findings are now skipped when the worktree path matches or is an ancestor of process.cwd(). (#1483)

  • frontmatter set / frontmatter merge no longer destroy must_haves object-lists — changing one frontmatter field (e.g. wave) silently dropped every provides: value and collapsed must_haves.artifacts/.prohibitions from a structured [{path, provides}] list into a malformed inline array, because the whole frontmatter was round-tripped through a lossy parse→serialize path that flattens object-list items to scalar strings. The write now preserves the original raw text for any structurally-unchanged top-level key and regenerates only the field that actually changed, so unrelated must_haves blocks survive verbatim. (#1572) (#1656)
  • Non-Claude runtime installs now resolve their own runtime and never attempt Claude-only worktree isolation — on any non-Claude install (Cursor, Gemini, Qwen, etc.) a runtime-neutral .planning/config.json previously resolved runtime=claude and enabled git worktree isolation, which only Claude Code's isolation="worktree" can honor — risking main-checkout edits while the workflow believed agents were isolated. Every non-Claude install now resolves its own runtime identity, defaults workflow.use_worktrees to false, fails closed if worktrees are forced on, and runs plan/execute inline in the manager/autonomous flows since only Codex can background-nest the pipeline's subagents. (#1521) (#1537)
  • Phase-aware commands now resolve project-code-prefixed ROADMAP headings such as MANIFOLD-117, while the roadmapper is instructed to keep project_code out of phase headings. (#1456)
  • workflow.mvp_mode now accepted by config-set; three undocumented workflow keys added to references — workflow.mvp_mode, workflow.code_review_command, and workflow.plan_chunked were consumed by planning-pipeline code but could not be set via config-set (they were missing from VALID_CONFIG_KEYS) or discovered via reference docs. All three are now in the schema and documented in references/planning-config.md. (#1500) (#1500)
  • Windsurf reinstall removes legacy .devin/skills/ artifacts — pre-#1615 installs wrote skills under .devin/skills/gsd-/ (Devin Desktop layout, #1085). #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Reinstalls now remove GSD-managed .devin/skills/gsd- dirs; user-owned content is preserved. (#1631)
  • adr-parser now classifies 9 previously-dropped punctuated ADR headers (Trade-offs, Non-Goals, Won't Do, Follow-up, How We'll Know, etc.) into their intended buckets instead of leaving them unmapped. (#1536)
  • Add prototype-pollution guard to the workstream/root config merge (_deepMergeConfig) so a config.json with a proto/constructor/prototype key can no longer spoof unset config flags. (#1534)
  • All GSD agents load on Gemini again — the Claude Skill/SlashCommand tools were converted to an invalid skill tool that Gemini rejects, aborting the load of 22 of 34 agents. They are now excluded from the Gemini and Gemini-backed Antigravity agent tools: frontmatter, the same way AskUserQuestion already is. (#1394) (#1418)
  • Antigravity config-dir resolution no longer shadows the active runtime — when more than one of ~/.gemini/antigravity, antigravity-ide, or antigravity-cli exists, GSD now resolves to the directory it actually installed into (marked by gsd-core/VERSION) instead of whichever directory existed first. Fixes silent misresolution where a CLI user who also had the Antigravity-IDE dir present was sent to the legacy dir (regression from #217). (#1442)
  • gsd-tools query agent-skills no longer silently drops a configured agent's skills under cwd or workstream drift — cmdAgentSkills now anchors to the project root via findProjectRoot before loading config, so invoking it from a descendant subdirectory or with a GSD_WORKSTREAM that has no scoped config correctly resolves the configured agent_skills block. A new loadConfigResolved(cwd, options) → { config, source, degraded } function reports provenance alongside the config object: source distinguishes 'root' | 'workstream' | 'builtin-defaults' | 'global-defaults'; degraded:true signals a workstream was requested but its config.json was absent. The --json IR of agent-skills gains four new fields — configured (bool), reason ('resolved' | 'not_configured' | 'configured_empty' | 'configured_unresolved'), source, and degraded — making silent failures visible and testable. A configured_empty or configured_unresolved agent emits a stderr WARNING; an unconfigured agent stays silent. (Closes #1366. Part of #1411, P2 / #1415.) (#1424)
  • findProjectRoot now respects explicit sub_repos config over implicit .git — when a parent workspace's .planning/config.json lists a child directory in sub_repos, that declaration takes precedence over the child's own .git/ directory. Previously, if the child had both .planning/ and .git/, the .git heuristic fired first and resolved to the child rather than the parent workspace, making the sub_repos declaration ineffective. (#1422)

phases clear now refuses to delete phase directories with uncommitted changes — cmdPhasesClear runs git status --porcelain over the phases directory before executing any deletion. If uncommitted or staged-but-not-committed files are found it aborts with a clear error message, preventing silent data loss at new-milestone time. Pass --force to bypass the guard when archival is already complete. Non-git projects are unaffected. (#1447, data-loss fix) (#1484)

999.x backlog phases are now excluded from total_phases, and total_phases can correct downward — deriveProgressFromRoadmap counted all progress-table rows whose phase cell started with a digit, so a 999.1 Backlog row inflated total_phases by one per entry (#1445). The same overcounting occurred in getMilestonePhaseFilter (which feeds isDirInMilestone and phaseDirs) and in the roadmapPhaseCount loop in buildStateFrontmatter. All three sites now filter phase tokens matching /^999\b/, consistent with the existing exclusion in init.cts. Additionally, shouldPreserveExistingProgress included total_phases in its ratchet check, preventing the counter from decreasing once set too high — e.g. after a 999.x fix or a ROADMAP correction (#1446). total_phases is now always taken from the freshly derived value; only completed_phases, total_plans, and completed_plans retain ratchet behaviour. (#1490)

  • Capability trust model was bypassable for project-scope third-party capabilities (#1459). The consent signal for a project-scope capability was its in-repo project ledger — but a project ledger is repo-plantable, so cloning or forging a repository activated that capability's executable surfaces (hooks, MCP servers) AND its declarative loop surfaces (steps, gates, contributions, federated config) AND its command dispatch with no user decision on the machine running it. The fix moves the authoritative consent signal off the repo tree into a new user-owned consent store at ${GSD_HOME||homedir()}/.gsd/consent.json (new leaf module src/capability-consent.cts): a bounded, non-throwing, atomically-written store keyed by (realpath(projectRoot), capability id). The security binding is a recomputed full-bundle content hash (bundleContentHash — a sha512 over every regular file in the bundle, manifest AND artifacts AND identity, symlinks and non-regular entries rejected, bounded), not the repo-plantable ledger integrity (which is '' for path/git/dir installs — a degenerate '' === '') and not the executable-only disclosure signature (which is constant for a declarative-only capability, so a repo-write attacker could swap capability.json for a malicious gate/contribution while consent still matched). The loader recomputes the bundle content hash at load and activates a project-scope overlay — declarative surfaces and command dispatch alike — only when it matches the consent record on this machine; any tamper (a swapped declarative manifest, an edited hook script, an empty-integrity local install) changes the hash and leaves the capability discovered but inactive (gsd capability list reports status: inactive with a reason). Global-scope overlays (under the user's own home) remain trusted without a per-project record. The lifecycle records consent (bound to the installed bundle's content hash) on a consented project install/upgrade and revokes it on remove, deriving the project root through one canonical helper (consentProjectRoot) shared by the install record site, the loader lookup, and trust revoke, so an install from a sub-directory is not immediately inactive. The disclosure signature now also covers each MCP server's transport/url/headers (non-stdio endpoints), env, cwd, and the raw args array, plus a command module's router, and every surface line is JSON-encoded (no delimiter-injection collisions) — so a swapped remote endpoint, header, environment (e.g. NODE_OPTIONS=--require evil.js), entry point, or non-string arg forces re-consent. The consent store serializes concurrent cross-project writes under a lockfile (no lost updates), enforces its record cap at write time, uses a collision-safe on-disk key for paths containing spaces, and tolerates a vanished directory on the durability fsync. New CLI: gsd capability trust list and gsd capability trust revoke <id> [--project <path>]. The loader's per-scope ledger read now goes through the shared bounded fd reader (a repo-planted FIFO ledger can no longer hang the loader) and reuses the ledger's shared isValidLedgerEntry validator for committed-entry parity. Integration hardening: the overlay consumers (capability-state, loop-resolver, the federated config-loader/config-schema) now thread the consent home (GSD_HOME) explicitly to the loader so a consented project capability is never looked up at the wrong home; the loader's discovered-but-inactive warning carries a structural kind: 'unconsented' discriminant that gsd capability list filters on (rather than matching the reason prose); a reconcile rollback that deletes a project-scope bundle also revokes its now-stale consent so an identical re-drop stays inactive; installCapability/upgradeCapability warn on stderr when a project install supplies no consent store and when the consent-store write fails (the install still succeeds — a consent-store IO error never fails an otherwise-successful install); gsd capability trust list now exposes the stored disclosureSignature and contentHash for diffing; and when GSD_HOME resolves equal to a genuine project root an in-repo bundle still requires a consent record (it is not deduped as trusted-global). Convergence hardening: the content-hash canonicalization — the security binding itself — is now injective and lossless. It length-FRAMES every component (an entry count, then per entry a type tag, the uint32 path length + path bytes, and for files the uint64 content length + the raw content bytes read via a new raw-Buffer reader, never a lossy UTF-8 decode) so neither a NUL embedded in file content can fake a file boundary (the old relpath + NUL + content + NUL framing was non-injective) nor can two binary artifacts that differ only in invalid-UTF-8 bytes collide on U+FFFD; empty directories are bound via typed directory markers so adding/removing one changes the hash. recordProjectConsent/revokeProjectConsent now throw rather than perform an unlocked read-modify-write when the consent-store lock cannot be acquired (the lifecycle already treats a consent-write failure as non-fatal and warns, so an install still succeeds). The consent lock and the lifecycle lock are now ONE shared hardened primitive (src/capability-lock.cts) — process-start-time liveness identity + a hard deadman — so the consent lock can no longer stale-steal a slow-but-live writer (the old mtime-only 60 s steal could). Finally, the MCP disclosure signature now folds in a stable hash of the full server config object the writer persists (not only the whitelisted fields), so an upgrade that changes any host-honored field outside the whitelist (a future envFile/workingDir/launch option) still forces re-consent. A final convergence pass closes four residual gaps: (1) the loader's overlay-root dedup and the CB-3 "project root == global home ⇒ require consent" comparison are now keyed on fs.realpathSync (fail-safe to path.resolve), so a symlinked GSD_HOME aliasing the project root can no longer slip an in-repo bundle into the trusted-global slot — it still requires a consent record; (2) the loader reads capability.json through the shared bounded fd reader (regular-file + size cap) instead of a raw fs.readFileSync, so a project-planted FIFO/device or oversized manifest skips the overlay (warning) rather than hanging or OOM-ing the loop; (3) the PATH component of the content hash is now hashed from raw directory-entry bytes (a { encoding: 'buffer' } walk, separator normalized at the byte level), so two bundle files whose names differ only in invalid-UTF-8 bytes (which a string decode would collapse to U+FFFD) no longer collide; and (4) the gsd capability trust revoke CLI now catches the consent-store lock-acquire failure and emits a clean, actionable error instead of surfacing a raw stack. A further convergence pass closes three more residual gaps and documents one irreducible limit: (1) the loader's user-owned consent gate now runs before the heavy pre-activation work (materializeHookFragments and cross-capability validation) for a project-scope overlay, so a forged in-repo bundle whose fragment.path points at an in-bundle FIFO/oversized file is skipped (unconsented → inactive) without ever reading that fragment — closing a pre-consent hang/OOM (the gate's decision is unchanged; only the work-ordering moved), and as defense-in-depth materializeHookFragments now reads each fragment body through the shared bounded fd reader (regular-file + size cap) so a FIFO/device/oversized fragment becomes an un-materializable-fragment validation error rather than a blocking read on any scope; (2) gsd capability list now reads each project capability.json through the same bounded reader instead of a raw fs.readFileSync, so a project-planted FIFO/device or oversized manifest omits that entry's metadata and exits cleanly rather than hanging/OOM-ing the list; (3) the loader's canonicalDir realpath failure is now strictly fail-safe — a candidate that would be classified trusted-global but whose realpathSync throws (a race/odd-FS, e.g. a symlinked GSD_HOME aliasing the project root) is reclassified conservatively to consent-required project, so it can no longer park an aliased project tree in the trusted-global slot (a non-existent global dir still resolves to a harmless no-op scan). Finally, an honest in-code comment at the loader consent gate documents the irreducible filesystem-primitive TOCTOU residual: the content hash binds the bundle at verification time, but a local writer racing between that verification and the capability's later execution can still mutate the bundle files — closing this window would require fd-pinned execution or an atomic content snapshot (native support not available at this layer); any persisted tamper is still caught on the next load (mirrors the #1462 lock-release residual — a documented real limit, not a dismissal). A final deep-convergence pass closes three more gaps: (1) the realpath fail-safe is now two-sided — a global overlay root is trusted (consent-free) ONLY when realpath(global) AND realpath(project) BOTH succeed AND resolve to DIFFERENT physical paths; the prior one-sided rule (demote only a realpath-failed global candidate) still let a symlinked GSD_HOME aliasing the project root bypass consent when the GLOBAL candidate realpathed fine but the PROJECT candidate's realpath failed (the keys never collided, so the in-repo bundle stayed in the no-consent global slot), so distinctness that cannot be proven (either side throws, or both resolve equal) now demotes the global to consent-required project whenever a genuine project root exists — while a genuinely non-existent project overlay (ENOENT) or a distinct real global root still stays trusted; (2) bundleContentHash now bounds the enumeration itself — it streams each directory via fs.opendirSync + readSync and throws the moment a cumulative entry counter exceeds the cap, BEFORE collecting/sorting a whole directory, so a malicious unconsented bundle with a huge single directory (or a deep tree) can no longer force unbounded memory/CPU before fail-closing (the cap is cumulative across the recursive walk; determinism is preserved by sorting the bounded set); and (3) a project remove no longer silently swallows the revoke-on-lock-failure throw — revokeProjectConsent throws on a consent-lock failure (a stale consent record a byte-identical re-drop could reactivate against), so removeCapability now surfaces it via a stderr warning naming the record AND a consentRevokeFailed/consentRevokeWarning flag on the result, which the CLI reports as a non-clean removal (telling the user to run gsd capability trust revoke). (#1473)

Capability --integrity is now verified or rejected per source, and hook commands are confined to the bundle — a supplied --integrity pin was silently dropped for npm, git, and local capability sources (only the tarball source verified it), so a user could believe content was pinned when it was not. npm now verifies the pin over the npm pack .tgz bytes; git and local sources, which have no single hashable artifact, now reject a supplied --integrity with an actionable error instead of ignoring it. Separately, a capability hook's relative script was written verbatim as the hook command, so it resolved against the working directory (not the capability bundle) at hook-exec time and a crafted relative path could escape the bundle; the command is now resolved against the capability's own install dir and realpath-confined to it, then written as an absolute path. That absolute command is consumed by a shell, which exposed two further problems now fixed: (1) a manifest could ship a file literally named run.sh; touch /tmp/pwn (filenames may legally contain ;, spaces, $, backtick, |, newline) and declare it as the hook script, so the emitted command injected a second shell command even though the file lived inside the bundle — the validator now rejects any hook script path outside a conservative [A-Za-z0-9._/-] allowlist (no whitespace, shell metacharacters, leading -, absolute path, or ..), failing the install/load loudly, and the confinement helper mirrors the same rejection defensively; (2) the absolute path begins with the install-home directory, which commonly contains spaces (e.g. /Users/Bob Smith/.claude/...) and word-split or broke when written unquoted — the emitted command is now POSIX single-quoted so the install prefix can neither break nor inject. (#1460) (#1481)

  • The capability loader never crashes on a single malformed overlay, and untrusted manifest/tar reads are size-bounded (ADR-1244 D2 invariant) — loadRegistry now makes the WHOLE per-candidate overlay-processing body total: ANY throw from ANY validator or step (including the committed validateCapability, which dereferences a malformed array entry such as gates: [null] / steps: [null] / contributions: [null] before its shape check) drops just that one overlay with a skip-warning instead of escaping the loader, while gatePointsOf is hardened to be total over null/non-array/malformed gates. The final canonical buildRegistry compose stays guarded: a throw on the merged set falls back to the frozen first-party registry plus a warning, records each dropped gate-declaring overlay's declared gate as blocked (incompatibleGateCapIds / blockedGates) so a dropped blocking gate FAILS CLOSED, AND now clears _overlay.commandRoots in the fallback so no dropped overlay retains a stale command root. Previously a malformed-array throw or a compose throw escaped the loader and crashed every consumer (loop-resolver, config-loader, surface, capability-state, gsd-tools). On the source side, the capability resolver/staging now reads every untrusted capability.json (tarball / npm / git / local) via the shared bounded fd-reader (regular-file + 8 MiB cap) instead of a raw fs.readFileSync, so an oversized or FIFO/non-regular extracted-or-local manifest can no longer OOM or hang the resolver; the fetch (realHttpsGet) bounds the downloaded response to 64 MiB; and stageValidated now enforces ONE uniform aggregate byte-budget (MAX_STAGED_BUNDLE_BYTES, 128 MiB) over the STAGED bundle directory via a bounded streaming walk (cumulative byte + entry counters; symlink / non-regular entries rejected) at the common staging chokepoint — AFTER staging and BEFORE validation/promotion — so a huge source tree, git repo, npm package, or gzip/tar bomb is rejected (and its staging dir cleaned up) before promotion, uniformly bounding the RESULT of copyDirRecursive / git clone / npm pack / tar -x that were previously only timeout-bounded. copyDirRecursive itself is now STREAMING and BUDGETED: it enumerates each directory via fs.opendirSync + dir.readSync() (one entry at a time) and threads CUMULATIVE entry (MAX_STAGED_BUNDLE_ENTRIES, 100k) and byte (MAX_STAGED_BUNDLE_BYTES) counters through the recursion, failing closed the MOMENT either cap is exceeded DURING the copy — closing a residual where the former fs.readdirSync(src, { withFileTypes: true }) materialized the ENTIRE directory-entry array into memory at staging time (BEFORE the post-copy budget walk), so a hostile source whose tree held a directory of millions of tiny files (fetch < 64 MiB, but a colossal dirent array) could OOM the process during the copy before the budget could fail closed; the post-copy walk is retained as a cheap belt-and-suspenders re-verification on what actually landed in staging. The spoofable per-member tar-header size parse (parseTarMemberSize, which mis-anchored on BSD tar -tv owner/group columns such as a Jan group → fail-open) was REMOVED in favor of that non-spoofable staged-dir budget; assertSafeTarMembers keeps its unambiguous traversal / symlink / hardlink NAME and TYPE guards. (#1461) (#1475)
  • Capability ledger: fail closed on corruption, with durable atomic writes and a race-safe install lock. A corrupt or unreadable .gsd-capabilities.json is now left in place and surfaced (not silently overwritten) — install/update/remove/list/reconcile fail closed and report it, so a corrupt ledger can no longer wipe prior capabilities' tracked files and shared-config fragments (which previously left unremovable orphans in settings.json/hooks.json). Ledger writes are atomic and crash-durable (exclusive temp file + fsync of file and directory + rename, with temp cleanup on failure). The per-capability lock is race-safe: a holder is identified by (pid, process start-time, hostname), so a reused PID cannot deadlock recovery and a verifiably-live holder is never stolen, with a hard deadman timeout for unverifiable or cross-host holders. Untrusted ledger and lock reads are bounded (regular-file + size caps; FIFOs/devices rejected) and validated through a single shared entry validator (prototype-safe ids, DoS length caps). (#1462, ADR-1244.) (#1469)

fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in planner verify blocks (#1482)

/gsd:pr-branch now handles sub-repos defined in config — when planning.sub_repos is set, the command scans each sub-repo for uncommitted changes and offers to create a branch, commit, push, and open a companion PR per sub-repo. Previously, sub-repos were silently ignored because all git commands ran against the shell's current directory instead of the intended repo path. All sub-repo git operations now use git -C <repo> so no shell-state assumptions are made. (#667)

  • /gsd-new-project AI Models prompt now exposes the adaptive model profile — both onboarding paths (auto-mode and interactive) listed only Balanced/Quality/Budget/Inherit, so the adaptive profile (role-based cost optimization across Claude/Codex/Gemini/OpenRouter/local) was unreachable through /gsd-new-project despite being a first-class catalog entry and documented in CONFIGURATION.md. Both prompts now use the proven two-question split (Q1: Adaptive / Standard tier / Inherit; Q2: Quality / Balanced / Budget) already shipped for /gsd:settings (#3784), and both config-new-project example payloads list adaptive. (#1516) (#1654)
  • --raw CLI commands no longer drop stdout on the error path — a command that emitted a JSON result/error envelope and then exited non-zero previously lost all of stdout (the output-capture wrapper discarded its buffer when the command threw to set a non-zero exit); the buffer is now flushed before the error propagates. (#1457) (#1457)
  • gsd install/upgrade now recovers a malformed ~/.gsd/defaults.json instead of leaving it broken — a defaults.json containing a valid-JSON-but-non-object value (null, [], a number, or a string) bypassed the parse catch and flowed through unrecovered: null threw a TypeError (swallowed by the outer guard, logging a confusing "Could not write" warning and leaving the file as null), while []/42/"str" silently kept their broken shape on every install. The non-Claude finishInstall step now resets any non-object parse result to a fresh {} before reading/writing it, so the file is repaired and resolve_model_ids defaults normally. (#1661)
  • Shipped milestones with a retired/folded phase now reach 100% — a phase struck through in ROADMAP (marked [x], with a directory but no completion artifact) was counted in progress.total_phases but could never be counted complete, freezing the milestone below 100% (e.g. 5/6 = 83%) with state sync --verify reporting no drift. Both STATE counting paths (state json and state sync) now exclude retired phases — detected from GFM strikethrough whose subject is the phase on a checklist/heading line — from both the phase-dir set and the heading count, via the canonical phase-id helpers so numeric, decimal, and project-code IDs match alike. (#1514) (#1568)
  • Codex installs no longer run with unsafe Claude-style worktree isolation — a Codex install with a runtime-neutral .planning/config.json was resolving its runtime as Claude and enabling git worktree isolation, which Codex's spawn_agent cannot honor; the Codex fail-closed guard was also silently dead because runtime/worktree config was read JSON-quoted and broke shell equality checks. Codex-emitted workflows now resolve runtime=codex, default workflow.use_worktrees to false, and fail closed when worktrees are forced on. (#1515) (#1519)
  • Windsurf installs expose /gsd- commands in Cascade again* — Windsurf runtime installs now emit workflow files under .windsurf/workflows instead of dead skills-only artifacts. (#1615) (#1622)
  • Capability settings.json hooks no longer fire on every tool and no longer fail when non-executable — installing a capability that declared a tool-scoped PreToolUse/PostToolUse hook wrote the entry with no matcher, so a guard intended for only Write|Edit fired on every tool call (including Bash) and a fail-closed guard could block the whole session; the emitted command was also a bare script path, so a .js-family hook delivered via git/tarball that lost the executable bit failed with Permission denied on every matching call. Install now honors a declared matcher (absent = match-all, unchanged for existing capabilities) and emits a node-prefixed command for .js/.cjs/.mjs hooks so they run regardless of file-mode bits. (#1634) (#1638)
  • /gsd:secure-phase now honors the configured ASVS level and block threshold — the security auditor previously received unsubstituted {SECURITY_ASVS} / {SECURITY_BLOCK_ON} placeholder text because secure-phase.md never assigned those variables. It now resolves workflow.security_asvs_level and workflow.security_block_on from config (--raw) before the auditor handoff. (#1625) (#1633)
  • Phase transitions now require fresh canonical verification - implementation-complete phases no longer advance when verification is missing, gap-bearing, human-pending, or stale relative to phase summaries. (#1548)
  • The security audit gate now respects workflow.security_block_on severity — /gsd:secure-phase previously blocked phase advancement on any open threat regardless of severity, so the documented security_block_on threshold had no effect (and the auditor's block vocabulary didn't even match the config enum). Threats now carry a per-threat Severity (critical|high|medium|low), and only open threats at or above the configured security_block_on severity count toward the blocking gate (SECURITY.md threats_open); none disables blocking, and a missing/unparseable severity fails closed as critical. (#1626) (#1635)
  • verify codebase-drift now reads workflow.drift_action and workflow.drift_threshold from the correct nested config shape — previously both keys silently no-oped because loadConfig() returns a flattened object and config?.workflow was always undefined. (#1504)
  • check.decision-coverage-plan no longer false-passes when CONTEXT.md decisions use the titled-colon bullet form — parseDecisions recognized the colon-immediate (- **D-NN:** text) and em-dash (- **D-NN — title** body) forms but dropped the titled-colon form (- **D-NN: Title.** body, where a title sits between the colon and the closing **) via the parse-miss guard. When all decisions used the titled convention, the parser returned 0 decisions and the coverage gate passed vacuously. A third per-form regex (checked last, a strict superset of the colon form) now parses the titled-colon form; id and [tags] trackability are honored. (#1665)
  • npm version no longer leaves capability-registry.cjs stale — the version npm lifecycle script now regenerates and stages the capability registry after stamping new version strings into all capability manifests, preventing the 1.6.0-rc regression where gen-capability-registry.cjs --check failed. (#1498) (#1499)
  • Antigravity installs all GSD slash-command skills where AGY can discover them — concrete skills such as /gsd-progress and /gsd-verify-work now land directly under the Antigravity skills directory instead of router-nested folders. (#1614) (#1616)
  • frontmatter set on an object-list field now fails closed instead of silently doing nothing — setting must_haves (or another object-list field) to a value whose lossy parse projection matched the original's was a silent no-op: the command reported {updated:true} but the change never applied (the writer's scalar-only parser had flattened both to the same shape). frontmatter set now detects a no-op write for dict-valued fields and surfaces a clear error directing the user to edit the file directly. Scalars and scalar arrays round-trip faithfully, so idempotent sets of those still report {updated:true} (no false positive). (#1664)
  • config-set now rejects invalid config values instead of storing them silently — out-of-enum strings, JSON array/object coercion (e.g. ["high"] stored as an array in a scalar key), and wrong-typed values for capability-registry-owned keys are validated against each key's declared schema at set time. Previously these were accepted and persisted, mis-configuring GSD. (#1628) (#1632)
  • OpenCode and other AGENTS-native runtimes now get a root AGENTS.md from /gsd:new-project — the workflow hardcoded a codex-only branch that sent every other runtime to .claude/CLAUDE.md, a location OpenCode never loads. A shared getProjectInstructionFile(runtime) policy (claude→.claude/CLAUDE.md, codex/opencode/kilo/kimi→AGENTS.md, copilot→.github/copilot-instructions.md, antigravity/gemini→GEMINI.md) is now the single source of truth consumed by both the new-project workflow and the generate-claude-md path, with a parity test guarding drift. (#1574)
  • roadmap upgrade now rejects an unsupported or malformed --convention value (including the --convention= form) instead of silently running the milestone-prefixed migration, and no longer hard-exits inside the command-routing hub. (#1539)
  • phase complete no longer duplicates a By-Phase row when the phase number's padding differs — completing a phase by its unpadded number (e.g. phase complete 5) against an existing zero-padded By-Phase row (| 05 |) appended a second | 5 | row instead of updating it, double-counting the phase in any column sum. The row matcher now canonicalizes a numeric phase to its integer form (matching 5, 05, 005 in either direction), so the existing row is upserted regardless of padding. (#1663)
  • Non-Claude installs no longer rewrite an explicit resolve_model_ids: true to "omit" — Codex, OpenCode, Gemini, and the other non-Claude runtimes were silently clobbering the deliberate opt-in to full materialized model IDs on every install/upgrade, so generated agent manifests inherited the active chat model instead of pinning the resolved model. The finish step now only defaults resolve_model_ids to "omit" when it is absent or falsy; an explicit true is preserved. (#1569) (#1653)
  • A failed roadmap upgrade --apply now actually rolls back .planning/ even when it is gitignored (commit_docs:false), instead of reporting a successful rollback while leaving the workspace half-migrated. Rollback is surgical and no longer runs a whole-repo git reset --hard. (#1543)
  • Codex runtime no longer crashes on startup — every gsd-tools command previously aborted with Cannot find module '../../../package.json' on Codex, whose runtime root has no package.json, because a module in the loader chain did a top-level require of it. The version emitted into Hermes skill frontmatter is now sourced lazily from the installed gsd-core/VERSION (validated semver), so gsd-tools loads on every runtime and never emits version: undefined. (#1383) (#1409)
  • /gsd-* commands in Windsurf Cascade resolve their command bodies — Windsurf slash-command workflows delegate to canonical command bodies at gsd-core/commands/gsd/X.md, but the install never copied those files. Commands appeared in the / menu yet silently failed when invoked because the LLM was told to read a missing file. Installs now copy commands/gsd/*.md into the workflow delegation target. (#1630)
  • query agent-skills no longer returns empty output on Windows — the plain (non---json) path wrote the <agent_skills> block then immediately called process.exit(0), which truncated the async stdout buffer on Windows pipes/files so every ${AGENT_SKILLS_*} workflow capture expanded empty and configured per-agent skills were silently dropped. It now flushes synchronously via the same writeAllSync helper the --json path uses. (#1400) (#1410)
  • phase complete now updates the By-Phase table on CRLF (Windows) STATE.md files — the By-Phase table matcher required bare \n line endings, so a STATE.md written or hand-edited with CRLF (\r\n) was treated as having no table: the completed phase's row was never upserted (and, with the velocity-from-table derivation, the total went stale). The matcher is now CRLF-tolerant (\r?\n) on the header/separator/lookahead, so CRLF STATE.md files are handled identically to LF. (#1662)
  • clean up stale get-shit-done paths in Codex and Kimi skill mirrors on upgrade (#1453) (#1453)
  • add phase.list-plans to gsd-tools — the command was referenced in agents/gsd-plan-checker.md but was missing from the router, causing 'Unknown phase subcommand' on every invocation (#1437)
  • roadmap analyze no longer reports phantom missing_phase_details for milestone-prefixed (M-NN) phase IDs (#1552)
  • workflow.security_asvs_level now actually scales security rigor — it was display-only (the planner hardcoded ASVS L1 and the auditor only echoed the level), so L2/L3 behaved identically to L1. The configured ASVS level now scales both planner threat-disposition rigor and auditor verification depth (L1 grep-presence → L2 boundary/vector checks → L3 end-to-end trace), defined in a new references/security-asvs-levels.md; the secure-phase clean-phase short-circuit now spawns the auditor at L2/L3 so deep verification runs even when the preliminary grep classification is clean. (#1627) (#1636)
  • Atomic file writes now retry a transient rename lock on Windows (a reader holding the target open) instead of falling back to a non-atomic write that could let a concurrent reader observe a truncated STATE.md/ROADMAP.md. (#1541)
  • Misconfigured agent skills no longer fail silently — when an agent's configured agent_skills paths all fail to resolve (e.g. a missing SKILL.md), gsd-tools query agent-skills now emits an aggregate warning to stderr and adds a warnings[] field to its --json output, instead of returning an empty block with no signal. (#1376) (#1376)
  • Decision-coverage gate now reads markdown-header and em-dash decisions, and fails loud when it can't parse them — check.decision-coverage-plan (a blocking gate) and gap-analysis previously extracted zero decisions from a populated CONTEXT.md that recorded its decisions under markdown headers (## Locked decisions) or with em-dash bullets (- **D-1 — title**), and silently reported a clean pass — so real decisions went un-checked. Decisions in those shapes are now recognized, and when decision-shaped content cannot be parsed (or a - **D-NN** bullet is malformed), the gate fails loud with a format-mismatch message instead of passing. (#1386) (#1386)
  • phase complete no longer double-counts Total plans completed velocity on re-run — re-running phase complete on an already-complete phase incremented the velocity total each time (2 -> 4 -> 6 ...), because the metric re-read the cumulative total and blind-added the phase's plan count on every invocation. The total is now derived from the By-Phase table's Plans column (the same source the table upserts against), so re-completing a phase upserts the same row and the sum stays stable — and a hand-edited inflated total self-heals to the true sum on the next completion. (#1582) (#1655)
  • verify schema-drift now resolves the target phase by its canonical token instead of substring containment, so a non-existent phase no longer silently matches a token-superstring phase (e.g. "1" matching "11-expansion") and runs the drift gate against the wrong phase. (#1640)

Security

  • Prompt-injection defence extended to the untrusted-input surface (LLM-playbook principle 12) — the read-injection scanner (a pattern-based pre-filter) now also scans WebFetch/WebSearch output (closing the largest untrusted channel at ingress), and the 10 research/doc-ingest agents (issue #1577 AC #2's named eight plus gsd-ai-researcher and gsd-domain-researcher, both web-ingress) isolate fetched/read content as data-not-instructions via a shared untrusted-input-boundary reference — this prompt-level boundary is what keeps an injection from being followed. An opt-in security.injection_blocking (registered config key; default advisory, unchanged) upgrades HIGH-confidence detections to a PostToolUse circuit-breaker: since the hook runs after the fetch, decision: "block" halts the agent's next step rather than redacting the already-fetched content (it is not a redactor). Based on arXiv 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472, 2503.00061. (#1585)
  • Third-party capability trust gate (ADR-1244 Phase 4) — installing a capability from a git/npm/tarball/local source now discloses every executable surface it ships (hooks, command modules, MCP servers, with the actual commands) and requires explicit consent before anything is promoted; integrity (sha512) and engines.gsd are verified before any code is staged, install never executes capability code, and reserved gsd-/gsd-core-/anthropic- namespaces are refused. capabilities.strict_known_registries gates which sources may be installed ([] = local-only lockdown; host-based allowlist otherwise) and capabilities.auto_update is off by default, re-prompting whenever a new version's executable set changes. An install ledger makes remove surgical (strips only the capability's own shared-config entries, preserving your hand-edits) and update an atomic, crash-safe stage-then-swap. (#1449) (#1449)

[1.5.0] - 2026-06-17

Added

  • gen-capability-registry now rejects duplicate artifact producers at the same Loop Extension Point — if two capability steps declare produces: [<same artifact>] at the same point, the generator throws at gen time naming the artifact, the point, and the producing capability ids, instead of letting the topological sort pick a winner silently (which left ADR-857 Decision #6's data-flow contract undefined). The check counts distinct (capId, stepIdx) producer steps, so a single step listing an artifact twice does not false-positive. ADR-894 §4's enumerated cross-capability invariant list gains the artifact-production-uniqueness rule. (#1123) (#1131)
  • gsd-tools drift-guard — deterministic plan-drift severity/authority decisions (ADR-22). The plan-review source-grounding pass now classifies cited-symbol drift through a tested seam (5-rung authority ladder, grep→intel auto-upgrade, severity mapping, rung≥3 hard-block) instead of re-deriving the rules from workflow prose on each run. (#1190) (#1242)
  • /gsd-progress --next --auto --converge now routes planning through plan-review convergence. ADR-15's designated primary convergence surface is wired into the progress/next workflow (previously only /gsd-autonomous --converge honored it; on /gsd-progress the flag was silently dropped). Accepts --cross-ai as an alias plus reviewer flags and --max-cycles N, and is gated on workflow.plan_review_convergence. (#1190) (#1237)

gsd-tools query teams-status + a plan-phase warning detect claude-code agent-teams — GSD's multi-agent orchestration can stall under claude-code's experimental agent-teams (a subagent's completion can fail to route back to the orchestrator). A new read-only query teams-status command reports { active, runtime, env_present, source } (and --active for a clean shell guard), and /gsd:plan-phase now emits a single non-fatal warning when agent-teams is detected, recommending you disable it for GSD workflows. The detector only activates on the claude runtime with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS strictly truthy — every other runtime and the teams-off path are completely unaffected. (#1355) (#1371)

  • spec-phase: prohibition probe — a prose-orchestrated Step 5.6 that surfaces the unwritten must-NOT constraints (values/safety/ethics) a feature could silently become but the spec never forbids. Two stages per requirement: an adversarial recall question ("what could this silently become that the author would NOT want?") then a one-pass precision classifier that drops routine engineering and keeps genuine prohibitions. Confirmed prohibitions become NEGATIVE SPEC acceptance criteria carrying a test/judgment verification tier, which plan-phase lifts into the must_haves.prohibitions sibling block (truths untouched). Judgment-tier items soft-gate at verify time (never silent, never hard-halt); unwired test-tier items fail closed. The recall stage is model-driven (no compiled engine); canon-bound concerns (OWASP/GDPR/fairness) are referred to /gsd:secure-phase. Additive and optional: existing SPECs without a Prohibitions section remain valid. Second adapter of the probe-core resolution model (ADR-550 Decision 7). (#1149)
  • Optional ## Business Context section in the PROJECT.md template — a four-field block (Customer, Revenue model, Success metric, Strategy notes) for monetized or customer-facing projects, positioned between Core Value and Requirements. Optional by default (an HTML comment tells non-business projects to delete it), capped at four one-line fields to stay a constraint reference rather than a business plan, and reviewed at each milestone by /gsd-complete-milestone when present. (#72) (#756)
  • Async external jobs can now defer an Execute step legally (external_job_waiting). An Execute step that dispatches a long-running external job and commits a .planning/async-jobs/<job>.json manifest — deferring SUMMARY.md — is now recognized as a legal deferred state, not an illegal partial-plan state. execute-phase safe-resume, resume-project, and pause-work reconcile against the manifest instead of re-dispatching (which would duplicate the external compute). This defines the versioned, scheduler-agnostic manifest stability contract consumed by the core loop; the scheduler adapter that produces manifests is the capability half (#1164). (#1165) (#1221)
  • MemPalace memory capability (opt-in) — adds cross-session/cross-project recall and verbatim+temporal-KG capture at GSD loop boundaries via the MemPalace MCP server and CLI; disabled by default, skip-on-error. (#1201) (#1201)
  • spec-phase: spec-completeness edge-probe — a taxonomy-driven Step 5.5 that walks each SPEC requirement against a closed 8-category edge taxonomy (boundary, adjacency, empty/degenerate, encoding, ordering, precision, idempotency, concurrency), proposes concrete candidate edges, and resolves each to covered/dismissed/backstop/unresolved. covered edges add acceptance criteria the planner lifts into must_haves.truths; a soft gate flags unresolved edges. Additive and optional: existing SPECs without an Edge Coverage section remain valid. (#584)
  • phase uat-passed predicate — new runtime-neutral command evaluates HUMAN-UAT results with markdown-aware parsing (ignores frontmatter, fenced code, blockquotes, and HTML comments) and reports pass only when every required check passes. (#1063) (#1063)
  • Bug-report issues that lack a valid GSD Version are now auto-closed on open by a new version-gate.yml GitHub Actions workflow. GitHub Issue Forms only enforce required: true in the web UI, so issues filed via the REST API, gh issue create, or AI reporters can arrive without a version; values like idk, _No response_, or an empty field are treated as missing. Affected issues receive a closing comment with instructions to add the version (e.g. 1.18.0) and reopen; maintainers can add the version-exempt label to opt an issue out. (#1181)
  • Kimi CLI runtime support is now documented and installable — users can install global GSD Agent Skills with --kimi --global, invoke them as /skill:gsd-*, and launch the generated custom agent explicitly with kimi --agent-file. The custom-agent (--agent-file) surface targets the legacy/Python kimi-cli contract; newer Kimi Code (@moonshot-ai/kimi-code) consumes the same /skill:gsd-* skills via --skills-dir instead. (#743)
  • agent_skills can now reference Claude-Code plugin-provided skills via the namespaced global:<plugin>:<skill> form (e.g. global:coderabbit:code-review). On the Claude runtime the agent's skills block emits a by-name Skill-tool load directive that resolves the plugin skill (no plugin-cache path is read); path-resolvable skills keep the existing @-include unchanged; on non-Claude runtimes a namespaced entry is skipped with a warning. The 22 agent_skills-consumer agents now carry the Skill tool so they can load plugin-provided skills. (#1261)
  • gsd-tools capability set — turn capabilities on/off and gate hooks from one command. Adds the write side of the capability system (ADR-857/ADR-1213): capability set <id> --on|--off toggles a capability through the runtime surface (the canonical on/off switch) and --gate <key>=<true|false> toggles a hook within an enabled capability, then re-resolves and reports — so disabling a capability is consistent across surface and config ("off means off") as a write-time invariant. /gsd:settings capability hook-gates now route through it. (#1213) (#1225)

Changed

  • Added an opt-in anthropic-fable model policy provider preset for Claude Fable 5 high-budget routing while preserving the existing Anthropic Opus 4.8 defaults and anthropic provider preset. (#1014) (#1015)
  • Windsurf/Devin workspace skills now install to the canonical .devin/skills/ directory — fresh workspace installs write skills under .devin/skills/ (Devin Desktop's documented preferred location) instead of .windsurf/skills/; the legacy .windsurf/skills/ layout is still recognized. The global ~/.codeium/windsurf/skills/ path is unchanged. (#1093) (#1093)

loadCentralConfigKeys now fails loud on a malformed central config-schema instead of silently returning an empty Set — ENOENT (the schema legitimately absent) still returns an empty Set silently, but a JSON parse error or any other read failure now writes a prominent stderr warning naming the schema file and throws ExitError(1). Previously a single catch (_) swallowed parse errors too, so a merge-conflict marker or truncated write in config-schema.manifest.json made every capability config key look non-central — the config-key collision / pending-migration gate fired zero warnings and --check passed clean, defeating the gate invisibly. (#1124) (#1131)

  • Capability hook rendering now consumes resolved Capability State — gsd-tools loop render-hooks uses the same installed/surfaced/configured state reported by gsd-tools capability state, so disabling a migrated capability at the runtime surface removes its workflow hooks even when config defaults are enabled. Migrated capability config keys remain accepted through the generated capability registry/federated config path instead of duplicated central VALID_CONFIG_KEYS entries. (#1136) (#1153)

Added ADR-857 Phase 6 capstone conformance coverage so migrated Capability activation keys cannot be read directly from host loop workflows, Capability-owned config keys stay out of the central schema, and the host loop workflow size budgets remain documented. The verify-work UI automation preflight now resolves UI activation through the Capability hook registry instead of reading workflow.ui_phase directly. (#1158)

  • ADR-857 phase 6 complete: optional features are now Capabilities, not inline loop branches. tdd, schema-gate, drift, gap-analysis, and profile-pipeline are migrated out of the five-step host loop into declarative Capabilities (loop hooks + a command family); their config keys are federated to capability ownership; and the plan-phase/execute-phase workflow bodies shrink accordingly. Two previously-declared-but-dead capability gates now actually fire — the security ship-time gate (ship:pre) and the UI safety gate (execute:wave:post) — and the phase-6 conformance gate is hardened to be un-gameable (rejects empty stubs, requires loop-body shrink, verifies hook dispatch and gate-result contracts). Behavior is preserved, verified across five adversarial review passes. (#1139, #1167, #1168, #1169) (#1183)

Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable gaps_found — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic check prohibition-enforcement sub-command (authored as src/prohibition-enforcement.cts, compiled by build:lib to the gitignored gsd-core/bin/lib/prohibition-enforcement.cjs) is the missing PRODUCER: it locates the wired mechanical check, runs it for a genuine non-vacuous pass, builds enforcementEvidence, and emits the dispositionForProhibition() verdict. The previously-unreachable green branch in dispositionForProhibition() is now reachable from the live pipeline — a test-tier prohibition with a genuinely-passing wired check disposes green and can reach passed, while a missing, non-attested, or non-passing check hard-gates (flagged, never green → gaps_found) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). verify-phase.md wires the consumer; the green/fail-closed policy in src/probe-core.cts is untouched. Both wired-check kinds are accepted (ADR-550 D2): a node --test negative test (requiring a real reported test — an empty file, which node --test counts as one passing "test", does NOT green) AND a lint/AST rule run as eslint --format json filtered by ruleId (so plugin rules like local/* load — bare --rule cannot), anchored on the in-tree local/no-source-grep rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in capabilities/ (D6). (#1259)

Honest scope — failFirst is caller-attested, not yet machine-proven. This lands the execution + non-vacuous-pass half: the producer requires the caller to attest failFirst: true and the check to genuinely run and pass. It does NOT yet independently prove the check fails-on-violation (the literal regression-must-fail-first property) — cheap proof of that at verify time needs running the check against a known violation fixture, which is a tracked follow-up (#1279). The red-first property currently rests on caller attestation, surfaced transparently in the evidence record.

Correction to the issue body (#1259): the issue's "96 invalid/error negative-proof cases" figure is wrong. For the no-source-grep anchor specifically, the genuine regression-must-fail-first proofs are its two invalid cases (the .includes() and .match() blocks) in tests/eslint-rules.test.cjs — not 96. The anchor argument is unaffected (those two cases ARE real fail-first proofs); only the count was off. (#1273)

  • The test-tier prohibition gate now has a deterministic SOURCE for its wired check — a resolved test-tier must_haves.prohibitions item MAY carry an optional check descriptor authored at spec-phase: the flat-scalar keys check_kind (node-test | lint-rule), check_target, and check_rule (lint-rule only). projectProhibitions projects these scalars deterministically and verify-phase reads them back (via descriptorFromProjection) to locate the check handed to check prohibition-enforcement — so a wired, passing test closes the gap with zero manual descriptor authoring (previously the verify-phase LLM had to invent {kind, target, rule} each run, #1259). This extends the ADR-550 Decision 3 prohibition-item shape (ratified in a dated 2026-06-15 ADR-550 addendum). The descriptor is optional and fully backward-compatible — a prohibition with no descriptor parses and disposes byte-identically to today — and fail-closed: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. The descriptor is represented as flat scalars (not a nested check:{} object) to keep the shared parseMustHavesBlock round-trip regression-free. Out of scope: machine-proven fail-first (#1279) and the dispositionForProhibition policy stay unchanged. (#1278) (#1301)

Test-tier prohibition fail-first is now MACHINE-PROVEN, not caller-attested — the deferred literal regression-must-fail-first property of ADR-550 Decision 4 (the gap #1259 / PR #1273 left as a tracked follow-up) has landed. The check prohibition-enforcement producer (src/prohibition-enforcement.cts, compiled by build:lib to the gitignored gsd-core/bin/lib/prohibition-enforcement.cjs) no longer trusts the caller's failFirst attestation: before a clean, non-vacuous pass can dispose a test-tier prohibition green, the new defaultProveFailFirst prover independently RUNS the wired check against a KNOWN VIOLATION and confirms it goes RED. Attestation is gone from the green AND (passed = proof.provenFailFirst === true && run.passed === true); any other outcome — passes-on-violation, can't-prove, throws, times out, or no violation source — hard-gates in BOTH interactive and autonomous modes (ADR-550 D4 / D3). The evidence record gains a failFirstProof field recording HOW fail-first was proven (FF-07). A caller can no longer green a toothless check.

The violation is sourced from a new descriptor field, CheckDescriptor.violationFixture — an author-supplied path to a known-bad subject. For a lint-rule the prover lints that fixture and requires the rule id to appear in the JSON report (the rule must have teeth); for a node-test the prover spawns the negative test with the subject injected through the GSD_PROHIB_SUBJECT env convention and requires a NON-VACUOUS red — # fail >= 1 AND a failing test named distinctly from the file (isNonVacuousNodeTestRed), so a violation fixture that merely CRASHES the test at load is not mistaken for the negative assertion firing red (symmetric with the clean-pass non-vacuity guard). The node-test prover also requires the violationFixture to EXIST (resolved against cwd) before spawning — a missing/typo'd path fail-CLOSES rather than letting an honest test's ENOENT crash forge a green (symmetric with the lint path's file-result guard). The deterministic spec→verify path composes end-to-end: a fourth flat scalar check_violation_fixture is projected by projectProhibitions and read back by descriptorFromProjection (rides both kinds), so a prohibition authored with all four check_* scalars machine-proves fail-first and greens through the projection alone — zero hand-authoring at verify time (#1278 + #1279 + #1346; round-trip pinned by a fast-check property + CHK-03(D) + an end-to-end COMPOSE capstone). One documented residual remains under #1346: the node-test proof confirms the fixture exists and the check reds, but cannot generically prove the red was caused by the subject's content rather than by the env merely being set. The lint-rule path is fully shippable now and is dogfooded against the in-tree local/no-source-grep rule; the node-test path ships its mechanism (a fixture-bearing descriptor IS machine-proven) and is exercised by SYNTHETIC temp fixtures — there is no live in-tree node --test prohibition to dogfood. CheckDescriptor.failFirst is DEMOTED, not removed (FF-08): it is kept as a non-authoritative hint so the #1259 route-JSON shape and the CheckDescriptor type stay backward-compatible mid-migration, but no path greens on it alone. The green/fail-closed policy in src/probe-core.cts (dispositionForProhibition, reads only evidence.length > 0) is untouched; the evidence array shape is additive. This closes ADR-550's D5d follow-up — see the dated 2026-06-15 ADR-550 addendum. (#1279)

PR-review flag — GSD_PROHIB_SUBJECT + violationFixture are PROPOSED, renamable conventions. Both are net-new surface with ZERO live in-tree consumers (no in-tree node-test prohibition yet; node-test proof runs only on synthetic test fixtures, the real dogfood stays the lint-rule). They are forward-looking scaffolding, so a later rename — or replacing the env var with an argv — is a mechanical, zero-migration find/replace. Surfacing them here so the maintainer can ratify, rename, or replace them at PR review with no migration cost, exactly as #1278's ADR addendum was reviewed at PR time. The failFirst demotion is likewise open to weighing outright removal; the keep-as-hint rationale is recorded in the ADR addendum. (#1314)

  • Read-only verifier/auditor agents now ship a Claude-Code disallowedTools deny-list — the installer injects a framework-level write-tool deny-list into the Claude copies of the read-only verifier/auditor agents (gsd-verifier, gsd-plan-checker, gsd-integration-checker, gsd-doc-verifier, gsd-eval-auditor, gsd-ui-auditor, gsd-ui-checker) so write actions are blocked even if a tool grant is inherited. Injected for Claude only; other runtimes are unaffected. (#1081) (#1081)

Antigravity workspace skills now install to the canonical .agents/ directory — fresh installs write workspace artifacts under .agents/ (the Google-Codelabs-documented base) instead of .agent/; the legacy .agent/ layout is still recognized so existing installs keep working. The global ~/.gemini/antigravity/ path is unchanged. (#1090) (#1090)

devin-desktop runtime alias for the Windsurf→Devin Desktop rebrand — the windsurf runtime now also answers to devin-desktop (CLI --devin-desktop), aiding discoverability after Cognition rebranded Windsurf as Devin Desktop. All paths are unchanged — global skills still install to ~/.codeium/windsurf/skills/. (#1086) (#1086)

  • Remove dead loadConfig export from configuration.cts — superseded by config-loader.cts (ADR-857 phase 2e, #885). All live callers already import loadConfig from config-loader.cjs or the core.cjs back-compat re-export; exhaustive grep confirms zero callers importing it from configuration.cjs. configuration.cjs now provides only the pure normalization and defaults primitives (normalizeLegacyKeys, mergeDefaults, migrateOnDisk, CONFIG_DEFAULTS) that config-loader.cjs depends on. (#893) (#893)
  • audit(#779): correct stale model-catalog IDs verified against live provider sources. The gemini opus default gemini-3-pro → gemini-3.1-pro-preview (the bare gemini-3-pro ID is undefined in gemini-cli source — only gemini-3-pro-preview/gemini-3.1-pro-preview exist) and the codex sonnet default gpt-5.3-codex → gpt-5.4 (deprecated per OpenAI's Codex models page); the same two IDs are also updated in the google/openai provider-preset entries. qwen3-coder-next was verified valid (callable on Alibaba Model Studio) and left unchanged. Adds a regression guard against the retired IDs and a sourcing/verification note in CONFIGURATION.md. Catalog IDs are internal defaults; users who pinned the old IDs must update their config. (#1047)
  • INVENTORY.md no longer carries (N shipped) count scalars — the hand-maintained absolute counts collided silently on merge (two branches each bumping the same integer to N+1 while the merged tree held N+2), red-flagging CI on the merge commit across all platforms. The manifest's name-set is now the sole registry, anchors are count-free and stable, and a guard test blocks re-adding a count. (#1179) (#1179)
  • Edge-probe precision probe text now names tie-breaking / rounding-mode (half-up vs half-to-even, ceil/floor/truncate), so a surfaced precision edge cues the most common rounding failure mode. Prose-only; firing rule and the 8-category core unchanged. (#1108)
  • Capability manifests now declare runtime compatibility through a validated runtimeCompat contract, and runtime descriptor interpreters now read artifact layout, skills-home, and hook-surface facts directly from runtime Capability descriptors instead of parallel runtime-name allowlists or fallbacks. This preserves existing supported runtime behavior while making future descriptor-backed runtimes additive. (#1157)
  • Planning-time research, AI integration, and pattern mapping now participate through Capability declarations and rendered plan:pre hooks, with developer documentation for building GSD capabilities. (#1141)
  • The planner now blocks plans that would self-trip their own verify gate — when an acceptance criterion negative-greps for a literal (grep -c 'LIT' file == 0) and that same literal appears verbatim in an <action> body, plan creation now fails at write time instead of letting the executor waste cycles on a comment-text echo at commit time. Unquoted/ambiguous grep targets warn instead of failing; add <!-- planner-discipline-allow: LIT --> to allowlist a legitimate occurrence. (#1062) (#1062)
  • Namespace router skills now nest their concrete sub-skills at install time (#69). On runtimes with non-recursive skill loaders (Claude global, Cline, Qwen, Hermes, Augment, Trae, Antigravity) the installer emits the 6 gsd-ns-* routers as the only top-level skill bundles and nests the ~61 concrete skills under <router>/skills/<name>/SKILL.md, cutting the eager skill-listing overhead to ≈6 entries. Concrete skills stay reachable via the router's Read skills/<name>/SKILL.md routing table. Breaking: on those runtimes the concrete skills are no longer invocable by bare name through the Skill tool / top-level listing — route via the namespace router (or the unchanged /gsd-* slash command where a commands surface exists). Legacy top-level gsd-<concrete>/ skill dirs are removed on upgrade. Recursive/unconfirmed loaders (Cursor, Codex, Copilot, Windsurf, CodeBuddy, OpenCode, Kilo) keep the flat layout. (#883)
  • Graphify now respects surface/profile state, not just graphify.enabled — gsd-tools graphify is off unless graphify is installed AND surfaced AND graphify.enabled is true (previously only the config key was checked). The gate is now runtime-aware: Codex/Cursor/etc. read their own runtime's surface instead of ~/.claude. (#1313) (#1313)
  • Isolated-executor recovery now fails safe — when an isolated (worktree) executor run is rejected (you decline to merge it) or over-reached the requested scope, /gsd:execute-phase and /gsd:quick no longer default or propose recovery by editing the primary checkout (main). The orchestrator halts safely and offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary checkout requires explicit, clearly-labeled confirmation. (#1292) (#1303)
  • Migrate code review, security, and Nyquist verification workflows to ADR-857 capability hooks. (#1147)
  • Intel and loop-hook rendering now honor the single capability active state — gsd-tools intel gates through the shared resolver (consistency; intel stays governed by intel.enabled), and loop-hook rendering now suppresses a config-disabled capability's hooks via the capability-level active gate (fail-closed), not just per-hook when. (#1315) (#1315)
  • Added no-drift guard tests (tests/issue-57-runtime-install-no-drift.test.cjs) that protect the Runtime Install Policy Module boundary (ADR-58) and the explicit Runtime Config Adapter Registry (#60). They fail loudly when supported-runtime metadata is added to an installer call site (allRuntimes, the interactive runtimeMap menu) without a matching registry adapter entry, or when config-mutation dispatch escapes the registry's declared install surfaces — catching reintroduction of the scattered per-runtime branching those seams removed. (#867)
  • Edge-probe now surfaces a zero-classification requirement (non-empty prose, no shape cue matched, no shapes override) as a single soft unclassified — review manually candidate instead of silently dropping it. Dismissible like any edge; shapes: [] opt-out stays silent; TAXONOMY unchanged. (#1117)
  • Capability state now reports a tri-state active — gsd-tools capability state adds an active field per capability (installed && surfaced && config-enabled), alongside the existing enabled (installed && surfaced). Internal isCapabilityActive(capId, cwd) lets consumers honor the single resolved on/off answer. (#1311) (#1311)
  • gsd-verifier no longer marks behavior-dependent must-haves VERIFIED on symbol presence alone — a truth that asserts a state transition or a cancellation/cleanup/ordering invariant is marked PRESENT_BEHAVIOR_UNVERIFIED when no test exercises it: excluded from the verified_truths score, reported as a behavior_unverified count, and routed to human verification, so a clean N/N now certifies behavioral evidence rather than mere symbol presence. (#966) (#1271)
  • verify plan-structure warns on cross-task region-scope conflicts (#968) — when a plan task's file-wide negative grep (! grep -Eq 'PAT' file / grep -c 'PAT' file == 0) bans a construct a sibling task legitimately requires elsewhere in the same file, plan validation now surfaces a warning pointing to the new region/function-scoped negative-gate idiom (documented in the gsd-planner guidance and the planner-antipatterns reference, with a worked banned-in-X / required-in-Y example). Warn-only: it never errors and never changes valid. (#1320) (#1320)

Fixed

  • gsd-intel-updater now writes the canonical intel filenames the gsd-tools intel CLI actually reads — the agent was instructed to emit short names (files.json, apis.json, deps.json) and a markdown arch.md, but the intel library reads only file-roles.json, api-map.json, dependency-graph.json, and arch-decisions.json (JSON). After /gsd:map-codebase --query refresh the output was orphaned, so intel status/validate reported the files missing and intel query returned nothing. The agent now emits the canonical long names and structured arch-decisions.json. (#1000) (#1037)
  • Installer no longer appends a duplicate managed hook when it is registered via an HTTP route — a hook re-registered as a type:"http" entry (local hook-server routing) carries its identity only in url, which the installer's presence check ignored, so a stock command duplicate was appended on every install/update and the hook ran twice per event. The presence check now also inspects h.url. (#1004) (#1032)
  • /gsd-code-review's fallow structural pre-pass now actually runs and delivers findings — it invoked fallow with flags no published fallow version accepts (--json, --profile, --stdin-files), so the pre-pass failed on every run and silently degraded (the structural-findings feature never delivered on any fallow version). It now uses fallow's real CLI (audit --format json --quiet, --changed-since for phase scope, and --max-crap mapped from the code_quality.fallow.profile preset: minimal→50, standard→30, strict→15), treats fallow's exit code 1 ("issues found") as a successful run instead of a crash (gating on a valid JSON report, not the exit code), and normalizes fallow's real audit --format json schema (dead_code.*, duplication.clone_groups) into the reviewer's <structural_findings> contract. The report normalizer — previously dead code parsing a schema fallow never shipped — is wired to the real schema and exercised against real fallow output. (#1012) (#1044)
  • worktree base-check now honors a user/global worktree.baseRef:"head" (and CLAUDE_CONFIG_DIR) — base-check resolved baseRef from the project checkout's .claude/ only, so a machine-wide head set via /config (the layer the harness itself honors) was invisible. On any phase/feature lane it returned shouldDegrade:true and execute-phase silently forced sequential execution, losing the parallel worktree execution the user configured. Resolution now falls back to the user/global settings.json (via getGlobalConfigDir('claude'), honoring CLAUDE_CONFIG_DIR) below the existing project-local and project-shared layers. (#1013) (#1038)
  • Agent SDK/state/commit steps now resolve gsd-tools on shim-only installs for every runtime — source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker, …) invoked bare gsd-tools …, which fails with command not found on shim-only installs where the binary is only reachable as <runtime-home>/gsd-core/bin/gsd-tools.cjs and is not on PATH. The agent then silently skipped init/state/validate/commit ceremony. #725 fixed only Codex's conversion layer; the source agents were never migrated, so the bug persisted on Claude Code and every other runtime that consumes the source agents directly. All 12 gsd-tools-calling agents now carry the canonical multi-runtime gsd_run resolver (the same preamble the workflow launchers use — covering claude/codex/cursor/gemini/copilot/windsurf/augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes), gsd-phase-researcher's stale claude-only resolver is upgraded to the canonical one, and the launcher-parity + bare-call regression guards are extended to agents/ so no runtime can silently regress. (#1041) (#1045)
  • gsd-tools generate-claude-md no longer clobbers a hand-crafted CLAUDE.md, and defaults the Claude-runtime output to ./.claude/CLAUDE.md — /gsd-new-project wrote a repo-root CLAUDE.md full of broad project documentation, overwriting/diluting an existing hand-authored instruction file. Now: (1) an existing instruction file that contains no GSD section markers (a hand-crafted file) is left untouched and the command reports action: "skipped" — pass --force to overwrite intentionally (the flag was already parsed but ignored); (2) the default output for Claude-family runtimes is ./.claude/CLAUDE.md (a valid project-scoped memory location) instead of repo-root ./CLAUDE.md, so generated content does not pollute a repo-root file. The config default (claude_md_path), the project config template, and the new-project workflow are aligned to the new location. Codex projects still write AGENTS.md. (#1098) (#1118)
  • state record-session updates an existing ## Session Continuity section in place instead of appending a duplicate ## Session block — on a freshly bootstrapped project (workstream / gsd2-import / new-project templates all emit ## Session Continuity), the auto-create path recognised only the normalized ## Session heading, so it appended a second session block. It now inserts only the missing canonical fields after the ## Session Continuity heading, preserving the heading and any existing prose, and the snapshot / frontmatter readers recognise that heading. (The originally reported recorded:false-yet-mutated symptom was already resolved by #944/#948.) (#1101) (#1113)
  • roadmap annotate-dependencies no longer fuses the preceding summary line onto the Plans: header — when the match regex's (?:^|\n) anchor consumed a leading newline (mid-string match), the replacement dropped it, producing corrupted output like **Plans:** 3 plansPlans:. The replacement now re-emits the leading newline when present. (#1103) (#1111)
  • /gsd-progress no longer reports a phase as complete (and routes to the next phase) when its verification ended human_needed or gaps_found — routing derived completeness from plan/summary counts only and never consulted the verification.status query (the seam built in #651). A new Step 1.7 consults it for the current phase, and the routing table sends gaps_found to /gsd:plan-phase {phase} --gaps (Route V.gaps) and human_needed to /gsd:verify-work {phase} (Route V.human) before the generic complete row. passed, missing (unverified), and unknown still route as complete, so unverified phases are not falsely blocked. (#1107) (#1116)
  • write-profile now writes USER-PROFILE.md to the active runtime's config home instead of always ~/.claude — under Codex, gsd-tools query write-profile wrote ~/.claude/gsd-core/USER-PROFILE.md while Codex discuss-phase advisor-mode (installed under ~/.codex) checked the Codex home and never found it, so advisor-mode silently stayed disabled. The default output path is now resolved via the runtime-aware getGlobalConfigDir (GSD_RUNTIME / config.runtime → e.g. ~/.codex for Codex), matching how the runtime's own workflows resolve it — mirroring generate-dev-preferences. Claude is unchanged (~/.claude); an explicit --output still wins. (#1114) (#1119)
  • /gsd:review no longer produces a silent empty Codex review on codex-cli < 0.137 — the codex exec invocation passed --dangerously-bypass-hook-trust (added in codex 0.137.0) unconditionally and discarded stderr, so on older CLIs codex exited with unexpected argument before reading the prompt and the empty output was treated as a completed review. The flag is now capability-probed (codex exec --help | grep) and applied via $CODEX_BYPASS_FLAG only when supported, codex stderr is captured to a .err file instead of /dev/null, and an empty Codex output is replaced with a diagnostic so a broken reviewer is surfaced rather than silently skipped. (#1115) (#1122)
  • sandbox_mode emission in Codex TOML is now gated on the runtime descriptor's sandboxTier axis — previously installCodexConfig emitted sandbox_mode unconditionally from a hardcoded policy map regardless of whether the runtime descriptor declared a sandbox tier, making the descriptor field cosmetic. resolveInstallPlan now projects sandboxTier from the capability registry, and generateCodexAgentToml / installCodexConfig gate emission on sandboxTier !== 'none'. The per-agent mode table CODEX_AGENT_SANDBOX remains GSD agent policy (not a runtime-descriptor property). For codex (sandboxTier === 'codex-agent-sandbox') output is byte-identical to before; for all other runtimes (sandboxTier === 'none') sandbox_mode is correctly omitted. resolveInstallPlan now fails loud (throws TypeError) on a missing or invalid sandboxTier descriptor axis rather than silently coercing garbage to 'none', preventing a corrupt/stale registry from silently dropping sandbox enforcement. Full removal of the per-agent registration-tax map remains tracked under #1138. (#1151) (#1152)
  • Installed runtimes no longer silently disable verify:post gates — in a global skills-runtime install (e.g. Codex at ~/.codex), the commands/gsd source tree is absent, so capability-state resolved an empty skill manifest. The full-profile * sentinel then materialized to an empty surfaced set, marking every capability surfaced=false → enabled=false. The result: gsd-tools loop render-hooks verify:post returned activeHooks: [] even with security_enforcement and nyquist_validation enabled, so the security and Nyquist gates never fired. Capability-state now falls back to the installed <configDir>/skills/gsd-*/SKILL.md layout when the source tree is unreachable, so verify:post again includes security -> secure-phase and nyquist -> validate-phase. (#1206) (#1206)
  • gsd install no longer warns that settings.local.json "may be malformed" when the file contains a valid JSON null. readSettings now treats a successfully-parsed null as empty settings ({}) instead of collapsing it into the parse-failure path, so a literal-null settings file is preserved silently; genuinely unparseable files still emit the warning. (#1191) (#1233)
  • gsd-tools no longer crashes at load on a fresh install — the installer omitted scripts/fix-slash-commands.cjs, which command-roster requires at module load, so every gsd-tools command failed with MODULE_NOT_FOUND. The installer now ships it (with a smoke assertion), and readCmdNames() tolerates a missing commands directory. (#1240) (#1240)
  • state begin-phase / complete-phase now advance the frontmatter status for pipe-table STATE.md, not only inline Status: files. The status update matched the YAML frontmatter status: line first and never updated a body | Status | … | cell, so the frontmatter status froze (e.g. stuck at planning); it now transitions correctly (planning → executing → completed) regardless of whether the body Status is inline or pipe-table. (#1255) (#1256)
  • state planned-phase now advances the pipe-table Status cell (and frontmatter status), and state begin-phase now updates the Current Position | Phase | / | Plan | cells instead of prepending stray inline lines. Systemic follow-up to #1255: planned-phase ran its body-field replacements on the full file content, so the YAML frontmatter status: line was matched before the body | Status | … | cell and the status never reached Ready to execute; and begin-phase had pipe-table branches only for Status/Last activity, so for pipe-table STATE.md the Phase/Plan rows were left stale while a spurious inline Phase: N — EXECUTING line was prepended. Both handlers now strip frontmatter before body-field replacement and update pipe-table cells in place, matching the inline-format behaviour. (#1257) (#1260)

Parallel worktree execution now has executor-authored cleanup metadata — executor agents capture their worktree path, branch, and expected base before task commits and return a parseable metadata block for execute-phase to prefer over runtime harness metadata. (#1297) (#1349)

UAT resume now accepts paused checkpoints — uat render-checkpoint treats a non-structured paused Current Test placeholder as a resume signal and derives the checkpoint from the first pending UAT test instead of failing as malformed. (#1300) (#1350)

phase complete now preserves prose-block STATE phase names — template-shaped Current Position prose now advances with the next phase name, avoids missing-field warnings, and keeps Last activity: on the template em-dash delimiter. (#1316) (#1351)

Claude skill installs now avoid rejected xhigh effort frontmatter — heavyweight GSD skills now ship with portable effort: max, and the Claude skill converter normalizes any remaining xhigh source effort before writing SKILL.md. (#1319) (#1352)

Glued letter-prefix phase directories now resolve correctly -- phase lookup now recognizes tokens like P0.3 and M1-2 from directory names, so phase commands can find their plans instead of reporting none found. (#1324) (#1353)

Update backups now ignore preserved shared skills and hooks -- /gsd-update custom-file detection now mirrors installer cleanup scope for shared runtime roots, so non-gsd-* skills and hooks are not copied into backup folders unnecessarily. (#1325) (#1354)

Codex skills no longer show up twice in autocomplete — GSD's Codex install wrote an agents/openai.yaml sidecar under every managed gsd-* skill directory, and recent Codex builds index both SKILL.md and the sidecar, so each skill appeared twice (once as gsd-foo, once as a humanized foo display name). The installer now stops emitting these sidecars and removes stale ones left by prior installs (pruning the empty agents/ directory), while preserving user-owned skill directories. Codex discovers GSD skills via SKILL.md alone. (#1326) (#1360)

The worktree path guard no longer blocks ordinary writes in non-GSD git worktrees — the gsd-worktree-path-guard PreToolUse hook fired for every Write/Edit in any linked git worktree, so Claude Code plan-mode writing its plan to ~/.claude/plans/<slug>.md from a manually-created worktree was hard-blocked. The hook now only enforces inside a GSD isolated-executor worktree (branch worktree-agent-*) and fails open when a target resolves to no git repository, while still blocking writes that escape to a different git root (the #260 protection) or into a repository's .git internals. (#1342) (#1361)

check.decision-coverage-plan no longer reports a false pass when a D-NN decision header has text before the colon — parseDecisions previously dropped any - **D-NN …:** bullet whose header contained a (parenthetical), em-dash, or other prose before the :**, silently narrowing the trackable set so the blocking coverage gate green-lit a phase whose dropped decisions were never checked. The parser now tolerates a freeform run before the colon (preserving [bracket] tags) and warns on any D-NN bullet it still cannot parse instead of dropping it. (#1343) (#1358)

Codex hooks.json is now always written in the nested { "hooks": { … } } shape Codex expects — the writer previously echoed back whatever shape it read, so an empty, absent, or legacy top-level hooks.json ({ "SessionStart": [...] }) stayed in the legacy shape that current Codex can reject or warn on. Every write now canonicalizes to the nested form, lifting any legacy top-level event entries (including mixed nested+top-level files) under hooks without dropping user-owned entries. Managed-hook dedup/removal is unchanged. (#1348) (#1363)

gsd install --cursor no longer leaves bare ~/.claude paths in installed artifacts — the Cursor install branch only rewrote the trailing-slash .claude forms, so bare ~/.claude / $HOME/.claude references survived into installed skills and workflows (e.g. gsd-surface, gsd-graphify, plan-phase, autonomous) and tripped the post-install "unreplaced .claude path reference(s)" warning, pointing at a directory that doesn't exist on a Cursor-only install. The Cursor branch now rewrites bare forms too (mirroring the Trae/Augment/Copilot branches), using a (?![\w-]) lookahead so .claude-plugin / .claudeignore are not corrupted. (#1356) (#1368)

  • /gsd-new-project and /gsd-new-milestone now self-heal when the research synthesizer returns SUMMARY.md inline instead of writing it — under some context loads the gsd-research-synthesizer agent hits an LLM false-refusal (fabricating a non-existent write restriction) and returns the SUMMARY.md content in its reply rather than writing .planning/research/SUMMARY.md. Prompt hardening (#240) reduced but did not eliminate this. Both workflows now verify the file exists after the synthesizer returns and, if it is missing but content came back inline, the orchestrator persists it before spawning gsd-roadmapper — so the roadmapper never fails with "SUMMARY.md not found". (#222) (#1042)
  • Codex agent TOML generation no longer pins model_reasoning_effort when the agent is intentionally inheriting the active Codex chat model. GSD still emits both model and model_reasoning_effort when a per-agent model override or runtime: "codex" resolver pins the model, avoiding the confusing partial state where the model followed Codex UI selection while effort followed GSD catalog defaults. (#838) (#842)
  • profile-pipeline temp output now lands under the reaped GSD temp root. cmdExtractMessages and cmdProfileSample previously created their output directories directly in os.tmpdir() root (gsd-pipeline-* / gsd-profile-*), which reapStaleTempFiles never scans (it only scans GSD_TEMP_DIR = os.tmpdir()/gsd). The directories accumulated forever. Both sites now call ensureGsdTempDir() and create under GSD_TEMP_DIR. Also adds missing after/afterEach teardown to four test fixtures that leaked gsd-* temp dirs on every npm test run. (#866) (#879)
  • getMilestonePhaseFilter now excludes phase headings inside fenced code blocks ( ``` or ~~~) — consistent with the fence-aware behavior of extractCurrentMilestone. Previously, a ### Phase N: line inside a fenced block was wrongly counted as a real phase. (#875) (#880)
  • gsd_run launcher shim now probes all non-Claude runtime homes before failing. The shim's last-resort detection previously stopped at $HOME/.claude, causing a false-positive fatal error on every non-Claude runtime (Hermes, Cursor, Codex, Copilot, Windsurf, Augment, Trae, Qwen, CodeBuddy, Cline, Grok, Antigravity, OpenCode, Kilo) when RUNTIME_DIR was unset and gsd-tools was not on PATH. The snippet now probes each runtime's config directory (respecting HERMES_HOME, CURSOR_CONFIG_DIR, CODEX_HOME, etc. with sensible $HOME-relative defaults) before emitting the install error. (#903)
  • validate health and validate consistency no longer emit false-positive W007 warnings for projects using checklist-style ROADMAP.md phases. buildRoadmapPhaseVariants() in src/validate.cts previously used only a heading-style regex (## Phase N: name), silently ignoring the supported checklist format (- [x] **Phase N: name**). This caused every on-disk phase directory to trigger W007 ("exists on disk but not in ROADMAP.md") when the project's ROADMAP used checklist-only notation. The fix adds a second regex pass mirroring the existing buildNotStartedPhaseVariants() approach. Additionally, cmdValidateConsistency() in src/verify.cts had a duplicate inline heading-only regex with the same gap — refactored to delegate to buildRoadmapPhaseVariants() (DRY). (#892) (#893)
  • init execute-phase and cmdCommit now produce correct branch_name when project_code is set — the {phase} substitution in phase_branch_template now calls normalizePhaseName(), stripping the project-code prefix and zero-padding the number, so the generated branch is e.g. gsd/phase-01-foundation instead of gsd/phase-CK-01-foundation. Both the execute-phase output path (src/init.cts) and the pre-execution commit path (src/commands.cts) are fixed. (#904) (#904)
  • syncStateFrontmatter no longer strips current_phase, current_phase_name, current_plan, and progress from STATE.md — when body annotations are absent (e.g. after an agent rewrites the body), the existing frontmatter values for those scalars are now preserved, mirroring the fallback already applied in cmdStateJson. (#905) (#905)
  • Top-level Claude Code /gsd-plan-phase now always spawns the researcher/planner/plan-checker agents instead of collapsing them inline — a <runtime_compatibility> block after </available_agent_types> makes the Agent-availability requirement explicit and documents that the workflow fails-closed (stops with a clear log message) in genuinely Agent-less contexts; seven "ORCHESTRATOR RULE — CODEX RUNTIME" labels are renamed to "ALL RUNTIMES" so the guard applies universally; execute-phase.md scopes its existing "Other runtimes" inline-fallback prose to non-Claude contexts, preserving the #853 backgrounded-agent behaviour. (#913) (#913)
  • /gsd-plan-phase, /gsd-execute-phase, /gsd-autonomous no longer carry context: fork — these are spawning orchestrators; a forked subagent context has no Agent tool, preventing them from spawning the subagents they require. effort: xhigh is preserved. Fixes /gsd:autonomous halting with "running as a forked subagent" on 1.4.1 (#921). Also replaces the introspection-based Agent-availability check in plan-phase's <runtime_compatibility> block with an attempt-based gate: the workflow now always attempts the Agent() call and only stops if a real tool-unavailable error is returned, eliminating false-negative aborts in top-level sessions (#922). (#921)
  • Claude global install reverted to flat skill layout so concrete skills are discoverable. PR #883 introduced nested skill layout for Claude at ~/.claude/skills/gsd-ns-<router>/skills/<stem>/SKILL.md, but Claude Code's skill discovery scans only one level under ~/.claude/skills/ — nested concrete skills were never listed in the Skill-tool available-skills list and direct Skill(skill="gsd-plan-phase") calls stopped working. This fix reverts Claude to the flat layout (~/.claude/skills/gsd-<name>/SKILL.md) so all ~61 concrete skills are top-level and immediately discoverable. The 6 other runtimes that confirmed non-recursive scanning (cline, qwen, hermes, augment, trae, antigravity) retain their nested layout. (#924) (#924)
  • gsd-context-monitor.js now echoes the actual invoking hook event name — instead of hardcoding hookEventName: "PostToolUse" (or "AfterTool" for Gemini), the hook reads data.hook_event_name from the stdin payload and falls back to the runtime heuristic only when the field is absent or blank; this fixes Claude Code rejecting hook output with "expected Stop but got PostToolUse" when the monitor is invoked by the Stop, SubagentStop, or PreCompact hooks registered in PR #821. (#925) (#926)
  • Fix --reapply verifier false-positives on post-#604-rename installs caused by two gaps in pristine-baseline handling:

Gap 1 (verify-reapply-patches.cjs): when backup-meta.json records a pristine_hash for a file but gsd-pristine/ has no corresponding snapshot on disk, the verifier fell to over-broad mode (every upstream-changed line treated as a user-added requirement) and produced FAIL_USER_LINES_MISSING false positives. Fix: return advisory OK_NO_BASELINE reason (non-blocking, exit 0) when a recorded hash is present but the pristine file is absent — the verifier cannot reason correctly without a baseline and must not block.

Gap 2 (new migration 004-prune-stale-pristine-get-shit-done): migration 003 removed legacy get-shit-done/ runtime files but left gsd-pristine/get-shit-done/ orphan snapshots in place. Those stale snapshots referenced get-shit-done/... key paths that no longer match the active gsd-core/... layout, contributing to FAIL_INSTALLED_MISSING false reports. Fix: add a new migration (not editing 003, to preserve its checksum) that removes all files under gsd-pristine/get-shit-done/. (#934) (#935)

  • /gsd-update changelog preview no longer silently fails — the installer now copies scripts/changeset/ and scripts/lib/ into the runtime config dir so $GSD_DIR/scripts/changeset/cli.cjs resolves at runtime; update.md was updated to use the correct installed path and to surface an explicit error if the CLI is missing rather than swallowing it. (#935)
  • plan-review-convergence now runs gsd-plan-phase inline instead of inside Agent() — both sites that previously wrapped gsd-plan-phase in Agent() (initial planning + replan loop) have been changed to bare Skill() calls at depth 0. On Claude Code, a depth-1 Agent has no Agent tool, so a wrapped plan-phase could never spawn gsd-planner or gsd-plan-checker — the replan loop silently failed to produce a revised plan whenever HIGH concerns were found. Running plan-phase inline from the depth-0 orchestrator (which retains the Agent tool) restores the full planner→checker sub-agent chain. A new structural guard test (bug-936-no-nested-spawner-wrap.test.cjs) statically scans all workflow files and fails if any workflow wraps a spawner orchestrator in Agent() without a RUNTIME != claude carve-out, preventing regression. (#936) (#939)
  • --json-errors now emits a structured error even when a handler throws unexpectedly — an unexpected (non-ExitError) throw fell through to a raw stack trace on stderr, breaking SDK structured-error parsing. (#965) (#987)
  • verify key-links docs now correctly state from:/to: are relative file paths — the reference implied component/endpoint values the verifier never supported, so locator-style links failed with a misleading 'Source file not found' and the author's pattern: was never evaluated. (#967) (#990)
  • Fixed a test-infrastructure regression (#996) where bug-969 hardening tests deleted the shared gsd-core/bin/lib/core.cjs during concurrent runs and the build tsbuildinfo lived inside the copied install tree, intermittently failing CI with MODULE_NOT_FOUND/ENOENT. The destructive tests now run hermetically against a temp project, and the tsbuildinfo moved out of gsd-core/bin/. (#969) (#1002)
  • gsd-planner now ships the Edit tool, so it can no longer destroy ROADMAP.md via a whole-file Write — the planner had Write but not Edit (the #571/#581 writer-agent gap), so an in-place ROADMAP edit fell back to a full overwrite that truncated committed milestone history. The update_roadmap step now directs scoped Edit calls and explicitly forbids passing the full file to Write. (#973) (#989)
  • graphify query --budget with no value now errors instead of silently ignoring the budget — a trailing --budget parsed as NaN and was treated as 'no budget', so the query ran unbounded with no warning. (#974) (#986)
  • The installer now resolves a stable fnm node path instead of the ephemeral multishell shim on Windows — managed .js hooks were pinned to fnm_multishells/<id>/node.exe, a per-shell-session path fnm later deletes, breaking every managed hook until reinstall. (#977) (#992)
  • gsd-tools milestone complete --force now actually overrides the unstarted-phase guard — the dispatcher never parsed --force, so the guard's own documented escape hatch was inert. (#978) (#982)

Trae and Windsurf installs no longer leak unreplaced ~/.claude / $HOME/.claude paths — both converters only rewrote trailing-slash .claude/ forms, so bare home-path references survived conversion and pointed users at the wrong config dir; bare forms are now rewritten (Codex/Cline #570/#782 parity) and CLAUDE_CONFIG_DIR maps to the runtime's own var, with .claude-plugin preserved. (#983) (#995)

  • Claude Code plugin installs no longer fail with empty @~/.claude/gsd-core/... includes — agents, commands, and templates @-include the canonical ~/.claude/gsd-core/ path, but a marketplace plugin install (claude plugin install) never creates that directory, so every include resolved to nothing and agents (e.g. the executor) failed. A new SessionStart hook (gsd-ensure-canonical-path.js) symlinks the canonical path's immutable subdirs (bin, contexts, references, templates, workflows) to the plugin's bundled tree. It is a no-op in classic bin/install.js installs, preserves user-generated files (e.g. USER-PROFILE.md), prunes stale links so it self-heals after claude plugin update, and uses Windows junctions. (#1207) (#1207)
  • /gsd-code-review, /gsd-code-review --fix, and /gsd-eval-review now inject configured agent_skills into their subagents — these review-family workflows previously spawned their reviewer/fixer/auditor agents (including the --auto re-review/re-fix loops) without the project-configured skill and rule context, so any agent_skills set for gsd-code-reviewer, gsd-code-fixer, or gsd-eval-auditor were silently ignored. They now query and inject those skills like the ~20 sibling workflows. (#1005)
  • phase complete no longer rewrites an existing roadmap completion date — repeat runs on an already-Complete phase preserve the recorded YYYY-MM-DD date (4- and 5-column layouts); empty/-/non-date cells are still stamped with the current date. (#1177)
  • Legacy ROADMAP projects no longer get deprecation-warning spam — the free-form ROADMAP warning fired on every command regardless of phase_id_convention; it now only warns when the milestone-prefixed convention is explicitly set and unmet. (#1218) (#1218)
  • Forking workflows target wrong base branch on master repos when origin/HEAD is unset — execute-phase, quick, ship, complete-milestone, and pr-branch detection bash fell through to a hardcoded main fallback whenever origin/HEAD was absent (common in git init + remote add + fetch without set-head, CI checkouts, and worktrees), causing GSD to fork phase branches off a non-existent main on master repos. Replaced with a single gsd_run query git.base-branch resolver that walks the full precedence ladder: config override → origin/HEAD symref → git remote show origin → local branch presence → "main". (#1198) (#1198)
  • query user-story.validate now works — mvp-phase and verify-work workflows both invoked this command to validate "As a / I want to / so that" user stories, but no CJS handler existed; every call errored with "Unknown command: user-story". (#1193) (#1193)
  • Context meter no longer sticks at 100% — the statusline reserved-buffer math was inverted, pinning usage at 100% whenever CLAUDE_CODE_AUTO_COMPACT_WINDOW equalled the total window. (#1194) (#1211)
  • Roadmapper honors phase_id_convention — new-project roadmaps now use milestone-prefixed phase IDs when phase_id_convention is set, instead of ignoring the default. (#1205) (#1215)

phase complete no longer emits false warnings from historical verification metadata or deferred requirement IDs — two distinct false-positive warning bugs: (A) the verification-status check used a full-text regex that matched previous_status: gaps_found in the file body, triggering an "unresolved gaps" warning even when the current frontmatter status: passed; the check now reads only the frontmatter status key via extractFrontmatter. (B) requirement IDs under explicitly deferred/backlog/future/v2 section headings in REQUIREMENTS.md were flagged as missing from the Traceability table; the check now skips any section whose heading matches those terms. (#1197) (#1197)

  • verify key-links no longer fails on planned future files — a from: link whose file is declared in a current/upcoming wave plan’s files_modified is now reported pending instead of a hard missing-file failure. (#1202) (#1219)
  • state patch and state record-session no longer corrupt STATE.md — a no-match patch no longer rewrites the file (was resetting milestone_name and resurrecting a stale stopped_at), and record-session now persists --stopped-at/--resume-file even when the body lacks the exact labels. (#952)
  • /gsd-update no longer flags managed-hooks-registry.cjs as a custom file — the shipped hook is now recorded in the file manifest, eliminating a perpetual false-positive custom-file warning. (#953)
  • gsd-tools no longer throws EAGAIN or truncates output under heavy load — the CLI's stdout/stderr writes now retry the transient EAGAIN/EINTR errnos and handle short writes when the output stream is a full non-blocking pipe (e.g. the parallel test runner), instead of throwing or silently dropping bytes. (#1009)
  • Quick worktree execution now accepts parent-or-plan bases for pre-dispatch plan commits — quick mode records the parent and plan commit around the pre-dispatch PLAN.md commit, lets the worktree guard accept either approved base, materializes the plan from git objects when a runtime forks from the parent, and teaches cleanup to validate the same allowed-base set. (#1265) (#1347)
  • phase add no longer reuses an existing phase number when that phase exists only as a roadmap bullet — the next-number scan now counts phases listed only as - [ ] **Phase N: ...** bullets (all checkbox variants, with or without a title), in addition to ### Phase N: section headers and on-disk phase directories, so a bullet-only phase is no longer shadowed and phase add appends after the highest used number. (#1249)
  • Preserve curated STATE.md progress frontmatter when state patch updates non-progress fields, while still allowing progress-related fields to resync from disk-derived project state. (#1345)
  • The installer no longer re-adds a duplicate managed hook when the user registered it in command+args (wrapped) form — the presence checks only inspected h.command, so an args-form wrapper (a common Windows windowless-launcher mitigation) was invisible and a stock entry was appended on every install/update, running the hook twice. (#976) (#994)
  • cmdSkillManifest now discovers concrete skills nested under gsd-ns-* routers (<root>/gsd-ns-<router>/skills/<stem>/SKILL.md), so gsd-health and gsd-settings report the correct count on nested-layout runtimes (cline, qwen, hermes, augment, trae, antigravity). The scan is scoped to gsd-ns-* router dirs only — unrelated user dirs that happen to have a skills/ subdirectory are not traversed. Dual-routed concretes (same skill installed under two routers) are deduped by name within each root. (#929) (#929)
  • state record-session no longer pins a CPU core forever — acquireStateLock busy-spun at 100% CPU when a recoverable errno (e.g. ENOENT from a removed worktree) persisted, because that retry path skipped the backoff sleep and the 30s time budget. Every retry path is now bounded and backed off. (#1236) (#1236)
  • /gsd-manager and /gsd-autonomous --interactive no longer silently skip worktree isolation and independent verification on Claude Code. They dispatched plan/execute as background agents, but a backgrounded Claude Code agent has no Agent/Task tool and cannot spawn the nested executors, plan-checker, or verifier — so isolation and verification silently never ran. Both workflows now resolve the runtime and run plan/execute inline on Claude Code; background dispatch is kept on runtimes that support nested subagents. (#863)
  • Researcher agents can now invoke Perplexity — gsd-phase-researcher and gsd-project-researcher referenced mcp__perplexity__* in their provider dispatch tables but never granted it in their tools: allowlist, so Perplexity web research silently fell through to the next provider. The grant is now generated from the researcher profiles, with a parity guard that fails if a future dispatch-table provider is added without its tool grant. (#1284) (#1288)
  • Init phase lookups now resolve active phases whose canonical details live in a flat Phase Details block outside the current milestone summary, restoring requirement coverage for plan/execute/phase-op flows. (#1344)
  • Installer no longer leaks gsd-cmd-rewrites-* temp directories. Each install that emitted slash commands left one fs.mkdtempSync directory under the system temp root; on tmpfs /tmp hosts these accumulated and consumed RAM-backed storage. installRuntimeArtifacts() now removes the temp copy in a finally once command files are copied. (#862)
  • validate agents (and validate health) now cross-reference the install manifest to detect manifest-backed Codex agent pair drift: when a generated agents/gsd-*.md / agents/gsd-*.toml pair has one side missing on disk, the agent is reported as incomplete and agents_found is false (previously a false-healthy agents_found: true, missing: []). validate health names the incomplete agents and recommends re-running the installer. The check no-ops when no manifest is present. (#1058) (#1079)
  • The map-codebase and docs-update workflows no longer collect background sub-agent results with the deprecated Claude Code TaskOutput tool — they keep run_in_background=true on the spawn and Read each agent's outputFile (from the async_launched result) once it reports completion, removing the TaskOutput(block=true) main-session hang surface (anthropics/claude-code#20236). Completion-marker contracts and on-disk verification are unchanged, and the non-Claude runtime fallbacks are preserved. (#1362)
  • model_policy is now honored on the default claude runtime — including the anthropic-fable Claude Fable 5 preset. Policy-resolved model IDs map to Claude Code agent aliases (e.g. claude-fable-5 → fable), and IDs without a Claude alias warn and fall back to the configured tier. Forward-port of #1133 (originally shipped on the 1.4.5 hotfix line). (#1133) (#1133)
  • /gsd-plan-review-convergence now blocks on actionable review findings outside PLAN.md (#724). The convergence summary contract includes current_actionable alongside current_high, and reviews-mode planning/checking requires actionable MEDIUM/LOW feedback to be incorporated or explicitly deferred in executable PLAN.md content. (#728)
  • Config docs/prompts now match the consumers — workflow.subagent_timeout is documented in milliseconds (default 300000), not "seconds (default 600)" (a user who entered 600 got a 600 ms timeout); review.models.<cli> is documented as a bare model id injected into --model/-m, not a shell command; and workflow.test_command / workflow.build_command (consumed by verify-phase, execute-phase, audit-fix, and the post-merge gate) are now accepted by config set and documented. (#1296) (#1299)
  • changeset new --pr 0 now accepted at creation — the required-field guard treated the integer 0 as a missing --pr flag, so the documented pr: 0 placeholder could not be authored via the CLI. (#1231) (#1231)
  • $gsd-quick Codex adapter no longer assumes typed spawn_agent(agent_type=...) — documents that typed planner/executor spawning needs the agent_type-capable Codex schema and provides a clearly-labeled generic-subagent fallback when only multi_agent_v1 is exposed. (#958)
  • state update and roadmap update-plan-progress now handle current Markdown artifact shapes — state field read/replace works on table-format STATE.md (| Status | … |), and roadmap update-plan-progress inserts missing per-plan checklist rows (filling partial gaps), tolerates Plans:/**Plans:**/**Plans**:, and scopes changes to the active milestone. (#1172)
  • state planned-phase now advances the Status field when the prior phase left a Complete ✓ (checkmark) or bare Complete terminal status. Previously such a status matched no known template default, so the transition was silently skipped and the state machine stayed stuck on the prior phase. Caveat-bearing statuses (e.g. Complete but needs manual QA) remain preserved. (#1070) (#1078)
  • state.* writes no longer silently revert the STATE.md frontmatter status/stopped_at — an incidental write (e.g. state record-session) that doesn't change the body's Status:/Stopped at: source field now preserves the existing frontmatter value instead of re-deriving it from possibly-stale body text. Legitimate transitions (e.g. begin-phase/complete-phase, which do update the body Status) still re-derive normally, so a verified-complete phase can no longer be flipped back to verifying by an unrelated write. (#1252)
  • audit-open no longer false-flags completed quick tasks — quick-task SUMMARYs now carry status: complete in frontmatter by construction, so the milestone-close auditor stops reporting finished quick tasks as [unknown]. (#951)
  • Workspace (local) Antigravity and Copilot skill installs no longer point at the global config home — a local install rewrote ~/.claude/ references in SKILL.md bodies to the global ~/.gemini/antigravity/ / ~/.copilot/ paths instead of the workspace-relative .agent/ / .github/, because the skills layout wrapper passed the runtime name into the converter's isGlobal parameter slot. (#1092) (#1092)
  • Fix the workflow gsd_run launcher being unreachable in later bash blocks on runtimes that run each fenced block in a fresh shell (e.g. Claude Code): ship a standalone gsd-core/bin/gsd_run executable and have the per-file preamble persist the launcher's bin dir onto PATH via CLAUDE_ENV_FILE, with the inline function definition kept as the fallback for all other runtimes. (#1084)
  • /gsd:phase insert and /gsd:phase --edit no longer dead-end recording Roadmap Evolution — query state.add-roadmap-evolution was rejected as "SDK-only" with an error that pointed back at the very command that just failed, and no CJS handler existed after the SDK retirement. The handler is now implemented in CJS, so the insert/edit phase workflows append the ### Roadmap Evolution entry under ## Accumulated Context (creating the subsection if missing, deduping identical entries) as documented. (#1148) (#1148)
  • Corrected the installer --help profile skill counts: core now shows 8 (was 7) and standard shows 14 (was 13), both derived from PROFILES so they can't drift again; the full line drops the stale hardcoded 66 for all skills. (#834) (#847)
  • Codex-installed GSD skills and agents no longer rely on a bare gsd-tools executable — generated Codex surfaces now call the bundled shim, and workflow launchers can resolve the Codex shim-only install path. (#731)

Wire the discuss loop step for capability hooks — capabilities can now register discuss:pre/discuss:post hooks (e.g. discuss-time context recall and CONTEXT capture); previously discuss was contract-declared but structurally unwireable. Also collapses the host-loop file set to a single source of truth and adds an authoring-time guard rejecting hooks at unwired extension points. (#1199) (#1199)

  • /gsd-autonomous --converge now routes phase planning through plan-review convergence instead of silently ignoring the flag. (#711) (#729)
  • Hermes skills now install at skills/gsd/gsd-/SKILL.md with name gsd-, restoring canonical /gsd- dispatch that was broken by the bare-stem prefix introduced in #3664. (#955)

[1.4.5] - 2026-06-12

Fixed

  • model_policy is now honored on the default claude runtime — including the anthropic-fable Claude Fable 5 preset. Policy-resolved model IDs map to Claude Code agent aliases (e.g. claude-fable-5 → fable), and IDs without a Claude alias warn and fall back to the configured tier. Previously the entire model_policy block was silently ignored on claude. (#1133) (#1133)

[1.4.4] - 2026-06-11

Changed

  • Added an opt-in anthropic-fable model policy provider preset for Claude Fable 5 high-budget routing while preserving the existing Anthropic Opus 4.8 defaults and anthropic provider preset. (#1014) (#1015)

[1.4.3] - 2026-06-09

Fixed

  • Fix --reapply verifier false-positives on post-#604-rename installs caused by two gaps in pristine-baseline handling:

Gap 1 (verify-reapply-patches.cjs): when backup-meta.json records a pristine_hash for a file but gsd-pristine/ has no corresponding snapshot on disk, the verifier fell to over-broad mode (every upstream-changed line treated as a user-added requirement) and produced FAIL_USER_LINES_MISSING false positives. Fix: return advisory OK_NO_BASELINE reason (non-blocking, exit 0) when a recorded hash is present but the pristine file is absent — the verifier cannot reason correctly without a baseline and must not block.

Gap 2 (new migration 004-prune-stale-pristine-snapshots): migration 003 removed legacy get-shit-done/ runtime files but left gsd-pristine/get-shit-done/ orphan snapshots in place. Those stale snapshots referenced get-shit-done/... key paths that no longer match the active gsd-core/... layout, contributing to FAIL_INSTALLED_MISSING false reports. Fix: add a new migration (not editing 003, to preserve its checksum) that removes all files under gsd-pristine/get-shit-done/. (#934) (#937)

  • /gsd-update changelog preview no longer silently fails — the installer now copies scripts/changeset/ and scripts/lib/ into the runtime config dir so $GSD_DIR/scripts/changeset/cli.cjs resolves at runtime; update.md was updated to use the correct installed path and to surface an explicit error if the CLI is missing rather than swallowing it. (#938)
  • plan-review-convergence now runs gsd-plan-phase inline instead of inside Agent() — both sites that previously wrapped gsd-plan-phase in Agent() (initial planning + replan loop) have been changed to bare Skill() calls at depth 0. On Claude Code, a depth-1 Agent has no Agent tool, so a wrapped plan-phase could never spawn gsd-planner or gsd-plan-checker — the replan loop silently failed to produce a revised plan whenever HIGH concerns were found. Running plan-phase inline from the depth-0 orchestrator (which retains the Agent tool) restores the full planner→checker sub-agent chain. A new structural guard test (bug-936-no-nested-spawner-wrap.test.cjs) statically scans all workflow files and fails if any workflow wraps a spawner orchestrator in Agent() without a RUNTIME != claude carve-out, preventing regression. (#936) (#939)

[1.4.2] - 2026-06-09

Fixed

  • /gsd-plan-phase, /gsd-execute-phase, /gsd-autonomous no longer carry context: fork — these are spawning orchestrators; a forked subagent context has no Agent tool, preventing them from spawning the subagents they require. effort: xhigh is preserved. Fixes /gsd:autonomous halting with "running as a forked subagent" on 1.4.1 (#921). Also replaces the introspection-based Agent-availability check in plan-phase's <runtime_compatibility> block with an attempt-based gate: the workflow now always attempts the Agent() call and only stops if a real tool-unavailable error is returned, eliminating false-negative aborts in top-level sessions (#922). (#921)
  • gsd-context-monitor.js now echoes the actual invoking hook event name — instead of hardcoding hookEventName: "PostToolUse" (or "AfterTool" for Gemini), the hook reads data.hook_event_name from the stdin payload and falls back to the runtime heuristic only when the field is absent or blank; this fixes Claude Code rejecting hook output with "expected Stop but got PostToolUse" when the monitor is invoked by the Stop, SubagentStop, or PreCompact hooks registered in PR #821. (#925) (#926)

[1.4.1] - 2026-06-09

Changed

  • Added no-drift guard tests (tests/issue-57-runtime-install-no-drift.test.cjs) that protect the Runtime Install Policy Module boundary (ADR-58) and the explicit Runtime Config Adapter Registry (#60). They fail loudly when supported-runtime metadata is added to an installer call site (allRuntimes, the interactive runtimeMap menu) without a matching registry adapter entry, or when config-mutation dispatch escapes the registry's declared install surfaces — catching reintroduction of the scattered per-runtime branching those seams removed. (#867)

Fixed

  • profile-pipeline temp output now lands under the reaped GSD temp root. cmdExtractMessages and cmdProfileSample previously created their output directories directly in os.tmpdir() root (gsd-pipeline-* / gsd-profile-*), which reapStaleTempFiles never scans (it only scans GSD_TEMP_DIR = os.tmpdir()/gsd). The directories accumulated forever. Both sites now call ensureGsdTempDir() and create under GSD_TEMP_DIR. Also adds missing after/afterEach teardown to four test fixtures that leaked gsd-* temp dirs on every npm test run. (#866) (#879)
  • gsd_run launcher shim now probes all non-Claude runtime homes before failing. The shim's last-resort detection previously stopped at $HOME/.claude, causing a false-positive fatal error on every non-Claude runtime (Hermes, Cursor, Codex, Copilot, Windsurf, Augment, Trae, Qwen, CodeBuddy, Cline, Grok, Antigravity, OpenCode, Kilo) when RUNTIME_DIR was unset and gsd-tools was not on PATH. The snippet now probes each runtime's config directory (respecting HERMES_HOME, CURSOR_CONFIG_DIR, CODEX_HOME, etc. with sensible $HOME-relative defaults) before emitting the install error. (#903)
  • validate health and validate consistency no longer emit false-positive W007 warnings for projects using checklist-style ROADMAP.md phases. buildRoadmapPhaseVariants() in src/validate.cts previously used only a heading-style regex (## Phase N: name), silently ignoring the supported checklist format (- [x] **Phase N: name**). This caused every on-disk phase directory to trigger W007 ("exists on disk but not in ROADMAP.md") when the project's ROADMAP used checklist-only notation. The fix adds a second regex pass mirroring the existing buildNotStartedPhaseVariants() approach. Additionally, cmdValidateConsistency() in src/verify.cts had a duplicate inline heading-only regex with the same gap — refactored to delegate to buildRoadmapPhaseVariants() (DRY). (#892) (#893)
  • init execute-phase and cmdCommit now produce correct branch_name when project_code is set — the {phase} substitution in phase_branch_template now calls normalizePhaseName(), stripping the project-code prefix and zero-padding the number, so the generated branch is e.g. gsd/phase-01-foundation instead of gsd/phase-CK-01-foundation. Both the execute-phase output path (src/init.cts) and the pre-execution commit path (src/commands.cts) are fixed. (#904) (#904)
  • syncStateFrontmatter no longer strips current_phase, current_phase_name, current_plan, and progress from STATE.md — when body annotations are absent (e.g. after an agent rewrites the body), the existing frontmatter values for those scalars are now preserved, mirroring the fallback already applied in cmdStateJson. (#905) (#905)
  • Top-level Claude Code /gsd-plan-phase now always spawns the researcher/planner/plan-checker agents instead of collapsing them inline — a <runtime_compatibility> block after </available_agent_types> makes the Agent-availability requirement explicit and documents that the workflow fails-closed (stops with a clear log message) in genuinely Agent-less contexts; seven "ORCHESTRATOR RULE — CODEX RUNTIME" labels are renamed to "ALL RUNTIMES" so the guard applies universally; execute-phase.md scopes its existing "Other runtimes" inline-fallback prose to non-Claude contexts, preserving the #853 backgrounded-agent behaviour. (#913) (#913)
  • /gsd-manager and /gsd-autonomous --interactive no longer silently skip worktree isolation and independent verification on Claude Code. They dispatched plan/execute as background agents, but a backgrounded Claude Code agent has no Agent/Task tool and cannot spawn the nested executors, plan-checker, or verifier — so isolation and verification silently never ran. Both workflows now resolve the runtime and run plan/execute inline on Claude Code; background dispatch is kept on runtimes that support nested subagents. (#863)
  • Installer no longer leaks gsd-cmd-rewrites-* temp directories. Each install that emitted slash commands left one fs.mkdtempSync directory under the system temp root; on tmpfs /tmp hosts these accumulated and consumed RAM-backed storage. installRuntimeArtifacts() now removes the temp copy in a finally once command files are copied. (#862)
  • Corrected the installer --help profile skill counts: core now shows 8 (was 7) and standard shows 14 (was 13), both derived from PROFILES so they can't drift again; the full line drops the stale hardcoded 66 for all skills. (#834) (#847)

[1.4.0] - 2026-06-08

Added

  • Research is now cached, curated-first, and code-governed — a content-addressed Research Store (per-source TTL), a single provider waterfall with confidence tiers, and registry-API package legitimacy replace the per-agent prose waterfall and the slopcheck bolt-on. (#664) Confidence is now verification-evidence-driven: provider identity alone no longer yields HIGH; HIGH requires ground-truth corroboration (e.g. legitimacyVerdict: 'OK'), authority alone caps at MEDIUM, and SLOP caps at LOW. (#664)
  • /gsd:plan-phase now accepts a --granularity <coarse|standard|fine> flag to override the configured planning granularity for a single invocation. The flag takes precedence over granularities.planning, top-level granularity, and planning.granularity config. Invalid values are rejected. (#703) (#750)
  • gsd-core can now be installed as a native Claude Code plugin — a new .claude-plugin/plugin.json manifest enables installing gsd-core via claude plugin install or the zero-friction ~/.claude/skills/ auto-load path (gsd-core@skills-dir), with slash commands auto-namespaced as /gsd-core:<command> (e.g. /gsd-core:plan-phase) and lifecycle management via claude plugin enable|disable|update. gsd-core's always-on guard and update hooks are wired for the plugin path through hooks/hooks.json using ${CLAUDE_PLUGIN_ROOT}. This is additive — the existing npm / file-copy installer is unchanged. (#797)
  • Installer pre-populates permissions.allow/deny for Claude Code — fresh Claude Code installs now receive GSD's known-safe tool-call patterns (Bash(npx gsd-core *), Read(.planning/*), Write(.planning/*), Read(STATE.md), Write(STATE.md)) in settings.json out of the box, eliminating first-run approval prompts. A deny block for credential files (Read(.env), Read(.env.*), Read(.secrets)) is also added for defense-in-depth. The merge is additive and idempotent; existing user-set entries are preserved. Uninstall removes only GSD-owned entries. (#768) (#819)

Added: register newly-available Claude Code lifecycle hooks — SubagentStop, Stop, PreCompact (all wired to gsd-context-monitor for context-headroom warnings), and FileChanged (matcher: config.json, wired to new gsd-config-reload.js hook that hot-reloads .planning/config.json context mid-session). Also updates hooks/hooks.json (plugin manifest) and managed-hooks-registry for drift-guard coverage (#770). (#821)

  • Gemini installs now register three additional hook events — BeforeAgent, AfterAgent, and BeforeModel — wired to gsd-context-monitor.js for per-turn context headroom tracking. Previously only SessionStart, BeforeTool, and AfterTool were registered. The installer also detects hooksConfig.enabled: false in the user's Gemini settings.json and emits a clear warning, surfacing the silent failure mode where all hooks are registered but never execute. (#776) (#829)
  • Cross-runtime command enrichment in the installer. Gemini CLI commands now use native {{args}} interpolation (translated from Claude's $ARGUMENTS) so typed arguments interpolate into the prompt body, and /gsd:progress injects live project state via a fixed, injection-safe !{cat .planning/STATE.md 2>/dev/null} shell block. Qwen Code skills now carry a numeric priority field so the most-used main-loop workflows (new-project, plan-phase, execute-phase, …) surface first in the /skills list. The OpenCode per-command model/agent/subtask enrichment was evaluated and intentionally not implemented — model would reintroduce the ProviderModelNotFoundError regression that the converter deliberately guards against for non-Anthropic providers (#1156), subtask/agent change execution semantics for GSD's interactive commands, and variant is not in the OpenCode command schema. (#778) (#825)
  • Emit native on-demand skills (skills/<name>/SKILL.md) for the OpenCode-family runtimes (OpenCode and Kilo) at install time, in addition to the existing flat command/ and file-based agents/ surfaces. OpenCode and Kilo share a config schema and both discover skills from skills/<name>/SKILL.md; the installer now stages each GSD command as a skill with minimal, spec-compliant frontmatter (name matching the directory, description 1–1024 chars) via a shared OpenCode-family skill writer. Skills respect the active install profile (core/minimal stage only their subset) and are removed on uninstall. (#784) (#810)
  • gsd install --cursor now writes .cursor/commands/gsd-<name>.md in addition to the existing .cursor/skills/ surface. Cursor 1.6 introduced plain-markdown slash commands (no frontmatter) in .cursor/commands/; they appear in the / menu in the Agent input. Each command file is generated from the same source as the skill but with frontmatter stripped and Cursor-specific content transforms applied (convertClaudeCommandToCursorCommand). The skills surface is unchanged — both surfaces are written on every install. (#805)
  • The GitHub Copilot installer now reaches lifecycle-hook and instruction parity with other first-class runtimes. It emits a self-contained sessionStart hook config (.github/hooks/gsd-session.json for local installs, ~/.copilot/hooks/gsd-session.json for global) and writes AGENTS.md at the repository root (which Copilot CLI reads as primary instructions) alongside copilot-instructions.md. The hook is an inline command hook with no separate script file, so it cannot dangle. Both artifacts are removed — with user-authored content preserved — on --uninstall. (#786) (#804)
  • Elevate the Cline runtime to hook parity. The installer now emits the Cline .clinerules/ directory form (.clinerules/gsd.md) instead of a single .clinerules file, adds a .clinerules/hooks/PreToolUse lifecycle hook (Cline v3.36+ JSON stdin → {cancel,errorMessage,contextModification} protocol; guards .planning/ artifacts and fails open), and merges GSD instructions into the cross-tool global ~/.agents/AGENTS.md target on global installs. A legacy single-file .clinerules is migrated to the directory form in place, and --uninstall removes the new artifacts and strips the GSD block from ~/.agents/AGENTS.md. (#787) (#803)
  • Qwen Code installs now register three additional hook events that Qwen Code supports beyond Claude Code: SubagentStop, Stop, and PreCompact — all wired to gsd-context-monitor.js for context headroom tracking at subagent completion, model stop, and pre-compaction. These events are Qwen-only; Claude Code installs are unchanged. UserPromptSubmit is deferred: gsd-prompt-guard exits unless tool_name is Write|Edit, making it a no-op for that payload shape. (#788) (#807)
  • CodeBuddy (Tencent) installs now emit /gsd-* slash commands. A --codebuddy install writes commands/gsd-<name>.md files to ~/.codebuddy/commands/ so GSD workflows are invokable from CodeBuddy's / menu (/gsd-phase, /gsd-ship, etc.), matching the integration depth of other fully-elevated runtimes (#789). The existing skills/gsd-<name>/SKILL.md files are now emitted with user-invocable: false so they stay out of the / menu — the commands surface is the single / entry point (no duplicate entries) and skills remain available for model invocation. Subagents (~/.codebuddy/agents/) were already emitted and are unchanged. Uninstall removes the gsd-* command files while preserving user-owned commands. No mcp.json is written — gsd ships no MCP server and CodeBuddy's mcp.json only registers external MCP servers.
(#830)
  • Augment installs now emit slash command definitions alongside skills. A global --augment install writes commands/gsd-<name>.md files to ~/.augment/commands/ in addition to the existing skills/gsd-<name>/SKILL.md files, matching the integration depth of other fully-elevated runtimes and allowing Auggie users to invoke GSD as slash commands (/gsd-phase, /gsd-ship, etc.) without manual configuration (#790). Content rewrites (path normalisation and Augment-specific branding) are applied at install time. Uninstall removes the gsd-* command files while preserving user-owned commands. mcpServers registration is explicitly excluded — gsd ships no MCP server and does not register third-party servers. (#801)
  • Issues are now checked for duplicates when opened: a no-LLM title-similarity check posts a challenge comment and applies a possible-duplicate label when a new issue closely matches existing open ones. Flagged issues that go unanswered for 24h are auto-closed as duplicates (reply, or react 👎 to the bot comment, to keep one open); a reply clears the label and routes to needs-maintainer-review. (#836) (#843)
  • Cursor now receives GSD lifecycle hooks via .cursor/hooks.json — a sessionStart hook injects the current workflow state as context at session start, and a postToolUse hook nudges the agent to update .planning/ after write-class operations, bringing Cursor to baseline hook parity with Gemini and Claude Code. (#777)
  • Gemini CLI extension package — gsd-core now ships a gemini-extension.json manifest (plus a GEMINI.md context payload) at the repository root, so Gemini CLI users can install, update, and remove GSD through Gemini's own extension lifecycle: gemini extensions install https://github.com/open-gsd/gsd-core, gemini extensions update gsd-core, gemini extensions uninstall gsd-core, and gemini extensions link <path> for local dev. The extension is discoverable in gemini extensions list and loads GSD's operating context into every session. Additive — the existing npx gsd-core --gemini installer (which provides the /gsd:* slash commands) is unchanged. (#775) (#775)
  • New agent_skills_security.trusted_global_roots config — opt-in allowlist of trusted root directories so symlinked global: agent skills whose real path resolves outside the default skills dir (e.g. ~/.claude/skills) are accepted; default [] is byte-identical and preserves the symlink-escape guard. (#754)
  • Added /gsd-update --next (alias --rc) to install or refresh from the @next RC dist-tag (ADR #660). A new parse_update_channel workflow step resolves the channel from $ARGUMENTS; the version check and all three npx install invocations thread $TAG instead of hardcoding @latest. When --next is used the version-comparison output gains a Channel: next (RC) banner so the user knows they are leaving the stable line; omitting the flag keeps @latest behavior byte-for-byte unchanged. check-latest-version.cjs gains ALLOWED_TAGS, buildViewArgs, and resolveTag exports, with an allowlist guard (enforced at both the CLI and function boundary) that rejects any dist-tag other than latest/next. (#815) (#839)

Changed

  • /gsd:plan-phase --research-phase <N> now auto-uses an existing RESEARCH.md instead of prompting update/view/skip. When research already exists and neither --research nor --view is passed, it emits a one-line notice and exits cleanly, matching the promptless behavior of standard /gsd:plan-phase <N>. Pass --research to force-refresh or --view to print the existing research. (#159) (#718)
  • Retire the installer's one-off runtime directory helpers (getGlobalDir/getOpencodeGlobalDir/getKiloGlobalDir) and consolidate per-runtime global config-dir resolution onto the single canonical projection runtime-homes:getGlobalConfigDir, extended with the --config-dir override and the opencode/kilo *_CONFIG file-path precedence. Behavior-preserving across all 15 install runtimes. (#56) (#802)
  • Make per-runtime config-mutation dispatch in the installer explicit: a new runtime config adapter registry maps each supported runtime to a typed config intent (install surface, shared-settings gate, finish-phase permission writer), and install()/finishInstall() dispatch by resolved intent instead of inline runtime === '...' branching. Behavior-preserving; unknown runtimes now fail loudly. (#60) (#795)
  • Verification status routing is now owned by a single queryable seam — ship.md and execute-phase.md both consume gsd_run query verification.status instead of re-deriving the passed/gaps_found/human_needed routing independently; the query returns next_action and next_command so per-status prose no longer needs to be kept in sync across files. This also fixes the broad-grep status misread in execute-phase.md where a body status: line (in a code block or copied artifact) could concatenate with the frontmatter value and misroute a valid passed phase; a parity test fails if a new verifier status value lacks a route. (#651) (#755)
  • Agent color: frontmatter now uses Claude Code's documented named colors (red/blue/green/yellow/purple/orange/pink/cyan) instead of hex values or the undocumented magenta, so the intended per-agent TUI color differentiation renders reliably across the Claude Code runtime. Display-only metadata; no behavior change. (#771) (#823)
  • Codex installs now register three additional stable hook events (SubagentStart, Stop, PostToolUse) wired to gsd-context-monitor.js, matching the full event coverage available since Codex CLI stabilised these hooks. The SessionStart hook entry gains a commandWindows field on Windows installs so the .cmd shim is used for native execution (Git Bash/MSYS cannot POSIX-exec node.exe directly). Both new-event registration and uninstall paths handle the flat { "EventName": [...] } and nested { "hooks": { "EventName": [...] } } hooks.json shapes. gsd-context-monitor.js and its Windows .cmd sibling are added to the managed-hook allowlist so idempotent re-runs de-duplicate entries correctly. (#772) (#827)
  • Codex CLI installs now emit two enrichments per agent and skill. Agent TOML enrichment: light-tier agents (haiku-equivalent, routingTier: "light" in model-catalog.json) get service_tier = "flex" and model_verbosity = "low" appended to their agent TOML, telling the Codex scheduler to use the flex tier (lower cost, background processing) and suppress verbose token output. Skill TUI chip: each installed gsd-* skill directory now receives an agents/openai.yaml file with interface.display_name and interface.short_description, making the skill appear in the Codex /skills picker with a human-readable name and description drawn from the skill's existing short-description frontmatter. Both enrichments are additive and backward-compatible with Codex CLI ≥ 0.130.0. (#774) (#828)
  • Cline global installs now emit skills, not just rules: gsd writes skills to ~/.cline/skills/<name>/SKILL.md for Cline ≥ v3.48.0 (see Cline skills docs), in addition to the existing .clinerules file. Each SKILL.md carries name/description frontmatter (agentskills.io) with paths rewritten to the .cline/ convention. Local installs remain .clinerules-only. The .clinerules rules file continues to be emitted for compatibility, and upgrading over an existing rules-only install emits the new skills on the next run. (#809)
  • Workflow size budget now measures bytes, not lines (#717). tests/workflow-size-budget.test.cjs re-bases its tier ceilings (XL/LARGE/DEFAULT) from line counts to byte counts — deterministic, no tokenizer, and matching the unit vendors bound on (Codex's 32,768-byte project_doc_max_bytes cap). The #597 tighten-only ratchet and per-file semantics are unchanged; the budget's caching-independent quality rationale (context rot / attention budget) is now documented. (#719)
  • The gsd-verifier agent no longer re-runs the full workspace test suite once per must-have during Step 7b spot-checks — it enumerates tests to prove existence and runs a single named test to prove a pass, invoking the full suite at most once per verification. (#753)
  • /gsd-plan-phase, /gsd-execute-phase, /gsd-autonomous now run in an isolated forked context on Claude Code — context: fork in skill frontmatter protects the main session's context budget. These three heavy skills also declare effort: xhigh; quick-status skills /gsd-progress and /gsd-stats declare effort: low. The installer preserves both fields when converting commands to Claude SKILL.md files. Runtimes that do not recognise these fields silently ignore them — no behaviour change on non-Claude runtimes. (#769)
  • /gsd:plan-phase and /gsd:execute-phase no longer eagerly load MVP-only guidance on non-MVP runs — the MVP planner rules, user-story template, Walking-Skeleton template, and MVP+TDD halt-report reference are now Read lazily by the planner/executor only when MVP / Walking-Skeleton / MVP+TDD mode is active, in both the workflow files and the gsd-planner/gsd-executor agent definitions, instead of being @-imported into every run. Behaviour is unchanged; non-MVP planning/execution simply carries less context. (#720) (#746)

Automated codex exec invocations in the review workflow now include --ephemeral (no session-state accumulation across automated/CI runs) and --dangerously-bypass-hook-trust (skip hook-trust prompts for hooks managed by gsd-core itself). These flags apply only to the non-interactive reviewer invocations in gsd-core/workflows/review.md. (#773) (#824)

  • Codex slash-command conversion no longer corrupts inline-wrapped /gsd-… file paths — the install-time converter now identifies a real /gsd-<command> mention by positive boundaries (opening delimiter + no path continuation) instead of an unbounded preceding-character denylist, closing the path-corruption class (#637 → #704) by construction while still converting legitimate backtick-wrapped mentions. (#747)
  • The release pipeline now automatically runs changeset render during the finalize job, promoting .changeset/ fragments into a dated CHANGELOG.md section before publishing — previously a manual step that was routinely skipped (leaving v1.3.0 and v1.3.1 unpromoted, #690). A new --allow-empty flag prevents the verify gate from hard-failing on no-change releases by emitting a dated heading with a _No notable changes._ placeholder when there are zero fragments. (#715)

Fixed

  • /gsd-review --cursor now actually invokes the Cursor agent. Detection probes the cursor-agent headless binary instead of the cursor IDE launcher, the invocation calls the single cursor-agent binary in print mode (not the two-token cursor agent, which the IDE treats as a file path), and the review prompt is passed as a file-path argument rather than piped to stdin (which cursor-agent -p ignores). On failure the captured stderr is surfaced instead of a silent empty result. (#686)
  • No more "gsd-core" console-window flash on Windows. Every gsd-core child process now passes windowsHide: true: the context monitor's record-session spawn, the execGit / execNpm / execTool helpers in shell-command-projection, the gsd-worktree-path-guard and gsd-workflow-guard hook git probes, check-command-router's git log call, and the roadmap-upgrade git status/rev-parse/reset/clean calls — matching the existing gsd-check-update spawn. execNpm (which uses shell: true → cmd.exe and runs on every SessionStart, i.e. every /clear) and the worktree-path guard (which runs on every Edit/Write in a worktree) were the most visible offenders. No behavior change on macOS/Linux, where the flag is ignored. (#688)
  • /gsd-review --agy no longer hangs the whole review on large prompts. On a big, file-path-rich prompt Antigravity's agy -p agentic Cascade can loop on its code_search/grep steps and never converge. The invocation now passes agy's own --print-timeout flag (its native print-mode cap) so a stalled run self-terminates through the tool's own mechanism; on a non-zero exit any partial output is discarded so the existing transcript fallback / "review failed" stub take over. (#689)
  • The roadmap parser now resolves fresh phases of the current milestone in multi-milestone roadmaps. extractCurrentMilestone() scoped the current-milestone window to its ## Phases checklist subsection and stopped at the milestone's own ## Milestone … (Phase Details) heading, so the ### Phase N: detail headers fell out of scope. Any command backed by the parser — init.phase-op (and therefore /gsd:discuss-phase and /gsd:plan-phase), state, roadmap list, and validate health (W006) — could not resolve phases of any milestone after the first until a .planning/phases/ directory already existed, blocking discuss/plan. The parser now also includes the current milestone's (Phase Details) section in scope, anchored to the selected milestone's version token so sibling sub-milestones do not cross-pollinate. (#730) (#748)
  • getGlobalSkillsBase('kilo') now resolves to ~/.kilo/skills — where Kilo Code actually discovers global skills — instead of ~/.config/kilo/skills. Per Kilo Code docs, global skills live in the .kilo directory within HOME (~/.kilo/skills/), independent of the XDG-based config dir at ~/.config/kilo. The kilo.jsonc config dir (~/.config/kilo) and the command/ path used by the installer are correct and unchanged. Blast radius: this corrects the resolved skills-base path used by doctor/status checks and agent-skills-block resolution (init.cjs); the installer writes commands (not skills) for Kilo, so no files were previously being written to the wrong location. (#806)
  • Honor the COPILOT_HOME environment variable when resolving the GitHub Copilot global config directory. Previously a global --copilot install ignored COPILOT_HOME and wrote all artifacts (skills, agents, copilot-instructions.md, the session hook) to ~/.copilot even when the user had relocated their Copilot home, making them undiscoverable by Copilot CLI. Resolution now follows --config-dir > COPILOT_CONFIG_DIR > COPILOT_HOME > ~/.copilot, mirroring the existing CODEX_HOME handling. Uninstall uses the same resolver and stays symmetric. (#812) (#814)
  • Release version bumps now keep runtime manifest versions in sync — .claude-plugin/plugin.json and gemini-extension.json are stamped to match package.json on every npm version, unblocking RC/finalize releases. New version-bearing manifests must be registered in scripts/sync-manifest-versions.cjs (enforced by a regression test). (#845)
  • npx @opengsd/gsd-core upgrades no longer abort with "applied migration checksum changed" — an already-applied installer migration whose recorded checksum drifted (e.g. a shipped body was edited) is now detected and reconciled automatically on the next install, instead of hard-failing the upgrade. Replaces the published-checksum allowlist with general self-healing recovery plus a CI baseline lock. (#675)
  • /gsd-import, /gsd-plan-review-convergence, and /gsd-spec-phase now run on global installs — these workflows resolve gsd-tools via the runtime launcher instead of a hardcoded $HOME path, so they no longer falsely report the tool as "not found" (and stop short) when only a global/shim install is present and no project-local runtime exists. (#642)
  • Worktree wave-cleanup no longer fails when the phase SUMMARY is committed — rescueSummaryArtifacts no longer copies an already-committed SUMMARY into the main checkout, which previously caused git merge --no-ff to abort with a permanent merge_failed (#706). (#709)
  • Phase execution no longer halts with exit 42 (worktree base mismatch) when run on a branch diverged from the default branch (#683). Claude Code forks worktree-isolated executors off the repository default branch (origin/HEAD), so running /gsd-execute-phase on an unmerged milestone/feature branch left every executor without the phase's plan files and tripped the worktree-branch-check guard (100% reproducible, all OSes). Execute-phase now detects this before dispatch and automatically degrades to sequential execution on the main working tree, recommending the permanent fix worktree.baseRef:"head". Both fresh installs and upgrades of GSD Core set worktree.baseRef:"head" in .claude/settings.local.json automatically (no-clobber) when workflow.use_worktrees is enabled (the default); gsd-tools worktree set-baseref remains available for manual use (e.g. after toggling worktrees on later). The exit 42 guard remains as a backstop. (#749)
  • Codex install no longer corrupts launcher paths — shell path segments like ${VAR}/gsd-core/ and $(cmd)/gsd-local-patches are no longer rewritten into a literal $gsd-core token during Codex markdown conversion (#704). (#710)
  • /gsd:surface no longer corrupts installed skill paths — re-surfacing (profile/enable/disable/reset) now applies the same per-runtime path rewrites as install, so SKILL.md bodies keep the correct install target instead of reverting to the converter's default ~/.claude paths. (#817)
  • /gsd:graphify, /gsd:import, and planning agents now resolve gsd-tools on global/shim-only installs — agent and command surfaces that invoked a hardcoded $HOME/.claude/...gsd-tools.cjs path now route through the resolved gsd_run launcher, so the step no longer reports the tool "not found" when there is no project-local runtime. (#707)
  • /gsd:surface no longer mis-names or orphans runtime command files — re-surfacing now writes the same gsd--prefixed command filenames as a fresh install for flat command dirs (Cursor, Augment, OpenCode, Kilo) and preserves user-authored command files instead of deleting them. (#822)
  • /gsd:update reliably previews release notes again — promotes the 1.3.x changelog into dated [1.3.0]/[1.3.1] sections, stops deleting the temp changelog before the human-readable render (no more (changelog unavailable)), and adds a release gate that blocks publishing a version whose CHANGELOG.md section was never promoted. (#694)

Security

  • gsd-tools config-set prototype-pollution guard hardened and regression-tested. The guard that blocks __proto__, prototype, and constructor segments in dotted config keys now uses inline literal comparisons at each property-write site (instead of a pre-loop Set check), so CodeQL's js/prototype-pollution-utility analysis recognises it as a sanitising barrier and code-scanning alert #26 clears. Runtime behaviour is unchanged from #663. Added regression tests that drive schema-valid dynamic-prefix keys (agent_skills.__proto__, agent_skills.constructor, features.__proto__, review.models.constructor) all the way to the guard — these reach setConfigValue past the schema gate and were previously the guard's only untested attack surface. (#751) (#752)
  • Hardened roadmap-phase parsing and config writes — resolved ReDoS in phase-heading/plan-filename regexes (validate/verify/commands/phase), blocked prototype-pollution through dotted config keys in config-set, and pinned qs >= 6.15.2 (DoS advisory). (#665)

1.3.1 - 2026-06-04

Security

  • Bumped hono to clear a moderate npm advisory carried transitively in the dependency tree. (#670)

Fixed

  • Installer-migration checksum drift no longer blocks upgrades — the updater now self-heals when a shipped migration's recorded checksum has drifted, reconciling the stored checksum instead of aborting. Restores upgrades across all OSes after shipped migration bodies were edited in a prior release. (#670)

1.3.0 - 2026-06-04

Added

  • Vertical MVP Slice mode — --mvp flag on /gsd-plan-phase switches the planner from horizontal layer decomposition to vertical feature-slice decomposition (UI→API→DB in one task sequence). On Phase 1 of a new project with no prior phase summaries, also emits SKELETON.md via Walking Skeleton mode. Composable with --tdd: --mvp --tdd produces vertical slices where every behavior-adding task starts with a failing test. Phase-level persistence via **Mode:** mvp in ROADMAP.md applies --mvp automatically without the flag. (#78)
  • /gsd-mvp-phase command — guided MVP planning: prompts for a user story (As a / I want to / So that), runs SPIDR story-splitting check (Spike/Paths/Interfaces/Data/Rules axes), writes **Mode:** mvp to ROADMAP.md, then delegates to /gsd-plan-phase. (#78)
  • MVP-aware UAT framing in verify-phase — when a phase has mode: mvp, the verifier generates a user-flow-first UAT script (walks the feature as a user would) before any technical checks. (#78)
  • MVP progress and stats display — progress and stats commands show Walking Skeleton completion status and per-feature-slice status lines for MVP-mode phases. (#78)
  • Six MVP reference files — planner-mvp-mode.md, skeleton-template.md, user-story-template.md, spidr-splitting.md, execute-mvp-tdd.md, verify-mvp-mode.md — loaded by the planner, executor, and verifier agents when MVP mode is active. (#78)
  • Milestone-prefixed phase ID convention (M-NN) for globally unique phase IDs within a project (#39)
  • getMilestoneFromPhaseId() and getPhaseDirFromPhaseId() helpers in core.cjs (#39)
  • W021 validation rule: fires when a phase ID's integer prefix mismatches its enclosing milestone section (#39)
  • gsd-tools roadmap validate subcommand for convention compliance checking (#39)
  • gsd-tools roadmap upgrade --convention milestone-prefixed migration tool (dry-run by default, --apply to mutate) (#39)
  • phase_id_convention config field (null | 'milestone-prefixed' | 'free-form'), defaults to null (legacy free-form, no breaking change) (#39)

Fixed

  • isDirInMilestone now correctly matches M-NN-style phase directories against milestone-prefixed ROADMAP headings (#39)
  • searchPhaseInContent heading regex now tolerates [bracket-token] scope prefix (e.g., ### [GSD] Phase 2-01:) (#39)
  • README version guidance now uses npm/package metadata as the source of truth — README, localized READMEs, and the docs index no longer present archived release-note or canary-stream numbers as the current GSD Core package version. (#545)

1.2.0 - 2026-05-31

1.2.0 is the current stable @opengsd/gsd-core release. It resumes the public package line after the release-version validation recovery documented in ADR 218 and makes @opengsd/gsd-core / gsd-core the canonical package and CLI identity.

Added

  • Plan-vs-codebase drift guard — plan review can verify generated plans against live source symbols before execution so hallucinated files, APIs, or commands are caught earlier. (#487)
  • Single Package Identity seam — package name, CLI identity, update checks, and installer identity are centralized so @opengsd/gsd-core stays consistent across runtime surfaces. (#499, #517, #521)
  • Cross-provider effort controls and fast-mode-aware routing — model-effort selection works across providers and can adjust routing for faster workflows. (#463)
  • Current public docs and install identity — README/docs now advertise GSD Core, @opengsd/gsd-core, and the gsd-core binary as the canonical user-facing surface. (#519, #523, #540)

Changed

  • SDK shim retired from installer/runtime docs — workflows now route through gsd-tools; dead SDK-shim verification and stale SDK-generated banners were removed. (#522, #515, #510)
  • Release numbering recovered at 1.2.0 — leading-zero release inputs are invalid and duplicate-version checks fail early before publish work begins. See ADR 218.
  • CI/test selection is more precise — affected-test selection now widens docs/test-impact correctly and avoids under-testing relevant PRs. (#495)

Fixed

  • Planning writes are more reliable — phase completion writes are transactional and no longer corrupt milestone progress counters. (#465, #514)
  • Roadmap and milestone parsing no longer leak stale phase details into active milestone state. (#513)
  • /gsd:update detects local Antigravity .agent installs and repo-local Claude installs correctly. (#512, #476)
  • Package identity registration no longer regresses update/runtime detection. (#521)

Legacy Release History

Release notes for every version published before the project was renamed to @opengsd/gsd-core — the retired get-shit-done-cc / get-shit-done-redux lineage, versions 1.0.0 → 1.42.x plus pre-release and canary builds — have been rolled up into a single archive:

➡️ docs/RELEASE-NOTES-LEGACY.md

Those legacy 1.x numbers belong to the previous package line and predate the current @opengsd/gsd-core versioning, which restarts at 1.0.0. They are preserved verbatim-in-spirit (condensed) in the archive and intentionally kept out of this file so the two version streams cannot collide.