chore: merge release v1.8.0 to main (#2500)

Reconciles main to the v1.8.0 tree while preserving the [1.7.0] CHANGELOG
section that release/1.8.0 omitted (root-caused in #2502: the main->next
auto-backmerge failed after 1.7.0, so next never received 1.7.0's promoted
CHANGELOG). Both parents kept so the v1.7.0 and v1.6.1 tags remain in
main's ancestry. This also delivers the fixed auto-backmerge.yml (build:lib
step, #2281) to main so the main->next backmerge stops failing.
This commit is contained in:
Tom Boucher
2026-07-21 20:40:16 -04:00
502 changed files with 39709 additions and 6263 deletions

View File

@@ -6,7 +6,7 @@ This directory holds **per-PR CHANGELOG fragments**. Every PR with user-facing c
Two PRs that both edit the `### Fixed` block of `CHANGELOG.md` always conflict on merge — git can't pick a serialization order without human input. Two PRs that each add a fresh `.changeset/<unique-name>.md` never conflict because they don't share lines.
See [#2975](https://github.com/open-gsd/get-shit-done-redux/issues/2975) for the full rationale.
See [#2975](https://github.com/open-gsd/gsd-core/issues/2975) for the full rationale.
## Adding a fragment

View File

@@ -9,7 +9,7 @@
{
"name": "gsd-core",
"description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.",
"version": "1.7.0",
"version": "1.8.0",
"source": "./",
"author": {
"name": "open-gsd",

View File

@@ -1,7 +1,7 @@
{
"name": "gsd-core",
"displayName": "GSD Core",
"version": "1.7.0",
"version": "1.8.0",
"description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.",
"author": {
"name": "open-gsd",

View File

@@ -141,6 +141,16 @@ jobs:
echo "dropped_oneline=" >> "$GITHUB_OUTPUT"
fi
# The version bump below fires the `version` npm lifecycle hook, which runs
# gen-capability-registry.cjs. That validator lazily require()s the built
# gsd-core/bin/lib/capability-ledger.cjs (a build:lib output, gitignored);
# when it is absent the bounded fragment reader falls back to a fail-closed
# stub and every capability fragment reports "could not be read", failing
# the sync. Build the ledger first so fragments materialize.
- name: Install dependencies and build (required by the version-sync hook)
if: steps.check.outputs.next_exists == 'true'
run: npm ci --silent && npm run build:lib
- name: Sync next's version to main's released version
if: steps.check.outputs.next_exists == 'true'
run: |

View File

@@ -18,9 +18,18 @@ jobs:
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5.0.1
with:
# Intentionally shallow — see tests/policy-lint-shallow-checkout.test.cjs.
# Depth 50 covers >99% of PRs; a deeper merge base fails CLOSED with a
# loud lint error, which is the accepted trade.
fetch-depth: 50
- name: Fetch base ref for diff
run: git fetch --depth=50 origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"
# The BASE REF itself must not be shallow (#2452): scripts/changeset/lint.cjs
# diffs with the three-dot form `origin/<base>...HEAD`, which needs a merge
# base. Truncating the base's ancestry makes git abort with
# `fatal: ...: no merge base` instead of reporting a real verdict.
# The shallow *checkout* above is the deliberate cost control; the shallow
# *base fetch* was not — it only shrank the window further.
run: git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"
env:
BASE_REF: ${{ github.event.pull_request.base.ref }}
- uses: actions/setup-node@a0853c24544627f65ddf259abe73b1d18a591444 # v5.0.0

View File

@@ -18,9 +18,18 @@ jobs:
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5.0.1
with:
# Intentionally shallow — see tests/policy-lint-shallow-checkout.test.cjs.
# Depth 50 covers >99% of PRs; a deeper merge base fails CLOSED with a
# loud lint error, which is the accepted trade.
fetch-depth: 50
- name: Fetch base ref for diff
run: git fetch --depth=50 origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"
# The BASE REF itself must not be shallow (#2452): scripts/lint-docs-required.cjs
# diffs with the three-dot form `origin/<base>...HEAD`, which needs a merge
# base. Truncating the base's ancestry makes git abort with
# `fatal: ...: no merge base` instead of reporting a real verdict.
# The shallow *checkout* above is the deliberate cost control; the shallow
# *base fetch* was not — it only shrank the window further.
run: git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"
env:
BASE_REF: ${{ github.event.pull_request.base.ref }}
- uses: actions/setup-node@a0853c24544627f65ddf259abe73b1d18a591444 # v5.0.0

View File

@@ -53,12 +53,28 @@ jobs:
# Ensure the base branch tip is available for the git diff below.
# fetch-depth: 0 above gets all history, but the remote ref name must exist.
# For workflow_dispatch (no base_ref) we fall back to `next`.
run: git fetch origin ${{ github.base_ref || 'next' }} --depth=1
#
# Do NOT pass --depth here (#2452): a shallow re-fetch truncates the
# freshly-fetched `next` history to a single commit, so the three-dot
# `origin/next...HEAD` diff in scripts/mutation-matrix.cjs can no longer
# compute a merge base and dies with `fatal: ... no merge base`. That
# only happens when the branch is BEHIND the base — an up-to-date branch
# incidentally passes because its merge base IS the fetched tip — which
# made the gate fail exactly on the PRs that most needed it, while the
# `mutate` shards never ran at all.
#
# BASE_NAME goes through `env:` rather than being interpolated straight
# into the shell, matching changeset-required.yml / docs-required.yml.
run: git fetch origin "${BASE_NAME}"
env:
BASE_NAME: ${{ github.base_ref || 'next' }}
- name: Compute mutation matrix
id: matrix
env:
BASE_NAME: ${{ github.base_ref || 'next' }}
run: |
BASE_REF="origin/${{ github.base_ref || 'next' }}"
BASE_REF="origin/${BASE_NAME}"
# Run the matrix script; capture the JSON output.
JSON=$(node scripts/mutation-matrix.cjs --base "${BASE_REF}")

View File

@@ -9,14 +9,29 @@ name: PR Target Validator
#
# See: docs/branching.md, docs/adr/230-introduce-next-integration-branch.md
# Trigger (#2331): pull_request_target, NOT pull_request. This job runs ONLY for
# non-OWNER/MEMBER/COLLABORATOR authors (see the `if:` below) — i.e. exactly the
# fork PRs whose `pull_request` GITHUB_TOKEN is downgraded to read-only
# regardless of the `permissions:` block. On that trigger the sticky-comment
# call below 403s, the unhandled rejection kills the step, and core.setFailed
# never runs: the author sees an API stack trace instead of "retarget to next".
# pull_request_target runs in the base-repo context with a write-capable token.
# Safe here (as in pr-template-format.yml): the only checkout is the BASE branch
# with persist-credentials: false, and the PR-controlled inputs (pr.base.ref /
# pr.head.ref) are read as data. No head code executes, so the "PR cannot edit
# the policy that judges it" property below is reinforced, not weakened.
on:
pull_request:
pull_request_target:
types: [opened, edited, reopened, synchronize]
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number }}
cancel-in-progress: true
# Scope unchanged by #2331 — see the note in pr-title-validator.yml. This
# workflow's `pull-requests: write` alone already authorizes the comment call
# (its sticky comment has posted 8 times on that scope); the 403 was the fork
# token downgrade, not a missing scope.
permissions:
contents: read
pull-requests: write
@@ -70,10 +85,21 @@ jobs:
}
// decision === 'blocked': base is main and head is not an allowed pattern.
//
// #2331: `head` is attacker-controlled (a fork author names their own
// branch) and the comment below is posted by github-actions[bot] with a
// write token. A backtick IS legal in a git ref name, so echoed raw into
// an inline-code span it closes the span early and the remainder renders
// as live Markdown. This is weaker than the pr-title-validator case —
// check-ref-format forbids space, ':', '[' and '*', so no bare URL, link
// or emphasis is expressible in a branch name — but it is the same class
// and is stripped identically rather than left to the charset to police.
const headForMarkdown = String(head).replace(/`/g, "'");
const msg = [
`### Wrong target branch`,
``,
`This PR targets \`main\` but the source branch \`${head}\` is not a release, hotfix, critical-fix, or back-merge branch.`,
`This PR targets \`main\` but the source branch \`${headForMarkdown}\` is not a release, hotfix, critical-fix, or back-merge branch.`,
``,
`**Most PRs should target \`next\`, not \`main\`.** See [docs/branching.md](../blob/main/docs/branching.md).`,
``,
@@ -91,28 +117,37 @@ jobs:
].join('\n');
// Post or update a sticky comment.
const { data: comments } = await github.rest.issues.listComments({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
});
const marker = '<!-- pr-target-validator -->';
const existing = comments.find(c => c.body && c.body.includes(marker));
const body = `${marker}\n${msg}`;
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
//
// #2331: the comment is a COURTESY, the verdict below is the GATE.
// A failure here must never suppress the verdict. pull_request_target
// should make the 403 impossible; this catch is defense in depth so a
// future permission change degrades the diagnostic, not the gate.
try {
const { data: comments } = await github.rest.issues.listComments({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
body,
});
const marker = '<!-- pr-target-validator -->';
const existing = comments.find(c => c.body && c.body.includes(marker));
const body = `${marker}\n${msg}`;
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
body,
});
}
} catch (err) {
core.warning(`Could not post the PR-target comment (${err.status || err.message}). Guidance follows:\n${msg}`);
}
if (warnOnly) {

View File

@@ -23,6 +23,18 @@ name: PR Title Validator
# matcher lands on the base branch it does not exist there — the introducing
# PR is skipped (bootstrap); every PR after merge is fully gated.
#
# Trigger (#2331): pull_request_target, NOT pull_request. A `pull_request`
# event raised from a fork hands the job a read-only GITHUB_TOKEN regardless of
# the `permissions:` block below, so the sticky-comment call 403s, the
# unhandled rejection kills this step, and core.setFailed never runs — the
# contributor sees an API stack trace instead of the retitle instructions.
# pull_request_target runs in the base-repo context with a write-capable token.
# This is safe here for the same reason it is safe in pr-template-format.yml:
# the only checkout is the BASE branch (persist-credentials: false) and the
# only PR-controlled input is `pr.title`, read as data. No head code executes.
# The trust boundary above is reinforced, not weakened — pull_request_target
# checks out base by definition.
#
# Unlike pr-target-validator, this runs for ALL authors (including members):
# the changelog drift that motivated #1549 came from member PRs.
#
@@ -32,13 +44,22 @@ name: PR Title Validator
# See: scripts/release-notes/conventional-title.cjs, CONTRIBUTING.md, issue #1549.
on:
pull_request:
pull_request_target:
types: [opened, edited, reopened, synchronize]
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number }}
cancel-in-progress: true
# Scope unchanged by #2331 — deliberately. `pull-requests: write` alone already
# authorizes github.rest.issues.createComment on a PR (GitHub accepts EITHER
# `issues` or `pull-requests` write for the issue-comments endpoint when the
# target is a PR). Verified against this repo's own history rather than the
# docs: this workflow has only ever declared `pull-requests: write` and its
# sticky comment has posted 26 times; require-issue-link.yml declares only
# `issues: write` and its comment posts too. The 403 was the fork token
# downgrade, NOT a missing scope — so widening the scope here would add
# privilege on a pull_request_target workflow while fixing nothing.
permissions:
contents: read
pull-requests: write
@@ -92,10 +113,24 @@ jobs:
return;
}
// #2331: `title` is attacker-controlled free text (a PR title has no
// charset restriction) and the comment below is posted by
// github-actions[bot] with a write token. Echoed raw into an inline-code
// span, a single backtick in the title closes the span early and the
// remainder renders as live Markdown — GFM autolinks a bare URL, so a
// fork author could make our own bot post an arbitrary clickable link
// into the PR thread (phishing that borrows the bot's credibility).
// Stripping the backtick is sufficient and complete: it is the only
// character that can break out of an inline-code span, and everything
// else is inert once it cannot. Only the RENDERED body needs this;
// core.warning/core.setFailed below go to the job log, not Markdown,
// and @actions/core already escapes workflow-command sequences there.
const titleForMarkdown = String(title).replace(/`/g, "'");
const msg = [
`### PR title needs the issue-ref convention`,
``,
`\`${title}\``,
`\`${titleForMarkdown}\``,
``,
result.message,
``,
@@ -113,28 +148,42 @@ jobs:
].join('\n');
// Post or update a sticky comment.
const { data: comments } = await github.rest.issues.listComments({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
});
const marker = '<!-- pr-title-validator -->';
const existing = comments.find(c => c.body && c.body.includes(marker));
const body = `${marker}\n${msg}`;
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
//
// #2331: the comment is a COURTESY, the verdict below is the GATE.
// Never let a failure here suppress the verdict — that inversion is
// the bug this guard exists to prevent (a 403 on the createComment
// call used to kill the step before core.setFailed ran, replacing
// "retitle as type(#issue): summary" with an API stack trace).
// pull_request_target should make the 403 impossible; this catch is
// defense in depth so a future permission change degrades the
// diagnostic rather than the gate.
try {
const { data: comments } = await github.rest.issues.listComments({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
body,
});
const marker = '<!-- pr-title-validator -->';
const existing = comments.find(c => c.body && c.body.includes(marker));
const body = `${marker}\n${msg}`;
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: pr.number,
body,
});
}
} catch (err) {
// Surface the guidance in the job log so it is not lost entirely.
core.warning(`Could not post the PR-title comment (${err.status || err.message}). Guidance follows:\n${msg}`);
}
if (warnOnly) {

View File

@@ -669,6 +669,21 @@ jobs:
VERSION: ${{ inputs.version }}
run: node scripts/verify-npm-publish.cjs --package @opengsd/gsd-core --version "$VERSION" --dist-tag latest
# Regression #2423: keep `next` at the last published release for final
# releases, not just rc/hotfix. Without this, `next` drifts to whatever
# rc.N the release branch forked from, and every npm script banner on
# `next` (and feature branches cut from it) reports the stale rc version
# — e.g. `lint:ci` reported `@opengsd/gsd-core@1.7.0-rc.6` after 1.7.0
# shipped. Mirrors the rc job's sync step at line ~479. Idempotent: a
# no-op when `next` is already at the target version.
- name: Sync next branch to the published release
if: ${{ !inputs.dry_run }}
continue-on-error: true
env:
GH_TOKEN: ${{ secrets.GSD_BOT_PR_TOKEN || secrets.GITHUB_TOKEN }}
VERSION: ${{ inputs.version }}
run: node scripts/sync-next-version.cjs "$VERSION"
- name: Summary
env:
VERSION: ${{ inputs.version }}

View File

@@ -1,13 +1,28 @@
name: Require Issue Link
# Trigger (#2331): pull_request_target, NOT pull_request. A `pull_request` event
# raised from a fork hands the job a read-only GITHUB_TOKEN regardless of the
# `permissions:` block, so the createComment call below 403s, the unhandled
# rejection kills the step, and the core.setFailed on the last line never runs —
# the contributor sees an API stack trace instead of "add Closes #NNN".
# pull_request_target runs in the base-repo context with a write-capable token.
# Safe here: this job performs NO checkout at all and reads the PR body only via
# an env var (never interpolated into a shell), so no head code executes. The
# #1389 fork-forgery carve-out below still holds — it is keyed on
# head.repo.full_name == github.repository, not on the branch name alone.
on:
pull_request:
pull_request_target:
types: [opened, edited, reopened, synchronize]
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
# Scope unchanged by #2331 — see the note in pr-title-validator.yml. `issues:
# write` alone already authorizes this workflow's issues.createComment call on a
# PR: its sticky comment has posted on same-repo PRs (#106, #164, #232, #259) on
# exactly this scope. The 403 was the fork token downgrade, not a missing scope,
# so no `pull-requests: write` is added — this job never calls a pulls.* API.
permissions:
issues: write
@@ -60,26 +75,34 @@ jobs:
'',
'Edit the PR description to add a valid `Closes #NNN`, `Fixes #NNN`, or `Resolves #NNN` line. This check will re-evaluate on the next PR update.',
].join('\n');
const comments = await github.paginate(github.rest.issues.listComments, {
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: prNumber,
per_page: 100,
});
const existing = comments.find(comment => comment.body && comment.body.includes(marker));
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
// #2331: the comment is a COURTESY, the setFailed below is the GATE.
// A failure here must never suppress the verdict. pull_request_target
// should make the 403 impossible; this catch is defense in depth so a
// future permission change degrades the diagnostic, not the gate.
try {
const comments = await github.paginate(github.rest.issues.listComments, {
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: prNumber,
body,
per_page: 100,
});
const existing = comments.find(comment => comment.body && comment.body.includes(marker));
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body,
});
} else {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: prNumber,
body,
});
}
} catch (err) {
core.warning(`Could not post the missing-issue-link comment (${err.status || err.message}). Guidance follows:\n${body}`);
}
core.setFailed('PR body must contain a closing issue reference (e.g. "Closes #123").');

View File

@@ -162,6 +162,13 @@ jobs:
if: github.event_name == 'pull_request'
env:
GITHUB_TOKEN: ${{ github.token }}
# Pin every job of this run to ONE base commit (#2472). Each job runs
# this step independently, minutes apart across a 12-job matrix, so
# merging the moving branch ref lets jobs see different trees when the
# base advances mid-run. The sharded lane needs all jobs to agree on a
# partition, and disagreement there drops a test file silently while
# CI stays green. base.sha is fixed for the life of the run.
CI_REBASE_BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: node scripts/ci-rebase-check.cjs
- name: Set up Node.js ${{ matrix.node-version }}
@@ -222,6 +229,9 @@ jobs:
!coverage/tmp
.nyc_output/
if-no-files-found: ignore
# ~440 MB per full run; same-run diagnostic never consumed by other
# jobs — keep short so it cannot accumulate against the org quota.
retention-days: 3
- name: Run integration tests
if: matrix.scope == 'full'
@@ -259,6 +269,13 @@ jobs:
if: github.event_name == 'pull_request'
env:
GITHUB_TOKEN: ${{ github.token }}
# Pin every job of this run to ONE base commit (#2472). Each job runs
# this step independently, minutes apart across a 12-job matrix, so
# merging the moving branch ref lets jobs see different trees when the
# base advances mid-run. The sharded lane needs all jobs to agree on a
# partition, and disagreement there drops a test file silently while
# CI stays green. base.sha is fixed for the life of the run.
CI_REBASE_BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: node scripts/ci-rebase-check.cjs
- name: Set up Node.js 22
uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
@@ -289,8 +306,9 @@ jobs:
run:
shell: ${{ matrix.shell }}
# The unit suite is sharded across 3 parallel runners per OS/node leg
# (#1212). Each shard runs a deterministic round-robin third of the sorted
# unit-file list via `run-tests.cjs --suite unit --shard i/3`, so per-job
# (#1212, cost-weighted in #2472). Each shard runs a deterministic
# cost-balanced third of the sorted unit-file list via
# `run-tests.cjs --suite unit --shard i/3`, so per-job
# wall-clock scales as O(total/3) and stays well under the cap as the suite
# grows — replacing the #869 timeout bump (15→20m) which only deferred the
# cliff. The cap stays at 20m as a generous backstop; a healthy shard now
@@ -369,6 +387,13 @@ jobs:
if: github.event_name == 'pull_request'
env:
GITHUB_TOKEN: ${{ github.token }}
# Pin every job of this run to ONE base commit (#2472). Each job runs
# this step independently, minutes apart across a 12-job matrix, so
# merging the moving branch ref lets jobs see different trees when the
# base advances mid-run. The sharded lane needs all jobs to agree on a
# partition, and disagreement there drops a test file silently while
# CI stays green. base.sha is fixed for the life of the run.
CI_REBASE_BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: node scripts/ci-rebase-check.cjs
- name: Set up Node.js ${{ matrix.node-version }}
@@ -387,7 +412,7 @@ jobs:
run: node scripts/check-npm-integrity.cjs
# The heavy unit suite is split across the 3 shards — each runs a
# deterministic round-robin third of the sorted unit-file list. The
# deterministic cost-balanced third of the sorted unit-file list (#2472). The
# union of shards 1/3 + 2/3 + 3/3 is the full unit suite, so coverage is
# unchanged; only wall-clock per job drops to ~total/3.
- name: Run unit tests (shard ${{ matrix.shard }}/3)

5
.gitignore vendored
View File

@@ -1,4 +1,4 @@
node_modules/
node_modules
.DS_Store
# ESLint cache
@@ -67,6 +67,7 @@ build/
# by `npm run build:lib`). Source of truth is src/; these are emitted, never edited.
# Published via prepublishOnly; built before test via pretest. Grows as modules migrate.
/tsconfig.build.tsbuildinfo
/gsd-core/bin/lib/broken-windows.cjs
/gsd-core/bin/lib/host-integration.cjs
/gsd-core/bin/lib/host-integration-sdk.cjs
/gsd-core/bin/lib/host-integration-adapters/imperative-hook-bus.cjs
@@ -150,6 +151,8 @@ build/
/gsd-core/bin/lib/installer-migrations/002-codex-legacy-hooks-json.cjs
/gsd-core/bin/lib/installer-migrations/003-rename-get-shit-done-to-gsd-core.cjs
/gsd-core/bin/lib/installer-migrations/004-prune-stale-pristine-snapshots.cjs
/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs
/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs
/gsd-core/bin/lib/observability/logger.cjs
/gsd-core/bin/lib/active-workstream-store.cjs
/gsd-core/bin/lib/adr-parser.cjs

View File

@@ -200,9 +200,23 @@ function mapToolInput(args) {
* @param {string} [opts.cwd] working directory for the child
* @returns {{ stdout: string, exitCode: number, timedOut: boolean }}
*/
const warnedMissingHooks = new Set();
function runHook(hookFile, payload, opts = {}) {
const hookPath = path.join(HOOKS_DIR, hookFile);
if (!fs.existsSync(hookPath)) {
// A missing guard script means the guard is silently NOT enforced — the
// exact failure mode of #2305 (plugin staged, hooks bundle not). Never
// break the tool call (the adapter's design contract), but never be
// silent about it either: warn loudly, once per hook file.
if (!warnedMissingHooks.has(hookFile)) {
warnedMissingHooks.add(hookFile);
console.error(
`[gsd-core] hook script missing: ${hookPath} — ${hookFile} is NOT ` +
"enforced. The GSD install may be incomplete; reinstall (or run " +
"/gsd-update) to restage the hooks/ bundle.",
);
}
return { stdout: "", exitCode: 0, timedOut: false };
}
const timeout = opts.timeout ?? 8000;

View File

@@ -200,9 +200,23 @@ function mapToolInput(args) {
* @param {string} [opts.cwd] working directory for the child
* @returns {{ stdout: string, exitCode: number, timedOut: boolean }}
*/
const warnedMissingHooks = new Set();
function runHook(hookFile, payload, opts = {}) {
const hookPath = path.join(HOOKS_DIR, hookFile);
if (!fs.existsSync(hookPath)) {
// A missing guard script means the guard is silently NOT enforced — the
// exact failure mode of #2305 (plugin staged, hooks bundle not). Never
// break the tool call (the adapter's design contract), but never be
// silent about it either: warn loudly, once per hook file.
if (!warnedMissingHooks.has(hookFile)) {
warnedMissingHooks.add(hookFile);
console.error(
`[gsd-core] hook script missing: ${hookPath} — ${hookFile} is NOT ` +
"enforced. The GSD install may be incomplete; reinstall (or run " +
"/gsd-update) to restage the hooks/ bundle.",
);
}
return { stdout: "", exitCode: 0, timedOut: false };
}
const timeout = opts.timeout ?? 8000;

View File

@@ -6,6 +6,281 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
## [Unreleased]
## [1.8.0] - 2026-07-22
### Added
- **A default-off, BETA, claude-only "Claude orchestration" capability** — adopts Claude Code's Workflow tool (`/effort ultracode`, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, restoring the wave parallelism + plan-checker + verifier that the #853 backgrounded-agent nesting limitation forces inline on Claude Code, and folding the existing `gsd-ultraplan-phase` plan-offload under the same runtime gate. When `claude_orchestration.enabled` is on AND the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets the floor (`claude_orchestration.min_agent_sdk_version`, default `0.3.149`), `execute-phase` emits a generated Workflow script (`waves → parallel() barriers`, `plans → agent({ agentType: 'gsd-executor', isolation: 'worktree' })`, `files_modified overlap → separate sequential stages`, `resumeFromRunId` wired to the phase run id, shared `budget` pool) that composes the SAME executor agent + worktree isolation the inline path uses, so artifacts/commits are produced identically. Detection is pure and fail-closed (any miss → inline), so on any runtime lacking the Workflow tool behaviour is byte-identical to today. Adds a pure module `gsd-core/bin/lib/claude-orchestration.cjs` (`detectWorkflowBackend`, `emitWorkflowScript`), the `capabilities/claude-orchestration/` declaration with two gated loop contributions (`execute:wave:post`, `plan:post`) and a `claude-orchestration` command family (`gsd-tools claude-orchestration detect-backend|emit-workflow`), federated config keys, and an ADR-1143 implementation amendment. (#1143) (#2044)
- **Phases that integrate an external API/SDK/service can no longer seal without a decided coverage matrix** — a new `api-coverage` gate on the `ai-integration` capability blocks `/gsd:verify-work` until the phase produces a `COVERAGE.md` enumerating the API's full capability surface, with every non-integrated capability an explicit, reasoned opt-out. Full coverage is the default; the matrix is the subtraction record, so "we integrated the API" can no longer silently mean "we integrated whatever the first use case exercised." Toggleable via `workflow.api_coverage_gate` (on by default). (#1562) (#2065)
- **OpenCode installs now auto-register the GSD companion MCP server (`mcp.gsd`)** — `--opencode` install writes a `mcp.gsd` entry (local stdio → `gsd-mcp-server`) into `opencode.json`, so OpenCode drives GSD's command + planning-state surface over MCP with no bespoke plugin (ADR-1239 Phase D / #1682). Idempotent and non-clobbering; a user-defined `mcp.gsd` is preserved. (#1682) (#1929)
- **OpenCode plugin handles `session.idle` + the `opencode-subset` hook dialect is implemented** — the GSD OpenCode plugin now recognizes `session.idle` (↔ Claude `Stop` lifecycle point), completing the compaction/idle pair (#1914 shipped compaction). The reserved `opencode-subset` dialect gains a consumer — `hookEventSurfaceFor()` in `host-integration.cts` — describing OpenCode's session/tool/file event subset (no workflow-phase events; the engine owns phase sequencing, ADR-1239 §OpenCode binding). Adds a Claude-parity test asserting the plugin covers the full declared subset. (#1682) (#1930)
- **GSD now warns when model config changed without re-running the installer on static-frontmatter runtimes** — on `codex` and `opencode`, editing `model_overrides` or `model_profile_overrides` or `model_policy.runtime_tiers` in `.planning/config.json` or `~/.gsd/defaults.json` previously had no effect until the user re-ran `gsd install <runtime>`, and the failure was silent: the sub-agent kept using the base model. Workflow entry points like `gsd-tools init *` now emit a one-line stderr warning naming the changed config file and the exact remediation command when they detect the config is newer than the baked agent files. The guard is read-only and warning-only by default, dedup'd per session, and skipped entirely on Claude Code because Claude Code resolves models at spawn time. Resolves #1688 as the structural follow-up to #1650. (#1692)
- **`gsd-tools state rebuild`** — new subcommand that re-derives STATE.md body structure from canonical sources (frontmatter + `.planning/phases/` disk scan), reconciling drifted `## Current Position` prose, dropping orphaned rows from the `**By Phase:**` table, clearing template-placeholder field values, and de-duplicating `## Session Continuity Archive` blocks. Every mutation is recorded in a `## Rebuild Log` audit section. Idempotent (running twice on a clean file is a no-op). Supports `--dry-run` (preview) and `--verbose` (tee log to stderr). Heavier, manual counterpart to the lightweight auto-triggered `state sync`. (#1830)
- **`graphify.graph_path` makes the knowledge-graph location configurable so one umbrella graph can serve multiple projects** — a new `.planning/config.json` key (path relative to project root, or absolute) overrides where `/gsd-graphify query|status|diff` read the graph, letting a single curated cross-repo umbrella graph serve every sibling sub-project without N drifting ~5 MB mirror copies. Previously the graph location was hardcoded to `<cwd>/.planning/graphs/` with no override; the only workaround was copying the umbrella `graph.json` into each project (which drifted, wasted disk, and could be silently overwritten by an in-project build). The diff snapshot travels with the configured graph; build stays project-scoped; unset → byte-identical default; a configured-but-missing file yields an actionable error naming the path. (#1825) (#2013)
- **Claude Sonnet 5 is now the `standard` (sonnet) tier model.** The model catalog and provider presets resolve the sonnet/standard tier to `claude-sonnet-5` (GA 2026-06-30) across the Anthropic-backed runtimes (`claude`, `copilot`, and the `anthropic`/`anthropic-fable` presets), plus the OpenRouter-style `anthropic/claude-sonnet-5` for `opencode`/`hermes`, replacing the superseded `claude-sonnet-4-6`. Opus and Haiku tier defaults are unchanged (the `haiku` high-effort preset's escalation slot tracks the current sonnet model). Shipped in 1.6.1. (#1847) (#1848)
- **`gsd-debugger` now guards fix acceptance with a multi-signal anti-overfitting gate** — a fix that greens the target test can no longer be silently accepted. The debugger now runs a five-signal guardrail before accepting a fix (target test, mutation check via Stryker, no-op/behavior-deleting diff detector, adjacent/held-out tests, and revert-and-reconfirm), degrades gracefully when Stryker or a test suite is absent (each skip is logged, never a silent pass), records every signal's result under `Resolution.verification` in the debug file, and returns a `FIX REJECTED BY GUARDRAIL` outcome that `gsd-debug-session-manager` surfaces for revise / accept-as-documented-debt / abandon. Full rules live in `gsd-core/references/debugger-fix-acceptance.md`. (#1958) (#2396)
- **`gsd-debugger` now ranks suspect code by Ochiai suspiciousness before forming hypotheses** — when a runnable test suite with per-test coverage exists (≥1 failing and ≥1 passing test), the debugger computes a spectrum-based fault-localization (Ochiai) ranking over the coverage and seeds the top-N suspicious locations into the Evidence section as first-class hypothesis candidates, narrowing the search space deterministically before any LLM reasoning. Tarantula is documented as a fallback formula. The step degrades cleanly (logged, never a silent pass) when there is no test suite, no failing tests, or no per-test coverage, and it is explicitly not trusted on flaky/Heisenbug spectra (pairs with the Phase 2B bug-taxonomy routing). Full rules live in `gsd-core/references/debugger-sbfl.md`. (#1959) (#2403)
- **`gsd-debugger` now branches root-cause analysis instead of chaining, guarding against 5-Whys single-cause bias** — before committing `root_cause`, the debugger enumerates candidate causes across ≥2 Ishikawa categories (code / config / environment / data) rather than a single linear "why" chain, and explicitly answers an AND-gate question ("could this failure require more than one contributing condition simultaneously?"). When the AND-gate fires, every contributing cause is recorded — so a multi-cause fix no longer recurs via the unaddressed second cause. `Resolution.root_cause` may now hold one OR a small set of contributing causes (additive; a single-cause session still records exactly one root_cause while the reasoning_checkpoint gains two RCA fields populated in every session). The Structured Reasoning Checkpoint gains `candidate_causes` + `and_gate` fields, and `debugger-philosophy.md` adds the single-cause-bias trap to its cognitive-bias table. Full rules live in `gsd-core/references/debugger-rca-branching.md`. (#1960) (#2405)
- **`gsd-debugger` now classifies each failure by bug class and routes the investigation technique accordingly, replacing the flat 11-technique menu with selection-by-class** — at a new Phase 1.75 the debugger assigns a `bug_class` (Bohrbug / Heisenbug-Mandelbug / Concurrency) and consults an explicit, inspectable routing table: Bohrbugs route to deterministic reproduction + SBFL (Phase 1.25) + git bisect; Heisenbugs/Mandelbugs route to record-replay (`rr`) + stability-stress + statistical sampling and **explicitly skip SBFL** (a flaky spectrum poisons the ranking); Concurrency bugs surface the atomicity/order/deadlock checklist before general techniques. The 11 techniques remain as routed targets, not an undifferentiated list (supersede, not append). `bug_class` + chosen strategy are written to the debug file; the common-bug-patterns catalog is cross-referenced to the taxonomy. Full rules live in `gsd-core/references/debugger-bug-taxonomy.md`. (#1961) (#2407)
- **`gsd-debugger` now hardens regression tests via PBT shrinking, explicit oracle classification, and boundary neighbors** — extending Minimal Reproduction and Test-First Debugging. When a bug triggers on a class of inputs, the debugger wraps the failing input in a property (fast-check for JS/TS, Hypothesis for Python) and lets the shrinker auto-minimize the counterexample, storing the **minimized** input as the regression seed; before writing the assertion it classifies the oracle as `specified` / `derived` (contract/model) / `metamorphic` / `implicit` (crash — weakest, never the silent default) and records it under `Resolution.oracle_type`; and it generates **boundary neighbors** (off-by-one, min/max, empty/singleton) around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — which is what the Phase 1A mutation guardrail needs to bite. Degrades gracefully to manual minimization when no PBT framework is present. Full rules live in `gsd-core/references/debugger-repro-hardening.md`. (#1962) (#2409)
- **`gsd-debugger` now emits a blameless-postmortem Prevention block at resolution, closing the loop on bug-class prevention** — at `archive_session` the debugger produces three blame-free components: a **branching 5-Whys** causal chain (branching per the Phase 2A RCA discipline, not a single linear chain; "agent error" prompts "why was that error possible?", never blame), a **"why wasn't this caught?"** answer naming the existing gate (test/typecheck/lint/review/verify) that missed it, and a **concrete recurrence guard** (a regression test / assertion / lint rule / knowledge-base pattern). The knowledge-base entry gains two structured fields — `why_not_caught` and `recurrence_guard` — so a future Phase-0 recall surfaces not just the prior fix but the prior *prevention* (additive; old entries without the fields still load). The session-manager's compact summary surfaces a one-line prevention summary. Full rules live in `gsd-core/references/debugger-prevention.md`; kept minimal — a block, not an incident-management subsystem. (#1963) (#2410)
- **Third-party capability gates now actually fire via a generic `command-exit-zero` predicate.** — a capability's declared `check.predicate` gate was rendered for display but never evaluated (only built-in `check.query` gates were enforced, and the `security` capability's gate worked solely via a hard-coded `ship.md` branch). A new generic evaluator (`gsd_run check predicate`) now evaluates `check.predicate` blocks by `kind`; the first built-in kind `command-exit-zero` runs a bounded `sh -c` command at the project root and blocks the loop on non-zero exit (timeout → block, fail-closed). The `execute:wave:post`, `execute:post`, and `plan:post` gate-dispatch sites route `predicate` gates to the new evaluator automatically. (#2008) (#2011)
- **GSD's lifecycle hooks now run under Kimi CLI** — installing GSD into Kimi wires its session-state, phase-boundary, graphify, and guard hooks into Kimi's own native `config.toml` `[[hooks]]` bus (Beta on Kimi's side) instead of silently no-op'ing, and GSD's Kimi subagents can now run in the background. Kimi's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2095) (#2159)
- **GSD is now installable on pi** — `npx @opengsd/gsd-core --pi` installs the GSD extension to `~/.pi/agent/extensions/gsd.cjs`, and `/gsd <family> <subcommand>` now dispatches real commands through the embedded engine (the reference binding previously could only run `query help`). Drives pi through the negotiated imperative Host-Integration adapter, with active-model steering and the full pi lifecycle-event surface. (#2102) (#2205)
- **The EoS Registry now lists GSD for Oh My Pi** — discover the independently maintained `tchivs/gsd-omp` protocol-v1 host integration, including exact install and uninstall commands, supported interface points, and negotiated host axes. (#2448)
- **Broken-windows ledger** — `/gsd:ship` now blocks (when `workflow.windows_enforce=true`, opt-in) while `.planning/WINDOWS.md` has any `open` entry, and the executor auto-populates the ledger with stubs, skipped tests, and unrun verifies as it works. Each window can be `waived` only with a recorded reason (auditable) or `fixed` (removed from the blocking set); `/gsd:progress` surfaces the open + waived counts. Backward-compatible: projects with no ledger ship cleanly (open_count starts at 0), and enforcement is off by default so tracking can precede the gate. Enable with `gsd config-set workflow.windows_enforce true`. (#1950) (#2441)
- **GSD now ships a pi extension** — a real, jiti-loadable ExtensionAPI module (`pi/gsd.cjs`) that registers `/gsd` (dispatches through the GSD command-routing hub) + `gsd_invoke` tool + `tool_call` event, installable at `~/.pi/agent/extensions/`. A reachability test proves the `/gsd` handler dispatches through the engine (keystone wired, not just registered on a mock). (#1965) (#1965)
- `plan-phase` now authors edge and prohibition predicates into PLAN.md `must_haves` when a phase SPEC omits `## Edge Coverage` / `## Prohibitions`, so goal-backward verification still has predicates to check on a spec-less phase (ADR-857 Phase 6). Gated by the new default-on `workflow.specless_probe_fallback` toggle — disable it to skip the fallback (the skip is recorded visibly in the plan). Spec-less prohibitions are authored descriptor-less and disposed flagged/unverified (honest verifier #1154), never a silent pass. (#1835)
- **Discover third-party GSD Capabilities in a new Community Capability Registry.** — A non-endorsing discoverability catalog where authors register a Capability via a documentation PR; each entry carries a live latest-release badge and a per-entry GitHub Discussion for community ranking and comments. (#2188) (#2188)
- **GSD now warns when a stale global CLI (e.g. a retired @gsd-build/sdk canary) shadows your project-local install** — the gsd-tools CLI startup detects when the running binary is outside the project root while a project-local install exists, and prints a remediation warning to stderr (non-blocking). (#1754) (#1755)
- **`gsd-mcp-server` — companion MCP server (interface points 1 + 5)** — a new bin command (`npx @opengsd/gsd-core gsd-mcp-server`) runs a stdio JSON-RPC 2.0 MCP server exposing `gsd_invoke_command` (→ the GSD command-routing hub) + `gsd_read_state` / `gsd_write_state` (→ `.planning/` state), so any MCP-consuming host (Claude Code, Codex, OpenCode, VS Code, Gemini CLI, Cursor, Cline, Hermes) can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681). Dependency-free (hand-rolled JSON-RPC). How-to: `docs/how-to/connect-gsd-mcp-server.md`. (#1810)
- **Opt-in absolute token count on the statusline context meter** — new `statusline.show_context_tokens` config (default `false`). When enabled, the meter shows the absolute context total after the percentage, e.g. "████░░░░░░ 46% (156k)", summing input, cache-creation, cache-read, and output tokens from the hook payload (a broader basis than the meter's percentage, which is derived from `used_percentage` and excludes output tokens — the two figures can diverge slightly). Default meter output is unchanged. (#2161) (#2174)
- **Long-running compute can now be externalized as async external jobs instead of blocking the agent turn** — a default-off external-job capability lets executors submit SLURM jobs, commit a .planning/async-jobs manifest, defer SUMMARY.md, and return external_job_waiting; the core loop already reconciles these manifests (#1165), so this adds the producer half (SLURM adapter, pure manifest module, planner/executor fragments, operation policy). (#1105) (#1998)
- **GSD now ships a repo-local VS Code extension** — a buildable extension (`vscode/extension.js` + `vscode/package.json`) that registers `gsd.invoke` (dispatches through the GSD command-routing hub) in the VS Code command palette. A reachability test proves the handler dispatches through the engine (keystone wired). Not Marketplace-published; mirrors the OpenCode plugin's bar. (#1966) (#1966)
- **Discover third-party GSD Embeddable Orchestration System (EoS) integrations in a new EoS Registry.** — A non-endorsing discoverability catalog where host-integration authors register via a documentation PR; each entry declares its Host-Integration interface points, negotiated axes, and protocol version, with a live release badge and a per-entry GitHub Discussion for ranking and comments. (#2193) (#2193)
- GSD Core ships a `.claude-plugin/marketplace.json` marketplace manifest so Claude-plugin-compatible runtimes (ZCODE et al.) can discover and install gsd-core from a custom marketplace source. Additive — the existing `.claude-plugin/plugin.json` and the Claude Code install path are unchanged. The catalog version (`plugins[0].version`) tracks `package.json` via the release version-sync. (#1861)
- **GSD now drives VS Code through the Embeddable Orchestration System** — the VS Code extension is rewired through the negotiated imperative Host-Integration adapter (active `vscode.lm` model, engine hook bus, sandboxed storage), gains native Language Model Tools (GSD skills as `#gsd-*` tools) and `#runSubagent` dispatch, and runs as a Web Extension (no Node APIs). (#2103) (#2210)
- **`/gsd:next` smart-entry workflow** — adds a state-aware entry point that classifies the current project situation (no-project, blocked, verify-failed, planning, executing, verify-pending, complete, and more) and recommends the right next GSD command. The `gsd-tools smart-entry [--json]` classifier handles phase ordering including decimal phase IDs; the `/gsd:next` skill surfaces the workflow with tiered fallback behavior. (#1798)
- OpenCode now runs GSD's lifecycle safety hooks (prompt-injection guard, read-before-edit guard, injection scanner, worktree/workflow guards, context monitor) via a native plugin installed to `~/.config/opencode/plugins/gsd-core.js`. OpenCode declares `hooksSurface: 'none'`, so these hooks were previously inert; the plugin bridges OpenCode's event bus onto GSD's existing hook scripts. Installed automatically by `npx @opengsd/gsd-core --opencode` and removed on uninstall. (#1923)
- **Opt-in compact GSD-state statusline format** — new `statusline.state_format` config, enum `full`|`compact` (default `full`, the existing rendering). `compact` renders "<version> · P<phase>/<total> · <status>" (e.g. "v1.12 · P7/12 · executing"), dropping the milestone name and progress bar and collapsing narrative statuses to the canonical vocabulary from `normalizeStateStatus()` — the canonical stuck state `paused` renders uppercase as `PAUSED`. Solves the unbounded-width problem where free-text status sentences push the context meter off the line. (#2162) (#2175)
- **`<precondition>` task element (Design by Contract)** — plans may now declare a runnable/checkable fact a task assumes (env var set, prior-phase artifact present, external-setup done) that plan ordering does not guarantee; the executor asserts it before running the task and halts with a checkpoint on unmet instead of building on a broken assumption. Plans that omit `<precondition>` behave exactly as today. (#1949) (#2422)
- **Config-gated provider escalation when a run hits a quota or rate limit** — an executor killed by a provider throttle stopped the phase and waited for a manual restart; escalating a tier did not help because the same throttled provider was still in play. Set `dynamic_routing.provider_escalation` to an ordered list of fallback model IDs and GSD now switches provider on a quota-exceeded failure, logs the swap (`sonnet → gpt-5`), honors the provider's `Retry-After`, caps the walk at `max_escalations`, and names every model tried once the list is spent. Opt-in — unset, quota failures keep today's manual recovery prompt. (#2296) (#2458)
- **Host-integration descriptors now carry an `extensionEvents` vocabulary** — the extension-system event surface (OpenCode, pi) is a separate descriptor field from managed `hookEvents`, so OpenCode declares `extensionEvents:opencode` without conflicting with the hooksSurface:none invariant. (#1946) (#1946)
- **`/gsd-review` now supports custom reviewer instances** — run one model-capable adapter (e.g. OpenCode) as several independent reviewer identities via a bounded `review.reviewer_instances` config, so two different models can review in a single pass without manually swapping config or hand-merging REVIEWS.md. (#1517) (#1766)
- **Opt-in git branch and working-state segment in the statusline** — the shell prompt's branch/dirty-state signal is hidden for the whole session under the Claude Code TUI, so wrong-branch commits and ship-time push rejections surface only after the fact. New `statusline.show_git` config (default `false`) renders the branch name plus staged/unstaged/untracked/ahead/behind markers (or ✓ when clean and in sync) after the directory segment. When disabled, no git subprocess is spawned and output is unchanged. (#2163) (#2183)
- **`/gsd:onboard` guides brownfield setup** — existing repos now have a top-level onboarding command that routes through codebase mapping, docs ingest, project initialization, and an onboarding summary without silently overwriting planning files. (#1994)
- **Plural/optional/chosen assumption-delta checkpoint during planning** — when a phase makes something plural, optional, or chosen that used to be singular, required, or derived, the planner is now prompted to re-ask whether the primary key / identity model still names the right thing, preventing silent architectural drift from accumulating into a later user-facing bug. Advisory (non-blocking); fires only on a detected signal. Toggle with workflow.assumption_delta. (#1561) (#1767)
- **`/gsd-ui-phase` now probes UI state coverage** — a new `ui-consideration-probe` (the third `probe-core` adapter) enumerates the shape-rooted UI states a UI-SPEC must resolve (empty/loading/error/populated/partial/overflow/zero-one-many/long-text). After the UI checker approves, the probe surfaces applicable considerations for each element, records a `## UI Considerations` section in the UI-SPEC, and plan-phase lifts each resolved consideration into `must_haves` — so a purely-visual state with no wired test routes to `insufficient_spec → human_needed` at verify rather than a silent pass. (#1979)
- **Host-Integration Interface (ADR-1239 Phase A)** — a versioned, negotiated capability contract (`runtime.hostIntegration`) over the six host-integration points (command, dispatch, model, hooks, state, artifact). Adds an in-process `negotiateHostCapabilities` handshake that fail-closes on undeclared/unknown/`undocumented` values (`effective ⊆ host-declared ∩ engine-known`), a typed degradation ladder, host-capability profiles, and a documentation-sourced per-CLI capability matrix for all 16 runtimes. Interface-definition only — no change to install behaviour. (#1690)
- **ZCode (Z.ai) is now an installable runtime** — a desktop Agentic Development Environment for the GLM-5.2 model can now be targeted with `--zcode`, landing GSD skills at `~/.zcode/skills/<name>/SKILL.md` plus slash commands and subagents. ZCode ships as a pure declarative capability descriptor (`capabilities/zcode/capability.json`) with zero hardcoded `runtime === 'zcode'` branches, reusing the Claude skill converter — the de-hardcoded, data-driven runtime path that 1.7.0 (ADR-1016 / ADR-1239) enables. (#1925) (#2039)
- **Reversibility tagging for planning decisions** — decisions can now be rated `reversible`, `costly`, or `one-way` by how expensive they are to undo. A `one-way` decision (one whose undo needs a data migration, breaks a published contract, or is impossible) earns a `checkpoint:decision` before the task that implements it, so an unattended run pauses for your sign-off instead of walking through the door. `costly` decisions are flagged in the plan without blocking; `reversible` ones flow as before. Pass `--no-reversibility-gates` to `/gsd:plan-phase` to suppress the checkpoint on runs you mean to leave unattended — ratings are still recorded either way. (#1951) (#2471)
### Changed
- **`gsd-debugger` now recalls prior resolved sessions semantically via MemPalace instead of keyword overlap** — at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions as candidate hypotheses, catching the same-root-cause / different-wording cases keyword overlap missed (a prior "requests hang under load" now surfaces for "API times out when many users connect"). Resolved sessions are indexed into MemPalace at archive (symptoms + root cause(s) + fix + recurrence guard). `knowledge-base.md` remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching against it (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Full rules live in `gsd-core/references/debugger-semantic-recall.md`. (#1964) (#2416)
- **The GSD CLI now self-heals a missing runtime build.** The compiled `gsd-core/bin/lib/*.cjs` modules are gitignored build artifacts (ADR-457) that ship prebuilt in the npm tarball but are absent on a Claude Code plugin-marketplace / git-clone install, which never runs `npm run build:lib`. Previously every command died at load with `Cannot find module './lib/cli-exit.cjs'`. The `gsd-tools` entrypoint now detects the missing output and compiles it once, on demand (lock-guarded so parallel invocations don't race), then proceeds — a single no-op check on the already-built npm path. When TypeScript is genuinely unavailable it prints an actionable `npm install && npm run build:lib` message instead of crashing. (#2036)
- **Internal: Claude Code's installer is now driven through the public Host-Integration Interface (ADR-1239 / EoS).** `bin/install.js` routes `claude` install/uninstall through the imperative adapter (`createImperativeAdapter`) instead of calling the engine directly, and its 13 hardcoded `runtime === 'claude'` / `runtime !== 'claude'` branches are folded into descriptor-driven `runtime.hostBehaviors` on `capabilities/claude/capability.json` (permission schema, `settings.local.json` scope routing, `.gsd-source` marker, effort frontmatter, canonical-workflow authorship, and more). Install/uninstall output is **byte-identical** for both the global skills layout and the local legacy layout (golden-parity asserted for both scopes); no other runtime changes. Removes the "add-a-host tax" of scattered string-equality checks for the tier-1 reference host. No user-facing change. (#2086) (#2106)
- **OpenCode is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS).** OpenCode and its Kilo sibling previously installed via a bespoke `runtime === 'opencode'`/`isOpencode` branch in `bin/install.js`; its commands+skills+plugin install now runs through the imperative adapter → the engine's combined-family install path (`installRuntimeArtifacts`), and every hardcoded `runtime === 'opencode'` branch is folded into descriptor-driven `runtime.hostBehaviors`. Install/uninstall output is **byte-identical** (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) **background dispatch** — OpenCode shipped experimental background subagents in v1.15 and made them default-on in v1.17, so `dispatch.background`/`backgroundDispatch` flip to `true`; GSD no longer force-flattens OpenCode-hosted wave dispatch (`shouldFlattenDispatch` now returns `false`), letting agents run concurrently where the host supports it. (2) **expanded event surface** — the OpenCode plugin now subscribes to `permission.asked`, `permission.replied`, and `session.error` (added to `EXTENSION_EVENT_SURFACES.opencode`), wiring the declared surface for future permission/error-aware bindings. (#2087) (#2108)
- **Codex is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS).** Codex previously installed via hardcoded `runtime === 'codex'`/`isCodex` projection in `bin/install.js`; its `config.toml` / agent-`.toml` / `hooks.json` install now runs through the declarative embedding adapter and descriptor-driven `runtime.hostBehaviors`, with **zero** positive `isCodex` gates and **zero** `runtime === 'codex'` branches remaining (source-guarded). Install/uninstall output stays byte-parity-gated (`tests/fixtures/golden-install-parity/codex.json`). Three Context7-verified upgrades land, each with a test driving the user-reachable surface: (1) **skill root** — GSD skills now install to Codex's canonical `$HOME/.agents/skills` (via a skills-kind `home` override) instead of the deprecated `$CODEX_HOME/skills` fallback, and pre-move installs are migrated (stale `~/.codex/skills/gsd-*` cleaned on both install and uninstall, user-owned content preserved); (2) **hook events** — GSD registers the six documented Codex lifecycle events it previously skipped (`PreToolUse`, `PermissionRequest`, `PreCompact`, `PostCompact`, `SubagentStop`, `UserPromptSubmit`, in addition to the existing `SessionStart`/`SubagentStart`/`Stop`/`PostToolUse`) in `hooks.json`, so `gsd-context-monitor` fires at the same points as in Claude Code, and the descriptor `extendedHookEvents` is reconciled from `[]` to the schema-valid wired subset; (3) **dispatch tuning** — `[agents] max_depth = 1` is written explicitly into the managed `config.toml` block to pin the negotiated `dispatch.maxDepth: 1` axis (`degradationFor` flattens GSD-hosted waves to single-level), and `validateCodexConfigSchema` now permits a known-scalar-only `[agents]` AgentsToml table (coexisting with the flattened `[agents.gsd-*]` role sub-tables) while still rejecting the `[[agents]]` and unknown-key break-forms from #2760. (#2088) (#2110)
- **Cursor is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS).** Cursor previously installed via hardcoded `runtime === 'cursor'`/`isCursor` branches in `bin/install.js`; its install/uninstall now runs through the imperative adapter, and every hardcoded cursor branch is folded into descriptor-driven `runtime.hostBehaviors` (reapplyCommand, frontmatterDialect, hooksJsonSurface, skipSharedHooksInstall, reportCommandsDir, managedHookEvents). Install/uninstall output is **byte-identical** (golden parity asserted for all 16 runtimes). Two Context7-verified upgrades land: (1) **expanded hook-bus coverage** — GSD registers all 6 managed lifecycle events in Cursor's `hooks.json` (`preToolUse`, `stop`, `subagentStart`, `subagentStop` in addition to the original `sessionStart`/`postToolUse`), driven by a new descriptor-driven adapter module (`src/host-integration-adapters/imperative-hook-bus.cts`) that reads `hostBehaviors.managedHookEvents` instead of a hardcoded event pair; cite https://cursor.com/docs/hooks. (2) **named/background nested subagent dispatch** — Cursor's `dispatch.background`/`backgroundDispatch`/`nested` are all `true` with `maxDepth: 2`, so `shouldFlattenDispatch(cursor)` returns `false` and GSD's wave-based execution drives Cursor's native background + depth-2 nested subagent invocation instead of flattening to inline sequential calls; cite https://cursor.com/docs/subagents + https://cursor.com/docs/sdk/typescript. (#2089) (#2120)
- **Cline is now driven through the public Host-Integration Interface, with two capability upgrades (ADR-1239 / EoS).** Cline previously installed via hardcoded `runtime === 'cline'`/`isCline` branches in `bin/install.js`; its install/uninstall now runs through the imperative adapter, and every hardcoded cline branch is folded into descriptor-driven `runtime.hostBehaviors` (reapplyCommand, frontmatterDialect, skipSharedHooksInstall, localTargetIsProjectRoot, clineRulesSurface, localCommandsViaRules). Install/uninstall output is **byte-identical** (golden parity asserted for cline + claude/cursor/codex/opencode). Two Context7-verified upgrades land: (1) **`AgentPlugin.hooks.beforeTool` planning guard** — the `.clinerules/hooks/PreToolUse` file-convention hook (#787) is re-implemented as a real Cline SDK `AgentPlugin` that cancels write-class calls targeting `.planning/` (same fail-open semantics), driven by a new descriptor-driven adapter module (`src/host-integration-adapters/cline-sdk-binding.cts`); cite https://github.com/cline/cline/blob/main/docs/sdk/plugins.mdx. (2) **`createAgentModel` model overrides** — `DefaultGateway.createAgentModel({providerId, modelId})` is wired so GSD's per-subagent `model_overrides`/`model_profile_overrides` resolution applies to Cline subagents (`modelMode: active`); cite https://github.com/cline/cline/blob/main/docs/sdk/reference/gateway.mdx. Cline's dispatch deliberately stays **degraded/flat** (`maxDepth: 1`, read-only, no nested spawning) per the documented host restriction — never silently upgraded to full nested/background. (#2090) (#2132)
- **Hermes Agent is now driven through the public Host-Integration Interface, with three capability upgrades (ADR-1239 / EoS).** Hermes previously installed via hardcoded `runtime === 'hermes'`/`isHermes` branches in `bin/install.js`; its install/uninstall now runs through the imperative adapter, and every hardcoded hermes branch is folded into descriptor-driven `runtime.hostBehaviors`. Three upgrades land: (1) **real plugin hook vocabulary** — GSD registers a new `extensionEvents: "hermes"` dialect carrying the 13 documented Hermes plugin events (`pre_tool_call`, `post_tool_call`, `pre_llm_call`, `post_llm_call`, `on_session_start`, `on_session_end`, `on_session_finalize`, `on_session_reset`, `subagent_start`, `subagent_stop`, `pre_gateway_dispatch`, `pre_approval_request`, `transform_tool_result`), replacing the borrowed `hookEvents: "claude"` 6-event surface that silently never fired; cite https://github.com/nousresearch/hermes-agent/blob/main/website/docs/user-guide/features/hooks.md. (2) **dispatch posture** — Hermes' `dispatch.nested: true` with `maxDepth: 1` is correctly negotiated (not silently flattened). (3) **branding/category metadata** — `DESCRIPTION.md` category descriptions, `version:` frontmatter, and branding rewrites are now descriptor-driven rather than hardcoded. Install/uninstall output is byte-identical (golden parity asserted for all runtimes). (#2091) (#2134)
- **Qwen Code now projects GSD's specialist agents as native subagents** — installing GSD into Qwen Code writes `~/.qwen/agents/gsd-*.md` files you can invoke directly (planner, executor, code-reviewer, …) instead of reaching them only through skill prose, and a `SubagentStart` hook now fires alongside `SubagentStop`. Qwen's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2092) (#2153)
- **Kilo Code now supports native hooks, active-model routing, and named subagent dispatch** — installing GSD into Kilo wires a lifecycle-hook plugin, keeps each agent's requested model instead of dropping it, projects GSD's specialist agents as invokable subagents, and documents the GSD MCP companion. Kilo's install is driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2093) (#2156)
- **GSD skills installed for Trae now carry SOLO stage metadata** — Trae's SOLO Agent can recognize GSD skills as workflow-stage skills for auto-invocation instead of requiring manual triggering. Several of Trae's install branches (shared-hooks gating, path rewrites) also move onto its capability descriptor. Note: the stage-metadata field is a best-effort/inferred shape — Trae publishes no formal schema. (#2094) (#2157)
- **Installing GSD into Antigravity now writes the `permissions.allow` rules its CLI documents** — so GSD's own reads and hooks aren't stuck on interactive prompts — and registers GSD's companion MCP server via a standalone `mcp_config.json` (best-effort: Antigravity's raw config schema isn't published, so this uses the Gemini-CLI-successor format). Antigravity's install is now driven by its negotiated capability descriptor instead of hardcoded runtime special-cases. (#2096) (#2165)
- **Augment Code now installs through its capability descriptor, with a native MCP companion** — installing GSD into Augment registers the GSD companion server in Augment's `settings.json` `mcpServers` and drives command/skill/agent conversion from Augment's negotiated descriptor instead of hardcoded runtime special-cases. (#2097) (#2166)
- **CodeBuddy now wires GSD's full extended lifecycle hook set and is driven by its capability descriptor** — installing GSD into CodeBuddy now registers `SubagentStart`, `SubagentStop`, `Stop`, and `PreCompact` hooks in its `settings.json` (it previously had none of these), matching the coverage Qwen/Kimi already ship, and CodeBuddy's install is fully descriptor-driven instead of via residual hardcoded runtime branches. (#2098) (#2169)
- **GitHub Copilot now wires GSD's full lifecycle hook bus and is driven by its capability descriptor** — installing GSD into Copilot registers `preToolUse`, `postToolUse`, `userPromptSubmitted`, and `sessionEnd` handlers in its `hooks/gsd-session.json` (beyond today's `sessionStart`-only advisory), and Copilot's residual hardcoded runtime branches are folded onto descriptor-driven `hostBehaviors`. (#2099) (#2172)
- **Windsurf now enforces GSD's write/command safety guards through Cascade's native hook bus** — installing GSD into Windsurf registers blocking `pre_write_code`/`pre_run_command` hooks in `.windsurf/hooks.json` (exit-code-2 blocking) and drives Windsurf's install from its capability descriptor instead of hardcoded runtime branches. (#2100) (#2190)
- **ZCode's install is now driven and regression-tested through its capability descriptor** — ZCode joins the dogfooded declarative-adapter reference hosts with a byte-identical install, and its shared-hooks exclusion is folded onto `hostBehaviors` instead of a hardcoded runtime branch. (Hook-automation and MCP upgrades remain blocked on ZCode publishing its on-disk config formats.) (#2101) (#2195)
- **Codex/OpenAI default models advance to the GPT-5.6 family (Sol/Terra/Luna)** — the Codex runtime tier defaults and the `openai` provider preset now resolve to current-generation model IDs instead of the superseded GPT-5.4/5.5 line, so Codex users on default profiles get improved agentic coding (Sol) and lower costs (Terra/Luna) without changing any config. (#2122) (#2146)
- **Internal: the installer's `program` (display-name) + `command` (slash-invocation) chains are now single-source lookups** — the 14-line `program` chain (an exact duplicate of `runtimeLabel`) → `getRuntimeLabel`, and the 14-line `command` chain (the per-runtime `/gsd-new-project` syntax: gemini `/gsd:`, codex `$`, cursor skill-mention, kimi `/skill:`, default `/gsd-new-project`) → new `getRuntimeNewProjectCommand(runtime)` helper (ADR-1239 Phase B / #1679 AC2 slice 4). `runtime ===` count in `bin/install.js`: 53 → 25 (cumulative this session: 129 → 25). Stdout strings preserved byte-for-byte; no install-output change (golden-parity 16/16). No user-facing change. (#1813)
- **Internal: the installer's per-function `is<Runtime>` flag-declaration blocks are now a single `runtimeFlags` lookup** — the four duplicated `const isX = runtime === 'x'` blocks in `bin/install.js` (uninstall / writeManager / install / a fourth helper — 48 branches) are collapsed into one `runtimeFlags(runtime)` helper in `runtime-name-policy.cts` (ADR-1239 Phase B / #1679 AC2 slice 3). The add-a-host tax for flags is removed (one `RUNTIME_FLAG_IDS` entry, not four declaration blocks). Install output is byte-identical for all 16 runtimes (golden-parity asserted); `runtime ===` count in `bin/install.js`: 101 → 53. No user-facing change. (#1811)
- **Internal: third-party descriptor loader enforces `configHome` write-confinement at load time** — `loadRegistry({includeInstalled:true, configHome})` now rejects (skip + warn, fail-closed) any installed third-party host-plugin descriptor whose declared `destSubpath` resolves outside the supplied `configHome`, before it is composed into the registry (ADR-1239 Phase C-2 / #1681 slice 2). The `configHome` option is optional and backward-compatible (omitted → no load-time check; install-time gate still bounds writes). No user-facing change for existing flows. (#1808)
- **Internal: agent install for cursor/windsurf/augment/trae/codebuddy now flows through the descriptor path** — ADR-1235 step 1 routes the trivial-converter runtime group's agents off the inline install() loop onto the descriptor-driven `installRuntimeArtifacts` path, applying the cross-cutting steps uniformly (pre-converter, no workflow-stamp). Agent output is byte-identical for all 16 runtimes (golden-parity asserted, global + local verified); no user-facing change. (#1764)
- **gsd-ui-checker gains an adversarial FORCE stance (LLM-playbook principle 16)** — the only verdict-producing critic that lacked one now resists rubber-stamping UI-SPEC contracts, with BLOCK/FLAG/PASS classification. Based on arXiv 2505.23840 (third-person objective persona), 2506.04975 (objective-not-hostile persona). (#1584)
- **Internal: the declarative embedding adapter is now named + bound behind a minimal `HostIntegrationInterface`** — `createDeclarativeAdapter({runtime})` (new `src/adapter-declarative.cts`) delegates in-process to `install-engine`'s `installRuntimeArtifacts`/`uninstallRuntimeArtifacts`, formalizing today's projection path as one of the two embedding adapters behind a common contract (ADR-1239 Phase C-1 / #1680 AC1). Output is byte-identical to today's install (gated by `golden-install-parity`). The full 6-point interface binding surface is deferred until the imperative adapter (AC2) fixes the shape (ADR-1239 open wire-shape question). No user-facing change — the adapter is not yet wired to any runtime path. (#1802)
- **Internal: getDirName is now derived from a documented `runtime.localConfigDir` descriptor field** — each runtime's local content-rewrite directory (e.g. `cursor`→`.cursor`, `copilot`→`.github`) moved from a hand-maintained if-chain into its capability descriptor (ADR-1239 Phase B), so it can no longer drift from the registry. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1757)
- **Internal: copyWithPathReplacement converter selection is now data-driven** — the installer's back-compat content-copy path replaced its 13 hardcoded `runtime === 'x'` flag chains with a single per-runtime dispatch table (ADR-1239 Phase B). Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing change. (#1759)
- **Phase-completion now writes `Status: All phases complete` instead of the overloaded bare `Milestone complete`** — the phase-level completion verb (`completePhaseCore`) was writing the same bare 'Milestone complete' string that the milestone-close verb uses for terminal state, causing a phase-level verb to own a milestone-level field. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts); milestone termination ('<version> milestone complete' / 'Awaiting next milestone') remains solely with the milestone-close verb. (#2204) (#2259)
- **#853 dispatch-flatten is now data-driven (ADR-1239 Phase B)** — whether GSD backgrounds the plan/execute orchestrator is decided from a documentation-sourced `backgroundDispatch` capability per host (via `gsd_run query dispatch-should-flatten`) instead of a hardcoded `runtime === 'codex'` check. **Cursor now backgrounds the orchestrator** (its docs document backgrounded subagent nesting); codex unchanged; all other hosts run inline. Fail-closed to inline on any uncertainty. (#1719)
- **Internal: companion MCP server module (interface points 1 + 5)** — `handleMessage`/`runServer` (new `src/mcp-server.cts`) is a minimal, dependency-free stdio JSON-RPC 2.0 server exposing `gsd_invoke_command` (→ the command-routing hub) + `gsd_read_state`/`gsd_write_state` (→ the Phase 3 stateIO seam), so any MCP-consuming host can drive GSD with no bespoke plugin (ADR-1239 Phase C-2 / #1681 slice 3a). Bin entry / packaging deferred to slice 3b. No user-facing change — the server is not yet wired to a bin entry. (#1809)
- **`requirements mark-complete` reports a per-surface write-set** — the command now returns a per-requirement `write_set` (checkbox + traceability surfaces) and a `write_set_complete` that is true only when every surface of every requirement applied, so a partial (checkbox-only) reconcile can no longer masquerade as full success even inside a multi-ID batch. Introduces the reusable ADR-2143 §5/§6 `Result` / `WriteSet` contract. (#2251) (#2251)
- **Internal: the imperative embedding adapter now composes the capability registry behind the same `HostIntegrationInterface`** — `createImperativeAdapter({runtime})` (new `src/adapter-imperative.cts`) calls `loadRegistry({includeInstalled:true})` (first-party-wins + consent + fail-closed — identical trust semantics to the CLI) and binds the engine surface behind the same contract the declarative adapter (AC1) satisfies, plus a `registry` accessor for an in-process host to bind its primitives to (ADR-1239 Phase C-1 / #1680 AC2). Concrete host binding is deferred to Phase 5. No user-facing change — the adapter is not yet wired to any runtime path. (#1803)
- **Internal: the model adapter seam exposes `passive` + `active` adapters selected by `modelMode`** — `createModelAdapter({modelMode})` (new `src/model-adapter.cts`): `passive` formalizes today's tier routing (delegates to `model-resolver.resolveModelForTier`), `active` is a host-supplied `sendRequest` seam (VS Code `vscode.lm` / pi providers), fail-closed until Phase 5 binds a concrete provider (ADR-1239 Phase C-1 / #1680 AC3). No user-facing change — the seam is not yet wired to any runtime path. (#1804)
- **Internal: derive the non-Claude runtime list from the capability registry** — `NON_CLAUDE_RUNTIMES` is now computed from the capability registry instead of a hand-maintained literal, so it can no longer drift from the per-runtime descriptors. No user-visible behavior change (the list is identical). (#1728)
- **Honest verifier — verify-phase now abstains on non-inferable `backstop` truths instead of confidently false-passing them (#1154).** When the spec's edge-probe marks a truth non-inferable (`verification: backstop`) and the verifier cannot confirm it with explicit evidence (a passing wired held-out/property test, or a directly-observed behavior), it now reports `human_needed` with reason `insufficient_spec` ("unverified — held-out test recommended") rather than a silent `passed`. Autonomous runs complete with "N unverified non-inferable checks"; interactive runs route to the end-of-phase human checkpoint. Inferable truths are never abstained (over-abstention guard); abstention is exogenous (driven by the tag, not self-judgment). Truth-axis mirror of the prohibition judgment-tier (ADR-550 D4). (#1738)
- Document Claude Code's advisor-tool inheritance in the model-profiles reference: the session-level advisor is inherited by all GSD subagents and composes with per-agent tiering, with candidate executor/advisor pairings, when it is worth enabling, and the session-level (no per-agent control) constraint. (#1922)
- **Extraction discipline for strict-format agents (LLM-playbook principle 8)** — gsd-doc-classifier and gsd-doc-synthesizer apply taxonomy/precedence rules directly without inventing content, reducing reasoning-induced format drift. Based on arXiv 2504.05081 (few-shot beats CoT for pattern tasks), 2506.00069 (terminal instruction placement), 2505.14810, 2505.11423. (#1584)
- **Internal: extracted the runtime-artifact install engine from `bin/install.js`** — `installRuntimeArtifacts`/`uninstallRuntimeArtifacts`/`installOpencodeFamilySkills` and their helpers now live in a dedicated `gsd-core/bin/lib/install-engine.cjs` module (ADR-1239 Phase B), so adapters can import the install pipeline instead of reaching into the 12k-line installer. Install output is byte-identical for all 16 runtimes (golden-parity asserted); no user-facing behaviour change. (#1735)
- **MemPalace `memory_mode` `kg_backend` and `replace` are now functional** — selecting either mode now routes recall through the palace instead of silently behaving like `augment`: `kg_backend` treats the palace temporal KG as the primary knowledge-graph source (native `.planning/graphs/` as fallback), and `replace` resolves recall through the palace as the source of truth. Every mode stays default-resilient — an unreachable palace falls back to native memory and no memory is lost. (#2010) (#2010)
- **`/gsd:surface` and `--materialize` now produce byte-identical agent output to a fresh install** — surface-path agents for descriptor-driven runtimes (cursor, windsurf, augment, trae, codebuddy, copilot, antigravity) now receive the same path-prefix rewrite, Co-Authored-By attribution, runtime-specific conversion, and body normalization as the install path. Copilot and Antigravity agents are now installed via the descriptor-driven path (copilot agents get the `.agent.md` filename rename). Cline remains on the inline loop (rules-only local branch). (#1575) (#2040)
- **Internal: hook-bus + stateIO adapter seams** — `createHookBus({bus})` (new `src/hook-bus.cts`, `host`/`engine`/`none` — engine is in-process pub/sub, host fail-closed, none silent) + `createStateIO({io})` (new `src/state-io.cts`, `filesystem`/`sandboxed-storage`/`session-log-append` — filesystem delegates to fs, the rest are fail-closed seams) (ADR-1239 Phase C-1 / #1680 AC4). Completes the Phase 3 adapter seam layer; concrete host binding is Phase 5. No user-facing change. (#1805)
- **Long-context model names render compactly in the statusline** — the verbose " (1M context)" suffix Claude Code appends to the model display name now collapses to a compact " (1M)" badge (tolerant of future window sizes and the abbreviated "ctx" variant: "(500K context)" → "(500K)", "(1M ctx)" → "(1M)"). Lossless — the long-context signal stays, the 12 characters of width don't. (#2160) (#2173)
- **Lazy-split `plan-phase.md` into a `steps/` directory** — ~4.7 KB lighter eager context per `/gsd-plan-phase` call via byte-invariant progressive disclosure (ADR-1610). (#1852) (#1934)
- **GSD subagents now self-load configured agent_skills regardless of orchestrator bash** — projects that map skills via `.planning/config.json` `agent_skills.<agent-type>` no longer silently lose them on `/gsd-autonomous` or Cursor, where `Skill()`-delegated workflow bash init did not reliably run. Each of the 22 consumer agents queries its own type at init and reads the listed skills, with a dedup guard so runtimes that also inject orchestrator-side (Claude Code) never carry two copies. (#1866) (#1868)
- **Internal: install/uninstall runtime labels are now sourced from a single `getRuntimeLabel` lookup** — the two duplicated `runtimeLabel` assignment chains in `bin/install.js` (uninstall + install) are collapsed into one curated label table in `runtime-name-policy.cts`, sibling to the registry-derived `getDirName` (ADR-1239 Phase B, #1679). Install output is byte-identical for all 16 runtimes (golden-parity asserted). Two console-label inconsistencies are normalized as a side effect: `kimi` shows 'Kimi CLI' in both sites, and `cline` uninstall no longer falls through to 'Claude Code'. (#1800)
- **Phase plans now lead with a verified end-to-end "tracer" slice by default** — every plan starts with one thin, production-quality slice wired through every layer, which the executor verifies before building out the remaining tasks, so an architectural dead-end surfaces after one commit instead of after ten. Pass `--no-tracer` to restore the previous horizontal-layer default; `--mvp` now layers user-story framing and the Walking Skeleton on top of the tracer-first ordering. (#1945) (#2294)
- **Internal: external-descriptor trust gate — load-time `configHome` confinement** — `assertDescriptorConfined(descriptor, configHome)` (new `src/external-descriptor-trust.cts`) fail-closed rejects any installed third-party host-plugin descriptor whose declared `destSubpath` resolves outside the user-approved `configHome`, before its install plan runs (ADR-1239 Phase C-2 / #1681 slice 1). Defense-in-depth load-time twin of Phase 2's install-time `assertDestWithinConfigHome`. Not yet wired into the loader (slice 2). No user-facing change. (#1806)
- **Internal: the installer's runtime → global-config-home hook-pathogen fragment is now a single `getGlobalConfigHomeFragment` lookup** — the 14-branch `if (runtime === 'x') return "'...'"` chain in `getConfigDirFromHome` (`bin/install.js`, the hook `path.join()` codegen mapping) is collapsed into one table in `runtime-name-policy.cts`, sibling to `getRuntimeLabel` (ADR-1239 Phase B, #1679 AC2 slice 2). Generated hook output is byte-identical for all 16 runtimes (golden-parity asserted); antigravity's dynamic env-overridable resolution is preserved in the caller. No user-facing change. (#1801)
### Removed
- **Removed the sunset Gemini CLI runtime — use Antigravity CLI instead** — Google discontinued Gemini CLI on 2026-06-18, so `npx gsd-core --gemini` now prints a deprecation notice and points you to Antigravity CLI (the official successor), which GSD already ships as a first-class runtime. (#1928) (#1996)
### Fixed
- The `verify-work` security-blocked presentation no longer offers next-phase planning. When security enforcement blocks phase advancement (no `SECURITY.md` produced), the workflow now routes only to the current-phase fix instead of competing `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}` options. (#1687)
- `milestone complete` and `roadmap analyze` now exclude the Phase 0 / Phase 999 backlog sentinels. A milestone whose only directory-less ROADMAP heading is a backlog sentinel can be completed without `--force`, and `roadmap analyze` no longer counts the sentinel in `phase_count` or routes `next_phase` into it. Completes the `^999` exclusion #1445 added to the progress denominators. (#1691)
- **`config-set` no longer silently coerces values into something the disk never sees** — `Number.isFinite` replaced `!isNaN` in the value parser so `Infinity`/`-Infinity` are no longer coerced to non-finite numbers that `JSON.stringify` then renders as `null` on disk while the CLI echoes `Infinity` (output ≠ disk). `context_window` now has a per-key validator requiring a finite positive integer (rejects `Infinity`, `0`, negatives, non-integers with a non-zero exit), and `project_code` is always persisted as a string so a leading-zero code like `007` survives verbatim instead of collapsing to `7`. Numeric coercion for genuine numeric keys (e.g. `granularity 42`) is unchanged. (#1581) (#2023)
- **`phase.complete` no longer reports a false `is_last_phase` on a `<details>`-wrapped checkbox checklist (#1591, #1752)** — when the active milestone's phase checklist was written as `- [ ] Phase N:` checkbox items inside a `<details>` block and the next phase had no directory on disk yet (still in planning), `phase.complete`'s `isLastPhase` roadmap-enumeration fallback used a heading-only pattern (`/#{2,4}\s*Phase…/`) that never matched checkbox items. It returned `is_last_phase: true, next_phase: null` on a mid-milestone phase and — via the milestone-complete cascade — wrongly flipped STATE.md to `Milestone complete` and decremented `progress.total_phases` (e.g. 8 → 7). The pattern now matches heading-style (`### Phase N:`), plain checkbox-list phases (`- [ ] Phase N:` / `- [x] Phase N:`), and the canonical **bold** checklist form the roadmap template emits (`- [ ] **Phase N: Name**`); `extractCurrentMilestone` already surfaces the `<details>`-wrapped checklist correctly, so no parser change was needed. Only the reproduced `phase.complete` fallback is changed; the heading-only sibling patterns elsewhere in `phase.cts` are untouched.
(#1819)
- The `<agent_skills>` block emitted by `gsd init` no longer leaks backslash paths into `@`-reference skill paths on Windows. The global skill directory (a native `path.join` result) was interpolated into the generated markdown without POSIX normalization, producing references like `@C:\…\skills\name/SKILL.md`; the reference is now normalized at the emit site so skill references use forward slashes on every platform. (#1736)
- **`/gsd-settings` no longer warns about four search-provider keys on fresh projects (#1747)** — `buildNewProjectConfig` emits seven search-provider availability flags and `research-provider.cts` `providerAvailability()` consumes all seven, but only three were registered in `VALID_CONFIG_KEYS` (`config-schema.manifest.json`). Running `/gsd-settings` on a freshly generated `.planning/config.json` printed `unknown config key(s) … tavily_search, ref_search, perplexity, jina — these will be ignored` even though the user never hand-edited the config. The four missing keys are now registered alongside `brave_search`/`firecrawl`/`exa_search` and documented in `docs/CONFIGURATION.md`; a drift guard in `tests/bug-2530-valid-config-keys.test.cjs` now requires every config-driven research-provider flag to be in the schema, so a future provider addition cannot reintroduce the drift. (#1814)
- **`gsd-tools state json` no longer reports conflated progress for an unversioned milestone (#1761)** — the ADR-1769 Phase 7 fix (#1794) taught `state sync` to leave Progress untouched when a milestone version is asserted but the ROADMAP has no versioned heading for it, but the `state json` **read** path still rebuilt progress via `buildStateFrontmatter`, whose phase-heading count fell back to the whole document and summed sibling milestones. `state json` therefore reported a conflated `total_phases` (e.g. 8 = 4+4 across two milestones) plus a derived `percent`, contradicting the sync guard on the very same project. The read path now mirrors the sync guard: when the asserted milestone cannot be bounded to a versioned ROADMAP heading, `total_phases` falls back to the on-disk phase-dir count and `percent` is omitted. Bounded milestones (versioned ROADMAP, or no milestone asserted) are unchanged; the signal rides on the existing `_diskScanCache` so `extractCurrentMilestone`'s return contract and its other callers are untouched. (#1818)
- **`gsd-graphify-update.sh` now reads the full multi-line command in Gate 2 (#1772)** — the PostToolUse auto-update hook joined `tool_name` + `\n` + `tool_input.command` and extracted the command with `sed -n '2p'` (line 2 only). Agent runtimes (Claude Code's Bash tool among them) routinely emit HEAD-advancing commits as multi-line scripts (`cd /path`, then `git add`, then `git commit …`), so line 2 was the `cd`, Gate 2's `*"git commit"*` match failed, and the rebuild silently no-op'd on real commits even with `graphify.auto_update: true`. The failure was invisible in manual probes because a single-line `git commit -m x` passes line 2 verbatim. The hook now captures line 2 through EOF (`sed -n '2,$p'`) so the `case` glob sees the full command string; single-line behavior is unchanged and multi-line commands without a HEAD-advancing op still no-op cleanly. (#1815)
- **`/gsd-thread close|resume` now writes the thread status/updated frontmatter (#1778)** — the thread workflow's CLOSE and RESUME branches invoked `frontmatter.set` with the pre-1.6 fully-positional shape (`frontmatter.set <file> <field> <value>`), but since 1.6 the dispatcher parses the file positionally and reads `field`/`value` from the named flags `--field`/`--value` via `parseNamedArgs`. The positional form left `field`/`value` undefined, `cmdFrontmatterSet` errored `file, field, and value required`, and the writes were silently skipped — so closing a thread never marked it `status: resolved` and resuming never marked it `status: in_progress`, with the error scrolling past on every thread command. All four sites (CLOSE `status`+`updated`, RESUME `status`+`updated`) now use the 1.6 hybrid form that `verify-work.md` already uses (`frontmatter.set <file> --field <field> --value <value>`). (#1816)
- **The installer no longer copies dead lifecycle hook scripts for ZCode** — it declares `hooksSurface: 'none'` and has no plugin surface, so the staged `hooks/*.js`, `hooks/*.sh`, `hooks/lib/` and the CommonJS `package.json` marker were dead weight in `~/.zcode/`. The hook-copy guards in `install.js` now exclude ZCode alongside the other no-hook runtimes. OpenCode, which also declares `hooksSurface: 'none'`, is deliberately kept: its native plugin adapter (#1914) spawns those staged hooks via OpenCode's event bus and needs both them and the marker. (This fix originally excluded Kilo too, on the premise that it had no plugin surface; that premise was wrong — Kilo's native plugin spawns the staged guard hooks, exactly like OpenCode's — and #2327 reverses the Kilo half.) (#2057)
- **Test gates can no longer hang forever on a watch-mode test runner.** vitest defaults to watch mode in an interactive terminal (exactly where `gsd-execute-phase` runs), so a resolved `npm test` / `pnpm test` that maps to vitest never exited and the orchestrator waited indefinitely until the user manually intervened. Every GSD test-command gate — the regression gate, the post-merge gate, the audit-fix gate, and the verify-phase gate — now routes the resolved command through a shared `normalize-test-command` helper that rewrites it to a one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script backed by watch-vitest → `CI=true` prefix; already-one-shot commands are left unchanged). The three gates that previously hung or silently continued — the regression, post-merge, and audit-fix gates — additionally bound execution with a configurable `workflow.test_gate_timeout` (default 600s), aborting or surfacing the cause on timeout instead of hanging; the verify-phase gate was already bounded (a fixed 5-minute limit) and keeps it, now naming watch mode on timeout. The normalizer only rewrites a runner named as a standalone command token (so paths/targets like `run-vitest.js` are never mangled), is length-capped and linear-time on adversarial input, and only reads a regular-file `package.json`. (#2060)
- **`settings-advanced.md` no longer has an orphan `</step>` around §8 Model Policy** — the §8 Model Policy block ended with a closing `</step>` but had no matching opening tag (5 opens / 6 closes), leaving its content as loose inter-step prose that could fail to execute reliably. Added the missing `<step name="model_policy">` opener so the section is a proper step. A new workflow `<step>`-tag-balance regression guard (fenced-code-stripped) now blocks any future orphan tag across all top-level workflows. (#1864) (#2014)
- **The runtime launcher now honors `CLAUDE_CONFIG_DIR`** — the `gsd_run` preamble embedded in every workflow/agent resolved the Claude global install only at `$HOME/.claude/gsd-core/bin/`, while the installer honored `CLAUDE_CONFIG_DIR`, so a global install redirected via `CLAUDE_CONFIG_DIR` was invisible to every `gsd_run` call (every GSD command failed with `gsd-tools.cjs not found`). The Claude resolver arm now uses `${CLAUDE_CONFIG_DIR:-$HOME/.claude}` — matching the installer and the other runtimes' `${VAR:-default}` pattern — so a custom `CLAUDE_CONFIG_DIR` is found and the default `$HOME/.claude` path is unchanged. Re-synced into all 95 workflows/agents; two capped workflows trimmed to stay under their byte budgets. (#1865) (#2024)
- **Node-test prohibition proofs now require a clean-fixture causation control** — a `node-test` prohibition's fail-first proof no longer accepts a deceptive content-independent negative test (one that reds merely because `GSD_PROHIB_SUBJECT` is *set*, ignoring the subject's content). The `check_clean_fixture` control is now **mandatory** for the `node-test` kind: a descriptor that omits it is un-provable and hard-gates, rather than greening on the violation alone. **Breaking (Hyrum):** a previously-green node-test prohibition with no clean fixture now hard-gates — blast radius is zero in-tree (no `node-test` prohibition ships today). The `lint-rule` kind is unchanged (its subject IS the linted file, no `GSD_PROHIB_SUBJECT` indirection). (#1906) (#2001)
- **Third-party capabilities now work on installed layouts.** `capability install` no longer rejects capabilities with a real `engines.gsd` range as "incompatible with GSD 0.0.0" — the host version is now read from the authoritative `gsd-core/VERSION` file across every runtime and the `capability install` CLI. The installer also now ships the registry generator scripts (`gen-capability-registry.cjs`, `gen-loop-host-contract.cjs`), so installed third-party capabilities actually compose into the loop instead of being silently discarded. (#1938)
- **`/gsd:verify-work` preserves verification state across gap-closure execution and no longer auto-promotes deferred follow-ups into blocking gaps** — resuming after `/gsd:execute-phase --gaps-only` used to lose the verification state: the UAT `## Gaps` still read `status: failed` even after their fix plans executed, so verify-work re-diagnosed them as fresh blockers, spawned a new gap plan, and reported only the new plan as verified. A state contract now links each gap to its fix plan: every UAT gap carries a stable `gap_id` (`G-{phase}-{N}`), gap-closure plans tag the ids they address in their frontmatter (`gap_ids: […]`), and a new `reconcile_gaps` step on resume marks a gap `status: resolved` when its plan has a matching `*-SUMMARY.md` — so fixed gaps aren't re-diagnosed and the phase can close. Separately, a deferred-follow-up branch captures future-work ideas (signals like "later", "next version", "out of scope") into a `## Deferred Follow-Ups` section instead of creating a blocking gap/plan. (#1921) (#2025)
- **`roadmap update-plan-progress` no longer counts stray non-plan `*-SUMMARY.md` files against phase completion** — remediation/gap-closure summaries (e.g. `30-FIX-CR02-SUMMARY.md`, `30-GAPCLOSURE-SUMMARY.md`) inflated `summary_count`, and once `summary_count >= plan_count` the phase silently flipped to `Complete` (checkbox checked, date stamped) even though several plans had no summary. A new `countMatchedSummaries` helper (core-utils) pairs summaries to plans via the `PLAN→SUMMARY` marker swap + the `<stem>-SUMMARY.md` form (layout-agnostic across root, bare, and nested layouts), so only a summary that corresponds to a real plan counts. Wired into `scanPhasePlans` (fixing roadmap listing, state sync, verification, workstream inventory at once) and `cmdRoadmapUpdatePlanProgress`. (#1988) (#2016)
- **`milestone complete --ws` requirements archive header now points at the workstream REQUIREMENTS.md** — the archive header string hardcoded the root path (`` `…see .planning/REQUIREMENTS.md` ``), so a workstream archive directed readers at the wrong file even though #1917 had already fixed the archive *locations* to land inside the workstream. The display path is now derived from the same workstream-aware `reqPath` the writer uses (`path.relative(cwd, reqPath)`), so root behavior is byte-identical and the workstream case correctly reads `.planning/workstreams/<ws>/REQUIREMENTS.md`. (#1993) (#2015)
- **Load-failed capability gates now fail open with a loud warning instead of blocking the whole project** — when an installed overlay (third-party) capability failed to load (e.g. an incompatible `engines.gsd` range) but had declared a `gate`-kind loop hook, the loop resolver injected a blocking synthetic gate (`blocking:true`, `onError:halt`) at every point where that capability declared a gate. A single incompatible capability therefore halted every `ship:pre` and `verify:post` in the project — unrelated to what the gate would have checked, and with no remediation surfaced. The resolver now injects no gate and instead emits a loud warning — to stderr and in the `loop render-hooks` envelope's `warnings` array — naming the load-failure reason and the exact `gsd capability remove <id>` remediation, and the loop proceeds (fail open). The capability id embedded in that remediation is validated against the canonical id shape first, so a malformed overlay directory name cannot inject shell metacharacters into the surfaced command. The loader still records `_overlay.blockedGates`; only the consequence changes from block to warn. `step`/`contribution` overlays were already skip-open. (#2009) (#2075)
- **`phase.complete` now updates the `## Progress` rollup row even when an earlier phase-numbered table precedes it** — the Progress-row writer used a non-global regex that matched *any* table row starting with the phase number, so it bound to the first such row (e.g. a `| Phase | Requirements | Count |` coverage table), no-op'd on the wrong 3-column row, and never reached the real Progress row. The regex is now scoped to the `## Progress` section so it binds to the correct table. The command still returned `roadmap_updated: true` (that field is `fs.existsSync(ROADMAP.md)`), masking the silent failure. (#2012) (#2032)
- **context7 now works for plugin-marketplace installs (8 agents regained doc lookup)** — the agents granted only `mcp__context7__*`, which matches a standalone context7 MCP server but not the official Claude Code plugin-marketplace install (`context7@claude-plugins-official`), whose tools are named `mcp__plugin_context7_context7__*`. The grant never matched, so advisor/ai/domain/phase/project/ui-researcher + planner + executor silently lost documentation lookup and fell back to WebSearch. All 8 agents now grant both forms, the researcher profile table is updated, and a parity guard asserts no agent grants the standalone form without the plugin form. (#2017) (#2029)
- **`applySurface` no longer deletes every `gsd-*` agent when the skills manifest resolves empty** — the agent-prune loop in `_syncGsdDir` deleted any `gsd-*.md` not in the staged set, and when the manifest was empty/unresolvable (null manifest, no array entries, no `files` key, or an unresolvable install source root), the staged set was empty → every agent was pruned. Skills were guarded by `pruneSkillDirs`'s manifest-membership check (conservative preservation on empty manifest); agents had no equivalent. The agent-prune loop is now skipped when the manifest is empty/absent, so agents are preserved while copy (adding genuinely new agents) still runs. (#2018) (#2031)
- **`planning-config.md` global-learnings path corrected to `~/.gsd/knowledge/`** — the `features.global_learnings` row directed users to `~/.gsd/learnings/`, but the implementation (`src/learnings.cts`, `execute-phase.md`) stores and reads global learnings from `~/.gsd/knowledge/`. Anyone following the docs to inspect, back up, or seed their global learnings looked in a directory the code never touches. (#2019) (#2026)
- **Removed dead SDK file references from runtime-loaded markdown that triggered an infinite `find.exe` storm on Windows** — `agents/gsd-executor.md` pointed at `sdk/src/query/QUERY-HANDLERS.md` and `gsd-core/workflows/reapply-patches.md` at `sdk/dist/cli.js`, both retired with the SDK package (ADR-0174). AI runtimes that resolve doc references by filesystem search ran `find / -iname …`; on Git Bash for Windows `/` maps to the drive root, so `find.exe` traversed the whole disk (14h+, orphaned processes, 4M+ open handles each, unkillable). The references now resolve to live paths, and a new regression guard asserts no `sdk/src|sdk/dist|sdk/handlers` file references remain in agents/workflows/references markdown. (#2020) (#2027)
- **`roadmap update-plan-progress` no longer checks the phase checkbox without verification** — the command stamped the phase-level ROADMAP checkbox and completion date the moment the last plan summary landed (called routinely after every wave and every plan), with **no verification gate** — unlike `phase.complete` which correctly requires `readVerificationStatus(...).status === 'passed'`. Now `isComplete` requires both all plan summaries AND a passed verification, matching the `cmdPhaseComplete` contract, so the checkbox only fires after `gsd-verifier` has confirmed the phase. (#2022) (#2030)
- **`phase complete` no longer marks a milestone done out of order, nor silently writes root state in workstream mode.** Completing the numerically-highest phase while an earlier phase was still outstanding wrongly flipped STATE.md to `Status: Milestone complete` (the milestone-end check only looked for higher-numbered phases, so an out-of-order completion — e.g. Phase 10 before Phase 9 — read as the end). It now reports milestone-end only when every lower-numbered phase in the milestone is checked complete. Separately, in workstream mode with no active workstream, `phase complete` previously fell back to root `.planning` and wrote STATE.md/ROADMAP.md (and the mislabel) into the shared root other workstreams read; it now fails safe — asking for `--ws <name>` or an active workstream — mirroring the existing `init progress` guard. (#2066) (#2066)
- **Phase directories whose slug begins with a single digit now resolve correctly.** A phase like `46-6-rs-pipeline-orchestrator` (roadmap name "6 Rs Pipeline Orchestrator") had its phase token over-collected as `46-6` instead of `46`, so `gsd-tools` phase-by-number lookups resolved `phase_dir=null` / `has_context=false` (breaking `init.plan-phase`, `init.phase-op`, and downstream execute/verify/ship). Numeric phase-token components must now be zero-padded (≥2 digits), so a single-digit slug word is no longer absorbed into the token. Fixed consistently across every same-class implementation — `extractPhaseToken`, `PHASE_TOKEN_FROM_DIR_RE` and `canonicalPlanStem` (health checks / plan pairing), `isDirInMilestone`'s numeric matcher (milestone filtering), and `extractCanonicalPlanId` — so the health-check and milestone-filter subsystems are fixed alongside phase resolution. (#2059)
- **`gsd-tools config-set <key> null` now clears (removes) the key instead of persisting the literal string `"null"`.** The documented "Clear" action previously fell through the value parser and stored `"null"` — a truthy value — so "cleared" keys stayed set and `config-get` returned `"null"`; for secret keys (`brave_search`/`firecrawl`/`exa_search`) a masked success line hid a truthy value on disk that integrations could pass along as a real credential. `config-set <key> null` now deletes the key (short-circuiting the typed per-key validators so clearing an enum/boolean/number key removes it rather than being rejected), making the "Clear" flows in `settings-integrations.md` / `settings-advanced.md` actually clear. (#2058)
- **`init plan-phase` no longer collapses foreign-prefixed task/workstream IDs into numeric phases** — a query like `MEM-01` (where `MEM` is not the configured `project_code`) used to have its prefix stripped and resolve to the unrelated numeric Phase 01; it now reports `phase_found: false` unless a phase directory or roadmap entry literally carries that prefix. The configured `project_code`'s own prefixed phases (e.g. `LKML-01` under `project_code: LKML`) continue to resolve as before. (#2056) (#2105)
- **`phase complete` no longer ticks the wrong phase's ROADMAP checkbox** — completing a phase whose number also appears in a later phase's description (e.g. an idempotent re-run of an already-complete phase) used to mark the *wrong* phase done, because the checkbox-matching regex greedily spanned from `]` to any later "Phase N" mention instead of only the immediately-following phase title. (#2067) (#2079)
- **`gsd-tools effort sync` no longer crashes in an installed runtime.** In any global install (e.g. `~/.claude/gsd-core/`), `effort sync` threw `Cannot find module '../../../bin/install.js'` — the command reached into the package-root `bin/install.js` for its install-time effort resolvers, but the installer only copies the `gsd-core/` subtree into a runtime home, so that file is never present there. As a result, `effort` config changes (`routing_tier_defaults` / `agent_overrides`) silently never reached installed agents without a full reinstall. The two resolvers (`readGsdEffectiveEffortConfig` + `resolveInstallTimeEffort`, with their helpers) are now extracted into a shipped `gsd-core/bin/lib/install-effort-resolver.cjs` that both `effort sync` and the installer import — a single source of truth that is always present in the installed tree. (#2076) (#2076)
- **`model_overrides` and per-phase-type models now actually apply to the assumptions-analyzer, code-reviewer, and code-fixer agents on Claude Code.** Previously `model_overrides["gsd-code-reviewer"]` / `["gsd-assumptions-analyzer"]` / `["gsd-code-fixer"]` (and `models.verification` / `models.discuss` / `models.execution`) were accepted and resolved but silently dropped — the workflows spawned these agents with no model, so they inherited the session model and the configured routing never took effect (no warning). Every spawn now threads its resolved model: `discuss-phase-assumptions`, `code-review`, and `code-review-fix` (both the re-review and the two fixer spawns) resolve it inline, and `quick`'s review step uses the code-reviewer's own resolved model instead of the executor's. The stale "`discuss` — reserved, no subagent" model-profile docs are corrected to list `gsd-assumptions-analyzer`, and the `verification` row now includes `gsd-code-reviewer`. (#2074) (#2074)
- **`/gsd-review`'s Antigravity CLI reviewer no longer fails silently on large prompts, unavailable pinned models, or pre-session stalls** — the `agy` invocation now uses a file-reference prompt to avoid exec arg-list overflow, is wrapped in an external wall-clock `timeout` paired with `--print-timeout` because `--print-timeout` cannot fire before `agy` creates a session, passes `--model` from `review.models.agy` when set as an escape hatch for a 404'd pinned model, and its empty-output stub now surfaces an `agy` cli.log diagnostic instead of a bare generic message. Supersedes the #687 "no external killer / inline `$(cat)`" contract, which predated `agy` gaining `--model` and predated its own guidance to pair `--print-timeout` with a terminal timeout. (#2073) (#2109)
- **`init execute-phase`, `init verify-work`, and `init phase-op` no longer collapse foreign-prefixed task IDs to numeric phases** — `MEM-01` under `project_code: LKML` was silently stripped to `01` and resolved to the unrelated numeric Phase 01, because the #2056 guard was applied only to `init plan-phase`. The guard is now extracted into shared helpers (`guardedFindPhase` / `guardedGetRoadmapPhase`) that delegate to the canonical `isForeignPrefixedPhaseQuery` from `phase-id.cts`, and all four init commands route through them. (#2104) (#2149)
- **`commit --files` now commits only the declared paths** — `gsd-tools commit --files A B` previously ran a bare `git commit` that absorbed the entire staged index, silently sweeping in unrelated files the caller never named. The commit now appends a pathspec (`-- <paths>`) so only the staged subset of `--files` lands in the commit; the no-`--files` default path is unchanged. Missing tracked files are still skipped (not committed as deletions, #2014), and when all declared files are missing the function short-circuits to `nothing_to_commit` instead of absorbing the index. (#2112) (#2148)
- **Fixed unresolvable bare `require('gsd-core/...')` in `gsd-surface` command doc** — the four `require()` examples now derive the engine path from `runtimeConfigDir` (resolvable at runtime), and the reinstall hint corrects `npm i -g gsd-core` to `npm i -g @opengsd/gsd-core`. (#2116) (#2213)
- **`milestone complete --dry-run` now prints a preview plan instead of silently mutating** — `gsd-tools milestone complete --dry-run` was neither parsed nor rejected, so a caller expecting a preview triggered the full destructive mutation (archive phases, move audit artifacts, rewrite STATE.md) with no way to back out. The `--dry-run` flag is now honored: it returns a JSON plan listing `would_archive` (roadmap, requirements, audit, phase dirs) and `would_update` (MILESTONES.md, STATE.md) targets with zero filesystem mutations. (#2118) (#2155)
- **`/gsd-secure-phase` now has a single SECURITY.md writer** — the `gsd-security-auditor` subagent previously held `Write`/`Edit` tools and was instructed to "write SECURITY.md" with no padded `<N>-` prefix and no template frontmatter, while the orchestrator's Step 6 also wrote the phase-scoped `<N>-SECURITY.md` from `templates/SECURITY.md`. The auditor is now return-only (drops `Write`/`Edit`, returns a structured verdict with `threats_open`); the orchestrator is the sole file writer. The workflow's Step 5 spawn constraints explicitly forbid the auditor from writing SECURITY.md. (#2119) (#2154)
- **Dead security scan exports removed; injection-scan docs corrected to match reality** — `scanEntropyAnomalies` and `shannonEntropy` were dead code with zero production callers (live hooks inline their own patterns for independence). REQ-SCAN-INJ-02/-03 now accurately describe what runs live (injection patterns, invisible Unicode) vs CI-only (base64-decode, codebase scan). (#2198) (#2211)
- **Post-merge, regression, and other GSD test/build gates no longer fail with a spurious "command not found" on stock macOS.** These gates hardcoded GNU coreutils' `timeout`, which stock macOS ships neither as `timeout` nor `gtimeout`; a passing build or test run now completes under a portable, coreutils-independent `run-with-timeout` wrapper instead of exiting 127 and being misreported as a failure. (#2351) (#2426)
- **Installed third-party capability skills now materialize on OpenCode and Kilo** — `capability install` + `capability set --runtime opencode` (or `kilo`) could report a capability as `installed: true, surfaced: true, active: true` while its skill was never written to `skills/gsd-<stem>/SKILL.md`: the OpenCode/Kilo combined-family install path never called the seam #2322 fixed for other runtimes. Installed capability skills now materialize the same way there too, bound to their declaring capability, with first-party skills always winning a name collision. (#2362) (#2434)
- **Shared requirement IDs across multiple plans no longer read `Complete` before every declaring plan (and phase verification) has finished** — `execute-plan.md` now gates completion on sibling plans' `SUMMARY.md` files via a new read-only `requirements ready-ids` check, and a `gaps_found` phase verification reverts any requirement ID this phase owns back out of `Complete` before the gap report renders. Single-plan requirement IDs are unaffected — no added latency. (#2388) (#2424)
- **`phase.add` no longer silently mistakes a goal-shaped description for a phase title** — a long or multi-sentence description used to land verbatim in the `### Phase N:` header with no signal anything was off; `phase.add` now returns a `warning` field when the description looks goal-shaped, and the phase-number auto-detect docs now correctly point callers at the orchestrating workflow instead of implying `gsd-tools.cjs` resolves it itself. (#2390) (#2425)
- **`response_language` now reaches orchestrator-owned prompts across most workflows and the UAT verification checkpoint frame** — previously only subagent prompts honored a configured `response_language`; the orchestrator's own questions (verify-work, new-project, new-milestone, quick, manager, and others) and the hardcoded English UAT checkpoint banner stayed in English regardless of configuration. Both now render in the configured language, with output byte-identical to before when unset. (#2402) (#2457)
- **Codex installer no longer double-registers each agent role in `config.toml`, eliminating one duplicate-role startup warning per agent** — `generateCodexConfigBlock` stopped emitting `[agents.gsd-*]` tables whose `config_file` pointed back at the same standalone TOMLs Codex already auto-discovers under `$CODEX_HOME/agents/`; reinstalling over an existing config also drops any legacy managed role tables left by a prior install while preserving unrelated user config and the user's own AgentsToml scalars. (#2406) (#2432)
- **Production dependency tree carries no known advisories** — five advisories disclosed against the transitive tree under `@anthropic-ai/claude-agent-sdk` → `@modelcontextprotocol/sdk` were cleared: `fast-uri` (GHSA-4c8g-83qw-93j6, high) and `hono` (GHSA-xgm2-5f3f-mvvc, GHSA-hvrm-45r6-mjfj, GHSA-w62v-xxxg-mg59) re-resolved to patched releases inside their already-declared ranges with no `package.json` change, and `@hono/node-server` (GHSA-frvp-7c67-39w9) pinned to `>=2.0.5` via `overrides` because `@modelcontextprotocol/sdk@1.29.0` — already the latest published version — still declares the vulnerable `^1.19.9` range. `npm audit --omit=dev` reports zero advisories. (#2496) (#2497)
- **Custom STATE.md frontmatter keys are no longer dropped on every mutating verb** — syncStateFrontmatter rebuilt the frontmatter from a fixed schema, silently dropping any custom key. It now carries forward existing keys the schema does not own. (#2202) (#2233)
- **Non-frontend phases with `UI hint: no` are no longer blocked by the UI-SPEC gate** — the UI safety gate's token list included the bare token `UI`, which matched GSD's own `**UI hint**: no` metadata line and false-detected a UI, blocking backend/infra phases at /gsd-plan-phase. An explicit `UI hint: yes|no` is now authoritative and the hint line is no longer token-sniffed. (#2150) (#2222)
- **OpenCode reviewer no longer silently yields an empty review on large prompts** — `/gsd-review --opencode` now invokes `opencode run --format json` and reconstructs the review from the assistant text parts, so a large-prompt run where the default `build` agent ends its turn with zero output tokens no longer produces an empty stub. When the agent genuinely emits no text, the stub now reports the stop reason, output-token count, and captured stderr instead of a generic message. (#1936) (#1992)
- **OpenCode's first-time install baseline now protects pre-existing files under the `commands/` directory, not just the legacy `command/` alias** — after #2329 moved OpenCode command materialization to `commands/`, the baseline scan that guards a machine's very first GSD-tracked install still only knew about the legacy `command/` directory, so a pre-existing, unrelated `commands/gsd-*.md` file was silently deleted by ordinary command materialization instead of blocking the install for an explicit keep/remove choice — the same protection `command/` already had. The scan now covers both directories. Kilo is unaffected and keeps using `command/`. (#2354)
- **api-coverage detector no longer false-positives non-API phases (and no longer fails open)** — the external-API-integration detector behind the blocking `verify:pre` seal gate required only same-line co-occurrence of an integration verb and an API noun, treated `/` as a word boundary (so first-party Next.js `src/app/api/…` route paths matched), and read any capitalized word before API/SDK/REST/GraphQL as a service name (so threat-model prose like "Resolver-only API" fired). It is now **fail-closed**: the compound rule requires the integration verb and API noun to share one clause (the clause boundary is the whole relationship test — no fragile word-gap cap that a genuine long integration clause would trip); fenced code, inline code spans, and path-shaped tokens are excluded before matching while external hosts like `api.stripe.com/v1` still count; and the `<Service> API` surface rule rejects stopwords, locality/protocol descriptors ("Internal API", "REST API"), compound modifiers, and first-party-qualified services, so a real vendor name (`Stripe API`) fires from any clause position. A phase that integrates no external API can declare it first-class in `COVERAGE.md` — `No external API integration: <reason>` — instead of fabricating a matrix row; when the detector still finds signals, the declaration overrides but the gate surfaces the overridden signals so the contradiction is visible. Because a false positive is cheaply dismissed by that declaration while a false negative silently slips a real API phase past the gate, the detector deliberately leans toward detecting. (#2365) (#2397)
- **`stale-bake-guard` hermeticity fix (test-isolation)** — the readGsdEffectiveModelOverrides subtest no longer reads the developer's real `~/.gsd/defaults.json`; the resolver now accepts a homedir seam so the test sandboxes HOME. (#2152) (#2223)
- **`/gsd-surface` (`list`/`status`) works on Claude Code global installs** — the installer now writes a `.gsd-source` marker pointing at its `commands/gsd` source, so `findInstallSourceRoot` resolves on the global skills layout (which ships no `commands/gsd` tree) instead of throwing `could not locate commands/gsd`. (#1487) (#1487)
- **Cursor no longer shows every `/gsd-*` command twice** — a `--cursor` install wrote both a skill and a slash command for each action, so every GSD entry appeared twice in Cursor's `/` menu. GSD now installs Cursor skills as `user-invocable: false` (matching the existing CodeBuddy behavior), so the slash command is the single `/` entry point while skills remain model-invocable. (#2341) (#2386)
- **`phase complete --phase N` now works alongside the positional form** — the phase verb family treated the first positional as the phase number, so `--phase 12` was passed as the literal phase name and failed with 'Phase --phase not found'. The phase family now accepts the --phase flag consistently with the state family, and unrecognized flags yield a usage error. (#2201) (#2231)
- **Third-party capability skills now surface correctly after install** — a skills-only `role: feature` capability installed `active` but its skills never reached the runtime surface, `capability enable`/`set` rejected it as `unknown capability`, and `capability list` disagreed with `capability state`. `resolveSurface` now unions the composed registry's `capabilityClusters` into the surfaced skill set (no on-disk linking), the writer validates against the composed overlay-aware registry, and `capability list` carries a `surfaced` field matching `capability state`. (#2054)
- **`/gsd-ship` no longer emits a 100%-missing TDD Audit noise table** — the TDD Audit PR-body section was always emitted, but the execute pipeline only writes `gate_status:` git trailers when TDD mode is active. Without TDD mode (the default), every commit was counted `missing` and the table was pure noise with no way to disable it. The section is now gated behind `workflow.tdd_mode`: when TDD mode is off, both the TDD Audit section and the aggregate `gate_status:` trailer are skipped entirely; when on, the existing behavior is preserved. (#2467)
- **`phases.clear` now archives phase history under the outgoing milestone version, not the newly-switched one** — because `new-milestone` advances the milestone before clearing leftover phases, the phase-history archive was silently misfiled under the new milestone's `<version>-phases/` directory. A new `--archive-version` override on `phases.clear` (threaded from the new-milestone workflow) files the archive under the previous milestone's version; without it, behavior is unchanged. (#2288) (#2323)
- Fixed: probe-core's runProbeCli now fails closed on per-item adapter garbage inside a well-shaped report envelope, matching its documented 'fails closed on adapter garbage' contract. (#1910)
- **Deferred out-of-scope findings logged to `deferred-items.md` are now surfaced** — the executor's SCOPE BOUNDARY convention writes discoveries to a phase directory's `deferred-items.md`, but nothing read it back, so those items were permanently invisible. `/gsd-progress`'s forensic audit and `audit-uat` now glob `.planning/phases/*/deferred-items.md` and surface unresolved entries. (#2287) (#2318)
- **`/gsd:verify-work` no longer silently terminates when all remaining UAT tests are blocked** — sessions with `blocked_count > 0` and `pending_count == 0` now route to `complete_session` as expected, enabling the zero-issues auto-transition path. (#1722)
- **state record-metric no longer appends per-plan rows into the By-Phase velocity table** — it now maintains its own Per-Plan Metrics table (self-created on first use), and its auto-create scaffold header is corrected. (#2253) (#2253)
- **Dynamic routing now escalates the model, not just effort** — with `dynamic_routing.enabled`, retry attempts advanced the reasoning effort but the model stayed pinned to the default tier because `resolve-execution` resolved the model without consulting `dynamic_routing`. `resolve-execution` now resolves the model per-attempt through the tier ladder (e.g. standard→heavy on attempt 1, capped at `max_escalations`); resolution is unchanged when dynamic routing is disabled. (#2068) (#2334)
- **`/gsd-next` no longer reports a project as complete while phases are still unchecked** — `smart-entry`'s completion check now grounds in ROADMAP.md's actual Progress table (global, authoritative) instead of STATE.md's stale milestone-scoped total_phases, and its status regex requires milestone-level language (`milestone complete` / `all phases complete` / `complete`) instead of matching any per-phase `shipped` or `done` substring. Together these fix the false-complete misclassification that could route `/gsd-next` toward `/gsd-new-milestone` — which archives still-pending phase directories. (#2466)
- Codex reviewer now captures the review via codex's --output-last-message flag instead of redirecting stdout, so Windows process-teardown output no longer pollutes the review file and slips past the empty-output guard. (#1709)
- **`last_activity` now shows your local calendar day** — the clock seam derived the date by slicing a UTC instant, so in negative-UTC-offset zones during UTC's early evening the date-only `last_activity` field jumped a day ahead of the operator's actual date (and of `last_updated`'s local date). Operator-facing date fields now use a host-local calendar day while internal/cosmetic stamps stay UTC. (#2136) (#2216)
- **A phase with a deliberately-unexecuted (superseded) plan no longer stays stuck below 100%** — a plan reassigned or dropped mid-phase can never gain a matching SUMMARY, yet plan-scan counted it forever, so the phase read In Progress and the milestone sat below 100% permanently — the plan-level analogue of the retired-phase bug (#1514). Mark such a plan `status: superseded` in its PLAN.md frontmatter and it is now excluded from both the plan and summary counts, so the phase completes honestly (a 13-plan phase with 2 superseded reads 11/11). Plans without the marker are unchanged. (#2349) (#2404)
- **`milestone_name` is no longer clobbered with a delimiter-led fragment** — getMilestoneInfo's `##` heading regex was unanchored, so it matched a heading quoted inside backticks in the Milestones bullet and wrote garbage like `— Active Milestone` over the curated milestone name on every phase transition. Now consults the 🚧 marker first, anchors the regex to line start, strips the leading delimiter, and widens the preserve guard so a bad derive keeps the existing name. (#2135) (#2215)
- **`init milestone-op` now counts project_code-prefixed phase directories correctly** — fully shipped milestones using the standard prefixed directory layout no longer report `completed_phases: 0` or stay falsely incomplete. (#1844) (#1844)
- **`/gsd-mempalace-capture` no longer crashes on first invocation** — the skill's own documented `rooms:` example wrote a flat list of bare strings, but mempalace's miner expects each entry as a dict with a `name` key, so following the example verbatim and running `mempalace mine` crashed with `TypeError: string indices must be integers, not 'str'`. Both `skills/gsd-mempalace-capture/SKILL.md` and `commands/gsd/mempalace-capture.md` now ship the corrected `- name: <room>` shape, so the documented example runs successfully end-to-end. (#2464)
- **`/gsd-quick` no longer halts with a stale-base worktree mismatch** — the worktree executor now degrades to sequential execution when its fork base has diverged from origin/HEAD, instead of spawning a worktree guaranteed to fail the base-mismatch guard. (#1991)
- **`GSD_ALLOW_SYMLINKED_DEST=1` lets users with intentional symlinked configHome layouts install/update again** — v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like `skills/` or `hooks/`) was a pre-existing symlink, with no opt-out. Three legitimate user-owned layouts were blocked: multi-account configs with symlinked shared skills/hooks (POSIX symlinks), Windows Junctions to shared skills dirs, and dotfiles-managed configHome (e.g. nix-darwin symlinking `~/.claude` itself to a version-controlled dir). The new env var follows user-owned symlinks instead of refusing them, while preserving the two load-bearing refusals from the original threat model: path-traversal in the destSubpath string itself (`../../etc`-style), and a symlink resolving to the install root itself (would let the prune pass wipe it). (#2393) (#2445)
- **`state record-session` no longer silently drops inserted fields on a CRLF `STATE.md`** — the section-rewrite regexes in `cmdStateRecordSession` used literal `\n` which couldn't match a CRLF STATE.md (`---\r\n`), so when a canonical session field (`Resume file` / `Stopped at` / `Last session`) was missing and had to be **inserted** via the section-rewrite path, the CRLF-tolerant detector entered the branch, the writer regex silently no-op'd, but `updated.push(...)` ran unconditionally. The command returned `{"recorded": true, "updated": ["Resume File"]}` while the field was never written to disk. With `core.autocrlf=input`, the CRLF working-tree file produced no `git diff`/`git status` change, so the bug was invisible. Both regexes now use the CRLF-tolerant `\r?\n` form (same canonical pattern already in use elsewhere), and a new defensive invariant gates `updated.push(...)` on the replace callback actually firing — so a future detector/writer drift will surface as missing `updated` entries rather than re-arming this silent-success class. (#2482)
- **`/code-review` no longer skips a phase whose SUMMARY.md records `~/`-prefixed file paths** — such a path was silently dropped as "deleted" (bash never tilde-expands a `~` that arrives as a variable's value), emptying the review scope and reporting "no source files changed" as a false success. Tilde paths are now expanded to `$HOME/…` before the deleted-file filter runs. (#2419)
- **Setting `external_job.submit_timeout_ms` / `poll_timeout_ms` / `artifact_dir` in `.planning/config.json` now actually configures the SLURM adapter** — the keys were declared by the external-job capability but the adapter only read env vars, so config edits silently had no effect. The adapter now resolves them through the canonical capability-config seam (env override > config > registry default), surfaces the resolved `artifact_dir` in `submit` output, documents why the contribution registers at `execute:wave:post` (#1164 asks for `wave:pre`, which `execute-phase.md` does not dispatch today; wiring it is a core-loop change #1164 explicitly defers), and gains unit coverage for the CLI surface (`parseFlags`, `findPlanningDir`, `resolveExternalJobSettings`, `formatShowReport`). (#1164) (#2006)
- **The Antigravity reviewer in `/gsd-review` no longer reviews blind** — `agy -p` never granted the agent the repo under review, so it frequently anchored on its own scratch directory and returned plan-text-only verdicts counted at full consensus weight. The reviewer is now granted the repo (capability-probed `--add-dir`) and anchored to the absolute repo root; a review that still runs without repo access is stamped `[reviewed-without-repo-access]` and down-weighted in the Consensus Summary. The cursor-agent prompt gains the same absolute-root anchor. (#2176) (#2184)
- **Non-Claude installs no longer brand all GSD output as Claude** — the installer never persisted `runtime: <id>` into `~/.gsd/defaults.json` for non-Claude runtimes, so `resolveRuntime()` (precedence: `GSD_RUNTIME` env > `config.runtime` > `'claude'`) fell through to the hard-coded `'claude'` default. A non-Claude install showed `agent_runtime: "claude"` and Claude-formatted `/gsd-*` slash hints with no env or config hand-set. The installer now persists `runtime: <runtime>` into `~/.gsd/defaults.json` for non-Claude runtimes, mirroring the existing `resolve_model_ids: "omit"` write at the same call site. Claude is the fallback so it needs no write; an explicit pre-existing `runtime` value is always preserved. (#2395) (#2446)
- **Autonomous reruns now skip phases with deferred verification until you resume them explicitly** — if a prior `/gsd-autonomous` run recorded `verification_deferred_human` or `verification_deferred_gaps`, later reruns no longer drop back into the same prompt loop and instead point you at the saved resume command. (#1846) (#1846)
- **`requirements mark-complete` no longer reports silent success when the traceability row is missing** — it OR-ed its checkbox and table-row writes into one flag, so a checkbox-only reconcile returned a payload byte-identical to a full reconcile while the traceability row stayed Pending (and re-run masked it as already-complete). It now surfaces `table_unmatched` for IDs whose checkbox reconciled but whose table row is absent, and treats a checked box with no table row as partial rather than done. (#2140) (#2219)
- state prune now resolves the current phase from the canonical location — frontmatter current_phase, the Current Phase field, or the prose Phase: line scoped to the ## Current Position section — instead of extracting Phase over the whole document, where stateExtractField's pipe-table fallback could latch onto an unrelated | Phase | N | row (e.g. a historical verification table) and compute a wrong prune cutoff. (#1832)
- **`model_overrides` Claude model IDs now resolve to Agent-tool aliases on the claude runtime** — a full Claude model ID (e.g. `claude-sonnet-5`) in `model_overrides` was returned verbatim and silently dropped by the Claude Agent tool (whose `model` parameter documents only tier aliases), causing the spawned subagent to inherit the parent session model instead of the configured one. It now maps to the tier alias (`sonnet`/`opus`/`haiku`/`fable`), consistent with the `model_policy` path (#1144). Bare aliases, non-Claude values, and non-Claude runtimes are unchanged; a Claude ID with no alias warns once and falls through to tier resolution. (#2041) (#2048)
- **`validate health` no longer false-flags the `adaptive` model profile, and now warns when a `models.<phase_type>` tier is invalid** — health reported `W004 invalid model_profile "adaptive"` for a profile that has been valid since v1.40, and a typo like `"planning": "opuss"` was accepted in silence while the resolver quietly ignored it. Health now sources its profile list from the model catalog and emits `W022` for unknown phase types and invalid tier values. (#2336)
- **Phase dirs whose slug leads with a multi-digit number (e.g. a year) resolve again** — a phase like `14-2026-photos-performance` (roadmap name "2026 Photos & Performance") had its phase token over-collected as `14-2026`, so `init.plan-phase`, `init.execute-phase`, `phase-plan-index`, `state.planned-phase`, and `roadmap.annotate-dependencies` reported `phase_dir=null` / `plan_count=0` while the directory existed. Continuation segments of a phase token are now capped at the exactly-2-digit zero-padded form the write side emits, via a single shared grammar source consumed by all five parsing sites (the residual case from #2043). (#2232) (#2254)
- Phase headers that place a parenthetical tag before the colon (`### Phase 26 (Cluster B): Title`) now resolve and enumerate the same as untagged headers. Previously the resolver returned not-found and `roadmap analyze`/listing silently dropped the phase (wrong phase_count, progress, and next_phase). Tag tolerance is applied at every phase-header read site; untagged and all existing header formats parse unchanged. (#1765)
- **`/gsd-stats` no longer misreports a phase as Not Started when two directories collide on the same phase key** — `cmdStats` now folds colliding statuses by precedence (Complete > Needs Review > Executed > In Progress > Planned > Not Started) instead of overwriting last-write-wins, so the furthest-along status wins regardless of `fs.readdirSync` order. Separately, `/gsd-health` now emits a new W023 warning whenever two or more real phase directories collide on the same normalized phase key, naming both directories and their independently-computed statuses (neutral wording — never guesses which is the real one). (#2461)
- Executor and milestone-summary/forensics workflows now call state.* commands with named flags so the named-only router records metrics, decisions, blockers, and session continuity instead of silently dropping positional args. (#1873)
- **bug-1367 install test no longer fails on Windows CI when hooks/dist isn't pre-built** — the test ran install.js without building its hooks/dist precondition (a gitignored build artifact the unit lane doesn't build), so on a lane without pre-built hooks the installer hit "Failed to install hooks: directory is empty" and the before-hook threw. The test now builds hooks in its own before() (mirroring golden-install-parity). (#1926) (#1927)
- **`/gsd-fast` now appends Quick Task rows to STATE.md again** — the log_to_state column-count guard used an off-by-one awk formula (`NF-1`) that was always one too high, so the schema gate rejected the very table quick.md creates and silently skipped the STATE.md update. Also now supports the 6-column validate-mode table. (#2133) (#2214)
- Build the gitignored `hooks/dist/` artifact once upfront in `scripts/run-tests.cjs` (the same chokepoint as `ensureBuiltArtifacts`), before any concurrent install test spawns `install.js`. Closes the scoped-CI first-build empty-dir race that intermittently failed install tests with `Failed to install hooks: directory is empty` (e.g. `bug-3683-workflow-colon-namespace-leak`). (#1967) (#1968)
- **workstream progress no longer reports shipped milestones as `executing`** — `gsd-tools workstream progress` now derives each workstream's status from authoritative shipped signals (an archived milestone snapshot under milestones/, or a SHIPPED marker in the workstream ROADMAP) instead of trusting the mutable STATE.md `Status` field, so a stale field can never hide a shipped/archived milestone. The output adds `status_source` (`field` | `derived`) and `status_conflict` (true when the derived value disagrees with the stale field). (#1913) (#1916)
- **Windows install/upgrade/state-write operations no longer fail on transient antivirus/indexer file locks** — the fs.renameSync atomic-publish sites (install state, hooks config, capability ledger/lifecycle, phase/workstream/milestone dirs, roadmap, planning/state locks) now retry EPERM/EBUSY/EACCES via retryRenameSync instead of propagating the transient lock; enforced by the new local/require-fs-op-fallback lint rule (ADR-1703 Phase 6). (#1740) (#1742)
- reconstructFrontmatter now emits valid YAML for scalars and block-array items that were previously serialized unescaped. Values carrying a YAML indicator plus a literal quote/backslash, embedded control characters, the empty string, a leading YAML indicator, or leading/trailing whitespace are now routed through a properly escaped double-quoted form, so frontmatter round-trips through strict parsers (js-yaml, PyYAML) instead of corrupting the block on the next state sync. (#1807)
- **`phase remove` no longer destroys the Progress table when removing the last phase** — deleting a phase used a whole-document regex whose scan, on the final phase, ran past the section and swept away the `## Progress` heading and its entire tracking table; the deletion is now structurally bounded to the phase’s own section. (#2253) (#2253)
- **Subagent prompts embedding orchestrator-relative planning paths now resolve correctly when the spawned subagent's own working directory differs from the orchestrator's (e.g. a git worktree)** — `init.*` (and `state.load`) command handlers now emit `state_path`, `roadmap_path`, `phase_dir`, `project_path`, `research_dir`, `codebase_dir`, `intel_dir`, `conflicts_path`, `debug_dir`, and similar fields as absolute paths anchored on the project root, and the planner/checker/verifier/synthesizer/roadmapper/debugger/mapper/classifier subagent-prompt blocks that previously hardcoded bare `.planning/...` literals now reference those fields instead; a subagent spawned into a different cwd would previously report real, already-committed files as missing. (#2376) (#2428)
- **`phases clear` archives phase directories instead of destroying them** — at a milestone switch, committed phase directories were hard-deleted (`rmSync`) with no archive, silently losing browsable phase history (the #1447 dirty-tree guard was a no-op for the common committed case). Phase directories are now moved to `milestones/<version>-phases/` (collision-safe; timestamp fallback when no version resolves), so history survives the switch. The #1447 uncommitted-changes guard is retained as a secondary backstop. (#1871) (#1919)
- **`/gsd-review` and `/gsd:ship` temp files are now scoped to a single per-run directory** — both workflows previously wrote prompt, section, and reviewer-output files to `/tmp/gsd-review-*-{phase}.*` keyed only on the bare phase number, so two projects sharing a phase number (or a crashed run's leftover file) could collide and silently feed a reviewer another project's stale content with no error; every temp path now lives under one `mktemp`-created run directory that's removed after the review completes. (#2358) (#2433)
- **Cross-AI review no longer silently drops the Codex/Claude/Gemini lanes on large plan sets** — the prompt-fed reviewer blocks in review.md invoked each CLI with no explicit timeout, so a slow source-grounded review was killed at the host default (~2 min) and the lane was silently lost. The workflow now directs a high Bash timeout and frames an empty output as a timeout (not the crash it was misdiagnosed as). (#2194) (#2226)
- **Runtime brand-swap no longer mislabels `<runtime_compatibility>` comparison tables** — every runtime installer that rebrands "Claude Code" to its own name (Cursor, Windsurf, Trae, Cline, CodeBuddy, Qwen, Hermes) also swapped it inside the runtime-comparison tables in shipped workflows, where "Claude Code" is a compared-runtime label, not a host self-reference — corrupting the comparison. Branding now protects `<runtime_compatibility>` regions while still rebranding genuine self-references. (#2284) (#2309)
- **`check tdd.review-checkpoint` no longer silently skips TDD plans with CRLF line endings** — the frontmatter regex at `src/check-command-router.cts:751` used literal `\n` which couldn't match a CRLF PLAN.md delimiter (`---\r\n`), so a Windows-authored `type: tdd` plan was silently classified as "no type:tdd plans found" and the advisory gate short-circuited to a confident pass with no violations table. The regex now uses the same CRLF-tolerant form (`/^---\r?\n([\s\S]*?)\r?\n---/`) already in use elsewhere in the same file (line 205, `extractPlanDesignatedSections`). With `core.autocrlf=input`, the triggering CRLF was invisible to `git diff`/`git status`, so the contributor had no way to tell their plan was being misclassified. (#2477)
- **Phase verification no longer reads `stale` from filesystem timestamps alone** — staleness is now derived from git commit times instead of file mtimes, so a phase whose report declares `status: passed` stays passed across a fresh `git clone`, `cp -R`, or an unrelated `touch`/reformat, instead of being silently downgraded to `stale` by a checkout-order mtime skew. (#2348) (#2394)
- **`/gsd-progress` no longer reports a stale root milestone in workstream mode** — in a multi-workstream project with no active workstream set, `gsd-tools query init.progress` silently fell back to root `.planning/STATE.md` (often stale) and reported it confidently. It now fails safe with an actionable error naming the available workstreams and the `--ws`/`workstream set` fix, so a stale root value is never reported. Flat mode and `--ws <name>` are unchanged. (#1912) (#1918)
- **Kilo installs now stage the shared PreToolUse guard hooks the native plugin spawns** — Kilo's capability descriptor declared both a `nativePlugin` (which spawns `gsd-prompt-guard`, `gsd-read-guard`, and `gsd-worktree-path-guard` as subprocesses) and `skipSharedHooksInstall: true` (which suppressed staging those scripts into the Kilo config dir), so every guard silently no-opped on every Kilo install. The skip flag is removed (Kilo now stages the same hooks bundle as OpenCode, whose byte-identical plugin was unaffected), and the plugin's `runHook` now warns loudly — once per hook file — when a guard script is missing instead of treating the absence as a silent allow. Resolves #2305. (#2327)
- **The decision-coverage gate no longer fails open on unrecognized decision-ID prefixes** — `check.decision-coverage-plan` classified a populated `<decisions>` block as "no trackable decisions" (a clean pass) whenever its IDs used a prefix the parser couldn't read (e.g. `D5-01` instead of `D-01`), silently skipping the gate on real decisions. The gate now recognizes any bold-lead-in decision bullet as evidence and fails loud (`could-not-parse`) when it can't read a populated block, instead of passing. (#2347) (#2389)
- Windows: stop double-quoting $CLAUDE_PROJECT_DIR-anchored managed node hook paths during the #2979 legacy rewrite, which produced "\"$CLAUDE_PROJECT_DIR\"/..." and broke every node managed hook with MODULE_NOT_FOUND (PreToolUse-guard deadlock). (#1746)
- **`/gsd-stats` and STATE.md progress no longer freeze stale `total_plans`** — the progress ratchet was applied to the whole progress record, so any single counter decreasing (e.g. `completed_plans`) froze every field including `total_plans`. Now `total_plans` always takes the freshly derived value (joining `total_phases` from #1446), so it corrects in both directions — upward when a new phase adds plans, downward when a milestone reorganization removes phases. The write-path `applyStatePreservation` also switched from wholesale block restore to per-field merge, so `state planned-phase` writes a consistent `total_plans` instead of the pre-transform stale value. (#2468)
- **phase complete now updates STATE progress on milestone-grouped roadmaps** — deriveProgressFromRoadmap parses the ## Progress table by header (column-by-name) instead of a fixed 4-column layout, so the 5-column milestone-grouped shape is no longer silently unparsed. (#2168)
- **Windows Claude Code hooks now work under PowerShell** — when Claude Code's hook runner resolves to PowerShell (not Git Bash), every GSD-installed hook failed with `Unexpected token` because the installer emitted bare quoted paths with no PowerShell call operator. The fix adds a `hookShell` parameter to the hook-command projection chain; when `hookShell='powershell'`, the `&` call operator is prepended. Default behavior (Git Bash, no prefix) is unchanged. (#2236) (#2261)
- **`/gsd-debug` now auto-resumes instead of stopping mid-investigation** — when the debug session-manager's own turn ended before the investigation was complete, the orchestrator treated the intermediate progress summary as completion and returned control to the user. It now recognizes a non-terminal `CONTINUE_REQUIRED` return, auto-resumes from the on-disk checkpoint, and only stops for genuine terminal conditions (with a no-progress anti-loop guard). (#2257) (#2300)
- **Installing a non-Claude runtime no longer breaks Claude's model resolution in no-project sessions** — the installer writes `resolve_model_ids:"omit"` for non-alias runtimes into the machine-wide `~/.gsd/defaults.json`, which any runtime read back, so install order silently flipped Claude's adaptive tier aliases (executor→sonnet, planner→opus) to an empty model string. Resolution is now scoped to the runtime actually resolving, via a per-install `.gsd-runtime` marker: Claude ignores a global-defaults omit and keeps its tier aliases, non-alias runtimes still omit, and an explicit project-level `omit`/`true` is always honored. (#2297) (#2332)
- **`check.decision-coverage-plan` no longer false-blocks on decisions cited in `<read_first>`/`<behavior>`/`<verify>`/`<acceptance_criteria>`/`<done>`** — the gate scanned only `<objective>`/`<tasks>`/`<task>`/`<action>` tag bodies while its remediation message claimed "(or body)". A decision faithfully cited in any of the five other planner-canonical tags (the natural place for "read this CONTEXT decision before editing" pointers, verification steps, acceptance criteria, etc.) was reported as uncovered with a misleading fix-hint that sent the fixer to "the body" — where a re-citation still failed. The scan now covers all nine planner-canonical tag bodies AND the message names the surfaces it actually scans, so message and behavior cannot drift apart again. (#2372) (#2443)
- **`capability state` and `loop render-hooks` now accept `--runtime` to override the auto-detected runtime** — previously both commands parsed only `--config-dir`, so the runtime config dir was derived from the persisted `.planning/config.json` runtime (precedence `GSD_RUNTIME` → `config.runtime` → `claude`). A repo that persisted `runtime:"codex"` resolved the config dir to `~/.codex`, where the Claude skill isn't installed, so every skill-bearing capability reported `surfaced:false` and `execute:post`/`verify:post` hooks silently no-op'd when the operator drove GSD from Claude Code. `--runtime <r>` (canonicalized, so aliases like `codex-app` work) now bypasses that fallback so the config dir resolves to the explicitly-named runtime's home. Behavior without the flag is unchanged. (#2003) (#2051)
- **`/gsd` now registers on pi** — installing GSD for pi wrote its extension as `gsd.cjs`, a suffix pi's extension auto-discovery skips silently, so `/gsd` never appeared and nothing reported an error. The extension now installs as `gsd.js`, and upgrading removes the stale `gsd.cjs`. (#2470) (#2478)
- **`phase complete` no longer false-reports REQ-IDs as missing when the traceability table leads with a status column** — the parser required the REQ-ID in the first column, so a table shaped `| ☐ | REQ-01 | …` matched zero rows and every body REQ-ID was reported missing. It now matches REQ-IDs in any column. (#2203) (#2234)
- **`init milestone-op` now ignores backlog `999.x` headings when counting milestone phases** — parked backlog items no longer inflate `phase_count` or pin `all_phases_complete` false for an otherwise finished milestone. (#1843) (#1843)
- **Phase archival is now wired end-to-end across the milestone lifecycle** — finishes the #1871 follow-up: `phases archive` is now a real command (the half-wired alias is routed, no longer errors Unknown), `milestone complete` archives phase dirs by default (`--no-archive-phases` opts out), and `new-milestone` §6 stages the archive move + source removal in the same commit so history is preserved atomically rather than left as orphaned uncommitted deletions. (#1871) (#1924)
- **`state update-progress` no longer mangles the frontmatter and discards the progress suffix** — its Progress: regex matched the raw STATE.md including frontmatter, so the YAML `progress:` key was hit first (corrupting the frontmatter) while the body line stayed stale and was silently reverted on the next write, and any descriptive suffix after the progress bar was destroyed. It now targets the body line only and preserves the suffix. (#2177) (#2224)
- **`/gsd-plan-review-convergence` no longer silently overrides configured reviewers with Codex** — a bare invocation (no reviewer flags) now respects `review.default_reviewers` (and, transitively, `review.reviewer_instances`) per ADR-0011/ADR-0015, instead of always injecting `--codex` and bypassing the configured default. Users without `review.default_reviewers` configured still get `--codex` as before. The startup banner now shows what will actually run. (#2451)
- **`/gsd-ship` no longer silently drops the ship-status note from STATE on merge** — the track_shipping step committed the STATE ship-note after creating the PR but never pushed it, so on a fast merge the note stayed local-only and never reached the default branch. The ship-note is now pushed onto the PR branch with a `[ci skip]` trailer so it lands on merge without a redundant pipeline. (#2138) (#2217)
- **`/gsd-debug` no longer stalls on a phantom background handoff** — the orchestrator treated the foreground session-manager spawn as a background task and queried its agent ID via TaskOutput (which needs a task ID), then waited on a handoff that was never queryable. The workflow now states the spawn is foreground/blocking, forbids passing an agent ID to TaskOutput, and gives a lost-handoff recovery path. (#2196) (#2227)
- **Roadmap phase lookup now ignores fenced examples and the backlog sentinel lane** — `roadmap get-phase` and `init plan-phase` no longer return fenced sample headings as real phases or treat `999.x` backlog items as active milestone work. (#1845) (#1845)
- **`phase complete` no longer checks the wrong ROADMAP checkbox or writes the plan count into a shipped milestone** — the roadmap mutators ran unanchored and un-milestone-scoped, so they could flip a bullet inside a backticked prose literal or a Backlog entry instead of the closing phase's, and write the plan count into a same-numbered phase in a shipped milestone. The checkbox flip is now line-anchored and both writers are scoped to the current milestone. (#2200) (#2229)
- **`audit-uat` no longer reports a false-clean `total_items: 0` when real items exist** — the parsers ignored two artifact shapes: a `## Gaps` section recording open findings, and verification items declared in frontmatter (`human_verification:` array) or as `### N.`+bold-paragraph blocks. audit-uat now surfaces unresolved `## Gaps` entries and reads the frontmatter array / heading shape, so a phase with outstanding UAT/verification work is no longer waved through as clean. (#2286) (#2317)
- **`claude_orchestration.enabled: true` now actually routes execute-phase waves through the Workflow backend** — the capability shipped registered-but-inert: nothing in `/gsd-execute-phase` ever called its backend detection, and the `execute:wave:pre` hook it needed was declared but never rendered, so enabling it had zero effect. execute-phase now renders `execute:wave:pre` before each wave and, when the capability is enabled and all gates pass, dispatches independent plans via the generated Workflow script; any gate miss or disabled config falls back to byte-identical inline dispatch. (#2285) (#2314)
- **`roadmap get-phase` resolves project-code-prefixed headings by bare number** — a bare-number query (e.g. `29`) now resolves a drifted `### Phase AB-29:` heading, matching the internal resolver used by `init.phase-op`; previously the CLI returned empty. A bare sibling (`### Phase 29:`) still takes precedence. A project-code-prefixed heading present only as a summary/checklist line (no matching detail section) now reports a `malformed_roadmap` diagnostic — for both prefixed and bare-number queries — instead of a silent empty result. (#2114) (#2139)
- **`query config-get` now returns capability-registry defaults for absent keys** — keys declared with a default in the capability registry (e.g. `workflow.security_enforcement`, which defaults to `true`) previously reported "Key not found" (exit 1) when missing from config.json, diverging from the runtime's own resolver and letting `... || echo false` guards silently read the security gate as disabled. config-get now resolves these through the same registry defaults the runtime uses. (#2256) (#2299)
- **`milestone complete --ws` now archives into the workstream instead of root** — the archive paths (MILESTONES.md, the milestones/ archive dir, and the per-version MILESTONE-AUDIT.md) were hardcoded to root `.planning/`, so a workstream milestone close scattered its artifacts into root and never produced a workstream-local archive. They now derive from the workstream-aware planning base (`planningPaths(cwd).planning`); flat-mode (no --ws) is unchanged. (#1911) (#1917)
- **`/gsd:new-milestone --ws <name>` no longer overwrites the shared PROJECT.md milestone heading** — in workstream mode the shared `.planning/PROJECT.md` had its `## Current Milestone` heading rewritten with one workstream's milestone, so with parallel workstreams whichever ran last silently won the shared heading. The milestone-state write in Step 4 is now skipped when a workstream is active, and the commit no longer stages PROJECT.md. The `--ws` flag is also now parsed into `${GSD_WS}`, which previously expanded to empty and silently dropped workstream scope from the suggested next-step routing hints. (#2338)
- **The context-monitor hook no longer fails Codex's Stop hook** — GSD wires `gsd-context-monitor` to Codex lifecycle events including `Stop`, but the hook emitted a `hookSpecificOutput.additionalContext` envelope that Codex's Stop schema rejects ("hook returned invalid stop hook JSON output") exactly when context was low. The hook now emits that envelope only for context-injection events (PostToolUse / AfterTool) and exits silently for Stop and every other lifecycle event, while its debounce and critical-session bookkeeping still run. (#2289) (#2324)
- **`phase complete` now reads milestone-grouped ROADMAP progress tables** — progress reported 0% on projects whose Progress table carries a Milestone column, because the reader assumed a fixed column position; it now resolves progress columns by name so both flat and milestone-grouped tables work (#2137). Quick Tasks logging via `/gsd:fast` also appends schema-correct, lock-safe rows instead of guessing the column count in shell (#2133). (#2248) (#2248)
- **Managed hooks no longer break after a volta node upgrade or prune** — on machines using volta to manage Node, the installer baked a version-pinned node path into every managed hook command. Once volta pruned that node version, every hook failed to spawn with `No such file or directory` at the start of each session, until the installer was re-run. Hook commands now resolve through volta's stable shim, which survives version changes. (#2335) (#2375)
- **Todo severity is now captured and surfaced end-to-end** — `/gsd-capture` (add-todo) now confirms a severity (blocker/major/minor/cosmetic) before writing a todo instead of silently omitting it, and `gsd-tools list-todos` / `init todos` now include the `severity` field in their JSON output (omitted for older todos that have none), so a backlog can be triaged by severity instead of by re-reading every file. (#2337) (#2381)
- **Skill-bearing capabilities now surface correctly on flat command-layout installs** — on an install using the flat `commands/gsd-<stem>.md` source layout (e.g. a Claude Code local project install with no `commands/gsd/` subdir), every skill-bearing capability (`nyquist`, `code-review`, `security`, `ui`, `mempalace`, `ai-integration`, `profile-pipeline`) was silently reported `surfaced:false`/`enabled:false`/`active:false`, so their loop hooks (`verify:post`, `execute:post`, etc.) never fired even with the corresponding `workflow.*` toggle on. The skill-manifest resolver now detects the flat layout and produces the same stems the nested `commands/gsd/*.md` loader does. (#1858) (#2049)
- **Claude Code installs now pre-approve `.planning/` and `STATE.md` writes** — the installer wrote `Write(.planning/*)`/`Write(STATE.md)` permission rules, but Claude Code has no standalone `Write` gate (file edits are gated via `Edit(pattern)`), so those rules never matched and every fresh install still hit first-run approval prompts (and a session-start warning). The installer now writes `Edit(...)` rules and migrates the stale `Write(...)` entries away on the next run. (#2278) (#2302)
- **Roadmap, requirements, and state table edits are confined to the right table** — the last ad-hoc table writers (phase completion updating roadmap progress, `requirements mark-complete`, and `state record-metric`/velocity) now route through the shared markdown-table seam, so a stray decoy table elsewhere in a document can no longer swallow a phase-progress update, a single ragged neighbouring row no longer silently aborts the whole edit, and per-plan metric recording no longer drops trailing section content or duplicates the section. (#2253) (#2253)
- **Installed third-party capability skills now materialize as real slash commands** — a capability could pass every check (`installed: true, surfaced: true, active: true`) and still never exist on disk: the registry layer counted the capability's skill as surfaced, but the file-copy step only ever scanned gsd-core's own bundled commands, so nothing was ever written to the runtime's `skills/` directory and the command was never invocable. Installed capability skills are now staged from where they live, bound to the capability that actually declared and registered them (never inferred from directory listing order), and are subject to the same runtime-targeted body rewrites as first-party skills — first-party skills still win any name collision. (#2340)
- **`/gsd:plan-review-convergence` can now use the Antigravity CLI reviewer** — its reviewer-flag whitelist predated the 1.7.0 Antigravity adapter and silently dropped `--agy`/`--antigravity`, so convergence fell back to `--codex` only and the working adapter was unreachable (especially after Gemini CLI's upstream shutdown). Both flags are now recognized and passed through to `/gsd-review` unchanged. (#2293) (#2325)
- **`npm run lint:ci` (and every npm script banner) on `next` and feature branches cut from `next` no longer reports a stale pre-release version after a final release** — the release pipeline's `finalize` job shipped `X.Y.0` to npm `latest` but never bumped `next` to match, so `next` carried the last `rc.N` placeholder indefinitely (observed: `1.7.0-rc.6` lingering after `1.7.0` shipped). The `finalize` job now runs `scripts/sync-next-version.cjs` — the same step the `rc` job already ran — keeping `next` at the last published release for every release type as `scripts/sync-next-version.cjs:6-9` always promised. (#2423) (#2437)
- **`verify plan-structure` no longer false-flags checkpoint tasks for missing `<action>`/`<verify>`/`<done>`** — every `<task type="checkpoint:*">` was reported as a structural error because the verifier unconditionally required the auto-task fields. It now branches on the task's `type` attribute: `checkpoint:human-verify` requires its canonical triple (`<what-built>`/`<how-to-verify>`/`<resume-signal>`), `checkpoint:decision` requires `<decision>`/`<options>`/`<resume-signal>`, `checkpoint:human-action` requires `<action>`/`<instructions>`/`<verification>`/`<resume-signal>` (per `gsd-core/references/checkpoints.md`), and unknown `checkpoint:*` subtypes require only the universal `<resume-signal>`. Non-checkpoint tasks keep the historical `<action>`/`<verify>`/`<done>`/`<files>` requirements unchanged. (#2473)
- **Hermes installs now project named-agent dispatch onto `delegate_task` instead of asserting a nonexistent `Agent` tool** — installed Hermes workflows brand-swapped "Claude Code"→"Hermes Agent" but kept literal `Agent(...)` calls and falsely claimed "The Agent tool IS available", which Hermes doesn't expose. A Hermes `.md` converter now rewrites named dispatch onto Hermes's `delegate_task` contract (embedding the resolved role prompt since Hermes has no named-agent lookup, mapping background dispatch, dropping unsupported per-call model), driven by the runtime's documented dispatch facts, and fails closed if a referenced role prompt is missing. (#2284) (#2309)
- **`phase complete` no longer silently drops requirement IDs the roadmap cites but REQUIREMENTS.md never defined** — completing a phase whose `**Requirements**:` line named an unregistered REQ-ID reported `requirements_updated: true` with zero warnings while the file was left byte-for-byte unchanged, indistinguishable from a run that wrote everything. Ghost IDs now raise a warning, `requirements_updated` reflects whether a write actually landed, an active heading like `## v1 Requirements` is no longer mistaken for a deferred section, and a phase whose every cited ID is unregistered still reports its missing-requirement rows instead of "No requirements or decisions to check." (#2339)
- **`~/.gsd/defaults.json` no longer silently drops `model_policy`, `model_profile_overrides`, and `runtime`** — the global-defaults path of config load now forwards these three keys identically to a project's `.planning/config.json`, so a machine-wide model policy / runtime / overrides specified globally is honored even outside a project. (#2069) (#2442)
- **ROADMAP phase edits can no longer escape their section** — completing a phase updated its plan count and per-plan checkboxes with whole-document regexes that could bleed into a neighbouring phase; those per-phase writes are now structurally bounded to the phase own section via a new `withSection` / `withPhaseSection` seam (#2130, #2067, #2080). (#2250) (#2250)
- **`close_phase_todos` no longer leaves moved todos as phantom unstaged deletions in `git status`** — the workflow step moved resolved todos from `.planning/todos/pending/` to `.planning/todos/completed/` with a plain `mv`, then committed by listing only the destination directory in `--files`. Git's index still tracked the moved file at its old `pending/` path, so the deletion was never staged and the moved-away file lingered as an unstaged deletion in `git status` until some later broad `git add -A` happened to catch it. The step's commit `--files` list now includes BOTH directories so `git add .planning/todos/pending/` stages the deletion atomically with the new `completed/` copy in the same commit. (#2415) (#2447)
- **STATE.md `## Session` fields now resolve on Windows** — the session-section reader used a `\n`-only heading regex that silently failed on a CRLF `## Session` heading, nulling all session state on Windows checkouts; it now reads through the CRLF-safe section seam. (#2253) (#2253)
- **Bullet/em-dash ROADMAP phases no longer resolve to `Phase null`** — the roadmap phase lookup matched only ATX headings with a colon, so a bullet entry like `- [ ] **Phase N — Name**` (which the roadmapper emits) failed to resolve and `Phase null` landed in STATE.md; a bullet-only ROADMAP also broke the milestone phase count. Phase lookup and the milestone filter now accept bullet/checkbox entries with an em-dash/en-dash/hyphen/colon separator. (#2199) (#2228)
- **Linuxbrew users no longer lose all GSD-managed hooks after `brew upgrade node`** — normalizeNodePath only recognized macOS Homebrew Cellar paths, so on Linux the version-pinned node path stayed baked into hook commands and 404'd after a node bump (and reinstall couldn't repair it). It now rewrites any Homebrew Cellar path — Intel, Apple Silicon, Linuxbrew, custom HOMEBREW_PREFIX — to the stable `<prefix>/bin/node` symlink. (#2185) (#2225)
- **`milestone complete` no longer corrupts the recorded phase** — closing a milestone (e.g. `v0.5`) previously overwrote `current_phase` in STATE.md with the version's minor digit, and a follow-up `state complete-phase` mined a bogus `0.5` token and rewrote the file; phase resolution is now anchored so the real phase is preserved and a milestone-closure line is rejected. (#2111) (#2131)
- **Headless MemPalace capture no longer fails silently** — the headless invocation `mempalace mine <path> --wing <wing> --room <room>` used a `--room` flag that does not exist on the `mine` subcommand (only `search` accepts `--room`), causing every headless/no-MCP capture run to fail with `unrecognized arguments: --room` and silently skip (onError: skip). The fix replaces the flag with MemPalace's documented room-assignment mechanism: stage the artifact under a room-named subfolder with a `mempalace.yaml` taxonomy so `detect_room()` assigns it via folder-path match. (#2220) (#2260)
- **Codex agents no longer fail to launch with an unsupported-model error** — GSD was writing an Anthropic tier name (`opus`/`sonnet`/`haiku`/`fable`) or a `claude-*` id into each Codex agent's `.toml` `model` field, which Codex rejects — fatally on a ChatGPT account (`The 'sonnet' model is not supported when using Codex with a ChatGPT account`). GSD now never writes an Anthropic-flavored model to a Codex agent: an explicit real-Codex model pin is kept, anything else is omitted so the agent inherits the working session model. (#2310) (#2312)
- Fixed: a hand-authored non-inferable backstop truth with a stray trailing space or surrounding quotes no longer silently grades green — it correctly abstains (insufficient_spec), restoring the #1154 honest-verifier guarantee. (#1909)
- **`commit_docs` no longer silently disables on CRLF `.gitignore` repos** — git check-ignore falsely reports a trailing-slash path (e.g. `.planning/`) as ignored when the .gitignore has CRLF line endings with blank lines. isGitIgnored now strips trailing slashes before querying, so the false positive cannot occur. (#2206) (#2235)
- **Phase-directory resolution fails loud on cross-project collisions** — when two unrelated GSD projects share a `.planning/phases/` tree, a bare phase number silently resolved to the first `0N-*` directory found. The fix detects multiple matches and surfaces an `ambiguous_matches` result. (#2237) (#2262)
- **Build/test gates no longer report a false failure on repos with no detectable build/test tooling** — the post-merge, regression, verify-phase, and audit-fix gates read `config-get workflow.build_command`/`workflow.test_command` without `--raw`, so an unset key returned the literal 2-byte string `""` rather than empty output. The `[ -z "$CMD" ]` guard then saw a non-empty value, skipped the Makefile/Cargo/go.mod/package.json auto-detection cascade, and executed the literal `""` as a command → exit 127, misread as a build/test failure (docs-only or planning-only repos, or any repo before its first build file). All of these reads now pass `--raw`, restoring the intended "no command detected — skip" no-op. (#2350) (#2399)
- **`scanPhasePlans` no longer counts PLAN-REVIEW artifacts as executable plans** — `*-PLAN-REVIEW.md` files were counted by the loose `/PLAN/i` fallback. The fix adds a `PLAN_REVIEW_RE` exclusion before the fallback. (#2252) (#2263)
- **Dependency tree no longer carries a known body-parser advisory** — GHSA-v422-hmwv-36x6 (low-severity DoS via invalid `limit` value, published 2026-07-20) in `body-parser@2.2.2` was pulled transitively via `@anthropic-ai/claude-agent-sdk` → `@modelcontextprotocol/sdk` → `express` and surfaced by `npm audit --omit=dev`. Re-resolved `body-parser` to 2.3.0 in `package-lock.json` within `express`'s already-declared `^2.2.1` range; no `overrides` block needed, `package.json` is unchanged. (#2473)
- **Milestone audit no longer flags a not-yet-validated phase as a Nyquist failure** — a phase that was planned but never run through `validate-phase` now reports as NOT-VALIDATED (a "run validate-phase" TODO) instead of collapsing into PARTIAL alongside phases whose validation genuinely failed. (#2117) (#2209)
- **CI gates no longer fail with `no merge base` on branches behind the base.** The mutation, changeset-required, and docs-required workflows shallow-fetched the base *ref*, truncating the ancestry their three-dot `origin/<base>...HEAD` diffs depend on — so the mutation gate reported failure and silently skipped its Stryker shards, leaving the 80% threshold unverified on any PR not already level with `next`. (#2452) (#2485)
- **OpenCode slash commands now install to the supported `commands/` directory instead of OpenCode's legacy `command/` alias** — GSD wrote all ~71 `/gsd-*` commands to `command/` (singular), which OpenCode's docs list only as a backwards-compatibility alias for the documented `commands/` (plural) convention. Commands now land in `~/.config/opencode/commands/` (global) and `.opencode/commands/` (local), and upgrading migrates the legacy directory, preserving any files you put there yourself. OpenCode currently resolves both names, so this is an alignment rather than a rescue — it takes GSD off a path the vendor may withdraw. Kilo is unaffected. (#2354)
### Security
- **`gate="blocking-human"` checkpoints are no longer auto-approved by the execute-phase orchestrator** — the package-legitimacy gate (#2827) spans two layers: `gsd-executor` refuses to auto-approve a `gate="blocking-human"` checkpoint and escalates it via `checkpoint_return_format` so a human can vet the package, and `execute-phase`'s `checkpoint_handling` step decides what happens next. That step dispatched purely on checkpoint *type* and never read `gate`, so under `--auto` / `--chain` it immediately auto-approved the very checkpoint the executor had just refused to auto-approve (`human-verify → {user_response} = "approved"`). The slopsquatting defence was therefore inert in exactly the unattended mode where nobody is watching: an `[ASSUMED]`/`[SUS]` package reached install with no human ever seeing the verification prompt. `checkpoint_handling` now carves out `gate="blocking-human"` (and the package-legitimacy `what-built` markers) ahead of every auto-mode branch, routing those checkpoints to the standard present-to-user flow regardless of type. `references/checkpoints.md` documents the `gate` attribute and its two values for the first time — previously `blocking-human` appeared nowhere outside `agents/gsd-executor.md`, so no planner had a documented way to author a checkpoint that auto-mode could not bypass. The existing regression test asserted the executor half only; it now asserts the orchestrator half too, which is why it stayed green while the gate was open. (#2107) (#2113)
- **Patched a transitive denial-of-service advisory in the production dependency tree** — `body-parser` reached GSD via the Claude Agent SDK's MCP dependency and, on versions through 2.2.2, silently stopped enforcing request size limits when given an invalid limit value (GHSA-v422-hmwv-36x6). Pinned to >=2.3.0. (#2470) (#2478)
- **`phases.clear --archive-version` and `milestone complete <version>` now reject version labels containing path separators or `..`** — the milestone version becomes a filesystem directory name that phase directories are moved into, so an unvalidated value could relocate phase history outside `.planning/milestones/`. Both now validate against a strict version-token pattern and fail loudly. (#2288) (#2323)
- **`query config-get` no longer leaks secret values or walks the prototype chain** — the `--default` fallback path printed secret-named keys (e.g. `brave_search`) in plaintext instead of masking them, and dotted-key traversal used raw property access so `config-get __proto__`/`constructor` resolved to JavaScript internals at exit 0 instead of erroring. Both absent-key resolution and traversal are now masked and own-property-gated. (#2256) (#2299)
- **Hardened phase/roadmap/plan markdown parsing against quadratic-time (ReDoS) CPU exhaustion** — a crafted `ROADMAP.md`, `STATE.md`, or `PLAN.md` with large runs of unclosed `(`, `[`, `<tag>`, `<!--`, or `<details>` could drive the phase-header, Plans-count, `files_modified`, and `<tag>`-block parsers into O(n²) scans (tens of seconds on a ~1.5 MB file). Every affected regex is now linear: header tag/bracket clauses are length-bounded, the Plans-count scan is section-local, and all `<tag>…</tag>` extraction routes through a single ReDoS-safe seam. (#2128) (#2141)
- **Installer writes are now confined to the declared config home** — the workflow/skill emit path (`copyWithPathReplacement`) and the Codex config writer (`installCodexConfig`) now reject any destination that escapes the install root: crafted or absolute paths, path-separator agent names, and pre-existing symlinks are refused before any delete or write. Fail-closed: an install write with no declared root is rejected rather than written unconfined. (#1725)
- **Install write-confinement (ADR-1239 Phase B)** — the installer now rejects any runtime-descriptor `destSubpath` that would write or delete outside the user's config home (path traversal, the config root itself, NUL bytes) and refuses to follow a pre-existing symlink that escapes it. Hardening only; no change to legitimate installs. (#1706)
## [1.7.0] - 2026-07-15
### Added

View File

@@ -53,7 +53,7 @@ Module owning projection from dispatch results/errors to CLI `{ exitCode, stdout
Module owning STATE.md parse, field extraction, field replacement, status normalization, and frontmatter reconstruction. It does not scan `.planning/phases` and does not own persistence or locking; phase/plan/summary counts arrive from inventory/progress Modules as inputs, and read-modify-write paths remain Adapters. Source of truth: `gsd-core/bin/lib/state-document.cjs`.
### STATE.md Transition Module
Module owning STATE.md lifecycle/maintenance transitions as intent-based methods (`beginPhase`, `advancePlan`, `completePhase`, `plannedPhase`, `milestoneSwitch`, `milestoneComplete`, `patch`, `sync`, `prune`, `update`, `rebuild`). Pure core `(content, intent, deps) → newContent` with injected I/O (file read/write, lock, disk scan); consults a field-classification table that names each STATE.md field's class (`derived-from-body` | `derived-from-disk` | `derived-from-external` | `curated` | `free`) and its preservation policy. Supersedes the 14 scattered RMW callbacks in `state.cts` and the direct `writeStateMd` callers in `milestone.cts:352` and `phase.cts:1770`; verify's `regenerateState` factory-reset primitive stays as a direct `writeStateMd` call. Absorbs `syncStateFrontmatter` + `readModifyWriteStateMd`'s post-sync preservation block; Encoding 3 (`cmdStateBuildFrontmatter`) stays separate — read path concern. Sibling/super-module of the STATE.md Document Module; consumes its `stateReplaceField`/`stateExtractField` primitives. Body section structure (`## Current Position`, `## Session`, etc.) lives as a constants block inside the Module. Append-only transitions (`addDecision`, `addBlocker`, etc.) stay on today's RMW seam for now. Targets the #1760/#1761/#1743/#1695/#1264/#1255/#1257/#3242 bug cluster. Migration per ADR-1372 §T6 sequenced as substrate + `beginPhase` first (PR1), then transition-by-transition with characterization tests first per transition. **ADR-1817 adds `rebuild` as the capstone 11th transition — the body-structure derivability contract.** Re-derives `## Current Position` prose from frontmatter and `## By-Phase Progress` table from phase dirs on disk; preserves `## Session` / `## Decisions` / unknown sections verbatim; de-duplicates `## Session Continuity Archive` (keep most-recent N, default 3); appends a structured audit entry to `## Rebuild Log` (`timestamp`, `kind`, `section`, `before`, `after`, `reason`) for every mutation. Hard idempotency guarantee: a no-mutation rebuild appends no log entry, so two successive invocations on a clean file are byte-identical. Non-overlapping with `sync` (3 lightweight frontmatter fields, auto-triggered) and orthogonal to `auto_prune_state` (age-based removal) — `rebuild` reconciles with current canonical sources, `prune` removes by retention policy, the two compose (rebuild first, then prune). Section ordering is invariant: rebuild rewrites content in place, never reorders. Targets the #1776/#1761/#1591 body-drift cluster that survived ADR-1769's per-field transitions. Phased per ADR-1817: Phase 0 = this ADR + predicates (closes #1817), Phase 1 = `rebuildCore` body + `rebuild` dispatch case + drift-class unit tests (#1827), Phase 2 = `cmdStateRebuild` CLI + `--dry-run`/`--verbose` + integration tests + docs + changeset (#1826). Source of truth: `gsd-core/bin/lib/state-transition.cjs` (generated from `src/state-transition.cts`).
Module owning STATE.md lifecycle/maintenance transitions as intent-based methods (`beginPhase`, `advancePlan`, `completePhase`, `plannedPhase`, `milestoneSwitch`, `milestoneComplete`, `patch`, `sync`, `prune`, `update`, `rebuild`). Pure core `(content, intent, deps) → newContent` with injected I/O (file read/write, lock, disk scan); consults a field-classification table that names each STATE.md field's class (`derived-from-body` | `derived-from-disk` | `derived-from-external` | `curated` | `free`) and its preservation policy. Supersedes the 14 scattered RMW callbacks in `state.cts` and the direct `writeStateMd` caller in `milestone.cts:552` (phase.cts's former direct caller has since been migrated away); verify's `regenerateState` factory-reset primitive stays as a direct `writeStateMd` call (`verify.cts:1925`). Absorbs `syncStateFrontmatter` + `readModifyWriteStateMd`'s post-sync preservation block; Encoding 3 (`cmdStateBuildFrontmatter`) stays separate — read path concern. Sibling/super-module of the STATE.md Document Module; consumes its `stateReplaceField`/`stateExtractField` primitives. Body section structure (`## Current Position`, `## Session`, etc.) lives as a constants block inside the Module. Append-only transitions (`addDecision`, `addBlocker`, etc.) stay on today's RMW seam for now. Targets the #1760/#1761/#1743/#1695/#1264/#1255/#1257/#3242 bug cluster. Migration per ADR-1372 §T6 sequenced as substrate + `beginPhase` first (PR1), then transition-by-transition with characterization tests first per transition. **ADR-1817 adds `rebuild` as the capstone 11th transition — the body-structure derivability contract.** Re-derives `## Current Position` prose from frontmatter and `## By-Phase Progress` table from phase dirs on disk; preserves `## Session` / `## Decisions` / unknown sections verbatim; de-duplicates `## Session Continuity Archive` (keep most-recent N, default 3); appends a structured audit entry to `## Rebuild Log` (`timestamp`, `kind`, `section`, `before`, `after`, `reason`) for every mutation. Hard idempotency guarantee: a no-mutation rebuild appends no log entry, so two successive invocations on a clean file are byte-identical. Non-overlapping with `sync` (3 lightweight frontmatter fields, auto-triggered) and orthogonal to `auto_prune_state` (age-based removal) — `rebuild` reconciles with current canonical sources, `prune` removes by retention policy, the two compose (rebuild first, then prune). Section ordering is invariant: rebuild rewrites content in place, never reorders. Targets the #1776/#1761/#1591 body-drift cluster that survived ADR-1769's per-field transitions. Phased per ADR-1817: Phase 0 = this ADR + predicates (closes #1817), Phase 1 = `rebuildCore` body + `rebuild` dispatch case + drift-class unit tests (#1827), Phase 2 = `cmdStateRebuild` CLI + `--dry-run`/`--verbose` + integration tests + docs + changeset (#1826). Source of truth: `gsd-core/bin/lib/state-transition.cjs` (generated from `src/state-transition.cts`).
### STATE.md Status Lifecycle (ADR-2207)
The `Status` field in STATE.md follows a strict lifecycle: `Ready to plan` → `All phases complete` (all phases done, milestone awaiting formal close) → `<version> milestone complete` (terminal, written only by the milestone-close verb `milestoneCompleteCore`) → `Awaiting next milestone` (archived). Phase-completion verbs write `All phases complete` on the last phase — never `Milestone complete` (the overloaded bare value was removed in #2204 per ADR-2207 to decouple phase-level writes from milestone termination). `normalizeStateStatus` maps any status containing "complete" → `completed`, so consumers using the normalized projection (workstream inventory's `status` field, statusline) recognize `All phases complete` without code changes. Note: `isCompletedInventory` (workstream-inventory-builder.cts) intentionally checks only for the terminal `\bmilestone\s+complete\b` / `\barchived\b` — `All phases complete` returns `false` (intermediate, not terminal).
@@ -76,6 +76,9 @@ Module owning the `init.*` family of query handlers that compose atomic queries
### Command Routing Hub
Single dispatch seam (`gsd-core/bin/lib/command-routing-hub.cjs`) that centralizes CJS routing, the no-throw pure-result contract, typed error variants, and dispatch-event emission for all command family adapters. Interface: `createHub({ cjsRegistry, manifest, logger }) → hub`; `hub.dispatch({ family, subcommand, args, cwd, raw, parentTraceId? }) → Result` where `Result = { ok: true, data } | { ok: false, kind, ...typedPayload }` and `kind ∈ { UnknownCommand, InvalidArgs, HandlerRefusal, HandlerFailure }`. The `InvalidArgs` variant carries an optional `exitReason?: string` field (amendment #1642 / #1644 Phase 1) holding the `ERROR_REASON` enum value, separate from `reason` (the explanation text); the `makeInvalidArgs(arg, reason, exitReason?)` factory omits the field when the third arg is absent, undefined, or empty — preserving the strict-keys invariant tested at `tests/command-routing-hub.test.cjs:444`. The Hub is single-runtime (no mode selection, no sdkLoader), never prints, never exits, never throws. Adapters call `createHub`, dispatch, then translate the pure Result to `output()`/`error()` calls; when an `InvalidArgs` Result carries `exitReason`, the adapter passes it as the second arg to `error(message, exitReason)` so the JSON-error envelope (`GSD_JSON_ERRORS=1`) preserves the typed reason. Source: `gsd-core/bin/lib/command-routing-hub.cjs`; ADR: `docs/adr/0174-retire-gsd-sdk-package-boundary.md` (§5 amended #1642).
### Command Dispatch Completion (ADR-2346, epic #2345)
Decision to dissolve `runCommand`'s 73-case ~2,338-line switch (the repo's #1 PageRank / #1 Tarjan-bridge / #1 complexity symbol) into a **two-layer dispatch**: families via the `commandFamilies` registry (ADR-959 mechanism, completed) and single-purpose **leaf verbs** via a dispatch table filling the prepared `_dispatchNonFamily` seam (today a dead shim returning `false`); `runCommand` collapses to a ~15-line `try registry → try leaf table → unknown`. Family/leaf rule: promote a cluster to a family iff ≥3 related subcommands + shared backing module + shared parse/return shape; lone verbs or pairs stay leaves. Nine families result (`state`/`phase`/`init`/`roadmap`/`validate`/`verify`/`capability` + promoted `config`/`research`/`resolve`/`git`); `worktree`+`workstream` stay leaves; ~40 remaining verbs rehome into ~4 themed leaf modules. A shared `parseFamilyArgs` (in `cjs-command-router-adapter.cts` beside `routeHubCommandFamily`) deletes the 4+ duplicated inline arg-parsers. The 706-line `capability` arm extracts to a thin `capability-command-router` (intel-shaped) + `capability-cli.cts` (CLI wiring), consolidating duplicated probes (`capHostVersion`→`readHostVersion()`, `capReadStrict`+drift-guard → one `readStrictKnownRegistries`). Behavior-preserving — each cutover proven equivalent by extending `tests/audit-command-cutover.test.cjs`'s 5-category template. Phased: P1 `parseFamilyArgs`+Tier-1 cutovers, P2 capability extraction, P3 new families, P4 leaf table + collapse. Source: `docs/adr/2346-command-dispatch-completion.md`; graduates ADR-959 `Proposed → Accepted`. _Avoid_: "the dispatch service" (when you mean the seam).
### Runtime Source Layout Module
Single-runtime seam layout for this repository after SDK retirement. Runtime execution paths live under `gsd-core/bin/lib/` and are grouped by seam concern (dispatch, manifest, handlers, runtime, observability, installer). ADR-0174 preserves the seam vocabulary and defines the canonical long-term shape as a seam-aligned TypeScript `src/` tree (`src/dispatch/`, `src/handlers/`, `src/errors/`, `src/manifest/`, `src/config/`, `src/state/`, `src/workstream/`, `src/runtime/`, `src/cli/`, `src/observability/`) compiled to CJS.
@@ -137,13 +140,13 @@ Module owning the layout-driven runtime-artifact install pipeline — `installRu
Module owning validation for Installer Migration Module records and planned actions. It enforces migration metadata, explicit install scopes, ownership evidence for destructive/config actions, and runtime contract citations for runtime config rewrites before a migration can enter planning or apply.
### Installer Module
Primary installer for all runtimes. Single production file: `bin/install.js` (generated). Exports: `install(isGlobal, runtime[, configDir])` → typed result `{ runtime, configDir, settingsPath, settings, statuslineCommand, updateBannerCommand }`; `uninstall(isGlobal, runtime[, configDir])`; `installRuntimeArtifacts(runtime, configDir, scope, resolvedProfile)`; `uninstallRuntimeArtifacts(runtime, configDir, scope)`; `writeManifest(configDir, runtime)`. Runtime enum: `allRuntimes` (15 values: claude, antigravity, augment, cline, codebuddy, codex, copilot, cursor, hermes, kilo, kimi, opencode, qwen, trae, windsurf). Directory helpers: `getDirName(runtime)` → local dir name; `getConfigDirFromHome(runtime, isGlobal)` → shell-quoted path fragment. Per-runtime global config-dir resolution is delegated to `gsd-core/bin/lib/runtime-homes.cjs:getGlobalConfigDir(runtime[, explicitDir])` — the canonical, env-var–aware projection (`explicitDir` override + opencode/kilo `*_CONFIG` file-path precedence); the legacy in-installer `getGlobalDir`/`getOpencodeGlobalDir`/`getKiloGlobalDir` were retired into it (#56). The same module exposes `detectAntigravityDirAmbiguity(opts)` — a side-effect-free probe reporting whether multiple `~/.gemini/antigravity{,-ide,-cli}` dirs coexist and which one GSD's `gsd-core/VERSION` marker (the `dot-home-nested` `probeExists`) resolves to, for installer / `/gsd-update` operator guidance when a pre-#217 install landed in the wrong sibling dir (#1441). Runtime-specific helpers: `resolveKiloConfigPath(configDir)`, `configureKiloPermissions(isGlobal[, explicitDir])`. Claude-specific permission helpers: `mergeClaudePermissions(settings)` — non-destructively appends GSD-owned allow/deny entries (see `GSD_CLAUDE_ALLOW_PERMISSIONS`, `GSD_CLAUDE_DENY_PERMISSIONS` constants) to a Claude Code settings object; called from `finishInstall` for `runtime === 'claude'` only; uninstall removes exactly these entries (#768). Layout-driven artifact copy/removal delegates to `gsd-core/bin/lib/runtime-artifact-layout.cjs:resolveRuntimeArtifactLayout` (throws `TypeError` for unknown runtimes). Seven runtimes with non-recursive skill loaders (claude global, cline, qwen, hermes, augment, trae, antigravity) use a nested router layout: 6 `gsd-ns-*` router bundles emitted as top-level skills, with concrete skills nested at `<router>/skills/<name>/SKILL.md` (hermes prefix='': `skills/gsd/ns-*/…`). The remaining skills-runtimes (cursor, codex, copilot, windsurf, codebuddy, opencode, kilo) use the flat `skills/gsd-<stem>/` layout unchanged. See Skill Surface Budget Module and Runtime Artifact Layout Module.
Primary installer for all runtimes. Single production file: `bin/install.js` (generated). Exports: `install(isGlobal, runtime[, configDir])` → typed result `{ runtime, configDir, settingsPath, settings, statuslineCommand, updateBannerCommand }`; `uninstall(isGlobal, runtime[, configDir])`; `installRuntimeArtifacts(runtime, configDir, scope, resolvedProfile)`; `uninstallRuntimeArtifacts(runtime, configDir, scope)`; `writeManifest(configDir, runtime)`. Runtime enum: `allRuntimes` (17 values: claude, antigravity, augment, cline, codebuddy, codex, copilot, cursor, hermes, kimi, kilo, opencode, pi, qwen, trae, windsurf, zcode). Directory helpers: `getDirName(runtime)` → local dir name; `getConfigDirFromHome(runtime, isGlobal)` → shell-quoted path fragment. Per-runtime global config-dir resolution is delegated to `gsd-core/bin/lib/runtime-homes.cjs:getGlobalConfigDir(runtime[, explicitDir])` — the canonical, env-var–aware projection (`explicitDir` override + opencode/kilo `*_CONFIG` file-path precedence); the legacy in-installer `getGlobalDir`/`getOpencodeGlobalDir`/`getKiloGlobalDir` were retired into it (#56). The same module exposes `detectAntigravityDirAmbiguity(opts)` — a side-effect-free probe reporting whether multiple `~/.gemini/antigravity{,-ide,-cli}` dirs coexist and which one GSD's `gsd-core/VERSION` marker (the `dot-home-nested` `probeExists`) resolves to, for installer / `/gsd-update` operator guidance when a pre-#217 install landed in the wrong sibling dir (#1441). Runtime-specific helpers: `resolveKiloConfigPath(configDir)`, `configureKiloPermissions(isGlobal[, explicitDir])`. Claude-specific permission helpers: `mergeClaudePermissions(settings)` — non-destructively appends GSD-owned allow/deny entries (see `GSD_CLAUDE_ALLOW_PERMISSIONS`, `GSD_CLAUDE_DENY_PERMISSIONS` constants) to a Claude Code settings object; called from `finishInstall` for `runtime === 'claude'` only; uninstall removes exactly these entries (#768). Layout-driven artifact copy/removal delegates to `gsd-core/bin/lib/runtime-artifact-layout.cjs:resolveRuntimeArtifactLayout` (throws `TypeError` for unknown runtimes). Five runtimes with non-recursive skill loaders (cline, qwen, hermes, augment, trae) use a nested router layout: 6 `gsd-ns-*` router bundles emitted as top-level skills, with concrete skills nested at `<router>/skills/<name>/SKILL.md` (hermes prefix='': `skills/gsd/ns-*/…`). claude (reverted from nested per #924 — the Skill tool errors on unrouted names) and antigravity (one-level scan, but concrete skills must be top-level discoverable) plus the remaining skills-runtimes (cursor, codex, copilot, windsurf, codebuddy, opencode, kilo) use the flat `skills/gsd-<stem>/` layout. See Skill Surface Budget Module and Runtime Artifact Layout Module.
### I/O Module
Module owning the tool's CLI I/O primitives: `output()` result emission (with large-payload temp-file spillover via `GSD_TEMP_DIR`/`ensureGsdTempDir`/`reapStaleTempFiles`), `error()` stderr emission with exit-code mapping, and the JSON-error-mode toggle (`setJsonErrorMode`/`getJsonErrorMode`, `ERROR_REASON`). Extracted from the Core module per ADR-857 rollout phase 1 (#859) so feature modules (`graphify`, `intel`, `audit`, `profile-pipeline`) depend on a small I/O seam instead of the core god-module; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/io.cjs` (generated from `src/io.cts`).
### Markdown Sectionizer
Canonical markdown-structure parsing seam (`gsd-core/bin/lib/markdown-sectionizer.cjs`, generated from `src/markdown-sectionizer.cts`). Pure functions, Node built-ins only. Exports: `stripFencedCode(content) → { text, unterminatedFence }` (CommonMark-correct state machine, CRLF-safe, signals unterminated fences); `tokenizeHeadings(content) → HeadingToken[]` (ATX headings outside fenced blocks, `{ level, text, line, offset }`); `collectSections(content, stopPredicate) → Section[]` (line-by-line section collection driven by a heading predicate); `collectSection(content, headingPredicate, { levelBounded, stripFences }) → Section | null` (single named section with level-bounded stop); `iterateBullets(sectionText) → BulletItem[]` (dash/checkbox/numbered markers with indented continuation); `extractTaggedBlocks(content, tagName) → string[]` (inner text of every `<tagName>…</tagName>` block in document order, tagName regex-escaped, caller decides fence-stripping — generalises `decisions.cts`'s bespoke extractor for T1); `replaceSection(content, section, newBody) → string` (pure character-offset splice using `Section.bodyStart`/`bodyEnd` for read-modify-write callers — eliminates T6 `state.cts`'s 7× inline `content.replace` pattern); `withSection(content, target, edit) → string` (resolve the section whose heading matches `target` — exact heading text or a `HeadingToken` predicate — and run `edit(body)` against ONLY that section's body before splicing the result back; bounded no-op when no heading matches or `edit` returns the same/non-string body; ADR-2143 §4 structurally retires the #2130/#2067/#2080 boundary-crossing class by confining any regex the caller runs to the matched section). `Section` carries `bodyStart`/`bodyEnd` offsets for `replaceSection`. ADR-1372 (epic #1372) establishes this seam and a tiered migration plan (T0–T7) to retire the 8+ ad-hoc markdown parsers and ~20 inline section-collects across `src/*.cts`. New `src/*.cts` modules must import this seam instead of hand-rolling fence strippers or heading-regex section walks (enforced by the `no-adhoc-markdown-parsing` ESLint rule landing in tier T7).
Canonical markdown-structure parsing seam (`gsd-core/bin/lib/markdown-sectionizer.cjs`, generated from `src/markdown-sectionizer.cts`). Pure functions, Node built-ins only. Exports: `stripFencedCode(content) → { text, unterminatedFence }` (CommonMark-correct state machine, CRLF-safe, signals unterminated fences); `stripInlineCode(content) → string` (per-line CommonMark inline-code-span stripper — removes `` `code` `` spans while leaving fenced blocks to `stripFencedCode`; #2365); `scanInlineCodeSpans(content) → InlineCodeSpan[]` (locates every inline code span as `{ start, end, content }`, offsets into the full string and spans never crossing a `\n`; callers that need the span CONTENT (e.g. api-coverage's package-name evidence, #2365) use this, callers that just want spans gone use `stripInlineCode`); `tokenizeHeadings(content) → HeadingToken[]` (ATX headings outside fenced blocks, `{ level, text, line, offset }`); `collectSections(content, stopPredicate) → Section[]` (line-by-line section collection driven by a heading predicate); `collectSection(content, headingPredicate, { levelBounded, stripFences }) → Section | null` (single named section with level-bounded stop); `iterateBullets(sectionText) → BulletItem[]` (dash/checkbox/numbered markers with indented continuation); `extractTaggedBlocks(content, tagName) → string[]` (inner text of every `<tagName>…</tagName>` block in document order, tagName regex-escaped, caller decides fence-stripping — generalises `decisions.cts`'s bespoke extractor for T1); `replaceSection(content, section, newBody) → string` (pure character-offset splice using `Section.bodyStart`/`bodyEnd` for read-modify-write callers — eliminates T6 `state.cts`'s 7× inline `content.replace` pattern); `withSection(content, target, edit) → string` (resolve the section whose heading matches `target` — exact heading text or a `HeadingToken` predicate — and run `edit(body)` against ONLY that section's body before splicing the result back; bounded no-op when no heading matches or `edit` returns the same/non-string body; ADR-2143 §4 structurally retires the #2130/#2067/#2080 boundary-crossing class by confining any regex the caller runs to the matched section). `Section` carries `bodyStart`/`bodyEnd` offsets for `replaceSection`. ADR-1372 (epic #1372) establishes this seam and a tiered migration plan (T0–T7) to retire the 8+ ad-hoc markdown parsers and ~20 inline section-collects across `src/*.cts`. New `src/*.cts` modules must import this seam instead of hand-rolling fence strippers or heading-regex section walks (enforced by the `no-adhoc-markdown-parsing` ESLint rule landing in tier T7).
### Markdown Table Model
Canonical GFM table parsing + schema registry seam (`gsd-core/bin/lib/markdown-table.cjs`, generated from `src/markdown-table.cts`; ADR-2143, epic #2143). Pure functions, Node built-ins only, string-in/value-out, no I/O. Exports: `parseMarkdownTable(sectionText) → Result<MarkdownTable>` (parses the first GFM pipe table found; typed `{ok:false,reason}` parse errors for no-table, missing/misaligned delimiter row, and ragged data rows — never silently drops or coerces a malformed row); `MarkdownTable` (`{columns: string[], rows: Record<string,string>[]}`, rows addressed by column name, not position); `Result<T>` (`{ok:true,value}\|{ok:false,reason}` — re-exported from the Write-Set Module, the ADR-2143 §5 single source of truth for this shape, so existing importers of `Result` from `markdown-table.cjs` are unaffected; deliberately distinct from command-routing-hub's dispatch `Result` `{ok,data\|kind}`; the two never mix); `TABLE_SCHEMAS` (`Record<string, CanonicalTableVariant[]>` — the canonical column-header variants for every GFM table GSD parses or generates: `RoadmapProgress` flat/milestone-grouped, `RequirementsTraceability`, `QuickTasks` no-status/with-status, `Security` trust-boundaries/threat-register/accepted-risks/audit-trail); `matchTableSchema(columns) → {id,label}\|null` (resolves a parsed header back to its canonical schema by exact column-name/order match). This registry is the single source of truth for ROADMAP/STATE/SECURITY canonical tables — a parity test (`tests/markdown-table.test.cjs`) asserts every variant's header appears verbatim in the template/workflow file that generates it, so the registry and templates can never silently drift (ADR-2143 §3 Generative-Fix-Divergence guard). `phase-lifecycle.cts`'s `deriveProgressFromRoadmap` is the first consumer: it locates the Progress section via the Markdown Sectionizer's `collectSection` and reads cells by column NAME through this seam, fixing #2137 (the prior position-anchored regex assumed `Status` was always the 3rd cell, which broke for the 5-column milestone-grouped `Milestone` variant).
@@ -167,7 +170,7 @@ Module owning project configuration loading: reads `.planning/config.json`, merg
Module owning model and effort resolution policy: resolves the model, runtime tier, planning granularity, reasoning effort, and fast-mode for a given agent by reading project config and resolving against the model profiles and catalog (`resolveModelInternal`, `resolveModelPolicy`, `resolveTierEntry`, `resolveModelForTier`, `resolveGranularityInternal`, `resolveEffortInternal`, `resolveFastModeInternal`, `resolveEffortForTier`, `nextEffort`, `assertValidGranularityOverride`). Depends only on leaf modules (`config-loader` for `loadConfig`, `configuration` for defaults, `model-profiles` and `model-catalog` for the static tables) — no other core dependency. Extracted from the Core module per ADR-857 rollout phase 2f (#888) — the final core.cts decomposition step; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/model-resolver.cjs` (generated from `src/model-resolver.cts`).
### Package Identity Module [Planned]
Single seam owning GSD's published-package coordinates so a repoint/rename is a one-line change instead of a tree-wide sweep. Source of truth is `package.json`; values are *derived*, not re-typed: `packageName` (`.name` → `@opengsd/get-shit-done-redux`), `binName` (`Object.keys(.bin)[0]` → `get-shit-done-redux`), `repoSlug` (parsed from `.repository.url` → `open-gsd/get-shit-done-redux`), plus derived `changelogRawUrl` and `manualInstallCommand({ scope, runtime })`. Generated `.cjs` per ADR-457 (generated-single-source); shipped under `gsd-core/bin/lib/`. Three consumer worlds: **Node** consumers `require()` it at runtime (worker, `check-latest-version.cjs`, `bin/install.js`); the **bash launcher** snippet receives the literal injected by `scripts/sync-runtime-launcher.cjs` at sync time; **prose/help** literals (`update.md`, installer help) carry a committed copy. A drift-guard lint (`scripts/lint-package-identity-drift.cjs`, sibling to `check:alias-drift`) fails CI on any raw package/repo literal outside `package.json`, the generated module, and the value-checked materialization sites — this is what keeps the seam real (`two adapters`, not one). Replaces the contradictory pair it consolidates: the runtime-broken `require('../package.json').name` in `hooks/gsd-check-update-worker.js` (#378, resolves to `undefined` post-install) and the hardcoded constant in `check-latest-version.cjs` (#2992). _Avoid_: "package name string", "the npm name" (when you mean the seam). See ADR-457 and Installer Module.
Single seam owning GSD's published-package coordinates so a repoint/rename is a one-line change instead of a tree-wide sweep. Source of truth is `package.json`; values are *derived*, not re-typed: `packageName` (`.name` → `@opengsd/gsd-core`), `binName` (`Object.keys(.bin)[0]` → `gsd-core`), `repoSlug` (parsed from `.repository.url` → `open-gsd/gsd-core`), plus derived `changelogRawUrl` and `manualInstallCommand({ scope, runtime })`. Generated `.cjs` per ADR-457 (generated-single-source); shipped under `gsd-core/bin/lib/`. Three consumer worlds: **Node** consumers `require()` it at runtime (worker, `check-latest-version.cjs`, `bin/install.js`); the **bash launcher** snippet receives the literal injected by `scripts/sync-runtime-launcher.cjs` at sync time; **prose/help** literals (`update.md`, installer help) carry a committed copy. A drift-guard lint (`scripts/lint-package-identity-drift.cjs`, sibling to `check:alias-drift`) fails CI on any raw package/repo literal outside `package.json`, the generated module, and the value-checked materialization sites — this is what keeps the seam real (`two adapters`, not one). Replaces the contradictory pair it consolidates: the runtime-broken `require('../package.json').name` in `hooks/gsd-check-update-worker.js` (#378, resolves to `undefined` post-install) and the hardcoded constant in `check-latest-version.cjs` (#2992). _Avoid_: "package name string", "the npm name" (when you mean the seam). See ADR-457 and Installer Module.
### Update Context Module [Planned]
Module owning install detection for `/gsd:update`. `resolveUpdateContext({ home, cwd, env, fs, preferredConfigDir, preferredRuntime })` is a pure, injected-fs port of update.md's former ~280-line `get_installed_version` bash; it reproduces the full precedence cascade — preferred-config-dir fast path, local-over-global probe with same-path dedup, env-var overrides (`CLAUDE_CONFIG_DIR`, `OPENCODE_CONFIG`, `KILO_CONFIG`, `XDG_CONFIG_HOME`, `CODEX_HOME`, …), and semver validation — and returns the 4-field contract `{ installedVersion, scope, runtime, gsdDir }` (scope ∈ `LOCAL`/`GLOBAL`/`UNKNOWN`). Antigravity is modelled first-class (its `.gemini/antigravity{,-ide,-cli}` dirs probe before bare `.gemini`; #3608). Exposed to the workflow as `gsd-tools update-context [--config-dir <d>] [--runtime <r>] --json`; `loadUpdateContext` wires the real fs. The workflow keeps only the execution_context path → `PREFERRED_*` derivation (the one input it alone knows). Source: `gsd-core/bin/lib/update-context.cjs`; tests: `tests/issue-498-update-context.test.cjs`. See Installer Module and Package Identity Module.
@@ -206,7 +209,7 @@ Generated central manifest projecting all co-located Capability declarations int
ADR-857 phase 3b seam that merges capability-declared config slices into the `loadConfig` return value. Implemented in `src/federated-config.cts` → `gsd-core/bin/lib/federated-config.cjs`. Exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig }) → { values, validKeys, warnings }`. Rules: central-schema keys are skipped with a `pending-migration` warning; malformed slices are skipped with a warning (never throws); valid federated keys (absent from the central schema) resolve to the user-supplied value (if type-matches) or the slice default. Object writes are guarded against prototype pollution with inline literal `__proto__`/`constructor`/`prototype` key checks. ADR-857 phase 6 made the channel live for migrated Capability keys: `config-schema.cjs` exposes `isCentralConfigKey()` for central ownership and `isValidConfigKey()` accepts central + runtime + dynamic + Capability-owned registry keys. `loadConfig` exposes `_setFederatedRegistryForTests`/`_resetFederatedRegistryForTests` seams for injecting a synthetic registry in tests.
### Capability Registry Overlay
Runtime seam (`gsd-core/bin/lib/capability-loader.cjs`, ADR-1244 D2) that composes the frozen first-party Capability Registry (`capability-registry.cjs`) with a validated installed overlay of third-party capability manifests discovered at load time. Install roots are global (`$GSD_HOME/.gsd/capabilities/<id>/capability.json`, where `GSD_HOME` defaults to `~`) and project (`<projectRoot>/.gsd/capabilities/<id>/capability.json`). Primary interface: `loadRegistry({ includeInstalled }) → registry` — when `includeInstalled` is true the overlay is merged via the canonical `buildRegistry` so all derived views (bySkill, byAgent, byLoopPoint, configKeys) cover first-party and overlay entries identically. First-party always wins: any overlay entry whose id, owned skill/agent stem, or federated config key collides with first-party, or whose id uses a reserved `gsd-`/`gsd-core-`/`anthropic-` prefix, is rejected at load time. Load-time re-gate: an overlay failing schema validation or whose `engines.gsd` semver range does not satisfy the running GSD version is skipped with a warning and never crashes the load loop. Per-hook-kind policy: a skipped capability that declared a `gate`-kind hook fails CLOSED (the loop resolver injects a blocking gate); skipped `step` or `contribution` capabilities skip open. A capability dir whose co-located ledger entry carries an in-flight `_pending` intent (a crashed/uncommitted install or upgrade, ADR-1244 Phase 4) is skipped OPEN (never activated until reconciliation commits or rolls it back). #1459 user-owned consent gate: a PROJECT-scope overlay is activated (declarative surfaces AND command dispatch) ONLY when the user-owned Capability Consent Store holds a record for `(realpath(projectRoot), id)` whose stored `contentHash` equals the bundle content hash the loader RECOMPUTES at load (`bundleContentHash(capDir)` over the whole on-disk bundle) — NOT the repo-plantable ledger integrity nor the executable-only disclosure signature — otherwise the cap is DISCOVERED-BUT-INACTIVE (a warning carrying `kind:'unconsented'`, no surfaces, empty commandRoots), so a forged/cloned in-repo project ledger or any post-consent tamper no longer activates anything; GLOBAL scope (under the user's own home) is trusted without a record, and the global-vs-project root dedup/escalation is realpath-keyed so a symlinked `GSD_HOME` aliasing the project root cannot bypass the gate (finding 1). The consent lookup is wrapped to fail CLOSED (inactive); both the per-scope ledger AND the `capability.json` manifest are read via the shared bounded `readSmallRegularFile` (a repo-planted FIFO/oversized ledger or manifest can no longer hang or OOM the loader — finding 2). The loader reuses the ledger's shared `isValidLedgerEntry` for committed-entry parity. Consumers wired to the overlay-aware registry: `config-loader.cjs`, `config-schema.cjs`, `capability-state.cjs`, `loop-resolver.cjs`.
Runtime seam (`gsd-core/bin/lib/capability-loader.cjs`, ADR-1244 D2) that composes the frozen first-party Capability Registry (`capability-registry.cjs`) with a validated installed overlay of third-party capability manifests discovered at load time. Install roots are global (`$GSD_HOME/.gsd/capabilities/<id>/capability.json`, where `GSD_HOME` defaults to `~`) and project (`<projectRoot>/.gsd/capabilities/<id>/capability.json`). Primary interface: `loadRegistry({ includeInstalled }) → registry` — when `includeInstalled` is true the overlay is merged via the canonical `buildRegistry` so all derived views (bySkill, byAgent, byLoopPoint, configKeys) cover first-party and overlay entries identically. First-party always wins: any overlay entry whose id, owned skill/agent stem, or federated config key collides with first-party, or whose id uses a reserved `gsd-`/`gsd-core-`/`anthropic-` prefix, is rejected at load time. Load-time re-gate: an overlay failing schema validation or whose `engines.gsd` semver range does not satisfy the running GSD version is skipped with a warning and never crashes the load loop. Per-hook-kind policy: a skipped capability that declared a `gate`-kind hook now fails OPEN (#2009) — the loop resolver injects no gate and instead emits a loud warning naming the load-failure reason and the exact `gsd capability remove <id>` remediation, and the loop proceeds (the loader still records `_overlay.blockedGates`; only the consequence changed from block to warn); skipped `step` or `contribution` capabilities skip open. A capability dir whose co-located ledger entry carries an in-flight `_pending` intent (a crashed/uncommitted install or upgrade, ADR-1244 Phase 4) is skipped OPEN (never activated until reconciliation commits or rolls it back). #1459 user-owned consent gate: a PROJECT-scope overlay is activated (declarative surfaces AND command dispatch) ONLY when the user-owned Capability Consent Store holds a record for `(realpath(projectRoot), id)` whose stored `contentHash` equals the bundle content hash the loader RECOMPUTES at load (`bundleContentHash(capDir)` over the whole on-disk bundle) — NOT the repo-plantable ledger integrity nor the executable-only disclosure signature — otherwise the cap is DISCOVERED-BUT-INACTIVE (a warning carrying `kind:'unconsented'`, no surfaces, empty commandRoots), so a forged/cloned in-repo project ledger or any post-consent tamper no longer activates anything; GLOBAL scope (under the user's own home) is trusted without a record, and the global-vs-project root dedup/escalation is realpath-keyed so a symlinked `GSD_HOME` aliasing the project root cannot bypass the gate (finding 1). The consent lookup is wrapped to fail CLOSED (inactive); both the per-scope ledger AND the `capability.json` manifest are read via the shared bounded `readSmallRegularFile` (a repo-planted FIFO/oversized ledger or manifest can no longer hang or OOM the loader — finding 2). The loader reuses the ledger's shared `isValidLedgerEntry` for committed-entry parity. Consumers wired to the overlay-aware registry: `config-loader.cjs`, `config-schema.cjs`, `capability-state.cjs`, `loop-resolver.cjs`.
### Community Capability Registry
Human-facing discoverability catalog (`docs/registries/capability-registry.md`, generated from `docs/registries/capabilities.json`; issue #2182) listing third-party Feature Capabilities registered by a docs PR so a solo developer can find one before installing it. Distinct from **Capability Registry** (the generated runtime manifest compiled from first-party `capability.json` declarations, ADR-894) and **Capability Registry Overlay** (the runtime seam that merges an installed third-party manifest into that generated registry at load time, ADR-1244 D2): this registry is a static document rendered by `scripts/gen-registry.cjs`, not a runtime data structure or loader. Each entry enumerates the capability's Loop Extension Points and hook kinds so a reader can judge blast radius before running `gsd capability install`, and declares its `engines.gsd` range. Inclusion is an explicit non-endorsement — a maintainer merged a link, nothing more — per `docs/registries/README.md`.
@@ -347,16 +350,25 @@ Second adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase pro
The SINGLE source of truth for "did the phase SPEC supply section X (with at least one resolved row)?" — the SPEC-section detection seam consumed by `plan-phase` Step 7.95 (the spec-less probe fallback) to decide, per section, whether to run the fallback. Replaces the ad-hoc `awk` that previously lived in the workflow body, which hard-coded the section header strings at the call site and hand-rolled markdown-table row counting — a brittleness that produced two bugs: an exact `^## Prohibitions$` anchor that missed the canonical `## Prohibitions (must-NOT)` heading, and a single-table row-counting assumption. **Suffix-tolerant header invariant:** `SECTION_HEADERS` regexes match a heading AND any parenthetical/whitespace suffix — `prohibitions` matches both `## Prohibitions` and `## Prohibitions (must-NOT)`; `edges` matches `## Edge Coverage` (and any future suffix); if spec-phase renames a heading, update HERE and the `templates/spec.md` heading together (the contract is pinned by `tests/spec-section.test.cjs`). **Supply rule:** `supplied = present AND dataRows > 0` — a present-but-empty section is NOT supplied (it triggers the fallback). **Multi-table robustness:** a blank or prose line resets the per-table state, so a section with multiple tables (or prose between them) counts every table's data rows without miscounting a second table's header row; the `|…|` line before a `|---|` separator is the table header row and is never counted. **Fail-safe:** a missing/unreadable SPEC file resolves to `present:false` / `supplied:false` (so the fallback fires) rather than throwing. Exports (locked surface): the `SpecSectionKey` type (`edges | prohibitions`), `SECTION_HEADERS` (the canonical header matchers), the `SectionStatus` shape (`{ key, present, dataRows, supplied }`), `countSectionDataRows` (pure `specText → { present, dataRows }`), and `specSectionStatus` (disk-reading wrapper). CLI: `node spec-section.cjs <specFile> <edges|prohibitions>` prints `SectionStatus` JSON — exit 0 on success (an absent file is a valid "not supplied" answer), exit 2 only on a usage error (missing args / bad key). Pure and dependency-free. Source of truth: `gsd-core/bin/lib/spec-section.cjs` (generated from `src/spec-section.cts`, gitignored per ADR-457). Tests: `tests/spec-section.test.cjs`. See Edge Probe Module, Prohibition Probe Module, and `references/specless-probe-fallback.md`.
### MVP Mode
Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`.
Phase-level planning **enrichment** layered on top of the default tracer-first decomposition (see Tracer Bullet): it frames the phase goal as a User Story and, on Phase 1 of a new project, emits a Walking Skeleton. Vertical slicing itself is now the default, so MVP Mode no longer *turns it on* — it adds the user-story framing + skeleton. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`.
### User Story
Phase-goal format under MVP Mode: `As a [role], I want to [capability], so that [outcome].` Required regex shape: `/^As a .+, I want to .+, so that .+\.$/`. Used as the framing input by `gsd-planner` (emits as bolded `## Phase Goal` header in PLAN.md) and as the verification target by `gsd-verifier` (the `[outcome]` clause is the goal-backward verification anchor). Authored interactively by `/gsd-mvp-phase`, validated by SPIDR Splitting when too large.
### Walking Skeleton
Phase 1 deliverable under `--mvp` on a new project: the thinnest end-to-end stack proving every layer (framework, DB, routing, deployment) works together. Emitted as `SKELETON.md` capturing the architectural decisions subsequent vertical slices inherit. Gate fires when `phase_number == "01"` AND `prior_summaries == 0` AND `MVP_MODE=true`. Scope intentionally narrow (PRD #2826 Q2) — does not retrofit existing projects.
Phase 1 deliverable under `--mvp` on a new project — the Phase-1 whole-application special case of a Tracer Bullet: the thinnest end-to-end stack proving every layer (framework, DB, routing, deployment) works together. Emitted as `SKELETON.md` capturing the architectural decisions subsequent vertical slices inherit. Gate fires when `phase_number == "01"` AND `prior_summaries == 0` AND `MVP_MODE=true`. Scope intentionally narrow (PRD #2826 Q2) — does not retrofit existing projects.
### Vertical Slice
Single-feature task that moves one user capability from open-to-close (happy path) end-to-end. Contrast with the horizontal layer (all models, then all APIs, then all UI). The MVP Mode planning unit; SPIDR Splitting axes (Spike, Paths, Interfaces, Data, Rules) are the canonical decomposition tools when a slice is too large for one phase.
Single-feature task that moves one user capability from open-to-close (happy path) end-to-end. Contrast with the horizontal layer (all models, then all APIs, then all UI). The default planning unit under tracer-first decomposition (the leading task is a Tracer Bullet); SPIDR Splitting axes (Spike, Paths, Interfaces, Data, Rules) are the canonical decomposition tools when a slice is too large for one phase.
### Tracer Bullet
The default GSD decomposition lead: a **permanent, production-quality, minimal end-to-end slice** that wires one path through every layer a phase touches and becomes part of the skeleton of the final system — written for keeps, not thrown away. Contrast with a **prototype** (throwaway reconnaissance code, deleted once its lesson is learned): a tracer's *functionality* gaps are acceptable but its *architectural* gaps are not; stubs are allowed only where they can later be filled without an architectural change. GSD ships tracers, never prototypes — which is why `gsd-planner` LEADS every plan with a `type="tracer"` task (default; `--no-tracer` / `TRACER_MODE=false` opts back into horizontal layers) and `gsd-executor` runs an early integration feedback gate on the tracer's `<verify>` before expansion tasks (autonomous: halt-on-fail; interactive: `checkpoint:human-verify`). Origin: *The Pragmatic Programmer* "Tracer Bullets" (#1945); the Walking Skeleton is the Phase-1 whole-application special case. See Vertical Slice, MVP Mode, Walking Skeleton, Precondition.
### Precondition
The front-of-task side of the GSD plan contract (issue #1949, *The Pragmatic Programmer* Topic 23 — Design by Contract). An optional `<precondition>` element on `<task>` stating, in a single line of runnable/checkable prose, what must already be true for the task to begin safely — e.g. "`OPENAI_API_KEY` is set", "`dist/schema.json` from Phase 02 exists", "server responds to GET /health". `gsd-executor` evaluates it before any other task work, using **read-only checks only** (file existence, env var presence with no value output, idempotent `GET /health`-style pings — no writes, no network POSTs, no secret emission; if a side-effecting check seems required, the executor halts and surfaces a checkpoint rather than running it): met-or-absent is a no-op (back-compat for every existing plan); unmet returns a `checkpoint:human-verify` with no partial commit, and is NEVER auto-approved under `AUTO_CFG=true` (a missing prerequisite is a fact the executor cannot establish on its own, not a verification step). `gsd-planner` emits `<precondition>` in exactly three cases — `user_setup` consumption, prior-phase artifact dependency, or env-var/runtime-config dependency — when the assumption is not already guaranteed by `depends_on` ordering. Closes the contract triad whose other two sides are postconditions (`<verify>`/`<done>`/`<acceptance_criteria>`) and invariants (`must_haves.truths`). The structural validator (`cmdVerifyPlanStructure`) does not reject unknown optional tags, so adding `<precondition>` passes plan-structure validation unchanged. Canonical schema reference: `docs/reference/plan-md.md` → Preconditions; emission rules + anti-patterns: `gsd-core/references/planner-preconditions.md`. The architectural-end companion is the Tracer Bullet (#1945); together they close both ends of the "outrunning your headlights" failure mode. See Tracer Bullet.
### Reversibility Rating
Classification of a planning decision by what undoing it would cost later (issue #1951, *The Pragmatic Programmer* Topic 15 — "Reversibility"; Bezos's one-way/two-way door framing). Three closed values: `reversible` (undo is local and cheap — one file, one function, an implementation swapped behind a stable interface), `costly` (undo touches many call sites or needs a coordinated change — shared interface shape, cross-module contract, dependency major bump), `one-way` (undo requires a data migration, breaks a published contract, or is impossible — on-disk/wire format, public API shape, external-service lock-in). Surfaced two places: `discuss-phase` records it inline on `<decisions>` entries in phase CONTEXT.md as `— **Reversibility:** <rating> — <rationale>` (optional; an unrated decision is treated as `reversible`), and `gsd-planner` carries it onto the implementing task as the optional `<reversibility rating="…">` element. Planner behavior is rating-dependent: `one-way` inserts a `checkpoint:decision` BEFORE the dependent task (and forces `autonomous: false`), `costly` is flagged in the plan but never blocks, `reversible` flows normally. Reuses the existing `checkpoint:decision` mechanism — no new checkpoint machinery. Override `--no-reversibility-gates` (`REVERSIBILITY_GATES=false`, `/gsd:plan-phase`) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings, so the signal survives a run that chose not to stop for it. The structural validator (`cmdVerifyPlanStructure`) does not reject unknown optional tags, so `<reversibility>` passes plan-structure validation unchanged. Canonical taxonomy owner: `gsd-core/references/planner-reversibility.md`; schema reference: `docs/reference/plan-md.md` → Reversibility; the Reversibility Test thinking model (`references/thinking-models-planning.md` #4) is the reasoning step that produces the rating and consumes this taxonomy rather than defining a second one. Primary anti-pattern: rating everything `one-way` (checkpoint fatigue) — default to `reversible` when unsure, and prefer *removing* irreversibility (writer seam, versioned contract, vendor adapter) over gating it. The decision-risk companion to the Precondition (#1949), which guards implementation assumptions. See Precondition, Tracer Bullet.
### Behavior-Adding Task
Predicate over a PLAN.md task: `tdd="true"` frontmatter AND `<behavior>` block names a user-visible outcome AND `<files>` includes at least one non-`*.md` / non-`*.json` / non-`*.test.*` source file. Pure doc/config/test-only tasks are exempt. The MVP+TDD Gate (in `references/execute-mvp-tdd.md`) only halts execution on this predicate; the gsd-executor agent applies all three checks at runtime. Currently a prose-only specification — no shared utility.
@@ -388,6 +400,9 @@ A legal deferred state of an Execute step (`external_job_waiting`): the executor
### External-job Capability
The producer half of the async external-job contract (#1164, part of #1105). Default-off Capability (`capabilities/external-job/capability.json`) that *writes* `.planning/async-jobs/<job>.json` manifests — the only thing that does; core never writes them. SLURM is the first backend (`sbatch --parsable` submit, `squeue` poll with `sacct` fallback, terminal-state mapping); the design stays scheduler-pluggable via the `backend` field (LSF/PBS/Kubernetes batch forward-declared, not built). Contributions inject at `execute:wave:post` into the executor (classify runtime budget → externalize `long_compute`, commit manifest + handoff, return `external_job_waiting`, defer SUMMARY.md) and at `plan:post` into the planner (emit `<runtime_budget>` quick|medium|unknown|long_compute per task). Activation key `external_job.enabled` (default `false`); sibling keys `external_job.backend`, `external_job.artifact_dir` (default `Artifacts/jobs`, per-job dirs — no fixed log paths, no hardcoded cluster/partition/account), `external_job.submit_timeout_ms` / `external_job.poll_timeout_ms` (bounded subprocesses per CLAUDE.md). Pure producer logic — SLURM state→manifest-status mapping, manifest build/validate, `sbatch`/`squeue`/`sacct` parsers, and the fail-closed manifest writer (refuses a second non-terminal job for a `plan_id` already in flight; refuses to clobber a malformed manifest) — lives in `gsd-core/src/external-job.cts` (generated to `gsd-core/bin/lib/external-job.cjs`); the operator CLI surface is `scripts/slurm-adapter.cjs` (`submit`/`poll`/`show`). Manifest commands are untrusted across the trust seam: `show` surfaces them for confirmation, never auto-runs `submit_command`/`verification_command`/`resume_command`. Test seam: `tests/external-job.test.cjs` (producer behavioral + fast-check property tests; the consumer invariant suite is `tests/external-job-waiting.test.cjs`).
### Broken Windows Ledger
The enforced cross-phase defect register operationalizing GSD's no-defer discipline as a tracked artifact (#1950). Markdown file at `.planning/WINDOWS.md` (project-level, cross-phase) with YAML frontmatter carrying scalar counts (`schema_version`, `open_count`, `waived_count`, `fixed_count`, `total_count`, `last_updated`) for the FAST path the gate reads via jq without parsing JSON, plus a JSON code block as the AUTHORITATIVE entries source; the two cross-check and fail closed on drift. Each entry: `{ id, kind, phase, file, line, description, status, reason, recorded_at, resolved_at }`; kinds are closed (`stub | todo | fixme | skipped-test | lint-warning | unmet-truth | unrun-verify | deviation`); statuses are closed (`open | waived | fixed`). The `broken-windows` Capability (`capabilities/broken-windows/capability.json`) registers one `ship:pre` gate with predicate `artifact-frontmatter-equals WINDOWS.md open_count == 0`; federated config key `workflow.windows_enforce` (default `false` — opt-in enforcement, tracking-only by default so a project can adopt the ledger before turning the gate on). Population is best-effort and never blocks execution: `agents/gsd-executor.md` appends stubs/skipped-tests/unrun-verifies via `gsd_run windows append` after writing SUMMARY.md. Source of truth: `src/broken-windows.cts` → `gsd-core/bin/lib/broken-windows.cjs` (pure `parseLedger`/`renderLedger`/`appendWindow`/`markWaived`/`markFixed` + I/O `cmdWindowsStatus`/`Append`/`Waive`/`MarkFixed`); CLI surface `gsd-tools windows status|append|waive|fixed`. Ship gate enforcement is a `capId == "broken-windows"` branch in `gsd-core/workflows/ship.md` preflight (sibling to the `security` branch); it reads `gsd_run windows status --raw` and fails closed on a non-zero/non-numeric `open_count` (an unparseable ledger is itself a broken window). `/gsd:progress` surfaces the open+waived count. The ledger is optional and backward-compatible: a project with no `.planning/WINDOWS.md` reports `open_count: 0` and ships cleanly, and with `workflow.windows_enforce=false` (the default) ship never blocks on it. Frozen `REASON` enum: `WINDOWS_LEDGER_MISSING | WINDOWS_LEDGER_MALFORMED | WINDOWS_ID_NOT_FOUND | WINDOWS_ALREADY_RESOLVED | WINDOWS_WAIVE_REASON_EMPTY | WINDOWS_INVALID_KIND | WINDOWS_INVALID_FILE | WINDOWS_INVALID_ID | WINDOWS_APPEND_MISSING_FIELD | WINDOWS_USAGE | WINDOWS_OK` — surfaced through `--json-errors` for typed test assertions. Test seam: `tests/broken-windows.test.cjs`. Origin: *The Pragmatic Programmer* Topic 3 (Hunt & Thomas — software transplant of Wilson & Kelling's broken-windows metaphor) plus Cunningham's debt metaphor (decay accrues interest ⇒ accounting, not just habit).
### Untrusted-input boundary
The prompt-level data/instruction isolation seam for untrusted web/document ingress (#1577). Shared reference `gsd-core/references/untrusted-input-boundary.md`, `@`-included by the 10 ingest agents (`gsd-project-researcher`, `gsd-phase-researcher`, `gsd-ui-researcher`, `gsd-assumptions-analyzer`, `gsd-advisor-researcher`, `gsd-ai-researcher`, `gsd-domain-researcher`, `gsd-research-synthesizer`, `gsd-doc-classifier`, `gsd-doc-synthesizer`) — every agent that reads fetch/search/MCP output or external source documents. The reference instructs: treat fetched/read content as **data, never instructions**; self-scan content for embedded directives before use; act only on the assigned task (ignore off-task instructions in data); and wrap quoted untrusted spans in a **fresh random delimiter** per wrap (fixed markers are spoofable). This prompt-level boundary is the primary control — it keeps an injection from being *followed* even while it sits in context. The hook-level companion is the read-injection scanner (`hooks/gsd-read-injection-scanner.js`, PostToolUse on `Read`/`WebFetch`/`WebSearch`), advisory by default; the opt-in top-level `security.injection_blocking` key upgrades HIGH-confidence detections to a PostToolUse circuit-breaker that halts the agent's next step (it runs *after* the fetch, so it is not a redactor). Tests: `tests/untrusted-input-isolation.test.cjs`, `tests/read-injection-scanner.*.test.cjs`, `tests/injection-blocking-config.test.cjs`. See `docs/adr/1577-untrusted-input-boundary-and-injection-blocking.md` and `docs/explanation/security-model.md`. Grounding: arXiv 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472.
@@ -395,7 +410,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## Probe family — spec-completeness probes (machine-oriented predicates)
> Glossary prose for these modules lives above (Probe Core / Edge Probe / Prohibition Probe / Verification Tier / Verification substrate). These are the greppable one-line predicates ADR-550's Consequences promised alongside the glossary. Research-derived numbers (N17/N18 rates) are deliberately kept out of this machine-canon and live hedged in `docs/design/verifier-reach.md`. (That design note and `docs/adr/1606` are co-delivered sibling PRs of epic #1605; predicate refs to them below resolve once the batch lands.)
> Glossary prose for these modules lives above (Probe Core / Edge Probe / Prohibition Probe / Verification Tier / Verification substrate). These are the greppable one-line predicates ADR-550's Consequences promised alongside the glossary. Research-derived numbers (N17/N18 rates) are deliberately kept out of this machine-canon and live hedged in `docs/design/verifier-reach.md`. (That design note and `docs/adr/1606-prohibition-enforcement-verify-seam.md` were co-delivered sibling PRs of epic #1605, now closed and landed; the predicate refs to them below are live.)
`PROBE.principle=verifier-reach-equals-spec-reach (a goal-backward verifier only checks assertions that exist; probes make omitted assertions exist before code) — ADR-857 verification-substrate boundary; docs/design/verifier-reach.md`
`PROBE.family=edge-probe(shape-axis)+prohibition-probe(must-NOT-axis)+ui-consideration-probe(UI-state-axis), shared probe-core, run as spec-phase/ui-phase soft gates (ADR-550 D7; #1867)`
@@ -423,8 +438,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## Test rules and lint
`RULESET.TESTS.no-source-grep=scripts/lint-no-source-grep.cjs rejects readFileSync source + .includes()/.match()/.startsWith() on the bound var; CI hard-fail`
`RULESET.TESTS.no-source-grep.stdout-extension=also flags assert.match/doesNotMatch on .stdout/.stderr — emit JSON from SUT, parse, assert on typed fields`
`RULESET.TESTS.no-source-grep=local/no-source-grep ESLint AST rule (eslint-rules/no-source-grep.cjs) rejects readFileSync of a source .cjs/.js/.ts path bound to a var later hit with .includes()/.match()/.startsWith()/.endsWith()/.indexOf()/.search(); error in tests/**/*.test.cjs, warn in gsd-core/bin/**/*.cjs + scripts/**/*.cjs (ADR 452 retired the old regex script, removed for good in #632)`
`RULESET.TESTS.no-source-grep.exemption=// allow-test-rule: <runtime-contract-is-the-product> with one-line justification; reserved for tests where the file content IS the product surface (STATE.md, config.toml, hooks.json, agent .md). Migration to typed-IR parser tracked in #2974.`
`RULESET.TESTS.no-source-grep.tmp-file-traps=reading tmp files written by the SUT in tests still trips lint; round-trip through CLI (e.g. frontmatter get) instead of readFileSync+.includes()`
@@ -438,12 +452,12 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`RULESET.TESTS.boundary-coverage.anti-pattern=test suites that pair budget:1_000_000 (trivially fits) with budget:1 (trivially overflows) and skip the boundary region; failure mode that shipped PR #3708 UNNEEDED_TRIM + FALSE_HARDFAIL regressions (commit 2df566ed, fixed bde1ae8f)`
`LEARNING.prompt-budget.boundary-gap=PR #3708 commit 2df566ed reserved NOTE_RESERVE_TOKENS in pressure-threshold AND in minSet pre-check; both buggy paths only fire when baseTokens ∈ (effectiveBudget - NOTE_RESERVE_TOKENS, effectiveBudget]; original test suite used budgets far from that band so neither path was exercised; fix bde1ae8f confines NOTE_RESERVE accounting to post-trim assembly path only; future budget/limit code MUST add boundary fixtures per RULESET.TESTS.boundary-coverage.fixtures`
`RULESET.TESTS.no-timing-assertion=do not assert on wall-clock elapsed time (Date.now() delta, performance.now(), process.hrtime() comparison); such assertions test the host machine not the SUT and flake on loaded CI runners; enforcement: local/no-elapsed-assertion ESLint rule (warn → error after #453); canonical replacement: clock-seam pattern with node:test mock.timers`
`RULESET.TESTS.no-timing-assertion=do not assert on wall-clock elapsed time (Date.now() delta, performance.now(), process.hrtime() comparison); such assertions test the host machine not the SUT and flake on loaded CI runners; enforcement: local/no-elapsed-assertion ESLint rule, currently warn (promotion to error tracked under open epic #1885, not #453 which already merged without completing it); canonical replacement: clock-seam pattern with node:test mock.timers`
`RULESET.TESTS.clock-seam=concurrency logic must accept an optional {clock=Date} parameter; tests control time via t.mock.timers.enable(['Date']) + t.mock.timers.setTime(0) + t.mock.timers.tick(N); real OS scheduler races are not a permitted test pattern after ADR 456 (2026-05-28); real-race tests are deleted once deterministic seam tests cover the same logical path; clock.cjs realClock adds nowIso() (→ new Date(this.now()).toISOString()) and today() (→ nowIso().split('T')[0]) so all date-stamping in state.cjs routes through the seam; subprocess time-pin adapter: set GSD_TEST_MODE=1 + GSD_NOW_MS=<epoch-ms> in runGsdTools env to pin the date written by the SUT without touching real wall-clock (issue #474)`
`RULESET.TESTS.property-based-testing=modules implementing parsing / transformation / budget-limit / bijective contracts must include at least one fast-check (fc) property test asserting a domain invariant; invariant categories: round-trip, monotonicity, boundary-containment, idempotency; property tests live in *.test.cjs alongside unit tests; CI signal: Stryker mutation score below 80% blocks merge`
`RULESET.TESTS.mutation-score=Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification`
`RULESET.TESTS.delete-bad-tests=pass-always / vacuous-truth / source-grep / elapsed-time / real-race / permanent-allow-test-rule tests are DELETED and replaced with compliant tests in the same PR; not skipped, not commented out, not permanently exempted; replacement must cover the same logical path via typed-surface assertion or clock-seam pattern`
`RULESET.TESTS.eslint-harness=ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at scripts/eslint-rules/; replaces scripts/lint-*.cjs regex scanners; three test-rigor rules (local/no-source-grep, local/no-magic-sleep-in-tests, local/no-elapsed-assertion) ship at warn, promoted to error after #453 cleanup sweep merges`
`RULESET.TESTS.eslint-harness=ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at eslint-rules/ (repo root, NOT scripts/eslint-rules/); replaces scripts/lint-*.cjs regex scanners (fully removed in #632); of the three test-rigor rules, local/no-source-grep and local/no-magic-sleep-in-tests are already promoted to error in tests/**/*.test.cjs scope (post-cleanup), local/no-elapsed-assertion remains at warn pending open epic #1885 (its dedicated ratchet issue #453 already merged without completing this promotion; follow-up #1888 was closed not-planned and folded into #1885)`
`RULESET.AUDIT.search-source-not-generated=verify an invariant/validation EXISTS by searching the AUTHORED source (src/*.cts OR the scripts/gen-*.cjs generator), never the generated bin/lib/*.cjs (gitignored, ADR-457); gen-time checks live in gen-*.cjs not the .cts it consumes → search BOTH before declaring absent; read generated .cjs only for output drift. Repro: grep src/*.cts for VALID_CONVERTER_NAMES → false "5e ConverterName unenforced"; actually enforced in gen-capability-registry.cjs. cf RULESET.TESTS.no-source-grep`
@@ -452,14 +466,14 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`RULESET.AGENT_SIZE_BUDGET=agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = per-file baseline (PRIMARY anti-creep: tests/agent-size-baseline.json pins each agents/gsd-*.md exact byte size) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). One 'npm run size:baseline' regenerates BOTH workflow and agent baselines via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter. A grown agent fails the baseline guard — regenerate + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes`
`RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; <step name="..."> XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name`
`RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/docs-update.test.cjs (folds former \`bug-3135-capture-backlog-workflow\`, consolidation epic #1969); INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill`
`RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets`
`RULESET.WORKFLOW_EXECUTE_END_TO_END=standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets — convention verified live across ~20 commands/gsd/*.md files; no ADR currently documents this specific phrasing rule (ADR-0002 covers the adjacent but distinct command-contract/@-ref-resolution seam, not this convention)`
`RULESET.WORKFLOW.COVERAGE-METADATA=#1602 SUMMARY frontmatter `coverage:` block (list of {id,description,requirement?,verification:[{kind∈unit|integration|e2e|automated_ui|manual_procedural|other, ref, status∈pass|fail|unknown}],human_judgment:bool,rationale?}) is the per-deliverable RTM consumed DETERMINISTICALLY by verify-work extract_tests via `gsd-tools uat classify-coverage --summary <f>` (src/coverage.cts → bin/lib/coverage.cjs). AUTHORING: execute-plan create_summary populates it from task <verify> results; every deliverable MUST be classified; fail-safe default = human_judgment:true + rationale. CLASSIFY CONTRACT: auto-pass (skip human) ONLY when human_judgment===false (strict boolean) AND verification non-empty AND every status==='pass' AND zero validation errors — else PRESENT to human. mode:legacy (no block) ⇒ byte-identical prose `## Accomplishments` fall-through; `coverage: []` ⇒ mode:coverage, zero entries (single-confirmation). Frozen IR: MODE/PRESENT_REASON/ERROR_CODE enums locked by tests/coverage-metadata-parser.test.cjs. extractFrontmatter CANNOT parse it (scalars-only `-` items) → dedicated parser, sibling of parseMustHavesBlock. Asymmetry by design: false-negative=redundant prompt (status quo); false-positive=shipped bug UAT existed to catch`
`RULESET.ALLOWED-TOOLS-FRONTMATTER=command's allowed-tools must cover every tool the workflow calls (including Write for file creation); thin-wrapper pattern makes this easy to miss`
`RULESET.ARGUMENTS-SANITIZE=any workflow step constructing .planning/.../{SLUG}.md path from user input ($ARGUMENTS, parsed remainder) must sanitize inline ([a-z0-9-] only, reject ..//\\, max-length) — "(already sanitized)" must trace back to explicit guard; RESUME/fallback modes need own guards`
`RULESET.SHARED-HELPERS-LINT-VS-TEST=when a lint script and test suite both implement same constant (CANONICAL_TOOLS) or parser (parseFrontmatter, executionContextRefs), extract to scripts/*-helpers.cjs required by both — silent divergence otherwise`
`RULESET.ADR-HEADER=every docs/adr/NNNN-*.md must open with - **Status:** Accepted|Proposed|Deprecated + - **Date:** YYYY-MM-DD immediately after title`
`RULESET.ADR-HEADER=every docs/adr/NNNN-*.md must open with - **Status:** Accepted|Proposed|Superseded (by [ADR-NNNN](file.md))|Legacy + - **Date:** YYYY-MM-DD immediately after title`
`RULESET.MANIFEST-CANONICAL-KEY=docs/INVENTORY-MANIFEST.json has a single top-level key: families; ALL SIX families.* arrays (agents/commands/workflows/references/cli_modules/hooks) are canonical, consumed by test suites — tests/inventory-manifest-sync.test.cjs reads all six, edit-phase/enh-2380/enh-2430 tests read commands+workflows; the old generated date field and the stale top-level workflows key are both gone; regen via node scripts/gen-inventory-manifest.cjs --write`
`RULESET.PR-SCOPE.one-concern-per-pr=split unrelated changes into separate PRs; cherry-pick doc changes to dedicated docs/ branch immediately, then force-push original to remove the commit`
@@ -473,7 +487,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## CodeRabbit + repo-process guards (machine-oriented predicates)
`RULESET.CONTRIB.GATE.ORDER=issue-first -> approval-label -> code -> PR-link -> changeset/no-changelog`
`RULESET.CONTRIB.CLASSIFY.fix=requires confirmed/confirmed-bug before implementation`
`RULESET.CONTRIB.CLASSIFY.fix=requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)`
`RULESET.CONTRIB.CLASSIFY.enhancement=requires approved-enhancement before implementation`
`RULESET.CONTRIB.CLASSIFY.feature=requires approved-feature before implementation`
@@ -505,7 +519,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`WORKSTREAM.INVARIANT.slug-contract=all .planning/workstreams/<name> must be addressable by set/get/status/complete`
`WORKSTREAM.REGRESSION.test-anchor=tests/workstream.test.cjs::normalizes --migrate-name to a valid workstream slug`
`ARCH.SKILL.improve-codebase.next-candidates=[Workstream Name Policy Module, Workstream Progress Projection Module, Active Workstream Pointer Store Module]`
`ARCH.SKILL.improve-codebase.next-candidates=[Workstream Progress Projection Module]`
`WORKTREE.SEAM.test-policy=cover all decision branches in policy module before changing prune behavior`
`WORKTREE.SEAM.test-anchors=[resolveWorktreeContext:has_local_planning|linked_worktree|not_git_repo|main_worktree, planWorktreePrune:git_list_failed|worktrees_present|no_worktrees|parser_throw_fallback, executeWorktreePrunePlan:missing_plan|skip_passthrough|unsupported_action|metadata_prune_only]`
@@ -525,7 +539,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## Release notes standard
`RELEASE-NOTES.SCOPE=GitHub Releases body for tags vX.Y.Z, vX.Y.Z-rcN; not CHANGELOG.md (changeset workflow owns that)`
`RELEASE-NOTES.SCOPE=GitHub Releases body for tags vX.Y.Z, vX.Y.Z-rc.N; not CHANGELOG.md (changeset workflow owns that)`
`RELEASE-NOTES.DEFAULT-STATE=auto-generated body is "What's Changed" PR list + Full Changelog link; treat as draft, not final`
`RELEASE-NOTES.GATE.hotfix=manual edit required; auto-generated body for vX.Y.{Z>0} is "Full Changelog only" and must be replaced with structured body`
`RELEASE-NOTES.GATE.rc=manual edit recommended; auto-generated PR list is acceptable for early RCs but final RC before vX.Y.0 should match standard`
@@ -547,15 +561,15 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`RELEASE-NOTES.WORKFLOW.edit=gh release edit <tag> --notes-file <path>`
`RELEASE-NOTES.WORKFLOW.view=gh release view <tag> --json body --jq .body`
`RELEASE-NOTES.WORKFLOW.token=must use .envrc GITHUB_TOKEN per project CLAUDE.md; never ambient gh auth`
`RELEASE-NOTES.WORKFLOW.token=must use .envrc GITHUB_TOKEN per RULESET.GH.AUTH.DEFAULT (this doc); never ambient gh auth`
`RELEASE-NOTES.WORKFLOW.idempotency=gh release edit overwrites body wholesale; safe to re-run after refining`
`RELEASE-NOTES.ANTI-PATTERN=raw "What's Changed" PR list as final body for hotfix or feature release; "Full Changelog only" body for tagged release with >0 user-facing fixes`
`RELEASE-NOTES.ANTI-PATTERN.implementation-first=do not lead bullet with file path or function name; lead with symptom/user-visible behavior`
`RELEASE-NOTES.ANTI-PATTERN.risk-commentary=do not include "may break", "be careful", "test thoroughly" - per global CLAUDE.md no-risk-commentary rule`
`RELEASE-NOTES.ANTI-PATTERN.risk-commentary=do not include "may break", "be careful", "test thoroughly" - release notes state what changed, not hedges about what might go wrong`
`RELEASE-NOTES.EXAMPLE.hotfix=v1.41.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.41.1) - 14 fixes grouped by 6 subgroups`
`RELEASE-NOTES.EXAMPLE.rc=v1.42.0-rc1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.42.0-rc1) - intro + Added/Changed/Fixed/Documentation taxonomy`
`RELEASE-NOTES.EXAMPLE.rc=v1.7.0-rc.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.7.0-rc.1) - intro + Added/Changed/Fixed/Documentation taxonomy`
`RELEASE-NOTES.EXAMPLE.minor-auto-acceptable=v1.41.0 - kept auto-generated body; many small fixes with clean conventional-commit titles`
`RELEASE-NOTES.TEMPLATE.hotfix=## Fixed\n\n### <subgroup>\n- **<bold change>** — <explanation>. (#<PR>)\n\n---\n\nInstall/upgrade: \`npx @opengsd/gsd-core@latest\`\n\n**Full Changelog**: <compare-url>`
@@ -574,14 +588,14 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`META.RULE.brief-no-paraphrase=writing "k040 — never leave changelog box unchecked" caused 5 of 8 agents to edit CHANGELOG.md in violation of CONTRIBUTING.md L110`
`PRED.k320.signal=changelog-direct-edit-forbidden`
`PRED.k320.canonical-source=CONTRIBUTING.md L110-123`
`PRED.k320.canonical-source=CONTRIBUTING.md L193-211`
`PRED.k320.rule=do not edit CHANGELOG.md in feature/fix/enhancement PRs`
`PRED.k320.cure=drop .changeset/<adj>-<noun>-<noun>.md fragment ONLY`
`PRED.k320.tool=npm run changeset -- --type <T> --pr <NNN> --body "..."`
`PRED.k320.types=Added|Changed|Deprecated|Removed|Fixed|Security`
`PRED.k320.opt-out-label=no-changelog`
`PRED.k320.ci-enforcement=scripts/changeset/lint.cjs`
`PRED.k320.ci-paths-monitored=bin/ gsd-core/ agents/ commands/ docs/ hooks/ tests/ scripts/`
`PRED.k320.ci-paths-monitored=bin/ gsd-core/ src/ agents/ commands/ hooks/ sdk/src/ sdk/prompts/`
`PRED.k320.recovery=open Removed-typed cleanup PR deleting only the redundant row`
`PRED.k320.evidence=PR #3302 merge-conflict against #3308 CHANGELOG.md row 2026-05-09`
@@ -631,12 +645,12 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`PRED.k327.cooldown-throttled=k322`
`PRED.k328.signal=pr-template-typed-heading-required`
`PRED.k328.canonical-source=CONTRIBUTING.md L101`
`PRED.k328.canonical-source=CONTRIBUTING.md L48,L64,L81 (template links) + .github/PULL_REQUEST_TEMPLATE/{fix,enhancement,feature}.md L1 (heading text)`
`PRED.k328.k100-restatement=heading must match issue class: bug→## Fix PR, enhancement→## Enhancement PR, feature→## Feature PR`
`PRED.k328.audit-list=[heading-matches-class, closing-keyword-present, changeset-fragment-or-no-changelog-label]`
`PRED.k329.signal=changeset-fragment-canonical-shape`
`PRED.k329.canonical-source=CONTRIBUTING.md L112-117 + .changeset/README.md`
`PRED.k329.canonical-source=CONTRIBUTING.md L196-202 + .changeset/README.md`
`PRED.k329.filename=.changeset/<adj>-<noun>-<noun>.md`
`PRED.k329.frontmatter=---\\ntype: <Added|Changed|Deprecated|Removed|Fixed|Security>\\npr: <NNN>\\n---`
`PRED.k329.body=**<Bold user-visible change>** — <symptom-led explanation>. (#<NNN>)`
@@ -665,7 +679,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## Triage and merge-wave lessons
`WAVE.LESSON.changelog-policy-violation-multiplier=brief contradicting CONTRIBUTING.md L110 produced violations on 5 of 8 PRs (#3300, #3302, #3304, #3305, #3308); k326 + k320 capture`
`WAVE.LESSON.changelog-policy-violation-multiplier=brief contradicting CONTRIBUTING.md's changelog-fragment policy ("CHANGELOG Entries — Drop a Fragment" section) produced violations on 5 of 8 PRs (#3300, #3302, #3304, #3305, #3308); k326 + k320 capture`
`WAVE.LESSON.cr-throttle-burst-correlation=8 PRs in <15min triggered k322 sustained-throttle on multiple PRs (#3306 worst case)`
`WAVE.LESSON.sibling-audit-overlap=k015-family parallel dispatch on #3297 + #3298 produced k323 add-backlog.md cross-PR overlap`
`WAVE.LESSON.agent-narrative-unreliable=k095/k324 confirmed at scale: 5 of 8 agents terminated mid-monitor with stale claims requiring direct verification`
@@ -686,13 +700,13 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`DEFECT.STATE-TRAMPLE.symptom=state-mutation paths overwrite curated values when body-derived computation is narrower than what's stored in frontmatter`
`DEFECT.STATE-TRAMPLE.examples=#3242 (Last Activity overwrote progress.completed_plans), #3257 (nested plans/ files uncounted), #3261 (buildStateFrontmatter), #3265 (canonical fields), #3286 (record-metric/add-decision sections)`
`DEFECT.STATE-TRAMPLE.detect=any state writer that calls buildStateFrontmatter without preserving existing progress.* keys; any mutation surface that does not honor shouldPreserveExistingProgress`
`DEFECT.STATE-TRAMPLE.fix-forward=route through state-document.cjs/.ts shouldPreserveExistingProgress + normalizeProgressNumbers (extracted in #3316 SDK-first seams)`
`DEFECT.STATE-TRAMPLE.fix-forward=route through state-document.cjs/.ts shouldPreserveExistingProgress + normalizeProgressNumbers (extracted in #3316; the sdk/ tree that PR originally targeted has since been fully retired per ADR-0174 — these functions now live solely in src/state-document.cts)`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.symptom=multiple workflow files independently construct .planning/phases/{NN}-{slug} paths; project_code prefix or slug normalization missing in some surfaces`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.examples=#3287 (init.phase-op + init.plan-phase first-touch), #3306/PRED.k015 (plan-milestone-gaps + import + add-backlog), #3297/#3298 (sibling reports)`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.detect=grep mkdir/touch/path.join with {NN}-{slug} or padded_phase + phase_slug; if not consuming expected_phase_dir from init.* JSON it is drifting`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.fix-forward=consume expected_phase_dir from init.phase-op / init.plan-phase output; never re-construct from padded_phase + slug in workflow steps`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.anchor=tests/bug-3298-phase-dir-prefix-drift-in-workflows.test.cjs (broad regression across workflow surfaces)`
`DEFECT.PHASE-DIR-PREFIX-DRIFT.anchor=tests/phase.test.cjs (expected_phase_dir assertions; consolidated from tests/bug-3298-phase-dir-prefix-drift-in-workflows.test.cjs into the Phase Lifecycle Module test suite in #3741)`
`DEFECT.STACKED-PR-AUTO-RETARGET.symptom=PR #N is stacked on branch B; branch B merges to main and is deleted; GitHub does not reliably auto-retarget #N to main; PR shows DIRTY/CONFLICTING with phantom conflicts`
`DEFECT.STACKED-PR-AUTO-RETARGET.examples=#3311 base fix/3255-add-json-errors-mode-gsd-tools deleted after #3304 merged`
@@ -720,12 +734,12 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`DEFECT.INVENTORY-DRIFT.fix-forward=update INVENTORY.md row entry; run node scripts/gen-inventory-manifest.cjs --write to regen INVENTORY-MANIFEST.json (all six families.* arrays are canonical — see RULESET.MANIFEST-CANONICAL-KEY)`
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH.symptom=adding to agents/gsd-planner.md (or other large agent files) exceeds the 45K char extraction-evidence threshold`
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH.state=gsd-planner.md is already 49,121 chars on main (over 45K); test fails on main; net-new content makes it strictly worse`
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH.state=gsd-planner.md is 49,125 chars on main, just under the test's actual PLANNER_EXTRACTED_LIMIT of 48K (49,152 chars — the test's own title still says "45K" but the enforced constant was raised in #2341); the test currently passes, but any further net-new content risks pushing it over`
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH.detect=tests/planner-decomposition.test.cjs ("planner is under 45K chars (proves mode sections were extracted)") and tests/reachability-check.test.cjs ("file stays under 50000 char limit")`
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH.fix-forward=mirror MVP mode pattern — extract full rules to gsd-core/references/planner-<mode>.md, leave a slim Detection section in the agent file with @-reference to the new file`
`DEFECT.CHANGESET-PR-FIELD-DRIFT.symptom=.changeset/*.md frontmatter pr: value is the issue number, a guess made before PR opened, or a stale stacked-PR number`
`DEFECT.CHANGESET-PR-FIELD-DRIFT.examples=#3316 (pr:3312 was the issue), #3325 (pr:3319 was a guess); already covered in CONTEXT.md L94 + L186 but recurs every cycle`
`DEFECT.CHANGESET-PR-FIELD-DRIFT.examples=#3316 (pr:3312 was the issue), #3325 (pr:3319 was a guess); recurs every cycle`
`DEFECT.CHANGESET-PR-FIELD-DRIFT.detect=changeset pr: value mismatches the actual PR number returned by gh api POST /pulls`
`DEFECT.CHANGESET-PR-FIELD-DRIFT.fix-forward=author changeset with placeholder pr:0; immediately after gh api POST /pulls returns the number, edit changeset and amend or follow-up commit; never guess`
@@ -764,8 +778,8 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`DEFECT.DEFAULT-FLIP-DOCUMENTATION.detect=any PR that changes a default value in CONFIG_DEFAULTS or buildNewProjectConfig; check that PR body Breaking Changes section explicitly covers (a) when the new default takes effect, (b) opt-back-in command, (c) effect on in-flight artifacts`
`DEFECT.DEFAULT-FLIP-DOCUMENTATION.fix-forward=template — "new default takes effect when .planning/config.json is rewritten (config-set, fresh project, regenerated config); existing artifacts continue to work; opt-back-in: gsd config-set <key> <old-value>"`
`DEFECT.SOURCE-GREP-IN-NEW-TESTS.symptom=new test file uses readFileSync + .includes() / .match() against source code (CONTEXT.md L82); contradicts the test rule lint script`
`DEFECT.SOURCE-GREP-IN-NEW-TESTS.detect=scripts/lint-no-source-grep.cjs (npm run lint:tests) fails with line-number-precise violation`
`DEFECT.SOURCE-GREP-IN-NEW-TESTS.symptom=new test file uses readFileSync + .includes() / .match() against source code (RULESET.TESTS.no-source-grep); contradicts the test rule lint script`
`DEFECT.SOURCE-GREP-IN-NEW-TESTS.detect=npm run lint (AST ESLint rule local/no-source-grep, eslint-rules/no-source-grep.cjs) fails with a line-number-precise violation`
`DEFECT.SOURCE-GREP-IN-NEW-TESTS.fix-forward=replace with runGsdTools(...) behavioral test capturing JSON; if asserting agent .md content (which IS the runtime contract) add // allow-test-rule: source-text-is-the-product with one-line justification`
`DEFECT.GENERATIVE-PRIORITY=these defect classes share a common root: parallel implementations diverge silently because no parity test enforces equality at the test layer`
@@ -773,7 +787,7 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
`DEFECT.GENERATIVE-EXEMPLAR=tests/runtime-launcher-parity.test.cjs (asserts every workflow bash block uses the canonical gsd_run launcher — the in-repo pattern for enforcing equality across parallel surfaces)`
`DEFECT.FRONTMATTER-SCALAR-BROAD-GREP.symptom=a YAML-frontmatter scalar (e.g. VERIFICATION.md status) read with grep "^key:" over the WHOLE markdown report instead of the frontmatter block; a key: line in the body (code block, copied artifact, example) returns extra matches that concatenate after cut|tr into a value matching no expected token, so a valid state is misrouted`
`DEFECT.FRONTMATTER-SCALAR-BROAD-GREP.examples=#586/PR #650 ship.md verification gate — grep "^status:" also matched body status: lines, yielding passed+gaps_found+human_needed instead of passed and blocking a passed phase; the same broad-grep still lives in execute-phase.md (consolidation tracked by #651)`
`DEFECT.FRONTMATTER-SCALAR-BROAD-GREP.examples=#586/PR #650 ship.md verification gate — grep "^status:" also matched body status: lines, yielding passed+gaps_found+human_needed instead of passed and blocking a passed phase; execute-phase.md has since been fixed to the frontmatter-scoped form (#651)`
`DEFECT.FRONTMATTER-SCALAR-BROAD-GREP.detect=grep "^<key>:" on a *.md whose result is compared to exact tokens, with no frontmatter scoping and no -m1; one body line beginning <key>: is enough to break it`
`DEFECT.FRONTMATTER-SCALAR-BROAD-GREP.fix-forward=scope to the leading frontmatter block and take the first match: sed -n '/^---$/,/^---$/p' "$f" | grep -m1 "^<key>:" | cut -d: -f2 | tr -d ' '; fix every parallel copy in the same change or consolidate behind one queryable seam (#651)`
`DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE.symptom=a test that parses a workflow bash block out of a *.md and runs it via execFileSync('bash',...) breaks on Windows two ways: the fence regex uses a literal \n after the bash fence that will not match CRLF and is flagged by local/no-crlf-fragile-split (the windows-test-parity-guard ratchet it formerly tripped was deleted in ADR-1703 Phase 4 #1726); and git-bash exists so a bash-presence probe is true, but an os.tmpdir() Windows path (C:\...) is un-globbable in bash so the pipeline returns empty and assertions fail`
@@ -824,16 +838,16 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
## Shell Command Projection Module (expanded glossary entry, 2026-05-13)
Module owning all OS-facing I/O for the tool: runtime-aware command-text rendering (hook commands, PATH action lines, shim scripts), subprocess dispatch (run-git, run-npm, run-tool, probeTty), and platform file I/O (platformWriteSync, platformReadSync, platformEnsureDir). Single seam for platform-conditional logic — one place to fix any shell or file write regression across Windows, macOS, and Linux. Lives in `gsd-core/bin/lib/shell-command-projection.cjs`. See ADR-0009 (superseded "does not execute" constraint) and ADR-0010 (superseded File Operation Engine).
Module owning all OS-facing I/O for the tool: runtime-aware command-text rendering (hook commands, PATH action lines, shim scripts), subprocess dispatch (execGit, execNpm, execTool, probeTty), and platform file I/O (platformWriteSync, platformReadSync, platformEnsureDir). Single seam for platform-conditional logic — one place to fix any shell or file write regression across Windows, macOS, and Linux. Lives in `gsd-core/bin/lib/shell-command-projection.cjs`. See ADR-0009 (superseded "does not execute" constraint) and ADR-0010 (superseded File Operation Engine).
Invariants:
- Result shape: all run-* return `{ exitCode, stdout, stderr }`; never throw on non-zero exit code.
- Platform policy owned at the seam: `shell: process.platform === 'win32'` lives only in run-npm; probeTty returns `null` on Windows.
- Result shape: all exec* functions return `{ exitCode, stdout, stderr }`; never throw on non-zero exit code.
- Platform policy owned at the seam: `shell: process.platform === 'win32'` lives only in execNpm; probeTty returns `null` on Windows.
- Normalization policy: platformWriteSync owns full `normalizeMd` for `.md`; CRLF-to-LF + trailing newline for all others; callers must NOT pre-call `normalizeMd`.
- `_normalizeMd` is re-implemented inline (not imported from `core.cjs`) to avoid circular dep.
- `atomicWriteFileSync`, `safeReadFile`, `normalizeMd` were in `core.cjs` exports (retired in epic #1267); callers now import these from their respective leaf modules directly.
- `atomicWriteFileSync`, `safeReadFile`, `normalizeMd` were duplicated wrappers in `core.cjs`, removed from there in this module's own Migration Phase 4 (`#3468`, see ADR-0009/ADR-0010) — a separate effort from epic #1267 (the core.cjs re-export-spine retirement); callers now import `platformWriteSync`/`platformReadSync`/`normalizeContent` from this module directly.
Migration plan: Phase 1 (#3465) seam additions complete; Phase 2 (#3466) targets 6 subprocess files; Phase 3 (#3467) targets 15 fs files (215 call sites); Phase 4 (#3468) removes compat exports.
Migration: Phases 1-4 (#3465-#3468) shipped 2026-05-13 — seam additions, subprocess-dispatch migration (6 files), fs migration (15 files / 215 call sites), and core.cjs compat-export removal are all complete; see ADR-0009's 2026-05-13 Update section.
---
@@ -892,7 +906,7 @@ Migration plan: Phase 1 (#3465) seam additions complete; Phase 2 (#3466) targets
`RULESET.HARNESS.test-memory-guard=~/.claude/hooks/test-memory-guard.sh fires on every Bash PreToolUse; if argv[0]∈{node|vitest|jest|mocha|tsx|ts-node|tap|ava|playwright|cypress} OR matches (npm|pnpm|yarn|bun) (run )?(t|test|tests|vitest|jest); blocks via hookSpecificOutput.permissionDecision=deny when sum(RSS of running matching procs, excluding tsserver|*-mcp|claude|Electron|...) ≥ 4 GiB OR when argv[0] basename matches a running process's argv[0]. Exception: node --version|-v|--help|-h|-p|-e are trivial probes and skip the check. Designed for a 24 GB Mac where prior accidental fan-out exhausted RAM`
`RULESET.PR-FLOW.docker-before-push=before ANY git push of any fix to any PR, run gsd-test (docker on the remote, mirrors ubuntu CI) and confirm exit 0. macOS-local node --test is NOT a substitute — many failures are platform-specific (path separators, case sensitivity, locale, fs semantics). Watchdog with Monitor on the output log; never set a sleep/timer and walk away. Source: user feedback 2026-05-16 — "we don't set a timer we actively watch and record results in real time as possible"`
`RULESET.PR-FLOW.docker-before-push=before ANY git push of any fix to any PR, run gsd-test (docker on the remote, mirrors ubuntu CI) and confirm exit 0. macOS-local node --test is NOT a substitute — many failures are platform-specific (path separators, case sensitivity, locale, fs semantics). Watchdog with Monitor on the output log; never set a sleep/timer and walk away. Source: user feedback 2026-05-16 — "we don't set a timer we actively watch and record results in real time as possible". SUPERSEDED 2026-07-17: 'confirm exit 0' is a false-green trap — piping/backgrounding can report exit 0 on a failed suite; gate on the verdict-line outcome:"passed" for the exact HEAD sha instead. See CLAUDE.md's gsd-test rule and the gsd-test-is-ref-based-commit-first predicate for the current, correct gating contract.`
`RULESET.PR-FLOW.templates-mandatory=every gh pr create|edit|gh issue create|edit MUST first invoke the gh-templates-first skill and Read (Read tool, not Bash cat — k321 read-tracking) the matching template in .github/. Apply ALL required sections; never write freeform bodies. Repo enforces this via gsd-pr-template-policy GitHub Action which flags any non-templated body — the bot allows the PR to stay open only because authors are contributors-or-higher, but the warning is a real complaint that must be cured. Source: user feedback 2026-05-16 (multi-message escalation) — "the whole reason i have that github action is because you fucking blow through and ignore using the templates"`
@@ -903,7 +917,7 @@ Migration plan: Phase 1 (#3465) seam additions complete; Phase 2 (#3466) targets
`EXEC.CLASSIFY.handler=gsd-core/bin/lib/agent-command-router.cjs:classifyAgentFailure (registered via command-aliases.cjs; mutation:false outputMode:json)`
`EXEC.CLASSIFY.workflow=gsd-core/workflows/execute-phase.md step 7; class-distinct prompts (quota-to-wait-for-reset; classify-handoff-bug-to-spot-check; unknown-to-continue/stop)`
`EXEC.CLASSIFY.classes={class:'quota-exceeded'|'classify-handoff-bug'|'unknown-failure', sentinel?, retryAfterSeconds?}`
`EXEC.CLASSIFY.sentinel-order=most specific first: 429 beats too-many-requests; quota beats resource_exhausted; case-insensitive; canonical sentinel value is lower-cased form`
`EXEC.CLASSIFY.sentinel-order=most specific first: 429 beats too-many-requests; resource_exhausted beats quota (array order in src/agent-command-router.cts QUOTA_SENTINELS checks resource_exhausted before quota); case-insensitive; canonical sentinel value is lower-cased form`
`EXEC.CLASSIFY.cross-runtime=Anthropic/CC: usage limit|rate limit|quota|429|retry-after; Copilot CLI: rate_limit (stem); Codex CLI: 429|usage_limit_reached|too many requests`
`EXEC.CLASSIFY.precedence=quota sentinel wins over classifyHandoffIfNeeded bug when both appear`
`EXEC.CLASSIFY.retry-after-parser=\bretry[-_ ]after[:\s]+(\d+)\b avoids embedded-word false matches like noretry-after`
@@ -920,11 +934,11 @@ Migration plan: Phase 1 (#3465) seam additions complete; Phase 2 (#3466) targets
`DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION.root-cause=gsd-test-summary lines 126-127 default LOCAL_OUT/DOCKER_OUT to fixed /tmp/gsd-test-{local,docker}.jsonl; concurrent line-buffered writers interleave bytes mid-multibyte → split UTF-8 sequence → decoder explodes on f.read()`
`DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION.detect=two gsd-test-summary --both runs in flight; UnicodeDecodeError in parse_events_from_string traceback; /tmp/gsd-test-*.jsonl size mismatch vs total events emitted`
`DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION.fix-forward=set per-invocation LOCAL_OUT=/tmp/gsd-test-<tag>-local.jsonl DOCKER_OUT=/tmp/gsd-test-<tag>-docker.jsonl env vars; or serialize the runs; upstream fix tracked in #3545 (default to tempfile.mkstemp + advisory flock)`
`DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION.upstream=open-gsd/gsd-core#3545`
`DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION.upstream=open-gsd/gsd-test-runner#4 (moved from #3545 in the predecessor repo, filed in the wrong repo; now CLOSED/COMPLETED — fix shipped)`
`DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL.symptom=spawned sub-agent kicks off gsd-test-summary --both via Bash run_in_background, then stops on the harness "you will be notified" message; never receives the notification because cross-turn task-notifications are only delivered to the top-level orchestrator`
`DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL.detect=sub-agent returns prematurely with text like "I should wait for the notification per CLAUDE.md" and incomplete work in its worktree (commits absent, push absent, PR absent)`
`DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL.fix-forward=keep gsd-test-summary --both at the top-level orchestrator; sub-agents either run it foreground with timeout: 1500000 (25min) and block, OR delegate the test step back to the orchestrator (write commits + return); never have a sub-agent fire-and-await a backgrounded long task`
`DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL.anchor=project CLAUDE.md "Top-level orchestrator (cross-turn notifications available) vs Sub-agent worker (no cross-turn notifications)" guidance — load-bearing for multi-worktree parallel fix dispatch`
`DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL.anchor=lesson: cross-turn task notifications are delivered only to the top-level orchestrator, never to a sub-agent — load-bearing for multi-worktree parallel fix dispatch (the CLAUDE.md passage this entry previously quoted verbatim has since been removed/rewritten; no live replacement citation exists)`
`DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT.symptom=sub-agent writes /gsd-<cmd> (legacy hyphen syntax) in code comments or doc strings while implementing a fix; lands as part of the implementation diff`
`DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT.examples=#3541 implementation included a typical /gsd-update path comment in installer-migration-report.cjs; caught by tests/slash-command-namespace.test.cjs (#3443 invariant)`
`DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT.detect=tests/slash-command-namespace.test.cjs prints "Found N retired /gsd-<cmd> reference(s) — use /gsd:<cmd> instead" with line-number-precise violations`
@@ -986,13 +1000,13 @@ Full detail in `~/.claude/skills/gsd-pr-fix-discipline/SKILL.md`. AI agents MUST
- **Symptom:** After base PR squash-merges, stacked PR shows conflicts or wrong diff; GitHub auto-retarget fails
- **Affected this session:** #158 stacked on #156
- **Fix:** `git rebase --onto main <old-base> <stacked-branch>` then force-push and `gh pr edit --base main`
- **Fix:** `git rebase --onto next <old-base> <stacked-branch>` then force-push and `gh pr edit --base next` (repo's integration branch as of the `next`-branch model introduced 2026-05-24; was `main` before that)
### `tee` pipe swallowing exit codes
- **Symptom:** `gsd-test-summary --both 2>&1 | tee /tmp/log` returns `0` even when Docker reports failures
- **Symptom:** piping any test-runner output through `tee` (e.g. `... 2>&1 | tee /tmp/log`) returns `0` even when the run reports failures — the pipeline exits on `tee`'s status, not the runner's
- **Affected this session:** Session-wide risk
- **Fix:** Run un-piped, or `set -o pipefail` before the pipe
- **Fix:** Run `gsd-test` (classic executor) UNPIPED per CLAUDE.md, or `set -o pipefail` before any pipe. Do not use the deprecated `gsd-test-summary` / `gsd-test-both` wrappers at all (see DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION / DEFECT.GSD-TEST-MIRROR-POISONED).
### Auto-merge disabled

View File

@@ -501,6 +501,14 @@ Required cases where relevant:
Property-style parser tests are encouraged for high-risk parsers. They must be deterministic: pin the seed, bound the iteration count, and print replay data on failure.
##### Fixture provenance (#2371)
**A gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples. A negative fixture must come from a source that does not know the gate exists.**
This is stricter than the adversarial-input rule above and exists because of it: `tests/fixtures/adversarial/` covers hostile input, but a fixture written by the parser's own author — even a deliberately "realistic" one — is still drawn from the author's mental model of the format. It can only ever confirm what the author already believed, never surface what they didn't anticipate. A property-test generator has the same failure mode one level up: seeding the generator from the writer/render function that produces the same format makes the document shape a constant, so the property can never explore a document the writer wouldn't produce (see the document-shaped vs. writer-seeded property tests in `tests/api-coverage.test.cjs` for a worked example — the writer-seeded one cannot fail against a decoy table; the document-shaped one can).
For a gate whose fixtures come from real user reports, put them under `tests/fixtures/representative/<gate>/` with a `MANIFEST.json` labeling each fixture's source issue and expected gate verdict, and drive them through the gate's real CLI entrypoint (gate-verdict altitude), not the parser function in isolation — see `tests/fixtures/representative/README.md` and `tests/representative-corpus.test.cjs`. If the gate is not yet fixed, do not mark the assertion `{ todo: true }` and do not skip it: this repo's test-runner (`gsd-test` / `gsd-test-runner`) has no concept of node:test's `todo` option — its JSONL result parser only recognizes `kind: "pass" | "fail"`, so a thrown todo-marked test is still counted as a real failure and blocks the push gate. Instead record BOTH the correct target verdict (`expected*`) and the exact current observed verdict (`currentBuggyOutput`) in the manifest, and assert against `currentBuggyOutput` — an honest, non-vacuous characterization of today's known-broken behavior that passes today and breaks loudly the moment the real fix changes the observed output, forcing the assertion to be flipped to `expected*`.
#### Filesystem writes and installers
Changes to install/uninstall flows, generated artifact writers, state/config writers, worktree safety, or any code that writes under `.planning`, runtime config dirs, `.claude`, `.codex`, `hooks`, or generated files must include fault-injection coverage where the seam allows it.

View File

@@ -60,6 +60,8 @@ New here? Follow [Your first project](docs/tutorials/your-first-project.md) for
## Documentation
**What's new in 1.7.0** → [docs/whats-new-1.7.0.md](docs/whats-new-1.7.0.md)
**Tutorials** — learning by doing:
- [Your first project](docs/tutorials/your-first-project.md)
- [Onboarding an existing codebase](docs/tutorials/onboarding-an-existing-codebase.md)

View File

@@ -270,30 +270,67 @@ If user selects 1 or 2: spawn continuation agent (with any additional context pr
If user selects 3: proceed to Step 4 with fix = "not applied".
### 3f. FIX REJECTED BY GUARDRAIL
When agent returns `## FIX REJECTED BY GUARDRAIL`:
Present the failing signal and evidence to the user via AskUserQuestion:
```
Fix rejected by the acceptance guardrail.
Failing signal: {failing signal}
Evidence: {why it failed}
Options:
1. Revise fix — spawn continuation agent to revise the fix so the signal passes
2. Accept as technical debt — record the unmet signal + justification (the fix lands without the gate passing; this is never silent)
3. Abandon — stop; session stays unresolved
```
If user selects 1: spawn continuation agent with `goal: find_and_fix` naming the failing signal to revise. Loop back to Step 3.
If user selects 2: spawn continuation agent instructed to record `guardrail_verdict: accepted_debt` + the justification in the debug file, then proceed to request_human_verification. Loop back to Step 3.
If user selects 3: proceed to Step 4 with fix = "not applied (guardrail rejected)".
## Step 4: Return Compact Summary
**Non-terminal early stop — check this FIRST.** Before returning any summary below, ask: is your own turn/context budget exhausted while the debugger (`gsd-debugger`) is still investigating — i.e. you have NOT reached `DEBUG COMPLETE`, a user-chosen `ABANDONED`, or exhausted the `INVESTIGATION INCONCLUSIVE` options? If so, do NOT fabricate a `DEBUG SESSION COMPLETE` or `ABANDONED` summary to fit this shape. Return the non-terminal marker instead:
```markdown
## CONTINUE_REQUIRED
**Session:** {debug_file_path}
**Status:** {status from frontmatter, e.g. investigating}
**Next action:** {next_action from Current Focus}
**Reason:** session-manager turn/context budget exhausted — investigation still in progress
```
`CONTINUE_REQUIRED` is distinct from both terminal shapes below AND from `## CHECKPOINT REACHED` (Step 3d): a `CHECKPOINT REACHED` is a genuine user-input/approval checkpoint that already correctly pauses via `AskUserQuestion` before looping back to Step 3 — it is not returned to the orchestrator. `CONTINUE_REQUIRED` is emitted only when no checkpoint is pending and the loop simply cannot proceed further in this turn. The orchestrator resumes by re-spawning this agent with the SAME `slug`/`debug_file_path` — the on-disk checkpoint at `.planning/debug/{slug}.md` (its `status` and `next_action`) is the source of truth for where to pick up. Never return control to the user as if the session were complete when it is not.
Read the resolved (or current) debug file to extract final Resolution values.
Return compact summary:
Return compact summary (terminal — investigation resolved):
```markdown
## DEBUG SESSION COMPLETE
**Session:** {final path — resolved/ if archived, otherwise debug_file_path}
**Root Cause:** {one sentence from Resolution.root_cause, or "not determined"}
**Root Cause:** {one sentence, or a '; '-joined list when the AND-gate identified multiple contributing causes, from Resolution.root_cause; or "not determined"}
**Fix:** {one sentence from Resolution.fix, or "not applied"}
**Cycles:** {N} (investigation) + {M} (fix)
**TDD:** {yes/no}
**Specialist review:** {specialist_hint used, or "none"}
**Prevention:** {one-line from the blameless postmortem — "why not caught: <gate, or 'none (no gate existed for this class)'>; guard: <artifact>"}
```
If the session was abandoned by user choice, return:
If the session was abandoned by user choice, return (terminal — user stopped):
```markdown
## DEBUG SESSION COMPLETE
**Session:** {debug_file_path}
**Root Cause:** {one sentence if found, or "not determined"}
**Root Cause:** {one sentence if found (or a '; '-joined list if the AND-gate identified multiple contributing causes), or "not determined"}
**Fix:** not applied
**Cycles:** {N}
**TDD:** {yes/no}
@@ -311,5 +348,6 @@ If the session was abandoned by user choice, return:
- [ ] Specialist dispatch executed when specialist_dispatch_enabled and hint maps to a skill
- [ ] TDD gate applied when tdd_mode=true and ROOT CAUSE FOUND
- [ ] Loop continues until DEBUG COMPLETE, ABANDONED, or user stops
- [ ] Non-terminal `CONTINUE_REQUIRED` (not a fabricated terminal summary) returned when the manager's own turn/context budget is exhausted mid-investigation
- [ ] Compact summary returned (at most 2K tokens)
</success_criteria>

View File

@@ -253,6 +253,10 @@ reasoning_checkpoint:
falsification_test: "[what specific observation would prove this hypothesis wrong]"
fix_rationale: "[why the proposed fix addresses the root cause — not just the symptom]"
blind_spots: "[what you haven't tested that could invalidate this hypothesis]"
candidate_causes:
- "[cause in category: code|config|environment|data]"
- "[cause in a DIFFERENT category — single-category is not a branch]"
and_gate: "[could this failure require >1 contributing condition simultaneously? yes/no + why — see RCA branching]"
```
**Check before proceeding:**
@@ -260,8 +264,9 @@ reasoning_checkpoint:
- Is the confirming evidence direct observation, not inference?
- Does the fix address the root cause or a symptom?
- Have you documented your blind spots honestly?
- **Did you branch across ≥2 categories and answer the AND-gate?** (Single-cause is fine when the AND-gate is no — but you must have checked.)
If you cannot fill all five fields with specific, concrete answers — you do not have a confirmed root cause yet. Return to investigation_loop.
If you cannot fill all seven fields with specific, concrete answers — you do not have a confirmed root cause yet. Return to investigation_loop.
## Minimal Reproduction
@@ -274,6 +279,7 @@ If you cannot fill all five fields with specific, concrete answers — you do no
3. Test: Does it still reproduce? YES = keep removed. NO = put back.
4. Repeat until bare minimum
5. Bug is now obvious in stripped-down code
6. **Shrinking (input-space bugs)** — when the bug triggers on a class of inputs, wrap it in a property (fast-check for JS/TS, Hypothesis for Python) and let the shrinker auto-minimize the counterexample; store the **minimized** input as the regression seed. See `gsd-core/references/debugger-repro-hardening.md`.
**Example:**
```jsx
@@ -443,18 +449,21 @@ MISMATCH: Checker looks in wrong directory → hooks "not found" → reported as
**The discipline:** Never assume a constructed path is correct. Resolve it to its actual value and verify the other side agrees. When two systems share a resource (file, directory, key), trace the full path in both.
## Technique Selection
## Technique Selection (routed by bug class)
| Situation | Technique |
|-----------|-----------|
| Large codebase, many files | Binary search |
| Confused about what's happening | Rubber duck, Observability first |
| Complex system, many interactions | Minimal reproduction |
| Know the desired output | Working backwards |
| Used to work, now doesn't | Differential debugging, Git bisect |
| Many possible causes | Comment out everything, Binary search |
| Paths, URLs, keys constructed from variables | Follow the indirection |
| Always | Observability first (before making changes) |
Classify the failure first (Phase 1.75), then route by class — not by ad-hoc
situation:
@~/.claude/gsd-core/references/debugger-bug-taxonomy.md
| bug_class | Route to | Revoke if already run |
|---|---|---|
| Bohrbug | deterministic reproduction → SBFL (Phase 1.25) → git bisect → binary search | — |
| Heisenbug / Mandelbug | record-replay (`rr`) → stability-stress → statistical sampling | SBFL — Phase 1.25 runs before classification; if it ran, mark its Evidence entry revoked (flaky spectrum poisons the ranking) |
| Concurrency | atomicity / order / deadlock checklist (see reference) FIRST | — |
| General (any class) | Binary search, Working backwards, Differential, Delta debugging, Comment-out-everything, Follow-the-indirection, Rubber duck, Observability first (always, before changes) | — |
The class rows pick the first move; the General lane holds situation-cued techniques that apply to any class. When the situation table and the class route disagree, the class route wins.
## Combining Techniques
@@ -590,6 +599,13 @@ function processUserData(user) {
// 5. Test is now regression protection forever
```
**Harden the regression test (so the Phase 1A mutation guardrail bites):**
@~/.claude/gsd-core/references/debugger-repro-hardening.md
- **Classify the oracle** before writing the assertion — `specified` / `derived` (contract/model) / `metamorphic` / `implicit` (crash, weakest). Record it under `Resolution.oracle_type`. Never default to implicit silently.
- **Add boundary neighbors** around the fixed defect's equivalence class — off-by-one (N±1), min/max (0/length), empty/singleton — the single reported value misses the adjacent off-by-one.
## Verification Checklist
```markdown
@@ -788,9 +804,11 @@ Each resolved session appends one entry:
## {slug} — {one-line description}
- **Date:** {ISO date}
- **Error patterns:** {comma-separated keywords extracted from symptoms.errors and symptoms.actual}
- **Root cause:** {from Resolution.root_cause}
- **Root cause(s):** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate fired}
- **Fix:** {from Resolution.fix}
- **Files changed:** {from Resolution.files_changed}
- **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
- **Recurrence guard:** {the concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / type refinement / config-default change / KB pattern}
---
```
@@ -804,9 +822,11 @@ At the **end of `archive_session`**, after the session file is moved to `resolve
## Matching Logic
Matching is keyword overlap, not semantic similarity. Extract nouns and error substrings from `Symptoms.errors` and `Symptoms.actual`. Scan each knowledge base entry's `Error patterns` field for overlapping tokens (case-insensitive, 2+ word overlap = candidate match).
**Semantic-first, keyword-fallback.** Query MemPalace with the current symptoms and surface the top-k meaning-similar prior resolutions — this catches same-root-cause/different-wording cases keyword overlap misses. Fall back to keyword overlap on `knowledge-base.md` when MemPalace is absent. See:
**Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis. Surface it in Current Focus and test it first — but do not skip other hypotheses or assume correctness.
@~/.claude/gsd-core/references/debugger-semantic-recall.md
**Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis — surface it in Current Focus and test it first; do not skip other hypotheses or assume correctness.
</knowledge_base_protocol>
@@ -966,12 +986,10 @@ At investigation decision points, apply structured reasoning:
**Autonomous investigation. Update file continuously.**
**Phase 0: Check knowledge base**
- If `.planning/debug/knowledge-base.md` exists, read it
- Extract keywords from `Symptoms.errors` and `Symptoms.actual` (nouns, error substrings, identifiers)
- Scan knowledge base entries for 2+ keyword overlap (case-insensitive)
- Query MemPalace semantically with the current symptoms (top-k meaning-similar prior resolutions); fall back to reading `.planning/debug/knowledge-base.md` and keyword overlap when MemPalace is absent
- If match found:
- Note in Current Focus: `known_pattern_candidate: "{matched slug} — {description}"`
- Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}.`
- Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}. Why not caught: {why_not_caught}. Recurrence guard: {recurrence_guard}.` (the last two are absent on old entries — that's fine; consume them when present)
- Test this hypothesis FIRST in Phase 2 — but treat it as one hypothesis, not a certainty
- If no match: proceed normally
@@ -983,14 +1001,32 @@ At investigation decision points, apply structured reasoning:
- Run app/tests to observe behavior
- APPEND to Evidence after each finding
**Phase 1.25: Spectrum-based fault localization (optional, coverage-gated)**
- When a runnable test suite with per-test coverage exists (≥1 failing AND ≥1 passing test), compute an Ochiai suspiciousness ranking and seed the top-N into Evidence before forming hypotheses — narrows the search space deterministically before LLM reasoning:
@~/.claude/gsd-core/references/debugger-sbfl.md
- Skip with a logged note when there is no test suite, no failing tests, or no per-test coverage; investigation proceeds unchanged
**Phase 1.5: Check common bug patterns**
- Read @~/.claude/gsd-core/references/common-bug-patterns.md
- Match symptoms to pattern categories using the Symptom-to-Category Quick Map
- Any matching patterns become hypothesis candidates for Phase 2
- If no patterns match, proceed to open-ended hypothesis formation
**Phase 1.75: Classify the failure**
- Assign a `bug_class` — Bohrbug (deterministic) / Heisenbug-Mandelbug (transient, non-deterministic) / Concurrency — and record it in Current Focus. The class routes which investigation technique to use:
@~/.claude/gsd-core/references/debugger-bug-taxonomy.md
- Bohrbug → reproduction + SBFL + bisect; Heisenbug/Mandelbug → record-replay/stability (skip SBFL — flaky spectra poison it); Concurrency → the atomicity/order/deadlock checklist first
**Phase 2: Form hypothesis**
- Based on evidence AND common pattern matches, form SPECIFIC, FALSIFIABLE hypothesis
- **Branch, don't chain** — at hypothesis formation (so it's done before the Phase 4 commit), enumerate candidate causes across ≥2 Ishikawa categories (code / config / environment / data) and answer the AND-gate check; `root_cause` may hold a set when the AND-gate fires:
@~/.claude/gsd-core/references/debugger-rca-branching.md
- Update Current Focus with hypothesis, test, expecting, next_action
**Phase 3: Test hypothesis**
@@ -1043,7 +1079,7 @@ Return structured diagnosis:
**Debug Session:** .planning/debug/{slug}.md
**Root Cause:** {from Resolution.root_cause}
**Root Cause:** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
**Evidence Summary:**
- {key finding 1}
@@ -1083,7 +1119,7 @@ Update status to "fixing".
**0. Structured Reasoning Checkpoint (MANDATORY)**
- Write the `reasoning_checkpoint` block to Current Focus (see Structured Reasoning Checkpoint in investigation_techniques)
- Verify all five fields can be filled with specific, concrete answers
- Verify every field can be filled with specific, concrete answers — including the RCA `candidate_causes` (≥2 categories) and `and_gate` fields
- If any field is vague or empty: return to investigation_loop — root cause is not confirmed
**1. Implement minimal fix**
@@ -1091,11 +1127,15 @@ Update status to "fixing".
- Make SMALLEST change that addresses root cause
- Update Resolution.fix and Resolution.files_changed
**2. Verify**
**2. Verify (Fix-Acceptance Guardrail)**
- Update status to "verifying"
- Test against original Symptoms
- If verification FAILS: status -> "investigating", return to investigation_loop
- If verification PASSES: Update Resolution.verification, proceed to request_human_verification
- Run the multi-signal guardrail before accepting the fix:
@~/.claude/gsd-core/references/debugger-fix-acceptance.md
- Record every signal's result under `Resolution.verification` (per-signal schema in the reference)
- If ANY applicable signal fails (and no documented technical-debt escape applies): return `## FIX REJECTED BY GUARDRAIL` (see structured_returns) — do NOT request human verification
- If all applicable signals pass: set `guardrail_verdict: accepted`, proceed to request_human_verification
</step>
<step name="request_human_verification">
@@ -1174,9 +1214,13 @@ Then commit planning docs via CLI (respects `commit_docs` config automatically):
gsd_run query commit "docs: resolve debug {slug}" --files .planning/debug/resolved/{slug}.md
```
**Append to knowledge base:**
**Append to knowledge base (with the Prevention block):**
Read `.planning/debug/resolved/{slug}.md` to extract final `Resolution` values. Then append to `.planning/debug/knowledge-base.md` (create file with header if it doesn't exist):
Read `.planning/debug/resolved/{slug}.md` to extract final `Resolution` values. Then produce the **Prevention block** — a blameless postmortem (branching 5-Whys per RCA, "why wasn't this caught?", and a concrete recurrence guard):
@~/.claude/gsd-core/references/debugger-prevention.md
Then append to `.planning/debug/knowledge-base.md` (create file with header if it doesn't exist):
If creating for the first time, write this header first:
```markdown
@@ -1193,9 +1237,11 @@ Then append the entry:
## {slug} — {one-line description of the bug}
- **Date:** {ISO date}
- **Error patterns:** {comma-separated keywords from Symptoms.errors + Symptoms.actual}
- **Root cause:** {Resolution.root_cause}
- **Root cause(s):** {Resolution.root_cause — joined as '; ' when multiple contributing causes were confirmed}
- **Fix:** {Resolution.fix}
- **Files changed:** {Resolution.files_changed joined as comma list}
- **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
- **Recurrence guard:** {concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / KB pattern / type refinement / config-default change}
---
```
@@ -1205,6 +1251,8 @@ Commit the knowledge base update alongside the resolved session:
gsd_run query commit "docs: update debug knowledge base with {slug}" --files .planning/debug/knowledge-base.md
```
**Index into MemPalace (when available)** per the semantic-recall reference — the Resolution summary (not raw symptoms), redacted — so a future Phase-0 query surfaces it by meaning. Skip with a logged note when MemPalace is absent or the KB write failed; `knowledge-base.md` is the durable fallback.
Report completion and offer next steps.
</step>
@@ -1298,7 +1346,7 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
**Debug Session:** .planning/debug/{slug}.md
**Root Cause:** {specific cause with evidence}
**Root Cause:** {specific cause with evidence — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
**Evidence Summary:**
- {key finding 1}
@@ -1334,6 +1382,16 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
Only return this after human verification confirms the fix.
## FIX REJECTED BY GUARDRAIL
Returned when a fix-acceptance guardrail signal fails (see `@~/.claude/gsd-core/references/debugger-fix-acceptance.md`). Do **not** mark the session resolved.
**Debug Session:** .planning/debug/{slug}.md
**Failing signal:** {signal 1–5 name}
**Evidence:** {why the signal failed — e.g. "mutant at fix site survived", "deletion-only diff with no RCA justification", "bug did not return on revert"}
The session-manager continuation surfaces this and offers revise / accept-as-debt / abandon.
## INVESTIGATION INCONCLUSIVE
```markdown

View File

@@ -144,6 +144,10 @@ At execution decision points, apply structured reasoning:
For each task:
0. **Precondition check (before any other task work):** If the task carries a `<precondition>` element, evaluate that single prose line first — it names a runnable/checkable fact the task assumes (env var set, prior-phase artifact present, server responding to `/health`, `user_setup` step done). Verify with **read-only checks only** — file existence, env var presence (no value output), idempotent `GET /health`-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.
- **Met OR absent:** continue with no visible change to execution flow. The precondition is a no-op for the rest of the task loop.
- **Unmet:** STOP — return a `checkpoint:human-verify` (use `checkpoint_return_format`) with `**Blocked by:** Precondition not met: <precondition text>`. Do NOT partial-commit the task. Unmet preconditions are NEVER auto-approved, even under `AUTO_CFG=true` — a missing prerequisite is not a verification step a human can rubber-stamp; it is a fact the executor cannot establish on its own. The human either satisfies the precondition (sets the env var, completes the `user_setup` step, regenerates the artifact) or reruns `/gsd:plan-phase` to restructure.
1. **If `type="auto"`:**
- Check for `tdd="true"` → follow TDD execution flow
- Execute task, apply deviation rules as needed
@@ -152,11 +156,17 @@ For each task:
- Commit (see task_commit_protocol)
- Track completion + commit hash for Summary
2. **If `type="checkpoint:*"`:**
2. **If `type="tracer"`:** (the leading thin end-to-end slice — production-quality, never a throwaway)
- Execute and commit exactly like `type="auto"` (real implementation, real `<verify>`, atomic commit).
- **Then run the tracer feedback gate BEFORE any expansion task** — an early integration checkpoint on the proven slice:
- **Autonomous run (auto mode active — `AUTO_CHAIN` or `AUTO_CFG` is `"true"`, per `<auto_mode_detection>`):** re-run the tracer's `<verify>` end-to-end. If it **fails**, HALT and surface it (deviation Rule 1) — do NOT proceed to expansion tasks. Pouring more layers onto a broken foundation is exactly the failure this gate prevents. If it passes, log `⚡ Tracer verified end-to-end — expanding` and continue.
- **Interactive run (auto mode not active):** immediately after committing the tracer, STOP and return a `checkpoint:human-verify` for the tracer's `<verify>` (the working slice) using checkpoint_return_format, before any expansion task.
3. **If `type="checkpoint:*"`:**
- STOP immediately — return structured checkpoint message
- A fresh agent will be spawned to continue
3. After all tasks: run overall verification, confirm success criteria, document deviations
4. After all tasks: run overall verification, confirm success criteria, document deviations
</step>
</execution_flow>
@@ -310,6 +320,8 @@ For full automation-first patterns, server lifecycle, CLI handling:
**Quick reference:** Users NEVER run CLI commands. Users ONLY visit URLs, click UI, evaluate visuals, provide secrets. Claude does all automation.
**Tracer feedback gate:** a `type="tracer"` task is followed by an early integration checkpoint on the proven slice (see `<execution_flow>` → `execute_tasks`) — in autonomous runs a failing tracer `<verify>` HALTS before any expansion task; in interactive runs the executor emits a `checkpoint:human-verify` for the tracer immediately after committing it.
---
**Auto-mode checkpoint behavior** (when `AUTO_CFG` is `"true"`):
@@ -660,6 +672,21 @@ Or: "None - plan executed exactly as written."
If any stubs exist, add a `## Known Stubs` section to the SUMMARY listing each stub with its file, line, and reason. These are tracked for the verifier to catch. Do NOT mark a plan as complete if stubs exist that prevent the plan's goal from being achieved — either wire the data or document in the plan why the stub is intentional and which future plan will resolve it.
**Broken-windows ledger (issue #1950).** For each stub, skipped test, or unrun `<verify>` recorded above, ALSO append it to the cross-phase defect register at `.planning/WINDOWS.md`. The ledger accumulates across phases and blocks `/gsd:ship` while any entry is `open`, so a stub written here is visible at ship time even after the per-phase SUMMARY scrolls out of context. Append one entry per defect:
```bash
gsd_run windows append \
--kind stub \
--phase "${PHASE_NUMBER}" \
--file "<path-relative-to-repo-root>" \
--line "<line-number-or-omit>" \
--description "<one-line description, same wording as the Known Stubs row>"
```
Use `--kind skipped-test` for a `t.skip(...)` / `test.todo(...)` you left behind, `--kind unrun-verify` for a `<verify>` you could not run, or `--kind deviation` for a documented plan deviation. The full kind vocabulary: `stub | todo | fixme | skipped-test | lint-warning | unmet-truth | unrun-verify | deviation`.
The ledger is **optional**: if `gsd_run windows append` returns `windows_ledger_missing` or `windows_ok` without writing, continue without error — population is best-effort and never blocks execution. Recording here is what makes the defect visible to the ship gate later; forgetting to record is the failure mode this ledger exists to prevent.
**Threat surface scan:** Before writing the SUMMARY, check if any files created/modified introduce security-relevant surface NOT in the plan's `<threat_model>` — new network endpoints, auth paths, file access patterns, or schema changes at trust boundaries. If found, add:
```markdown

View File

@@ -66,7 +66,7 @@ The orchestrator provides user decisions in `<user_decisions>` tags from `/gsd:d
**Self-check before returning:** For each plan, verify:
- [ ] Every locked decision (D-01, D-02, etc.) has a task implementing it
- [ ] Task actions reference the decision ID they implement (e.g., "per D-03")
(The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, and `<action>` tag bodies, as well as markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
(The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, `<action>`, `<read_first>`, `<behavior>`, `<verify>`, `<acceptance_criteria>`, and `<done>` tag bodies, as well as `## must_haves`/`truths`/`tasks`/`objective` markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
- [ ] No task implements a deferred idea
- [ ] Discretion areas are handled reasonably
@@ -193,23 +193,21 @@ Every task has four required fields:
**Grep gate hygiene:** `grep -c` counts comments, so header prose can be self-invalidating. Use `grep -v '^#' | grep -c token`. Bare `== 0` gates on unfiltered files are forbidden.
<comment_text_discipline>
**Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for (`grep -c 'LIT' file == 0`) must NOT appear verbatim in any `<action>` body — JSDoc samples, head-comment references, or "what NOT to do" snippets echo into the written file and trip the executor's commit-time gate. `validate_plan` (`verify.plan-structure`) fails plan creation on violation. Rephrase the literal by concept, or — when it must legitimately appear — add an allowlist marker on its own line:
`<!-- planner-discipline-allow: LIT -->`
Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
**Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for must NOT appear verbatim in any `<action>` body. Full rules + `<!-- planner-discipline-allow: LIT -->` allowlist + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
</comment_text_discipline>
<region_scoped_negative_gate>
**Region-scoped negative gates (WARN, #968):** Region-scope a file-wide negative grep when a sibling task needs that construct elsewhere in the same file; `validate_plan` WARNS. See: @gsd-core/references/planner-antipatterns.md ("Region-Scoped Negative Gates").
**Verify-gate hygiene (#1478/#1479):** See @gsd-core/references/planner-antipatterns.md.
**Region-scoped negative gates (WARN, #968)** and **Verify-gate hygiene (#1478/#1479):** @gsd-core/references/planner-antipatterns.md.
</region_scoped_negative_gate>
**<done>:** Acceptance criteria - measurable state of completion.
- Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
- Bad: "Authentication is complete"
**<precondition>** (optional, one prose line): a runnable/checkable fact the task assumes that plan ordering does not guarantee — external setup (`user_setup`), a prior-phase artifact, or an env var. The executor asserts it before running the task and halts on unmet. Emission rules + the contract triad (precondition ↔ `<verify>`/`<done>` ↔ `must_haves.truths`): @~/.claude/gsd-core/references/planner-preconditions.md.
**<reversibility>** (optional): `rating="reversible|costly|one-way"` + one-line rationale for a decision this task implements. `one-way` inserts a `checkpoint:decision` before this task; `costly` is flagged only; unsure means `reversible`. Rules: @~/.claude/gsd-core/references/planner-reversibility.md
See @~/.claude/gsd-core/references/planner-guidance.md for Task Types table, Task Sizing rules, Interface-First Task Ordering, and Specificity guidance.
## TDD Detection
@@ -250,34 +248,33 @@ Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, config
`workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
## MVP Mode Detection
## Tracer-First Decomposition (default)
**When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: Read `~/.claude/gsd-core/references/planner-mvp-mode.md` for the vertical-slice rules (lazy — only on MVP runs).
**Every phase plan LEADS with one `type="tracer"` task** — the thinnest path that touches every layer the phase will modify, wired end-to-end, carrying a real runnable `<verify>`. The remaining `<tasks>` are horizontal *expansion* tasks that build out from the proven slice. This is the default for **every** phase; it is not gated behind a flag. Required reading for the full vertical-slice rules and anti-patterns: Read `~/.claude/gsd-core/references/planner-mvp-mode.md`.
**Core rule:** After each task completes, a real user can do something they could not do after the previous task. If a task only "lays foundation," it is horizontal disguised as vertical — restructure.
**Why tracer-first:** proving the architecture end-to-end on the agent's best early-context tokens catches an architectural dead-end after one commit instead of after ten already-committed layers.
**Plan structure under MVP_MODE:**
**A tracer is production-quality, not a prototype.** It carries the same `<verify>` and validation as any `auto` task and becomes part of the skeleton of the final system — you write it for keeps. Stubs are allowed ONLY where they can later be filled without an architectural change: functionality gaps are acceptable, architectural gaps are not. (Glossary: `tracer bullet` vs `prototype` in `CONTEXT.md` — GSD ships tracers, never prototypes.)
1. Frame the phase goal as a user story at the top of `PLAN.md`. The user story is sourced from the `**Goal:**` line in ROADMAP.md (set by `mvp-phase`). Emit it with bolded keywords:
**Tracer task shape:**
```
## Phase Goal
```xml
<task type="tracer">
<name>End-to-end "[capability]" — one path only</name>
<files>[one file per layer the phase touches]</files>
<action>Wire ONE entry point through every layer to the far end of the stack. No other call sites, no batching. Real error handling on the single path.</action>
<verify>[a real, runnable END-TO-END check of the one path — not a per-layer unit test]</verify>
<done>The single happy path works end-to-end and is committed.</done>
</task>
```
**As a** [user role], **I want to** [capability], **so that** [outcome].
```
**Core rule (expansion tasks):** after each task a real user can do something they could not before. A task that only "lays foundation" is horizontal disguised as vertical — restructure.
Format rules (Read `~/.claude/gsd-core/references/user-story-template.md`):
- All three slots required. If the ROADMAP `**Goal:**` line is not in user-story format, surface the discrepancy and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story.
- Bold the three keywords (`**As a**`, `**I want to**`, `**so that**`) when emitting to PLAN.md. The ROADMAP form does not use bolded keywords; the PLAN form does.
2. First task: failing end-to-end test for the happy path.
3. Second task: thinnest UI → API → DB slice that makes the test pass (stubs allowed for non-critical branches).
4. Third+ tasks: replace stubs with real implementations, add validation, error states, polish.
**`--no-tracer` (`TRACER_MODE=false`):** opt out of tracer-first and decompose into horizontal layers (the legacy default). Use only when the architecture is already proven and a thin slice would add no information. Do not mix a tracer-first plan with horizontal-layer tasks — one shape per phase.
**Mode is all-or-nothing per phase** (PRD decision Q1). Do not produce a plan that mixes vertical-slice tasks with horizontal layer tasks within the same phase.
**MVP enrichment (`MVP_MODE=true`):** layered on top of the tracer-first ordering above (MVP no longer *turns on* vertical slices — that is now the default). It adds: (1) frame the phase goal as a user story at the top of `PLAN.md`, sourced from the ROADMAP `**Goal:**` line, bolding `**As a**` / `**I want to**` / `**so that**` (Read `~/.claude/gsd-core/references/user-story-template.md`; if the Goal line is not in user-story format, surface it and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story); and (2) **Walking Skeleton mode** (`WALKING_SKELETON=true`, Phase 1 of a new project) — emit `SKELETON.md` from `~/.claude/gsd-core/references/skeleton-template.md` alongside `PLAN.md`. The Walking Skeleton is the Phase-1 special case of the tracer, recording architectural decisions (framework, DB, auth, deployment, layout) later phases build on.
**Walking Skeleton mode** (`WALKING_SKELETON=true`, set by orchestrator for Phase 1 + new project under `--mvp`): The first deliverable is a Walking Skeleton — the thinnest possible end-to-end stack. In addition to `PLAN.md`, produce `SKELETON.md` using the template at `~/.claude/gsd-core/references/skeleton-template.md` (Read it now). `SKELETON.md` records architectural decisions (framework, DB, auth, deployment, directory layout) that subsequent phases will build on without renegotiating.
**Compatibility with TDD detection:** When both `MVP_MODE=true` and `workflow.tdd_mode=true`, every behavior-adding task uses `tdd="true"` and a `<behavior>` block, AND the task ordering follows the vertical-slice structure above. The first task is always a failing end-to-end test.
**TDD composition (`workflow.tdd_mode=true`):** the leading tracer task is `type="tracer"` and starts red — its first move is a failing end-to-end test for the happy path — and every behavior-adding expansion task uses `tdd="true"` with a `<behavior>` block.
See @~/.claude/gsd-core/references/planner-guidance.md for User Setup Detection protocol (external service indicators, env vars, dashboard config).
@@ -542,15 +539,9 @@ Do NOT use for: Deploying (use CLI), creating webhooks (use API), creating datab
When Claude tries CLI/API and gets auth error → creates checkpoint → user authenticates → Claude retries. Auth gates are created dynamically, NOT pre-planned.
## Writing Guidelines
## Writing Guidelines, Anti-Patterns, and Extended Examples
**DO:** Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.
**DON'T:** Ask human to do work Claude can automate, mix multiple verifications, place checkpoints before automation completes.
## Anti-Patterns and Extended Examples
For checkpoint anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
For checkpoint writing guidelines (DO/DON'T), anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
@~/.claude/gsd-core/references/planner-antipatterns.md
</checkpoints>
@@ -768,6 +759,8 @@ At decision points during plan creation, apply structured reasoning:
Decompose phase into tasks. **Think dependencies first, not sequence.**
**Lead with the tracer.** Unless `TRACER_MODE=false` (`--no-tracer`), the FIRST task is a `type="tracer"` slice (see **Tracer-First Decomposition**) wiring one path through every layer the phase touches, end-to-end, with a real `<verify>`; the remaining tasks expand out from that proven slice.
For each task:
1. What does it NEED? (files, types, APIs that must exist)
2. What does it CREATE? (files, types, APIs others might need)

View File

@@ -537,11 +537,11 @@ grep -R -n -E 'probe-[^[:space:]]+\.sh|scripts/.*/tests/probe-.*\.sh' "$PHASE_DI
1. Build the `PROBES` list from explicit PLAN declarations first; include conventional `scripts/*/tests/probe-*.sh` when the phase is a migration/tooling phase or the success criteria mention probes.
2. For every documented probe path, if the file is missing or unreadable, mark `MISSING_PROBE` and set `status: gaps_found`. Do not require the executable bit because probes run through `bash "$probe"`.
3. Run each probe from the built `PROBES` list (declared + conventional) from the repository root:
3. Run each probe from the built `PROBES` list from the repository root:
```bash
for probe in "${PROBES[@]}"; do
timeout 30s bash "$probe"
gsd_run run-with-timeout 30 -- bash "$probe"
done
```

File diff suppressed because it is too large Load Diff

View File

@@ -1,7 +1,7 @@
{
"id": "ai-integration",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "AI design contract",
"description": "AI-SPEC design contract workflow for phases that build AI systems; owns the AI integration command, agents, and workflow.ai_integration_phase activation key.",
"tier": "full",

View File

@@ -34,6 +34,19 @@ the checkpoint entirely and continue planning. Do not raise it with the user.
**If `detected` is `true`:** an external-API integration is in scope. You MUST
produce a **coverage matrix** before the plan is finalized.
**If `detected` is `true` but the phase genuinely integrates no external API**
(the detector is deterministic, not infallible — confirm by re-reading the phase
scope, not by preference): do NOT fabricate a matrix row for a capability that
does not exist. Write a reasoned declaration to `${PHASE_DIR}/COVERAGE.md`
instead:
```markdown
No external API integration: <one-line reason — what the phase touches instead>.
```
The reason is required, exactly like an `OPT-OUT` reason. The seal-time gate
accepts this declaration in place of a matrix.
## Produce the coverage matrix
Enumerate the external API's full **capability surface** — the verb/endpoint/method
@@ -81,7 +94,8 @@ This checkpoint is enforced. At `verify:pre` the `api-coverage.verify-pre` gate
runs `check api-coverage.verify-pre <phase-dir>`:
- If `COVERAGE.md` exists, it is validated — every row needs a valid decision and
every `OPT-OUT` a reason. A malformed/partial matrix **blocks the seal**.
every `OPT-OUT` a reason. A malformed/partial matrix **blocks the seal**. A
reasoned `No external API integration: …` declaration (and no rows) passes.
- If `COVERAGE.md` is absent, the detector runs again over the phase scope. If a
strong external-API-integration signal is found, the seal is **blocked** until a
matrix is produced. If no signal is found, the phase is treated as a non-API

View File

@@ -1,7 +1,7 @@
{
"id": "antigravity",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Antigravity",
"description": "Google Antigravity IDE — nested under ~/.gemini/antigravity; probed across 1.x and 2.x layouts; Gemini hook event dialect; flat skill layout; tier-1 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "assumption-delta",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Assumption-delta architecture checkpoint",
"description": "Rarely-firing advisory checkpoint that triggers when a phase makes something plural, optional, or chosen that used to be singular, required, or derived. Surfaces one identity-model question (promote the new general representation to primary, or add it alongside?) so a silent primary-key drift does not accumulate into a later user-facing bug. Non-blocking; fires only on a detected signal.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "audit",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Audit",
"description": "Open-artifact audit and UAT-gap audit for milestone close gates; exposes `gsd-tools audit-uat` (cross-phase UAT outstanding items) and `gsd-tools audit-open` (structured open-artifact scan across debug, tasks, threads, todos, seeds, UAT, verification, context-questions).",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "augment",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Augment Code",
"description": "Augment Code CLI — commands + nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.",
"tier": "core",

View File

@@ -0,0 +1,46 @@
{
"id": "broken-windows",
"role": "feature",
"version": "1.8.0",
"title": "Broken-windows ledger",
"description": "Cross-phase defect register accumulating stubs, TODOs, skipped tests, unrun verifies, and unmet truths into .planning/WINDOWS.md. Blocks /gsd-ship while any window is open unless explicitly waived with a recorded reason. Operationalizes GSD's no-defer discipline as a tracked, enforced artifact (issue #1950).",
"tier": "full",
"requires": [],
"engines": {
"gsd": ">=1.7.0"
},
"runtimeCompat": {
"supported": [
"*"
],
"unsupported": []
},
"skills": [],
"agents": [],
"hooks": [],
"config": {
"workflow.windows_enforce": {
"type": "boolean",
"default": false,
"description": "Enable the blocking ship:pre gate for the broken-windows ledger. When true (opt-in), /gsd-ship blocks while .planning/WINDOWS.md has any open entry. When false (default), windows are still tracked (the executor and verifier still populate WINDOWS.md via gsd-tools windows append) but ship does not block — teams can adopt tracking before enforcement. Issue #1950."
}
},
"steps": [],
"contributions": [],
"gates": [
{
"point": "ship:pre",
"check": {
"predicate": {
"kind": "artifact-frontmatter-equals",
"artifact": "WINDOWS.md",
"field": "open_count",
"equals": 0
}
},
"when": "workflow.windows_enforce",
"blocking": true,
"onError": "halt"
}
]
}

View File

@@ -1,7 +1,7 @@
{
"id": "claude-orchestration",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Claude orchestration (Workflow backend)",
"description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.",
"tier": "full",
@@ -25,7 +25,8 @@
"router": "routeClaudeOrchestrationCommand",
"subcommands": [
"detect-backend",
"emit-workflow"
"emit-workflow",
"resolve-wave-dispatch"
]
}
],
@@ -55,10 +56,10 @@
"steps": [],
"contributions": [
{
"point": "execute:wave:post",
"point": "execute:wave:pre",
"into": "executor",
"fragment": {
"path": "fragments/execute-wave-post.md"
"path": "fragments/execute-wave-pre.md"
},
"produces": [],
"consumes": [

View File

@@ -1,64 +0,0 @@
# Claude orchestration — Workflow execution backend (BETA)
> Injected at `execute:wave:post` `into: executor` only when
> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.
## When this contribution is active
The Claude orchestration capability is **default-off and BETA**. It activates only
when ALL of the following hold:
1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND
2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent
SDK-specific), AND
3. `claude_orchestration.execution_backend` resolves to `workflow` — either
explicitly, or via `auto` — **and** the Agent SDK version is
`>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK
floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release
or older SDK never activates the preview backend).
Detection is fail-closed: any miss degrades to **inline, manual, one-agent-per-
message dispatch** — exactly today's behaviour. On a non-Claude runtime this
contribution is a no-op.
## What the executor does when the Workflow backend is active
Instead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor,
isolation=worktree, run_in_background=true)` per message (which on Claude Code
cannot nest further subagents — #853 — and so degrades to sequential inline
execution), execute-phase **emits a generated Workflow script** and lets the main
loop orchestrate it:
- **waves → one or more sequential `parallel()` barriers** — each wave is a
barrier group; when plans within a wave share `files_modified`, they are split
into separate sequential stages within that wave's barrier (the next wave
still waits for the previous wave to complete).
- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**
— the SAME executor agent and worktree isolation the inline path uses, so the
produced `SUMMARY.md` and commits are identical.
- **`files_modified` overlap → separate sequential stages** — two plans that
touch the same file are placed in different stages within the wave (the same
overlap rule execute-phase already applies inline).
- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase
resumes without re-running completed plans.
- **`budget(tokens)`** — a shared token pool across the whole phase when the
orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a
function parameter, not a config key; the orchestrator decides the budget).
The emitter is a pure function exposed through the capability command surface:
`gsd-tools claude-orchestration emit-workflow --waves <manifest.json> --run-id <id>
[--phase-dir <dir>] [--budget <n>]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript`
directly). It maps the phase's wave/plan manifest to the Workflow script string
and never invokes the Workflow tool itself; the orchestrator runs the emitted
script. Detection is resolved by the orchestrator calling the pure
`detectWorkflowBackend` with the LIVE host descriptor (the CLI
`gsd-tools claude-orchestration detect-backend` is a simulation harness that
assumes a capable host unless `--no-nested-dispatch` is passed — it does not probe
the real runtime; the orchestrator supplies the real descriptor).
## Fallback contract
If detection resolves to `inline` (tool absent, SDK too old, runtime not Claude,
or the capability disabled), execute-phase MUST proceed with the standard inline
wave dispatch. The executor MUST NOT assume parallelism, a shared budget, or
resume-from-run-id semantics in that mode.

View File

@@ -0,0 +1,167 @@
# Claude orchestration — Workflow execution backend (BETA)
> Injected at `execute:wave:pre` `into: executor` only when
> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.
## When this contribution is active
The Claude orchestration capability is **default-off and BETA**. It activates only
when ALL of the following hold:
1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND
2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent
SDK-specific), AND
3. `claude_orchestration.execution_backend` resolves to `workflow` — either
explicitly, or via `auto` — **and** the Agent SDK version is
`>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK
floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release
or older SDK never activates the preview backend).
Detection is fail-closed: any miss degrades to **inline, manual, one-agent-per-
message dispatch** — exactly today's behaviour. On a non-Claude runtime this
contribution is a no-op.
## Why `execute:wave:pre` (not `execute:wave:post`)
This is a **dispatch-backend selector** — it decides HOW a wave's executor agents
are spawned. That decision has to be made BEFORE the wave's `Agent()` calls in
`execute-phase.md` step 3, not after the wave has already finished (#2285). The
capability previously registered at `execute:wave:post`, which fires only after
worktree merge/post-merge tests/tracking updates — by then the wave was already
dispatched inline, so the contribution was structurally unable to change how
dispatch happened. This fragment is injected at the point that actually precedes
dispatch.
## What the orchestrator does when the Workflow backend is active
Before spawning executor agents for the current wave (execute-phase.md step 3),
resolve the dispatch backend through the single composed CLI seam:
```bash
gsd-tools claude-orchestration resolve-wave-dispatch \
--waves "$WAVE_MANIFEST_PATH" --run-id "$PHASE_RUN_ID" \
--runtime "$RUNTIME" \
${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"} \
--phase-dir "$PHASE_DIR" --raw
```
This composes `detectWorkflowBackend` (the gate ladder above) with
`emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure
function backing it is `resolveWaveDispatch` in
`gsd-core/bin/lib/claude-orchestration.cjs`. Response shape:
`{ backend: 'inline'|'workflow', reason, script?, summary? }`.
### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`, `$AGENT_SDK_VERSION`)
These are NOT pre-existing execute-phase.md variables — the orchestrator builds
them at this step, from data it already has in-context from `discover_and_group_plans`
(the `PLAN_INDEX` JSON) and step 2.5 (the per-plan `USE_WORKTREES_FOR_PLAN` decision):
1. **`$PHASE_DIR`** — reuse `{phase_dir}` from the `INIT` bundle (already loaded
in the `initialize` step). No new value needed.
2. **`$PHASE_RUN_ID`** — a stable identifier for THIS phase-execution attempt, so
`resumeFromRunId` can resume an interrupted run without re-dispatching plans
the Workflow tool already completed. Construct it deterministically —
`execute-{phase_number}-{phase_slug}` — from `INIT`'s `phase_number`/`phase_slug`
(both are already validated identifiers used elsewhere in this workflow, so
they satisfy `emitWorkflowScript`'s `isScriptableIdentifier` check). Do NOT
mint a new random id per wave — the SAME `$PHASE_RUN_ID` is reused for every
wave in the phase so the Workflow tool can correctly track cross-wave resume
state.
3. **`$WAVE_MANIFEST_PATH`** — a fresh temp file for THIS wave's manifest (one
wave = one `waves` array with a single entry, matching the wave-by-wave
dispatch loop; do not batch multiple waves into one manifest — waves are
dispatched in wave order, not all at once):
```bash
WAVE_MANIFEST_PATH=$(mktemp "${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX") && mv "$WAVE_MANIFEST_PATH" "$WAVE_MANIFEST_PATH.json" && WAVE_MANIFEST_PATH="$WAVE_MANIFEST_PATH.json"
```
Then **use the Write tool** (not a bash/jq pipeline — the orchestrator already
has every field parsed in-context) to write the manifest JSON to
`$WAVE_MANIFEST_PATH`:
```json
{
"waves": [
{
"id": "wave-{N}",
"plans": [
{
"id": "{plan_id}",
"brief": "{the SAME <objective>...<success_criteria> prompt block step 3 builds for this plan's inline Agent() call}",
"files_modified": ["{from PLAN_INDEX.plans[].files_modified for this plan}"],
"use_worktree": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}
}
]
}
]
}
```
- **`id`** — the plan id from `PLAN_INDEX`, e.g. `"01-01"`.
- **`brief`** — MUST carry the same task content as step 3's inline `Agent()`
prompt (the `<objective>`/`<execution_context>`/`<files_to_read>`/
`<success_criteria>` block, with `{plan_number}`/`{phase_number}`/
`{phase_name}` substituted) — a short summary here would NOT reproduce
step 3's behavior and would violate the "identical artifacts" contract.
- **`files_modified`** — copy verbatim from the plan's `PLAN_INDEX` entry.
- **`use_worktree`** — `true` for every plan UNLESS step 2.5's per-plan
worktree gate (`execute-phase/steps/per-plan-worktree-gate.md`) set
`USE_WORKTREES_FOR_PLAN=false` for that plan (submodule-touching plan, or
project-level `USE_WORKTREES=false`) — in which case pass `false` here so
`emitWorkflowScript` omits `isolation: "worktree"` for that plan (#2772 /
#2285 finding 1). **Never** hardcode `true` — that would force worktree
isolation on a plan the inline path explicitly keeps out of worktrees.
4. **`$AGENT_SDK_VERSION`** — see below; OMIT when unknown (fails closed).
**Agent SDK version:** the orchestrator has no scriptable (bash-computable) way
to introspect the live Agent SDK version. When it can determine the version
(e.g. from a host-exposed value it can read directly), pass
`--agent-sdk-version`. When it cannot, OMIT the flag — `resolveWaveDispatch`'s
gate 5 (`agent_sdk_version_unknown`) then fails closed to `inline` by design;
this is not a bug, it is the same fail-closed posture documented above applied
to a real absence of information.
**If `backend == "workflow"`:** run the emitted `script` via the Workflow tool
for THIS wave instead of the per-message `Agent()` loop in step 3. The script
composes the SAME `gsd-executor` agent type the inline path uses, with
worktree isolation applied PER PLAN from the manifest's `use_worktree` field
(see `emitWorkflowScript`):
- **waves → one or more sequential `parallel()` barriers** — each wave is a
barrier group; when plans within a wave share `files_modified`, they are split
into separate sequential stages within that wave's barrier.
- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**
when `use_worktree` is not `false`, or `agent(brief, { agentType: 'gsd-executor' })`
(no isolation) when it is — so the produced `SUMMARY.md` and commits are
identical to inline dispatch, INCLUDING the inline path's submodule safety
gate (#2772 / #2285 finding 1).
- **`files_modified` overlap → separate sequential stages** — the same overlap
rule execute-phase already applies inline (step 1 of the wave loop).
- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase
resumes without re-running completed plans.
The orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup,
post-merge gate, tracking update) exactly as it does for inline dispatch — the
Workflow backend only replaces HOW agents are spawned for this wave, not what
happens after they return.
**If `backend == "inline"`** (any gate miss, or `resolve-wave-dispatch` itself
unavailable/erroring): proceed to step 3's standard per-message `Agent()`
dispatch — the default, byte-identical-to-today path. `onError: skip` on this
contribution means a `resolve-wave-dispatch` command failure is treated exactly
like an `inline` result, never as a fatal wave error.
## Fallback contract
Detection is fail-closed end-to-end: capability disabled, non-Claude runtime,
`execution_backend:"inline"`, missing/incapable host descriptor, unknown or
below-floor Agent SDK version, or an `emitWorkflowScript` failure on a malformed
wave manifest — ANY of these degrades to `backend:"inline"` and execute-phase's
standard inline dispatch (step 3) runs unmodified. The Workflow backend never
partially activates; the executor MUST NOT assume parallelism, a shared budget,
or resume-from-run-id semantics when `backend == "inline"`.

View File

@@ -1,7 +1,7 @@
{
"id": "claude",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Claude Code",
"description": "Anthropic Claude Code — primary development runtime; tier-1 support with full hook surface and skills-based global install.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "cline",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Cline",
"description": "Cline (VS Code extension) — global-only nested-skill layout; cline-rules hook surface (.clinerules); no hook events emitted; tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "code-review",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Code review",
"description": "Source-file code review and review-fix workflow support for completed execution work.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "codebuddy",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "CodeBuddy",
"description": "CodeBuddy (Tencent) — converted commands + skills artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "codex",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "OpenAI Codex CLI",
"description": "OpenAI Codex CLI — shell-var command style; per-agent sandbox tiers; config.toml + hooks.json hook surface; tier-1 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "copilot",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "GitHub Copilot",
"description": "GitHub Copilot (VS Code) — markdown config format; copilot-inline hook surface; no hook events emitted; flat skill nesting (unconfirmed recursive loader); tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "cursor",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Cursor",
"description": "Cursor IDE — skills + converted commands artifact layout; hooks.json surface; Claude hook event dialect; recursive skill loader (flat nesting); tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "drift",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Drift detection gates",
"description": "Drift detection gates for the planning loop. At execute:wave:post: a blocking schema drift gate (detects schema files changed without a database push) and a non-blocking codebase drift gate (detects structural additions not reflected in STRUCTURE.md). At plan:pre: a non-blocking, warn-only codebase drift gate (gated on workflow.plan_drift_precheck) that flags a stale codebase map before planning, so plans are authored against a fresh STRUCTURE.md instead of discovering drift mid-execution.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "external-job",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Async external-job scheduler adapter",
"description": "Default-off producer of the async external-job manifest (#1164). At execute:wave:post an executor can externalize long-running compute (SLURM first, scheduler-pluggable), commit a .planning/async-jobs/<job>.json manifest, defer SUMMARY.md, and return external_job_waiting. The core loop (#1165) consumes the manifest; this capability is the only thing that writes it. NOTE on contribution point: #1164 specifies execute:wave:pre, but execute-phase.md only dispatches execute:wave:post today (wave:pre is declared in the loop host contract but not rendered); wiring wave:pre dispatch is a core-loop change #1164 explicitly puts out of scope, so this capability registers at wave:post and the executor honors the runtime_budget classification guidance before running any tagged task. The adapter (scripts/slurm-adapter.cjs) reads external_job.submit_timeout_ms / poll_timeout_ms / artifact_dir through the canonical capability-config seam (env override > config > registry default).",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "gap-analysis",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Post-planning gap analysis",
"description": "Proactive, non-blocking post-planning coverage report. After all PLAN.md files are generated, cross-references every REQ-ID and D-ID from REQUIREMENTS.md and CONTEXT.md against plan bodies. Emits a Source | Item | Status table. Does not block phase advancement.",
"tier": "standard",

View File

@@ -1,7 +1,7 @@
{
"id": "graphify",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Knowledge graph",
"description": "Build, query, and inspect the project knowledge graph in `.planning/graphs/`; exposes graphify CLI subcommands (build, query, status, diff) and the /gsd-graphify skill.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "hermes",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Hermes Agent",
"description": "Hermes Agent (NousResearch) — skills nest under skills/gsd/ category bucket; nested skill layout; settings-json hook surface; Claude hook event dialect; tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "intel",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Codebase intelligence",
"description": "Code-intelligence store for codebase querying, diff, snapshot, and API-surface extraction; exposes `gsd-tools intel` subcommands (query, status, update, diff, snapshot, patch-meta, validate, extract-exports, api-surface) and backs `/gsd-map-codebase` and `gsd-intel-updater`.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "kilo",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Kilo Code",
"description": "Kilo Code — XDG-based config dir; global skills at ~/.kilo/skills (separate from XDG config); flat command/ + skills artifact layout; no lifecycle hook registration; tier-2 support.",
"tier": "core",
@@ -101,8 +101,7 @@
"file": "gsd-core.js",
"source": ".kilo/plugins/gsd-core.js"
},
"skipUpdateBannerCommand": true,
"skipSharedHooksInstall": true
"skipUpdateBannerCommand": true
}
}
}

View File

@@ -1,7 +1,7 @@
{
"id": "kimi",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Kimi CLI",
"description": "Kimi CLI (Moonshot AI) — generic agents root at ~/.config/agents; skills + kimi-agents artifact layout; native config.toml [[hooks]] bus at ~/.kimi/config.toml; background dispatch; tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "mempalace",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "MemPalace memory",
"description": "Cross-session, cross-project memory: deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries, via the MemPalace MCP server and CLI.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "nyquist",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Nyquist validation",
"description": "Validation coverage audit that maps executed work back to tests and manual-only evidence.",
"tier": "full",

View File

@@ -1,9 +1,9 @@
{
"id": "opencode",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "OpenCode",
"description": "OpenCode — XDG-based config dir; flat command/ + skills artifact layout; settings-json config format; no lifecycle hook registration; tier-2 support.",
"description": "OpenCode — XDG-based config dir; flat commands/ + skills artifact layout; settings-json config format; no lifecycle hook registration; tier-2 support.",
"tier": "core",
"requires": [],
"engines": {
@@ -25,7 +25,7 @@
"global": [
{
"kind": "commands",
"destSubpath": "command",
"destSubpath": "commands",
"prefix": "gsd-",
"nesting": "flat",
"recursive": false,
@@ -43,7 +43,7 @@
"local": [
{
"kind": "commands",
"destSubpath": "command",
"destSubpath": "commands",
"prefix": "gsd-",
"nesting": "flat",
"recursive": false,
@@ -88,7 +88,7 @@
"hostBehaviors": {
"reapplyCommand": "/gsd-update --reapply",
"attributionConfigResolver": "opencode",
"flatCommandDir": "command",
"flatCommandDir": "commands",
"combinedFamilyInstall": true,
"frontmatterDialect": "opencode",
"nativePlugin": {

View File

@@ -1,7 +1,7 @@
{
"id": "pattern-mapper",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Pattern mapping",
"description": "Optional codebase-pattern mapping before planning; owns the pattern mapper agent and workflow.pattern_mapper activation key.",
"tier": "full",

View File

@@ -1,9 +1,9 @@
{
"id": "pi",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "pi",
"description": "pi (pi.dev) — bun-runtime programmatic-CLI; TS ExtensionAPI (registerCommand/registerTool/registerProvider/pi.on); single native-extension file at ~/.pi/agent/extensions/gsd.cjs; no shared-settings hook surface; tier-2 support.",
"description": "pi (pi.dev) — bun-runtime programmatic-CLI; TS ExtensionAPI (registerCommand/registerTool/registerProvider/pi.on); single native-extension file at ~/.pi/agent/extensions/gsd.js (.js, not .cjs — pi's extension auto-discovery accepts only .ts/.js, #2470); no shared-settings hook surface; tier-2 support.",
"tier": "core",
"requires": [],
"engines": {
@@ -51,7 +51,7 @@
"hostBehaviors": {
"nativePlugin": {
"dir": "extensions",
"file": "gsd.cjs",
"file": "gsd.js",
"source": "pi/gsd.cjs"
},
"pluginOnlyInstall": true

View File

@@ -1,7 +1,7 @@
{
"id": "profile-pipeline",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Developer profiling pipeline",
"description": "Developer behavioral profiling from Claude Code session history; scans session JSONL files, extracts and samples user messages, and generates profile artifacts (USER-PROFILE.md, dev-preferences.md, CLAUDE.md sections). Exposes eight `gsd-tools` commands: scan-sessions, extract-messages, profile-sample (pipeline phase) and write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md (output phase). Backs the /gsd-profile-user skill and gsd-user-profiler agent.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "qwen",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Qwen Code",
"description": "Qwen Code (Alibaba) — nested-skill artifact layout; settings-json hook surface; Claude hook event dialect; tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "research",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Phase research",
"description": "Optional phase research before planning; owns the phase researcher agent and workflow.research activation key.",
"tier": "standard",

View File

@@ -1,7 +1,7 @@
{
"id": "schema-gate",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Schema push detection gate",
"description": "Detects ORM schema-relevant files in the phase scope during planning and injects a mandatory [BLOCKING] schema push task into the plan. Prevents false-positive verification where build/types pass because TypeScript types come from config, not the live database.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "security",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Security enforcement",
"description": "Threat mitigation verification and ship-time security blocking for phases with security enforcement enabled.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "tdd",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "Test-driven development",
"description": "Injects TDD heuristics into the planner and enforces RED/GREEN gate compliance on type:tdd plans after execution. Owns workflow.tdd_mode; the --tdd CLI flag is the ephemeral override.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "trae",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Trae IDE",
"description": "Trae IDE — nested-skill artifact layout; no hook surface (profile-marker-only config); tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "ui",
"role": "feature",
"version": "1.7.0",
"version": "1.8.0",
"title": "UI design contracts",
"description": "UI-SPEC design contract + retrospective UI audit for frontend phases.",
"tier": "full",

View File

@@ -1,7 +1,7 @@
{
"id": "vscode",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "VS Code",
"description": "VS Code — Marketplace/VSIX extension; no file-projected config directory; IDE-profile reference host (active vscode.lm model, engine-owned hook bus, sandboxed globalState/workspaceState stateIO).",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "windsurf",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "Windsurf",
"description": "Windsurf (Codeium) — workspace workflow artifact layout for slash commands; Cascade native hooks.json blocking hook bus (pre_write_code, pre_run_command); tier-2 support.",
"tier": "core",

View File

@@ -1,7 +1,7 @@
{
"id": "zcode",
"role": "runtime",
"version": "1.7.0",
"version": "1.8.0",
"title": "ZCode",
"description": "ZCode (Z.ai) — desktop Agentic Development Environment for GLM-5.2; Claude-shaped nested skills at ~/.zcode/skills/<name>/SKILL.md, slash commands, named subagents, native MCP; declarative plugin surface; profile-marker install; tier-2 community support.",
"tier": "core",

View File

@@ -28,7 +28,7 @@ Flow: Select Framework → Research Docs → Research Domain → Design Eval Str
</execution_context>
<context>
Phase number: $ARGUMENTS — optional, auto-detects next unplanned phase if omitted.
Phase number: $ARGUMENTS — optional; when omitted, the orchestrating workflow reads ROADMAP.md and selects the next unplanned phase. This is not a `gsd-tools.cjs` CLI feature — the CLI's phase-lookup primitives require an explicit phase number.
</context>
<process>

View File

@@ -64,12 +64,16 @@ On any error or timeout, stop and let the phase continue -- capture is best-effo
# One-time: declare the GSD room taxonomy so detect_room() recognizes these folders
mkdir -p "$STAGE"
[ -f "$STAGE/mempalace.yaml" ] || cat > "$STAGE/mempalace.yaml" <<'YAML'
# Each entry MUST be a dict with a `name` key (the miner's detect_room()
# indexes room["name"] — a bare-string list crashes _mine_impl with
# TypeError: string indices must be integers, not 'str'). Optional fields:
# `description`, `keywords` (matched against folder-path segments).
rooms:
- decisions
- planning
- milestones
- problems
- general
- name: decisions
- name: planning
- name: milestones
- name: problems
- name: general
YAML
# Suppress MemPalace cache artifacts written into the scanned tree
[ -f "$STAGE/.gitignore" ] || echo "mempalace_embedder.json" > "$STAGE/.gitignore"

View File

@@ -1,7 +1,7 @@
---
name: gsd:new-milestone
description: Start a new milestone cycle — update PROJECT.md and route to requirements
argument-hint: "[milestone name, e.g., 'v1.1 Notifications']"
argument-hint: "[milestone name, e.g., 'v1.1 Notifications'] [--ws <name>]"
allowed-tools:
- Read
- Write

View File

@@ -1,7 +1,7 @@
---
name: gsd:plan-phase
description: Create detailed phase plan (PLAN.md) with verification loop
argument-hint: "[phase] [--auto] [--research] [--skip-research] [--research-phase <N>] [--view] [--gaps] [--skip-verify] [--prd <file>] [--ingest <path-or-glob>] [--ingest-format <auto|nygard|madr|narrative>] [--reviews] [--text] [--tdd] [--mvp]"
argument-hint: "[phase] [--auto] [--research] [--skip-research] [--research-phase <N>] [--view] [--gaps] [--skip-verify] [--prd <file>] [--ingest <path-or-glob>] [--ingest-format <auto|nygard|madr|narrative>] [--reviews] [--text] [--tdd] [--mvp] [--no-tracer] [--no-reversibility-gates]"
effort: max
allowed-tools:
- Read
@@ -40,7 +40,7 @@ Create executable phase prompts (PLAN.md files) for a roadmap phase with integra
</runtime_note>
<context>
Phase number: $ARGUMENTS (optional — auto-detects next unplanned phase if omitted)
Phase number: $ARGUMENTS (optional — when omitted, the orchestrating workflow reads ROADMAP.md and selects the next unplanned phase; `gsd-tools.cjs` itself has no auto-detect feature and requires an explicit phase number)
**Flags:**
- `--research` — Force re-research even if RESEARCH.md exists
@@ -52,7 +52,9 @@ Phase number: $ARGUMENTS (optional — auto-detects next unplanned phase if omit
- `--ingest-format <auto|nygard|madr|narrative>` — Optional ADR parser format override (`auto` default).
- `--reviews` — Replan incorporating cross-AI review feedback from REVIEWS.md (produced by `/gsd:review`)
- `--text` — Use plain-text numbered lists instead of TUI menus (required for `/rc` remote sessions)
- `--mvp` — Vertical MVP mode. Planner organizes tasks as feature slices (UI→API→DB) instead of horizontal layers. On Phase 1 of a new project, also emits `SKELETON.md` (Walking Skeleton). Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md.
- `--mvp` — MVP enrichment on top of the default tracer-first ordering: frames the phase goal as a user story and, on Phase 1 of a new project, also emits `SKELETON.md` (Walking Skeleton). Vertical slicing itself is now the default (see `--no-tracer`); `--mvp` no longer *turns it on*. Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md.
- `--no-tracer` — Opt out of the default **tracer-first** decomposition and plan horizontal layers (the legacy default). By default every plan LEADS with one production-quality end-to-end `tracer` slice that is verified before any expansion task.
- `--no-reversibility-gates` — Suppress the human checkpoint that a **one-way-door** decision normally earns, for runs you intend to leave unattended. By default a decision rated `one-way` (undo needs a migration, breaks a published contract, or is impossible) gets a `checkpoint:decision` before the task implementing it. Ratings are still recorded on tasks and `costly` items still flagged — the flag changes what stops the run, not what the plan remembers.
Normalize phase input in step 2 before any directory lookups.
</context>

View File

@@ -1,7 +1,7 @@
---
name: gsd:plan-review-convergence
description: "Cross-AI plan convergence - replan until review concerns are resolved."
argument-hint: "<phase> [--codex] [--gemini] [--claude] [--opencode] [--ollama] [--lm-studio] [--llama-cpp] [--text] [--ws <name>] [--all] [--max-cycles N]"
argument-hint: "<phase> [--codex] [--gemini] [--claude] [--opencode] [--ollama] [--lm-studio] [--llama-cpp] [--agy] [--text] [--ws <name>] [--all] [--max-cycles N]"
allowed-tools:
- Read
- Write
@@ -40,8 +40,9 @@ Replaces gsd-plan-phase's internal gsd-plan-checker with external AI reviewers (
Phase number: extracted from $ARGUMENTS (required)
**Flags:**
- `--codex` — Use Codex CLI as reviewer (default if no reviewer specified)
- `--codex` — Use Codex CLI as reviewer (default if no reviewer flag given AND `review.default_reviewers` is unset; otherwise `review.default_reviewers` wins per ADR-0011 — #2315)
- `--gemini` — Use Gemini CLI as reviewer
- `--agy` / `--antigravity` — Use Antigravity CLI as reviewer (successor to the discontinued Gemini CLI)
- `--claude` — Use Claude CLI as reviewer (separate session)
- `--opencode` — Use OpenCode as reviewer
- `--ollama` — Use local Ollama server as reviewer (OpenAI-compatible, default host `http://localhost:11434`; configure model via `review.models.ollama`)

View File

@@ -215,7 +215,8 @@ GSD uses a multi-agent architecture where thin orchestrators (workflow files) sp
- Fresh 200K context window per plan
- Follows XML task instructions precisely
- Atomic git commit per completed task
- Handles checkpoint types: auto, human-verify, decision, human-action
- Handles task types: auto, tracer, checkpoint (human-verify, decision, human-action)
- Tracer feedback gate: after a `tracer` slice, verifies it end-to-end before expansion tasks — autonomous runs halt on failure; interactive runs emit a human-verify checkpoint
- Reports deviations from plan in SUMMARY.md
- Invokes node repair on verification failure
@@ -399,6 +400,13 @@ runs its default whole-repo scan.
- Tracks hypotheses, evidence, and eliminated theories
- State persists across context resets
- Requires human verification before marking resolved
- Runs a multi-signal fix-acceptance guardrail (mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm) before accepting a fix; degrades gracefully when Stryker or a test suite is absent
- Ranks suspect code by Ochiai suspiciousness from test pass/fail coverage (spectrum-based fault localization) before forming hypotheses; skips cleanly when no coverage exists
- Branches root-cause analysis across ≥2 Ishikawa categories and applies an AND-gate check before committing root_cause (guards against 5-Whys single-cause bias); root_cause may hold a set when the AND-gate fires
- Classifies each failure as Bohrbug / Heisenbug-Mandelbug / Concurrency at Phase 1.75 and routes the investigation technique accordingly (routes Bohrbugs to SBFL+bisect, Heisenbugs to record-replay/stability with SBFL skipped, Concurrency to the atomicity/order/deadlock checklist)
- Hardens regression tests via PBT shrinking (minimized counterexample as the seed), explicit oracle classification (specified/derived/metamorphic/implicit), and boundary neighbors around the fixed equivalence class
- Emits a blameless-postmortem Prevention block at resolution (branching 5-Whys, why-wasn't-this-caught, a concrete recurrence guard) and records `why_not_caught` + `recurrence_guard` in the knowledge base so the same bug class is prevented, not just fixed
- Recalls prior resolved sessions semantically via MemPalace at Phase 0 (top-k meaning-similar), catching same-root-cause/different-wording cases keyword overlap misses; falls back to keyword matching when MemPalace is absent
- Appends to persistent knowledge base on resolution
- Consults knowledge base on new sessions

View File

@@ -305,6 +305,8 @@ See [`docs/INVENTORY.md`](INVENTORY.md#hooks) for the authoritative hook roster.
CJS command family routers dispatch through `CommandRoutingHub`. The hub owns the no-throw pure-result contract (`hub.dispatch()` catches internal exceptions and returns `{ ok: false, kind, ...typedPayload }`) and the closed runtime error taxonomy (`UnknownCommand`, `InvalidArgs`, `HandlerRefusal`, `HandlerFailure`). Router adapters remain thin CLI translators — they build the hub, call `dispatch`, then map the Result to `output()`/`error()` calls. The runtime is single-path (no dual-runtime mode selection). See `docs/adr/0174-retire-gsd-sdk-package-boundary.md`.
> **Planned (ADR-2346 / epic #2345):** the `runCommand` 73-case switch is being dissolved into a two-layer dispatch — families via the `commandFamilies` registry (ADR-959 mechanism, completed) and single-purpose leaf verbs via a table filling the prepared `_dispatchNonFamily` seam — collapsing `runCommand` to a ~15-line dispatcher. Behavior-preserving; tracked phase-by-phase under epic #2345. The current-state description above holds until each phase lands.
### Capability Command Dispatch (`gsd-core/bin/gsd-tools.cjs`, ADR-1244 D7)
Command families declared by capabilities (`commands: [{ family, module, router }]`) are dispatched from the registry rather than a hardcoded switch. The `runCommand` default arm tries, in order:
@@ -832,10 +834,10 @@ The migration-specific ownership and source snapshots live in
| Runtime | Global root | Local root | Invocation surface | Agent surface | Config and hooks |
| --- | --- | --- | --- | --- | --- |
| Claude Code | `~/.claude` | `./.claude` | Global `skills/gsd-*/SKILL.md` (flat, #924); local `commands/gsd/*.md` | `agents/gsd-*.md` | `settings.json` hook and statusLine entries |
| OpenCode | `~/.config/opencode` | `./.opencode` | `command/gsd-*.md` | `agents/gsd-*.md` | `opencode.json` or `opencode.jsonc`; no GSD hooks |
| OpenCode | `~/.config/opencode` | `./.opencode` | `commands/gsd-*.md` | `agents/gsd-*.md` | `opencode.json` or `opencode.jsonc`; no GSD hooks |
| Kilo | `~/.config/kilo` | `./.kilo` | `command/gsd-*.md` | `agents/gsd-*.md` | `kilo.json` or `kilo.jsonc`; no GSD hooks |
| Kimi CLI | First-existing generic root: `~/.config/agents` recommended, then `~/.agents` when `~/.agents/skills` exists and `~/.config/agents/skills` does not | Deferred and guarded | `skills/gsd-*/SKILL.md` (flat) invoked as `/skill:gsd-*` | `agents/gsd.yaml`, `agents/gsd.md`, and `agents/subagents/gsd-*` YAML/prompt pairs | Explicit `kimi --agent-file <configRoot>/agents/gsd.yaml`; no GSD hooks or statusline |
| Codex | `~/.codex` | `./.codex` | `skills/gsd-*/SKILL.md` (flat) | `agents/` source markdown plus per-agent TOML | `config.toml` `[agents.gsd-*]`, `[features].hooks` (canonical; legacy alias `codex_hooks` is recognized and migrated forward on reinstall, #3566), and hook tables |
| Codex | `~/.codex` | `./.codex` | `skills/gsd-*/SKILL.md` (flat) | `agents/` source markdown plus per-agent TOML (Codex auto-discovers each `agents/gsd-*.toml`; this is the sole canonical role registration, #2406) | `config.toml` bare `[agents]` dispatch-tuning scalar (`max_depth`, no per-role `[agents.gsd-*]` tables), `[features].hooks` (canonical; legacy alias `codex_hooks` is recognized and migrated forward on reinstall, #3566), and hook tables |
| GitHub Copilot | `~/.copilot` | `./.github` | `skills/gsd-*/SKILL.md` (flat), `copilot-instructions.md`, and `AGENTS.md` (repo root, local) | `.agent.md` files | Self-contained `sessionStart` hook (`hooks/gsd-session.json`, inline `command` type); no statusline |
| Antigravity | auto-detected: `~/.gemini/antigravity`, `~/.gemini/antigravity-ide`, or `~/.gemini/antigravity-cli` | `./.agent` | `skills/gsd-*/SKILL.md` (flat, #1614) | `agents/gsd-*.md` | Gemini-style `settings.json` hook entries when installed by GSD |
| Cursor | `~/.cursor` | `./.cursor` | `skills/gsd-*/SKILL.md` (flat) | `agents/gsd-*.md` | Rule references under `rules/`; `hooks.json` with sessionStart context injection and postToolUse STATE.md monitor (#777) |

View File

@@ -193,7 +193,7 @@ Research, plan, and verify a phase.
| Argument | Required | Description |
|----------|----------|-------------|
| `N` | No | Phase number (defaults to next unplanned phase) |
| `N` | No | Phase number (if omitted, the orchestrating workflow reads ROADMAP.md and targets the next unplanned phase — not a `gsd-tools.cjs` CLI feature) |
| Flag | Description |
|------|-------------|
@@ -211,8 +211,10 @@ Research, plan, and verify a phase.
| `--validate` | Run state validation before planning begins |
| `--bounce` | Run external plan bounce validation after planning (uses `workflow.plan_bounce_script`) |
| `--skip-bounce` | Skip plan bounce even if enabled in config |
| `--mvp` | Vertical MVP mode — planner organizes tasks as feature slices (UI→API→DB) instead of horizontal layers. On Phase 1 of a new project with no prior phase summaries, also emits `SKELETON.md` (Walking Skeleton). Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md, which applies `--mvp` automatically without the flag. |
| `--tdd` | TDD mode — planner applies `type: tdd` to eligible behavior-adding tasks so each begins with a failing test. Composable with `--mvp`: `--mvp --tdd` produces vertical slices where every behavior-adding task starts red-green. |
| `--mvp` | MVP enrichment on top of the default tracer-first ordering — frames the phase goal as a user story and, on Phase 1 of a new project with no prior phase summaries, also emits `SKELETON.md` (Walking Skeleton). Vertical slicing is now the default (see `--no-tracer`); `--mvp` no longer turns it on. Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md, which applies `--mvp` automatically without the flag. |
| `--no-tracer` | Opt out of the default **tracer-first** decomposition and plan horizontal layers (the legacy default). By default every plan leads with one production-quality end-to-end `tracer` slice that the executor verifies before any expansion task. |
| `--no-reversibility-gates` | Suppress the human checkpoint that a **one-way-door** decision normally earns, for runs you intend to leave unattended. By default a decision rated `one-way` — undoing it needs a data migration, breaks a published contract, or is impossible — gets a `checkpoint:decision` inserted before the task that implements it. Ratings are still recorded on tasks and `costly` decisions are still flagged, so the flag changes what stops the run, not what the plan remembers. |
| `--tdd` | TDD mode — planner applies `type: tdd` to eligible behavior-adding tasks so each begins with a failing test. Composable with `--mvp`: `--mvp --tdd` produces vertical slices where every behavior-adding task starts red-green. The leading `tracer` task also starts red under `--tdd`. |
| `--granularity <coarse\|standard\|fine>` | Override the planning granularity for this invocation, ignoring config. Valid values: `coarse`, `standard`, `fine`. Takes precedence over `granularities.planning`, top-level `granularity`, and `planning.granularity` config. |
**Prerequisites:** `.planning/ROADMAP.md` exists
@@ -385,6 +387,11 @@ Create PR from completed phase work with auto-generated body.
- Key decisions
- Optional configured PRD-style sections from `ship.pr_body_sections`
**Ship gates (capability-driven):** `/gsd:ship` runs every active `ship:pre` gate from the capability registry. Two are on by default:
- **Security** (`security` capability): blocks while `SECURITY.md` reports `threats_open > 0`. Resolve via `/gsd:secure-phase {n}`.
- **Broken-windows ledger** (`broken-windows` capability, issue #1950): when `workflow.windows_enforce=true` is set, blocks while `.planning/WINDOWS.md` reports any `open` entry. The ledger accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases. Resolve an entry with `gsd-tools windows fixed <id>` (defect resolved) or `gsd-tools windows waive <id> "<reason>"` (justified deferral — reason is required and recorded). Inspect via `gsd-tools windows status`. Enforcement is **opt-in** (default `workflow.windows_enforce=false`): enable with `gsd config-set workflow.windows_enforce true`; tracking continues regardless.
See [Custom PR Body Sections](ship-pr-body-sections.md) for onboarding, examples, and validation rules.
---
@@ -482,6 +489,7 @@ Start next version cycle.
|----------|----------|-------------|
| `name` | No | Milestone name |
| `--reset-phase-numbers` | No | Restart the new milestone at Phase 1 and archive old phase dirs before roadmapping |
| `--ws <name>` | No | Scope the milestone to a workstream; skips the shared `PROJECT.md` write |
**Prerequisites:** Previous milestone completed
**Produces:** Updated `PROJECT.md`, new `REQUIREMENTS.md`, new `ROADMAP.md`
@@ -490,6 +498,7 @@ Start next version cycle.
/gsd-new-milestone # Interactive
/gsd-new-milestone "v2.0 Mobile" # Named milestone
/gsd-new-milestone --reset-phase-numbers "v2.0 Mobile" # Restart milestone numbering at 1
/gsd-new-milestone --ws search "v2.0 Search" # Scope to a workstream
```
---

View File

@@ -160,7 +160,8 @@ GSD stores project settings in `.planning/config.json`. Created during `/gsd-new
| `dynamic_routing.enabled` | boolean | `true`, `false` | `false` | Master switch for [dynamic routing with failure-tier escalation](#dynamic-routing-with-failure-tier-escalation-dynamic_routing--added-in-v140). When `true`, agents resolve to `tier_models[default_tier]` and escalate one tier up on orchestrator-detected soft failure. Added in v1.40 ([#3024](https://github.com/open-gsd/gsd-core/pull/3031)) |
| `dynamic_routing.tier_models.<tier>` | enum | `opus`, `sonnet`, `haiku` | (none) | Tier alias for `light`, `standard`, or `heavy`. Used when `dynamic_routing.enabled: true`. Added in v1.40 |
| `dynamic_routing.escalate_on_failure` | boolean | `true`, `false` | `true` | When `false`, escalation is disabled even if `enabled: true` — every attempt uses the default tier. Added in v1.40 |
| `dynamic_routing.max_escalations` | integer | `0`, `1`, `2`, … | `1` | Hard cap on retries per agent invocation. Beyond the cap the resolver returns the cap-tier model. Added in v1.40 |
| `dynamic_routing.max_escalations` | integer | `0`, `1`, `2`, … | `1` | Hard cap on retries per agent invocation. Beyond the cap the resolver returns the cap-tier model. Also caps `provider_escalation`. Added in v1.40 |
| `dynamic_routing.provider_escalation` | string[] | ordered model IDs | (none) | Opt-in fallback providers tried when a run dies on a quota / rate limit — see [provider escalation](#provider-escalation-on-quota-exceeded--added-in-v143). Added in v1.43 ([#2296](https://github.com/open-gsd/gsd-core/issues/2296)) |
| `project_code` | string | any short string | (none) | Prefix for phase directory names (e.g., `"ABC"` produces `ABC-01-setup/`). Added in v1.31 |
| `phase_id_convention` | enum | `"milestone-prefixed"`, `null` | `null` | Phase ID naming convention. `null` = legacy numeric IDs (`Phase 1`, `Phase 2`). `"milestone-prefixed"` = globally unique IDs that encode the enclosing milestone (`Phase 1-01`, `Phase 1-02`). Run `gsd-tools roadmap upgrade --convention milestone-prefixed` to migrate an existing ROADMAP.md. |
| `response_language` | string | language code | (none) | Language for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`). Propagates to all spawned agents for cross-phase language consistency. Added in v1.32 |
@@ -453,6 +454,7 @@ If `.planning/` is in `.gitignore`, `commit_docs` is automatically `false` regar
| `statusline.show_last_command` | boolean | `false` | Append `last: /<cmd>` suffix to the statusline showing the most recently invoked slash command. Opt-in; reads the active session transcript to extract the latest `<command-name>` tag (closes #2538) |
| `statusline.context_position` | string | `"end"` | Position of the context-window meter. `"end"` (default) renders at line tail; `"front"` renders immediately after the model name so the meter stays visible in narrow terminals. Closes #2937 |
| `statusline.show_context_tokens` | boolean | `false` | Append the absolute token count (e.g. `(156k)`) after the context meter's percentage. Sums input, cache-creation, cache-read, and output tokens from the hook payload — a broader basis than the meter's percentage (which excludes output tokens), so the two figures can diverge slightly. Opt-in; the meter is unchanged when the flag is absent |
| `statusline.state_format` | string | `"full"` | Format of the GSD-state segment. `"full"` (default) is the existing rendering with milestone name and progress bar. `"compact"` renders `<version> · P<phase>/<total> · <status>` (e.g. `v1.12 · P7/12 · executing`) — drops the milestone name and bar, and collapses narrative statuses to the canonical keyword set from `normalizeStateStatus()` (`paused` — the canonical stuck state — renders uppercase as `PAUSED`) |
| `statusline.show_git` | boolean | `false` | Append a git segment after the directory: current branch plus compact work-state markers (`+staged` `~unstaged` `?untracked` `↑ahead` `↓behind`, or `✓` when clean and in sync). One `git status --porcelain=v2` call per render; the segment is absent outside a git repo or when git is unavailable |
The prompt injection guard hook (`gsd-prompt-guard.js`) is always active and cannot be disabled — it's a security feature, not a workflow toggle.
@@ -926,7 +928,9 @@ plans and shipped code (issue #2492).
existing requirements coverage gate, before plans are committed. For each
trackable decision in `<decisions>`, it checks that the decision id
(`D-NN`) or its text appears in at least one plan's `must_haves`,
`truths`, or body. A miss surfaces the missing decision by id and refuses
`truths`, or `objective` (front-matter), a `## must_haves`/`truths`/`tasks`/`objective`
heading, or an `<objective>`/`<tasks>`/`<task>`/`<action>`/`<read_first>`/`<behavior>`/`<verify>`/`<acceptance_criteria>`/`<done>`
tag body. A miss surfaces the missing decision by id and refuses
to mark the phase planned.
**Verify-phase validation gate (NON-BLOCKING).** Runs alongside the other
@@ -1239,7 +1243,40 @@ The `dynamic_routing` block is **disabled by default** — `enabled: false` (or
| `dynamic_routing.tier_models.standard` | enum | (none) | Tier alias for standard. Typically `sonnet`. |
| `dynamic_routing.tier_models.heavy` | enum | (none) | Tier alias for heavy. Typically `opus`. |
| `dynamic_routing.escalate_on_failure` | boolean | `true` | When false, escalation is disabled (every attempt uses the default tier). |
| `dynamic_routing.max_escalations` | integer | `1` | Hard cap on retries per agent invocation. Prevents runaway loops. |
| `dynamic_routing.max_escalations` | integer | `1` | Hard cap on retries per agent invocation. Prevents runaway loops. Also caps the provider ladder below. |
| `dynamic_routing.provider_escalation` | string[] | (none) | Ordered fallback model IDs tried when a run dies on a provider **quota / rate limit**. Added in v1.43 ([#2296](https://github.com/open-gsd/gsd-core/issues/2296)) |
#### Provider escalation on quota-exceeded — added in v1.43
The tier ladder above escalates *within one provider*. That does not help when the
provider itself is what ran out: a heavier tier on the same throttled account is still
throttled. `provider_escalation` is a separate, opt-in ladder for exactly that case.
```json
{
"dynamic_routing": {
"enabled": true,
"tier_models": { "light": "haiku", "standard": "sonnet", "heavy": "opus" },
"provider_escalation": ["gpt-5", "nvidia/llama-3.3"],
"max_escalations": 2
}
}
```
When an executor dies and `gsd-tools agent classify-failure` classifies the error body as
`quota-exceeded`, `execute-phase` re-resolves the model from this list instead of waiting
for a quota reset, logs the switch (`sonnet → gpt-5`), and honors any `Retry-After` the
provider sent. The ladder is capped at `min(max_escalations, provider_escalation.length)`;
once spent, GSD reports every model it tried and falls back to the manual recovery prompt
rather than silently retrying the last one.
- **Opt-in.** With no `provider_escalation` configured, quota failures keep today's manual
wait-for-reset prompt exactly as before.
- **Quota only.** Other failure classes (`classify-handoff-bug`, `unknown-failure`) never
consult this ladder — they keep the tier ladder.
- **`escalate_on_failure: false`** disables this ladder too.
- Entries are opaque model IDs passed to the runtime. Blank and non-string entries are
dropped; the surviving order is preserved.
#### When to use which
@@ -1578,6 +1615,7 @@ Use `provider: "generic"` (or `"custom"`) for OpenRouter, LiteLLM, local gateway
| `GSD_AUDIT_ARGS` | Set to `1` to include command args in audit/error events (omitted by default) |
| `GSD_PROJECT` | Override project root for multi-project workspace support (v1.32) |
| `GSD_SKIP_SCHEMA_CHECK` | Skip schema drift detection during execute-phase (v1.31) |
| `GSD_ALLOW_SYMLINKED_DEST` | Set to `1` (or `true`) to permit install/update when `CLAUDE_CONFIG_DIR` (or any artifact-kind child like `skills/`, `hooks/`) is an **intentional, user-owned symlink** pointing outside the install root. v1.7.x write-confinement (ADR-1239 Phase B) refuses such layouts by default to prevent untrusted `destSubpath` traversal. Opt in only if you manage configHome via symlinked external dirs, multi-account config layouts (`~/.claude-personal`, `~/.claude-team`), or dotfiles-managed configHome (nix-darwin, etc.). Two refusals remain load-bearing even with opt-in: path-traversal in `destSubpath` (`../../etc`-style), and a symlink whose resolved target equals the install root itself (would let the prune pass wipe it). |
| `WSL_DISTRO_NAME` | Detected by installer for WSL path handling |
---

View File

@@ -171,6 +171,17 @@
- [MemPalace Memory Capability](#145-mempalace-memory-capability)
- [Spec-Phase Prohibition Probe](#146-spec-phase-prohibition-probe)
- [Capability Management Command](#147-capability-management-command)
- [Smart Entry Launcher](#148-smart-entry-launcher)
- [v1.7.0 Features](#v170-features)
- [Embeddable Orchestration System (Host-Integration Interface)](#149-embeddable-orchestration-system-host-integration-interface)
- [Discoverability Registries](#150-discoverability-registries)
- [Companion MCP Server](#151-companion-mcp-server)
- [Statusline Token Count & Git Segment](#152-statusline-token-count--git-segment)
- [Model Catalog Advances](#153-model-catalog-advances)
- [Claude Orchestration Capability (BETA)](#154-claude-orchestration-capability-beta)
- [External-Job Capability](#155-external-job-capability)
- [API-Coverage Gate](#156-api-coverage-gate)
- [State Rebuild & Configurable Graph Path](#157-state-rebuild--configurable-graph-path)
---
@@ -298,6 +309,7 @@
- REQ-PLAN-07: System MUST prompt user to run `/gsd-ui-phase` if frontend phase detected and no UI-SPEC.md exists (UI safety gate)
- REQ-PLAN-08: System MUST include Nyquist validation mapping when `workflow.nyquist_validation` is enabled
- REQ-PLAN-09: System MUST verify all phase requirements are covered by at least one plan before planning completes (requirements coverage gate)
- REQ-PLAN-10: System MUST support an optional `<reversibility rating="reversible|costly|one-way">` element recording how costly a decision would be to undo, and MUST insert a `checkpoint:decision` before the task implementing a `one-way` decision unless `--no-reversibility-gates` is set (`costly` is flagged without blocking; `reversible` and unrated flow normally)
**Produces:**
| Artifact | Description |
@@ -3256,3 +3268,101 @@ The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, s
**Reference:** [Smart Entry Design](superpowers/specs/2026-06-27-gsd-smart-entry-design.md)
---
## v1.7.0 Features
> These are features new to **@opengsd/gsd-core 1.7.0** (the current release line: 1.0.0 → 1.2.0 → … → 1.6.1 → 1.7.0). The preceding `v1.27`–`v1.43.0` sections use the retired get-shit-done-cc / get-shit-done-redux feature numbering and are not gsd-core releases — see [Legacy Release Notes](RELEASE-NOTES-LEGACY.md).
### 149. Embeddable Orchestration System (Host-Integration Interface)
**Purpose:** Express every host integration against one public, versioned contract (ADR-1239 Phase A, #1690) instead of bespoke per-host wiring, so onboarding a new host becomes additive descriptor work.
**Behavior:** The interface exposes six interface points (`command`, `dispatch`, `model`, `hooks`, `state`, `artifact`), eight negotiated axes, and a `PROTOCOL_VERSION` handshake that negotiates down to `min(host, engine)`. In 1.7.0, 14 runtimes were migrated onto the interface via imperative adapters (OpenCode #2087, Cursor #2089, Cline #2090, Hermes #2091, Qwen #2092, Kilo #2093, Trae #2094, Kimi #2095, Antigravity #2096, Augment #2097), a declarative adapter (Codex #2088), plus full lifecycle-hook wiring for CodeBuddy (#2098), GitHub Copilot (#2099), and Windsurf (#2100). Descriptors gained an `extensionEvents` vocabulary (#1946), and `/gsd:surface` now reproduces a runtime's agent output byte-for-byte from the installer's descriptors (#1575).
**New runtimes:** ZCode (Z.ai — Agentic Development Environment for GLM-5.2, #1925), pi (`npx @opengsd/gsd-core --pi`, #2102), and a repo-local VS Code extension driven through the adapter (#2103). The retired Gemini CLI now redirects to Antigravity CLI, its official successor (#1928).
**Reference:** [The Embeddable Orchestration System](explanation/embeddable-orchestration-system.md) · [Host-Integration Interface](reference/host-integration-interface.md) · [Interface versioning policy](explanation/interface-versioning-policy.md)
---
### 150. Discoverability Registries
**Purpose:** Two non-endorsing catalogs for third-party extensions (#2182).
**Behavior:** The **Community Capability Registry** (#2188) lists third-party Feature Capabilities installed with `gsd capability install`; the **EoS Registry** (#2193) lists third-party host integrations built on the ADR-1239 interface. Every entry embeds a live release badge and links to a GitHub Discussion. Registration is a documentation PR, regenerated with `npm run gen:registry`.
**Reference:** [GSD Registries](registries/README.md)
---
### 151. Companion MCP Server
**Command:** `gsd-mcp-server`
**Purpose:** A companion MCP server exposing GSD over stdio JSON-RPC 2.0, covering interface points 1 and 5 (#1681).
**Behavior:** OpenCode installs auto-register it as `mcp.gsd` (#1682). OpenCode also gained the `opencode-subset` hook dialect plus `session.idle` handling (#1682) and now runs GSD's lifecycle safety hooks — prompt-injection guard, read-before-edit guard, and injection scanner (#1923).
---
### 152. Statusline Token Count & Git Segment
**Purpose:** Opt-in statusline additions surfacing more session context.
**Behavior:** An absolute token count on the context meter (#2161) and a git branch + working-state segment (#2163), both opt-in. A companion opt-in **compact GSD-state format** condenses the GSD state segment (#2162).
**Configuration:** `statusline.*`
---
### 153. Model Catalog Advances
**Purpose:** Refresh the default model tiers and how models are surfaced.
**Behavior:** Codex/OpenAI defaults advance to the **GPT-5.6 family (Sol / Terra / Luna)** (#2122); the verbose `(1M context)` model suffix collapses to a compact `(1M)` badge (#2160). GSD warns when model config changes without re-running the installer on static-frontmatter runtimes such as Codex and OpenCode (#1688).
**Reference:** [Configuration](CONFIGURATION.md) · [Configure model profiles](how-to/configure-model-profiles.md)
---
### 154. Claude Orchestration Capability (BETA)
**Purpose:** A default-off, BETA, Claude-only capability that adopts Claude Code's Workflow tool for parallel sub-agent orchestration (#1143).
**Reference:** [The Claude orchestration capability](explanation/claude-orchestration-capability.md)
---
### 155. External-Job Capability
**Purpose:** A default-off capability that externalizes long-running compute as asynchronous external jobs, e.g. SLURM submission (#1165).
**Configuration:** `external_job.submit_timeout_ms`, `external_job.poll_timeout_ms`, `external_job.artifact_dir` (#1164)
---
### 156. API-Coverage Gate
**Command:** `/gsd:verify-work`
**Purpose:** A phase that integrates an external API, SDK, or service can no longer seal verification without a decided coverage matrix (#1562).
---
### 157. State Rebuild & Configurable Graph Path
**Behavior:** A new `gsd-tools state rebuild` subcommand re-derives `STATE.md` from source (#1830). The new `graphify.graph_path` setting makes the knowledge-graph location configurable, so a single umbrella graph can serve several projects (#1825).
---
### 158. Broken-Windows Ledger
**Behavior:** A cross-phase defect register at `.planning/WINDOWS.md` accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths (#1950). `/gsd:ship` blocks while any entry is `open`; an entry can be `waived` only with a recorded reason (auditable) or marked `fixed` (removed from the blocking set). `/gsd:progress` surfaces the open + waived counts.
**Commands:** `gsd-tools windows status | append | waive | fixed`.
**Config:** `workflow.windows_enforce` (gate active, default `false` — opt-in enforcement). Enable with `gsd config-set workflow.windows_enforce true`. Tracking (the ledger itself, populated by the executor) is always on; only the ship gate is opt-in.
**Backward compatibility:** A project with no `.planning/WINDOWS.md` reports `open_count: 0` and ships cleanly; the gate only activates once windows are recorded.
**Configuration:** `graphify.graph_path`

View File

@@ -214,7 +214,14 @@
"common-bug-patterns.md",
"context-budget.md",
"continuation-format.md",
"debugger-bug-taxonomy.md",
"debugger-fix-acceptance.md",
"debugger-philosophy.md",
"debugger-prevention.md",
"debugger-rca-branching.md",
"debugger-repro-hardening.md",
"debugger-sbfl.md",
"debugger-semantic-recall.md",
"decimal-phase-calculation.md",
"doc-conflict-engine.md",
"domain-probes.md",
@@ -222,6 +229,9 @@
"execute-mvp-tdd.md",
"execute-phase-between-wave-reset.md",
"execute-phase-context-guard.md",
"execute-phase-quota-recovery.md",
"execute-phase-requirement-revert.md",
"execute-phase-response-language.md",
"execute-phase-wave-guard.md",
"executor-examples.md",
"gate-prompts.md",
@@ -246,6 +256,8 @@
"planner-interface-context.md",
"planner-load-graph-context.md",
"planner-mvp-mode.md",
"planner-preconditions.md",
"planner-reversibility.md",
"planner-reviews.md",
"planner-revision.md",
"planner-source-audit.md",
@@ -299,7 +311,9 @@
"assumption-delta.cjs",
"audit-command-router.cjs",
"audit.cjs",
"broken-windows.cjs",
"capability-activation.cjs",
"capability-command-router.cjs",
"capability-consent.cjs",
"capability-ledger.cjs",
"capability-lifecycle.cjs",

View File

@@ -294,6 +294,13 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum
| `ui-brand.md` | Visual output formatting patterns. |
| `common-bug-patterns.md` | Common bug patterns for code review and verification. |
| `debugger-philosophy.md` | Evergreen debugging disciplines loaded by `gsd-debugger`. |
| `debugger-fix-acceptance.md` | Multi-signal fix-acceptance guardrail (anti-overfitting) loaded by `gsd-debugger`. |
| `debugger-sbfl.md` | Spectrum-based fault localization (Ochiai) pre-filter loaded by `gsd-debugger`. |
| `debugger-rca-branching.md` | RCA branching (fishbone + AND-gate) anti-single-cause discipline loaded by `gsd-debugger`. |
| `debugger-bug-taxonomy.md` | Bug-taxonomy classification (Bohrbug/Heisenbug/Concurrency) + technique routing table loaded by `gsd-debugger`. |
| `debugger-repro-hardening.md` | Regression-test hardening (PBT shrinking + oracle classification + boundary neighbors) loaded by `gsd-debugger`. |
| `debugger-prevention.md` | Prevention / blameless-postmortem output (5-Whys + why-not-caught + recurrence guard) loaded by `gsd-debugger`. |
| `debugger-semantic-recall.md` | Semantic knowledge-base recall via MemPalace (keyword-fallback) loaded by `gsd-debugger`. |
| `mandatory-initial-read.md` | Shared required-reading boilerplate injected into agent prompts. |
| `agent-skills-bootstrap.md` | Shared agent_skills self-load contract (query + Read + dedup guard) injected into all 22 consumer agents. |
| `project-skills-discovery.md` | Shared project-skills-discovery boilerplate injected into agent prompts. |
@@ -308,6 +315,9 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum
| `agent-contracts.md` | Formal interface between orchestrators and agents. |
| `context-budget.md` | Context window budget allocation rules. |
| `execute-phase-context-guard.md` | Context exhaustion guard step for `execute-phase` wave loop — `workflow.context_guard_mode` dispatch table (warn/auto/off) and POOR-tier pause-work trigger (#1452). |
| `execute-phase-requirement-revert.md` | Gap-report step for `execute-phase` — reverts this phase's own shared requirement IDs out of `Complete` in REQUIREMENTS.md before rendering a `gaps_found` report, scoped to `PHASE_REQ_IDS` so other phases' rows are untouched (#2388). |
| `execute-phase-response-language.md` | Response-language directive for `execute-phase` orchestrator output (questions, narration, report-template prose); extracted to keep the workflow under the frozen pre-phase-6 byte ceiling (#2402). |
| `execute-phase-quota-recovery.md` | Step 7.1 detail for `execute-phase` — `quota-exceeded` recovery: the opt-in `dynamic_routing.provider_escalation` ladder (swap provider, honor `Retry-After`, fail loudly when spent) and the default manual wait-for-reset prompt (#2296). |
| `continuation-format.md` | Session continuation/resume format. |
| `domain-probes.md` | Domain-specific probing questions for discuss-phase. |
| `edge-probe.md` | Spec-phase edge-completeness probe — 8-category edge taxonomy, shape classification, and the `requirements → checks → verifier` resolution model (Step 5.5). |
@@ -377,6 +387,8 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
| `planner-revision.md` | Plan revision patterns for iterative refinement. |
| `planner-source-audit.md` | Planner source-audit and authority-limit rules. |
| `planner-mvp-mode.md` | Vertical-slice planning rules for MVP mode. |
| `planner-preconditions.md` | Emission rules for the optional `<precondition>` task element (issue #1949, Design by Contract): when to emit, the three cases (user_setup / prior-phase artifact / env-var), format, anti-patterns, and the contract triad mapping. |
| `planner-reversibility.md` | Canonical reversibility taxonomy for the optional `<reversibility>` task element (issue #1951): the three ratings (`reversible` / `costly` / `one-way`), the `checkpoint:decision` insertion rule for one-way doors, the `--no-reversibility-gates` override, and the checkpoint-fatigue anti-patterns. |
| `planner-human-verify-mode.md` | Rules for `workflow.human_verify_mode = end-of-phase`: suppress `checkpoint:human-verify` task emission and route deferred items via `<verify><human-check>`. |
| `planner-graphify-auto-update.md` | How `load_graph_context` surfaces `.last-build-status.json` auto-update state (running / failed / stale head) alongside the existing staleness annotation. Opt-in via `graphify.auto_update` (#3347). |
| `planner-interface-context.md` | Interface context rules for executors — how to extract key interfaces/types/exports from existing code and document new interfaces that downstream plans will consume. |
@@ -399,11 +411,12 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `active-workstream-store.cjs` | Workstream source precedence and selection (CLI `--ws` > `GSD_WORKSTREAM` env > stored pointer); name validation and environment propagation |
| `adr-parser.cjs` | ADR decision parser for plan-phase ingest express path; normalizes section synonyms, parses status/decision/scope fences, and enforces status rejection gates |
| `agent-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools agent` |
| `api-coverage.cjs` | API-coverage detector + matrix validator (#1562) — pure `detectApiIntegration` (compound verb+noun signal + `<Service> API/SDK` surface; strips fenced code) and `validateCoverageMatrix`/`parseCoverageMatrix`/`renderCoverageMatrix` for the COVERAGE.md artifact; STDIN CLI (`echo "$SCOPE" \| node .../api-coverage.cjs [--json]`, exit 0=detected/1=none/2=error); consumed by the `ai-integration` capability's `plan:pre` contribution and blocking `verify:pre` gate (`check api-coverage.verify-pre`) |
| `api-coverage.cjs` | API-coverage detector + matrix validator (#1562, #2365) — pure `detectApiIntegration` (fail-closed: same-clause verb+noun signal + `<Service> API/SDK` surface naming a real service; strips fenced code, inline code, and path-shaped tokens; external hosts count, first-party route paths do not) and `validateCoverageMatrix`/`parseCoverageMatrix`/`renderCoverageMatrix` for the COVERAGE.md artifact (incl. the `No external API integration: <reason>` declaration); STDIN CLI (`echo "$SCOPE" \| node .../api-coverage.cjs [--json]`, exit 0=detected/1=none/2=error); consumed by the `ai-integration` capability's `plan:pre` contribution and blocking `verify:pre` gate (`check api-coverage.verify-pre`) |
| `artifacts.cjs` | Canonical artifact registry — known `.planning/` root file names; used by `gsd-health` W019 lint |
| `audit-command-router.cjs` | ADR-959 capability command router for `gsd-tools audit-uat` and `gsd-tools audit-open` — extracted from hardcoded cases in `gsd-tools.cjs`; dispatches to `uat.cjs:cmdAuditUat` and `audit.cjs:{auditOpenArtifacts,formatAuditReport}`; phase 4d-impl-3 |
| `audit.cjs` | Audit dispatch, audit open sessions, audit storage helpers |
| `capability-activation.cjs` | Capability activation resolver shared by config validation and capability-state consumers — resolves registry-owned config keys from raw runtime config without re-centralizing migrated settings |
| `capability-command-router.cjs` | ADR-2346 P2 host command router for `gsd-tools capability` — relocated verbatim from the former 706-line `case 'capability':` arm in `gsd-tools.cjs`; dispatched via `HOST_COMMAND_ROUTERS` in `runCommand`'s default case; wires `capability-lifecycle`/`-trust`/`-consent`/`-state`/`-writer`; hand-authored CJS (sibling of `ensure-runtime-build.cjs`) |
| `capability-consent.cjs` | User-owned capability consent store (#1459) — bounded, non-throwing JSON store at `${GSD_HOME\|\|homedir()}/.gsd/consent.json` (NEVER under a repo) keyed by `${realpath(projectRoot)} <id>`; exports `consentStorePath`/`readConsentStore`/`hasProjectConsent` (matches iff integrity AND disclosureSignature both match)/`recordProjectConsent` (atomic+durable write)/`revokeProjectConsent`; the authoritative consent signal that gates PROJECT-scope third-party capability activation so a forged/cloned project ledger no longer activates anything until the user consents on THIS machine |
| `capability-lock.cjs` | Shared cross-process lock primitive (#1459 finding 4) — the SINGLE hardened lockfile protocol used by BOTH capability-lifecycle (`.gsd/capabilities/.lock`) and capability-consent (`.consent.lock`); exports `acquireLock(lockPath, opts?)`/`releaseLock(handle)` with pid + process-start-time liveness identity, a hard deadman, and token+inode owner-safe release — NEVER stale-steals a verified-live same-host holder, reclaims only a provably-dead/unverifiable holder, never deadlocks; `opts.maxAttempts`/`opts.waitForFresh` let the consent store serialize genuinely-contended writers; `_setLockProbes`/`_resetLockProbes` are test seams |
| `capability-ledger.cjs` | Per-runtime install ledger (ADR-1244 D4) — atomic read/write of `.gsd-capabilities.json` recording `{ id, version, source, integrity, files[], sharedEdits[] }` per installed capability; exports `readLedger`/`writeLedger`/`recordInstall`/`removeEntry`/`reconcile` (orphan detection)/`readSmallRegularFile` (utf8) + `readSmallRegularFileBuffer` (raw bytes, the byte-exact consent-hash reader, #1459 finding 1); atomic commit point and reconciliation basis for Phase-4 upgrade/remove |
@@ -414,6 +427,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `capability-state.cjs` | Unified capability-state resolver (ADR-857 phase 4b/6) — composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; exports pure `resolveCapabilityState`, reusable `resolveCapabilityRuntimeState`, and I/O handler `cmdCapabilityState`; command surface: `gsd-tools capability state [--config-dir <path>]` emitting `{ runtimeConfigDir, capabilities[] }` |
| `capability-trust.cjs` | Capability trust gate (ADR-1244 Phase 4, D5 + compatibility half of D6) — PURE policy module: `discloseExecutableSurfaces` (hooks/command modules/mcpServers), `evaluateInstallTrust` (compose source policy + reserved-namespace + engines gate + disclosure → allowed/requiresConsent/blockReasons), `evaluateSourceAllowed` (`strict_known_registries`: permissive/lockdown/host-allowlist), `checkEngines` (engines.gsd hard gate + `compatVersions` graceful-downgrade), `executableSetChanged` (auto-update re-consent trigger); no sandbox — see `docs/explanation/capability-trust-model.md` |
| `capability-validator.cjs` | Shared runtime-callable capability validator (ADR-1244 D2) — extracted from `scripts/gen-capability-registry.cjs` so the build-time generator and the runtime overlay loader share ONE validation implementation (generative-parity guarded); exports `validateCapability`/`validateCrossCapability`/`validateVersionEnvelope`/`validateConsumesGlobal`/… plus the closed-vocabulary sets and `SEMVER_RE` |
| `broken-windows.cjs` | Broken-windows ledger library (issue #1950) — typed IR + I/O for `.planning/WINDOWS.md` (cross-phase defect register); pure `parseLedger`/`renderLedger`/`appendWindow`/`markWaived`/`markFixed`/`openCount` + I/O `cmdWindowsStatus`/`cmdWindowsAppend`/`cmdWindowsWaive`/`cmdWindowsMarkFixed`; frozen `REASON` enum for typed-error assertions; CLI surface `gsd-tools windows status\|append\|waive\|fixed`. Generated from `src/broken-windows.cts` |
| `capability-writer.cjs` | Capability State Writer (ADR-1213) — write-side inverse of the resolver; projects desired per-capability enabled/gates onto surface + config substrates, then re-resolves (assert-and-report); exports `setCapabilityState` and I/O handler `cmdCapabilitySet`; command surface: `gsd-tools capability set <id> [--on\|--off] [--gate <key>=<true\|false>]` |
| `check-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools check` |
| `cli-exit.cjs` | `ExitError` class and `runMain()` helper — CLI entrypoints throw `ExitError` instead of calling `process.exit()`; `runMain()` translates the outcome into `process.exitCode` so output flushes cleanly |
@@ -465,7 +479,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `legacy-cleanup.cjs` | Detect and remove leftover get-shit-done-cc artifacts; exports `planLegacyCleanup` (pure scan) and `applyLegacyCleanup` (thin IO applier) that root out stale files from the old package across every GSD-managed runtime config directory (#607) |
| `loop-host-contract.cjs` | Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts for the five-step pipeline (discuss/plan/execute/verify/ship); emitted by `scripts/gen-loop-host-contract.cjs --write` (ADR-894 §3); consumed by `gen-capability-registry.cjs` |
| `loop-resolver.cjs` | Loop Extension Point resolver — ADR-857 phase 3c/6 registry-consuming query; given a canonical loop point, filters `byLoopPoint` by resolved Capability State plus config activation (`when` key traversal with prototype-pollution guard), returns `{ point, activeHooks, rendered }` envelope; `resolveLoopHooks` and `renderLoopHooks` are pure (no I/O); command surface: `gsd-tools loop render-hooks <point> [--config-dir <path>]` |
| `markdown-sectionizer.cjs` | Canonical markdown-structure parsing seam (ADR-1372, epic #1372) — pure, Node built-ins only; exports `stripFencedCode` (CommonMark-correct fence stripper, CRLF-safe), `tokenizeHeadings` (ATX headings outside fenced blocks), `collectSections`/`collectSection` (line-by-line section collection with `bodyStart`/`bodyEnd` offsets), `iterateBullets` (dash/checkbox/numbered markers), `extractTaggedBlocks` (inner text of `<tag>…</tag>` blocks, caller decides fence-stripping), `replaceSection` (pure character-offset body splice for read-modify-write callers), and `withSection` (resolve a section by heading/predicate and run an edit callback against ONLY its body, splicing the result back — ADR-2143 §4 bounded mutation); foundation for T0–T7 migration tiers retiring 8+ ad-hoc parsers |
| `markdown-sectionizer.cjs` | Canonical markdown-structure parsing seam (ADR-1372, epic #1372) — pure, Node built-ins only; exports `stripFencedCode` (CommonMark-correct fence stripper, CRLF-safe), `stripInlineCode` (per-line CommonMark inline-code-span stripper, #2365), `tokenizeHeadings` (ATX headings outside fenced blocks), `collectSections`/`collectSection` (line-by-line section collection with `bodyStart`/`bodyEnd` offsets), `iterateBullets` (dash/checkbox/numbered markers), `extractTaggedBlocks` (inner text of `<tag>…</tag>` blocks, caller decides fence-stripping), `replaceSection` (pure character-offset body splice for read-modify-write callers), and `withSection` (resolve a section by heading/predicate and run an edit callback against ONLY its body, splicing the result back — ADR-2143 §4 bounded mutation); foundation for T0–T7 migration tiers retiring 8+ ad-hoc parsers |
| `markdown-table.cjs` | Canonical GFM table model + `TABLE_SCHEMAS` registry seam (ADR-2143, epic #2143) — pure, Node built-ins only; exports `parseMarkdownTable(sectionText) → Result<MarkdownTable>` (parses the first GFM pipe table, typed parse errors for ragged/malformed rows rather than silent coercion), `MarkdownTable` (`{columns, rows}`, rows addressed by column name), `Result<T>` (`{ok:true,value}\|{ok:false,reason}` — distinct from command-routing-hub's dispatch `Result`), `TABLE_SCHEMAS` (canonical column-header variants for `RoadmapProgress`/`RequirementsTraceability`/`QuickTasks`/`Security` tables), and `matchTableSchema(columns) → {id,label}\|null` (resolves parsed headers back to a canonical schema); consumed by `phase-lifecycle.cts`'s `deriveProgressFromRoadmap` (fixes #2137, the 5-column milestone-grouped Progress table) |
| `milestone.cjs` | Milestone archival, requirements marking |
| `model-catalog.cjs` | CJS adapter over the shared model catalog JSON; exports canonical runtime tier defaults, agent profile maps, alias maps, and routing metadata for all CLI consumers |

View File

@@ -74,6 +74,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [The capability trust model](explanation/capability-trust-model.md) — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- [How overlay capabilities compose](explanation/capability-overlay-model.md) — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- [Architecture](ARCHITECTURE.md) — system architecture, agent model, and data flow
- [The Embeddable Orchestration System](explanation/embeddable-orchestration-system.md) — one public, versioned contract for embedding GSD across many hosts
- [Discuss modes](workflow-discuss-mode.md) — assumptions mode vs interview mode for `/gsd-discuss-phase`
- [Context monitoring](context-monitor.md) — context window monitoring hook architecture
- [Issue-driven orchestration](issue-driven-orchestration.md) — recipe for driving GSD from a tracker issue using existing primitives
@@ -82,5 +83,6 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
## Related
- [What's new in 1.7.0](whats-new-1.7.0.md) — curated highlights of the 1.7.0 release
- [Root README](../README.md) — landing page, quickstart, and documentation overview
- [Changelog](../CHANGELOG.md) — release history

View File

@@ -209,6 +209,55 @@ npm run ci:test-scope -- --files "commands/gsd/plan-phase.md"
node scripts/ci-test-scope.cjs --base origin/next --head HEAD
```
## Chunk packing and the test timing table
`scripts/run-tests.cjs` does not hand the whole selected file list to one
`node --test` process. It packs the files into **chunks**, each spawned
separately, because Windows caps a command line at 32,767 characters and because
each chunk gets its own 600s timeout (`RUN_TESTS_CHUNK_TIMEOUT_MS`) and a fresh
process, which bounds memory pressure.
How files are distributed across those chunks decides whether the slowest chunk
sits near that timeout while the others idle. The packer weights each file by its
**measured duration**, read from `tests/test-timings.json`, and places files with
LPT (longest-processing-time-first: heaviest file first, each into the currently
lightest chunk). Before #2456 the weight was guessed from the filename, which
mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.
### Reference
| Knob | Default | Meaning |
|---|---|---|
| `RUN_TESTS_MAX_FILES_PER_CHUNK` | `60` | Per-chunk weight budget. Weights are normalized so an **average-cost** file weighs 1, so this still reads as "about 60 average files". |
| `RUN_TESTS_MAX_CMDLINE_CHARS` | `28000` | argv ceiling per chunk, with headroom under the Windows 32,767 limit. |
| `RUN_TESTS_TIMINGS_FILE` | `tests/test-timings.json` | Path to the timing table. Tests override it to inject a synthetic cost profile. |
| `RUN_TESTS_CHUNK_TIMEOUT_MS` | `600000` | Per-chunk timeout. |
The timing table is **advisory and deliberately un-gated**. There is no `--check`
mode and no CI lint that fails on staleness, because timing data legitimately
varies run to run. A file missing from the table falls back to the table's median
weight, and a missing or unparseable table falls back to uniform weight — so
drift costs chunk *balance*, never a red build. A count-based floor additionally
guarantees the packer never produces fewer chunks than plain count-based packing
would, so a badly stale table cannot collapse the suite into a few fat chunks.
### How-to: regenerate the timing table
Regenerate when the suite's cost profile has visibly drifted — after adding or
removing expensive tests, not on a schedule. The input is a `node:test` reporter
event stream from a `gsd-test` run:
```bash
node scripts/gen-test-timings.cjs \
~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node22.jsonl \
~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node24.jsonl
```
Pass every lane you have. A file's recorded time is the **max** across the
supplied streams, not the mean: the packer exists to keep the *slowest* lane's
slowest chunk away from the timeout, so the conservative bound is the right one.
Keys are sorted so a regeneration diff shows only the files whose cost moved.
## Best practices for forward-compat (Node 24/26)
- Use `process.execPath` when spawning Node in tests so each matrix lane exercises the lane's Node version.

View File

@@ -231,7 +231,7 @@ See [docs/workflow-discuss-mode.md](workflow-discuss-mode.md) for the full discu
The discuss-phase captures implementation decisions in CONTEXT.md under a `<decisions>` block as numbered bullets (`- **D-01:** …`). Two gates ensure those decisions survive into plans and shipped code.
**Plan-phase translation gate (blocking).** After planning, GSD refuses to mark the phase planned until every trackable decision appears in at least one plan's `must_haves`, `truths`, or body.
**Plan-phase translation gate (blocking).** After planning, GSD refuses to mark the phase planned until every trackable decision appears in at least one plan's scanned surfaces: front-matter `must_haves`/`truths`/`objective`, a `## must_haves`/`truths`/`tasks`/`objective` heading, or an `<objective>`/`<tasks>`/`<task>`/`<action>`/`<read_first>`/`<behavior>`/`<verify>`/`<acceptance_criteria>`/`<done>` tag body.
**Verify-phase validation gate (non-blocking).** During verification, GSD searches plans, SUMMARY.md, modified files, and recent commit messages for each trackable decision. Misses are logged to VERIFICATION.md as a warning section; verification status is unchanged.
@@ -886,9 +886,9 @@ Since v1.3.1, the installer pre-populates `~/.claude/settings.json` (or
"allow": [
"Bash(npx gsd-core *)",
"Read(.planning/*)",
"Write(.planning/*)",
"Edit(.planning/*)",
"Read(STATE.md)",
"Write(STATE.md)"
"Edit(STATE.md)"
],
"deny": [
"Read(.env)",

View File

@@ -1,6 +1,6 @@
# SDK Architecture seam map for query/runtime surfaces
- **Status:** Superseded by ADR-0174 (2026-05-23); originally Accepted (2026-05-09)
- **Status:** Superseded by [ADR-0174](0174-retire-gsd-sdk-package-boundary.md) (2026-05-23); originally Accepted (2026-05-09)
- **Date:** 2026-05-09
We decided to keep SDK architecture explicitly module-seamed rather than allow feature logic to spread across query handlers, runtime adapters, and compatibility shims. This ADR is the top-level map for SDK seams and their ownership boundaries.

View File

@@ -1,6 +1,6 @@
# SDK Package Seam Module owns SDK-to-get-shit-done-redux compatibility
- **Status:** Superseded by ADR-0174 (2026-05-23); originally Accepted (2026-05-07)
- **Status:** Superseded by [ADR-0174](0174-retire-gsd-sdk-package-boundary.md) (2026-05-23); originally Accepted (2026-05-07)
- **Date:** 2026-05-07
We decided to define one explicit SDK Package Seam Module for the `@opengsd/gsd-sdk` → `@opengsd/get-shit-done-redux` transition. During this transition, install-layout probing, legacy `gsd-tools.cjs` discovery, legacy `core.cjs` discovery, and compatibility-only missing-asset diagnostics must live behind one seam instead of leaking across SDK Modules. This keeps callers thin, raises leverage for standalone-SDK testing, and improves locality by making package-readiness bugs land in one place. First tracer-bullet slice: add one compatibility Adapter Module at this seam and migrate current legacy asset callers onto it before broader native replacement work.

View File

@@ -1,6 +1,7 @@
# Shell Command Projection Module owns runtime-aware OS command rendering
- **Status:** Accepted
- **Supersedes:** [ADR-0010](0010-file-operation-engine-module.md) (File Operation Engine Module) — absorbed into this seam's Phases 3–4 (`#3467`–`#3468`), 2026-05-13
- **Date:** 2026-05-12
We propose introducing a Shell Command Projection Module that owns projection from typed command intent to concrete shell/runtime-specific command text. GSD currently hand-builds hook commands, PATH repair commands, shim scripts, and other serialized OS-facing command strings across installer call sites. That drift has repeatedly produced cross-shell regressions (`#2376`, `#2979`, `#3002`, `#3011`, `#3181`, `#3393`, `#3413`). The proposed seam concentrates quoting, path-style, and runtime-wrapper policy in one module while keeping real subprocess execution on array-arg/non-shell paths.

View File

@@ -1,6 +1,6 @@
# File Operation Engine Module owns safe runtime/config file mutations
- **Status:** Superseded by ADR-0009 (Shell Command Projection Module expansion, Phases 3–4, `#3467`–`#3468`)
- **Status:** Superseded by [ADR-0009](0009-shell-command-projection-module.md) (Shell Command Projection Module expansion, Phases 3–4, `#3467`–`#3468`)
- **Date:** 2026-05-12
- **Superseded:** 2026-05-13

View File

@@ -1,8 +1,10 @@
# Skill Surface Budget Module owns install-time skill listing curation
- **Status:** Proposed
- **Status:** Superseded by [ADR-0011](0011-skill-surface-budget-module.md) (Skill Surface Budget Module — install-time profile staging and runtime surface control); originally Proposed (2026-05-12)
- **Date:** 2026-05-12
> **Provenance of this status (2026-07-16).** This file said `Proposed` while the hand-maintained index in `README.md` recorded it as *"Skill Surface Budget Module — earlier draft superseded by ADR-0011"*, status *"Superseded by 0011"*. The index was right and the file was stale. When the index became a generated artifact (derived from these files), that assertion would have been silently dropped and this superseded draft would have reappeared as a live `Proposed` decision — so it is recorded here, at its source, instead. This is the one status corrected from the old index rather than left for ratification, because leaving it would have *lost* a decision the maintainer had already made.
We propose extending the existing install profile seam (`gsd-core/bin/lib/install-profiles.cjs`) into a **Skill Surface Budget Module** that owns which subset of GSD's 66 skills is written to the runtime config dirs, and that owns the per-skill `requires:` dependency manifest used to keep that subset closed under cross-skill references. GSD currently ships a binary `--minimal` / full toggle; runtimes that enumerate skills (Claude Code, OpenCode, etc.) cap the `<available_skills>` system-prompt block at `skillListingBudgetFraction` of the context window (default 1% = ~2k tokens at 200k), and GSD alone consumes ~60% of that cap (#3408). Further description shrinkage is unavailable — `scripts/lint-descriptions.cjs` already enforces a hard 100-char ceiling and the mean is 72.5 chars. The remaining lever is surfacing fewer skills, which requires a typed profile model plus a dependency manifest, not more ad-hoc allowlists.
## Decision

View File

@@ -1,11 +1,11 @@
# PRD — `review.default_reviewers` config key for `/gsd-review` reviewer selection
- **Status:** Draft
- **Status:** Legacy — frozen historical record; not a pattern to follow (see the note below)
- **Date:** 2026-05-13
- **Issue:** `#3079`
- **Related ADR:** `0011-review-default-reviewers.md`
- **Related ADR:** [`0011-review-default-reviewers.md`](0011-review-default-reviewers.md)
> This PRD is filed alongside its ADR under `docs/adr/` for co-location. The repo does not yet have a `docs/prd/` directory; if maintainers prefer one, this file can move there with the `0011-` prefix preserved.
> **Note (2026-07-16).** This PRD's original note said "the repo does not yet have a `docs/prd/` directory; if maintainers prefer one, this file can move there." That directory **now exists**, and [`docs/prd/README.md`](../prd/README.md) records this file's disposition: it *"predates this directory and is preserved as immutable historical record. It is not a pattern to follow. New PRDs live here."* It is therefore kept in place, and its status is `Legacy` — the decision is frozen for provenance, not superseded by a specific successor. New PRDs go in `docs/prd/`.
## TL;DR

View File

@@ -1,10 +1,28 @@
# `review.default_reviewers` config key scopes the no-flag `/gsd-review` fan-out
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-05-13); see "Ratification" below
- **Date:** 2026-05-13
We propose adding a `review.default_reviewers` key to `.planning/config.json` that scopes the no-flag default of `/gsd-review` to a user-chosen subset of detected CLI reviewers. Today the no-flag branch of `workflows/review.md` (line 52) invokes **all available** CLIs, which for multi-CLI users plus local model servers (ollama, lm-studio, llama.cpp) means probing up to ~10 backends per review, paying timeout costs on servers that aren't running and burning tokens on reviewers the user doesn't want for routine work (`#3079`). The only workaround today is patching `workflows/review.md` in place; that patch is wiped on every `/gsd-update` and requires `/gsd-update --reapply` to restore, with no machine-readable record of intent. The proposed key sits inside the existing `review.*` namespace (alongside `review.models.<cli>` and `review.*_host`), follows GSD's **absent = enabled** config philosophy, and is implementable as a one-line config read plus an intersection on the detected reviewer set.
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive; the Status field had sat stale at "Proposed" for roughly 65 days after the decision actually shipped.
**Evidence the decision shipped:**
- Landing commit `245d5f66a` ("feat: add review.default_reviewers config for /gsd-review defaults (#3464)", 2026-05-13) added the schema, resolution logic, workflow wiring, docs, and three test files in one change.
- `src/review-reviewer-selection.cts` (329 lines) exports `KNOWN_REVIEWER_SLUGS` (line 51) and `normalizeConfiguredDefaultReviewers` (line 105), implementing the ADR's precedence order (explicit flags > `--all` > `review.default_reviewers` > all detected).
- `src/config.cts:878` handles `kp === 'review.default_reviewers'` for `config-get`/`config-set`, running values through `normalizeConfiguredDefaultReviewers` and surfacing schema errors.
- `gsd-core/workflows/review.md` (no-flag branch, ~lines 55-70) intersects detected reviewers with `review.default_reviewers` exactly as specified, including unknown-slug warnings and undetected-slug info notes.
- `docs/CONFIGURATION.md:219-225` documents the key, type, default, and precedence; `docs/COMMANDS.md:1451-1461` documents usage with a `gsd config-set` example.
- Four test files are present and current: `tests/review-default-reviewers-config.test.cjs`, `tests/review-default-reviewers-resolution.test.cjs`, `tests/review-default-reviewers-workflow.test.cjs`, `tests/review-reviewer-instances.test.cjs`.
- `.changeset/archived/daring-badgers-munch.md` (type: Added, pr: 3464) is archived, confirming release tooling already processed it.
Governance state: the owning issue (`#3079`, referenced above) and its landing PR (`#3464`) both 404 against the current `open-gsd/gsd-core` tracker — their numbering belongs to a predecessor repo whose issue space predates this repo's 2026-05 range (which topped out near `#540`), consistent with known predecessor-repo numbering rather than a fabricated reference. No in-tracker close event is directly checkable; the shipped-code evidence above substitutes for it.
**Known gaps at ratification:** two of the ADR's own non-blocking open questions remain genuinely unresolved — Q-2 (`--no-default` flag) and Q-3 (`review.profiles.*` namespace) — exactly as the ADR itself scoped them as future/non-blocking, so this is expected rather than a regression.
## Decision
- Add **`review.default_reviewers`** to the `config.json` schema as `string[]`, validated against the existing CLI slug pattern `^[a-zA-Z0-9_-]+$` (the same pattern used for `review.models.<cli>` slugs).

View File

@@ -3,6 +3,8 @@
- **Status:** Accepted
- **Date:** 2026-05-12
- **Decision date:** 2026-05-12
- **Supersedes:** [ADR-0010](0010-skill-surface-budget-module.md) (Skill Surface Budget Module — earlier draft, install-time skill listing curation)
- **Subsumed by:** [ADR-857](857-capability-system.md) (Capability system) — generalizes this module; this seam remains live at `src/surface.cts:348` (`applySurface`)
- **Implementation:** feat/3408-skills-description-dropped-due-to-size, PR <TBD>
Every installed `gsd-*` skill costs eager system-prompt tokens: runtimes (Claude Code, opencode, and others) enumerate all skill descriptions in `<available_skills>` on every turn. With 66 skills and 33 agents, GSD alone consumes roughly 60% of the default 1%-of-context skill-listing budget, causing descriptions to drop when users stack multiple plugins (#3408).

View File

@@ -1,6 +1,6 @@
# CommandRoutingHub as single dispatch seam for CJS command families
- **Status:** Superseded by ADR-0174 (2026-05-23); originally Accepted (2026-05-20)
- **Status:** Superseded by [ADR-0174](0174-retire-gsd-sdk-package-boundary.md) (2026-05-23); originally Accepted (2026-05-20)
- **Date:** 2026-05-20
## Context

View File

@@ -7,6 +7,20 @@
- **Realizes:** [ADR-857](857-capability-system.md) Branch 8 (host-CLI support as `role: runtime` Capabilities)
- **Materializes:** [ADR-58](58-runtime-install-policy-module.md) (the typed `InstallPlan` projection)
- **Builds on:** [ADR-3660](3660-runtime-artifact-layout-module.md) (artifact layout), [ADR-894](894-capability-declaration-format.md) (the `role: runtime` body, already validated)
- **Subsumed by:** [ADR-1239](1239-gsd-embeddable-orchestration-engine.md) (GSD as an Embeddable Orchestration Engine) — read it first; see the amendment below
## Amendment (2026-07-16): subsumed by ADR-1239 (EoS) — this ADR is the *declarative adapter*, not the whole architecture
[ADR-1239](1239-gsd-embeddable-orchestration-engine.md) — **GSD as an Embeddable Orchestration Engine** (EoS), Accepted — subsumes this ADR **as the declarative adapter** in a larger frame, and inverts its direction of travel:
- This ADR answers *"how does GSD project its files onto a host CLI we already know?"* — GSD reaches into the host.
- ADR-1239 inverts that: **GSD is the engine; the host loads it through a negotiated Host-Integration Interface**, and a third party writes the thin host-plugin. ADR-1239 calls this "flips *projection* to *embedding*, and **unifies** them."
**This ADR is not superseded and its status is unchanged.** The runtime descriptor is real, live, and load-bearing: it remains the *declarative* adapter within EoS. But it is a **component of** the current architecture, not the statement of it. A reader who takes this ADR as the top-level answer to "how does GSD meet a host?" will reach the wrong conclusion for any new host.
**Read [ADR-1239](1239-gsd-embeddable-orchestration-engine.md) first.**
Recorded because ADR-1239 declared this subsumption while this file recorded nothing, leaving the pointer one-way and EoS undiscoverable from here.
## Context

View File

@@ -7,6 +7,14 @@
- **Blocked by:** [#857](https://github.com/open-gsd/gsd-core/issues/857) being **released** (Proposed → Accepted + capability infrastructure shipped). Not actionable until then.
- **Relates to:** [#853](https://github.com/open-gsd/gsd-core/issues/853) (Claude Code backgrounded agents cannot nest subagents), existing BETA skill `gsd-ultraplan-phase`
## Why this is still `Proposed` (audited 2026-07-17)
Confirmed shipped, on-tree: the capability is real and registered, not vaporware. `capabilities/claude-orchestration/capability.json` exists with detection + emission (`detectWorkflowBackend` / `emitWorkflowScript`) in `src/claude-orchestration.cts` (compiled to `gsd-core/bin/lib/claude-orchestration.cjs`), federated config (`claude_orchestration.enabled` / `execution_backend` / `min_agent_sdk_version`), and 1,552 lines of tests across `tests/claude-orchestration.test.cjs`, `tests/claude-orchestration-command-router.test.cjs`, and `tests/fix-2285-claude-orchestration-wiring.test.cjs`. The previously-fatal wiring bug, #2285 ("claude-orchestration capability (#1143) registered as active but never wired into execute-phase orchestrator prompt"), is closed COMPLETED (2026-07-15) — one day before this audit — and the owning feature issue #1143 is also closed COMPLETED.
**The blocker.** The ADR sets its own bar for ratification in its own Amendment (above): "flipping to Accepted follows maintainer sign-off on the E2E behaviour once exercised on Claude Code with the Workflow tool present." No such exercise is recorded anywhere in issues, PRs, or tests. Every test in the three files above operates at the contract or CLI-subprocess layer — asserting the *shape* of an emitted script or the return value of `resolve-wave-dispatch` — none constructs or executes an actual Workflow-tool run (`grep -rn "Workflow(" tests/claude-orchestration*.test.cjs tests/fix-2285-*.test.cjs` returns no hits). Two further gaps sit inside the ADR's own Decision section: (1) Decision §1's claimed net effect — "wave parallelism, the plan-checker, and the verifier are restored" — is narrower than what shipped: `capability.json`'s own description says "the plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates"; (2) Decision §3's fold-in of the `gsd-ultraplan-phase` skill into the capability's `skills[]` has not happened — `capability.json` still shows `"skills": []`, and no follow-up issue for the migration the Amendment promises exists (searched via `gh issue list --search`, no result).
**Unblock condition.** Ratify once: (a) a real Claude Code session with the Workflow tool present and `claude_orchestration.enabled=true` drives an `execute-phase` wave through the Workflow backend, and the result is recorded (issue comment, PR, or a test that actually builds/executes a `Workflow` script rather than asserting emitted-script shape) — that is the maintainer sign-off the ADR itself asks for; and (b) the Decision section's "plan-checker and verifier restored" language is reconciled with the shipped scope (either corrected to match, or backed by a tracked issue for the deferred wiring `capability.json` already discloses). The `skills[]` migration (item 3) is lower priority since it is openly disclosed as deferred rather than silently dropped, but should carry a tracked issue number before ratification so it doesn't quietly vanish.
## Context
Claude Code ships two orchestration primitives GSD does not yet treat as first-class:

View File

@@ -6,6 +6,16 @@
- **Completes:** Capability system (ADR-857) — the write half of the phase-4 "Wire" step
- **Builds on:** Capability declaration format (ADR-894), Capability command contribution (ADR-959), Skill Surface Budget Module (ADR-0011)
## Why this is still `Proposed` (audited 2026-07-17)
**What shipped.** The module this ADR decided is real and in production use: `setCapabilityState` / `cmdCapabilitySet` are implemented at `src/capability-writer.cts:140` and `:446`, wired into the CLI at `gsd-core/bin/gsd-tools.cjs:1941` and `:2314`, and `gsd-core/workflows/settings.md:448` routes gate writes through `capability set --gate`. `CONTEXT.md:249` carries the glossary entry, and the landing commit (`bf634b95c`, "feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)") is a confirmed ancestor of `origin/next`. The write-side invariant this ADR set out to build — off means off, enforced at write time — is in force.
**The blocker.** This ADR's own Decision section (lines 27–29, as written above) declares the writer's return shape as `{ capabilities: CapabilityStateEntry[]; warnings: string[] }`, and decision item 4 says post-write divergence "is returned as warnings, not silently swallowed" — a warnings-only channel, no separate hard-failure signal. The shipped code does not match that: `src/capability-writer.cts:113-117` defines `SetCapabilityStateResult` as `{ capabilities, warnings, errors }`, and `errors[]` is populated both by pre-write validation rejections (e.g. `"unknown capability"`, `"cannot enable ... not in the install profile"`, which abort with zero writes) and by post-write assert failures (e.g. `"failed to disable ... still surfaced after write"`), which the CLI (`cmdCapabilitySet`, lines 483–493) turns into a non-zero exit — a real hard-failure channel this ADR's Decision section does not describe. This is not drift or a bug: it is a later, deliberate redesign. ADR-1411 (Accepted; 2026-06-18 Amendment) states explicitly that `capability-writer`'s "`errors[]` (operation-not-applied) is load-bearing and cannot fold into `warnings[]` (advisory)" and records the mutation-verb shape as `{ capabilities, warnings, errors }` (ADR-1411 lines 81, 86) — superseding the two-field interface this ADR decided. ADR-1213's own text has never been updated to note the amendment or to revise the signature, so as written it misdescribes the interface actually shipped.
**Dropped claim.** A second refutation argument held that the parent ADR-857 carried an explicit governance caveat reserving any Proposed→Accepted flip in this ADR family for a maintainer, and that flipping ADR-1213 on shipped-code evidence alone would repeat a move ADR-857 itself refused to make unilaterally. That premise no longer holds: `docs/adr/857-capability-system.md` now reads "Status: Accepted — ratified 2026-07-17" with a "## Ratification (2026-07-17)" section, and the caveat text this argument quoted is no longer present anywhere in that file (confirmed by direct search). ADR-857 was ratified in the same 2026-07-17 audit pass that reviewed this ADR, so this argument is dropped rather than carried forward as a live blocker.
**Unblock condition.** Revise this ADR's Decision section — the return-shape signature and item 4's assert-and-report description — to match what shipped: `{ capabilities: CapabilityStateEntry[]; warnings: string[]; errors: string[] }`, with `errors` describing operation-not-applied hard failures (pre-write validation rejects, post-write assert failures) distinct from advisory `warnings`. Either fold in a one-line "Amended by ADR-1411" pointer or edit the signature directly. Once the Decision section states the interface actually in the tree, this ADR is ready to ratify — the underlying mechanism is already proven in production.
## Context
ADR-857 promised: *"one resolved capability state replaces three contradicting toggle systems; 'off' means off."* The **read** side delivers it. The **Capability State Resolver** (`src/capability-state.cts`) collapses three substrates into one resolved state:

View File

@@ -1,12 +1,33 @@
# ADR-1244 — Capability Ecosystem: third-party authoring, versioned manifests, and URL import/upgrade/remove
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-06-14); see "Ratification" below
- **Date:** 2026-06-14
> **Relationship to other ADRs.** This ADR **amends and extends ADR-857 Decisions 7 and 8** — it does not reverse them. ADR-857 D7 deferred third-party code-loading "to its own ADR"; D8 deferred third-party CLI support "to an external loader + trust/validation gate, no rework because runtimes are already descriptors." This *is* that ADR, and it *delivers* that gate. It builds on **ADR-894** (capability declaration format), **ADR-1016** (runtime capability descriptor), and **ADR-58** (InstallPlan seam). Tracked by [#1244](https://github.com/open-gsd/gsd-core/issues/1244). Target release: **1.6.0**.
---
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive; the Status field sat at Proposed for 33 days after the owning issue and all six phase sub-issues had already closed as shipped.
**Evidence the decision shipped:**
- Issue #1244 and all six phase sub-issues (#1430–#1435, Phase 1 through Phase 6) are CLOSED / `stateReason: COMPLETED`.
- **D1** (versioned manifest): `capabilities/*/capability.json` carry `version` + `engines` (confirmed in `ai-integration`, `antigravity`, `claude-orchestration`).
- **D2** (runtime overlay): `src/capability-loader.cts:486` exports `loadRegistry({ includeInstalled })`.
- **D3** (source resolver): `src/capability-source.cts:1080` exports `resolveCapabilitySource`, backed by the four adapters `resolveLocal` (861), `resolveGit` (892), `resolveNpm` (944), `resolveTarball` (1017).
- **D4** (ledger): `src/capability-ledger.cts` (42.4K) exists with `tests/capability-ledger.test.cjs` (111.0K) covering it.
- **D5** (trust model): `src/capability-trust.cts` and `src/capability-consent.cts` exist; `strictKnownRegistries` is threaded through `src/capability-lifecycle.cts` at lines 171, 881, 956, and 1079, each backed by `tests/capability-trust.test.cjs` and `tests/capability-consent.test.cjs`.
- **D6** (upgrade/compat): `src/capability-lifecycle.cts:1078` implements `upgradeCapability` under the documented atomic stage-then-swap (comment header at line 1056); `compatVersions` downgrade handling is present at lines 126, 920, and 1111.
- **D9** (capability matrix): `docs/reference/capability-matrix.md` (9.4K) exists and is generated from the registry.
Governance: owning issue #1244, `stateReason: COMPLETED`, closed 2026-07-07.
**Known gaps at ratification:** the D8 cross-reference promised back into ADR-857 ("D7 and D8... extended by ADR-1244") was never written — `docs/adr/857-capability-system.md` has no mention of ADR-1244. And epic #1900 (ADR-1244 edge hardening: MCP arg/cwd confinement, tarball/registry SSRF denylist, duplicated injection patterns) remains OPEN with all three of its filed children (#1901, #1902, #1903) closed `NOT_PLANNED` — the epic's own text scopes this as post-ship hardening on an already fail-closed pipeline, not a reversal of any D1–D9 decision, but the hardening itself is not yet scheduled.
---
## Context
ADR-857 turned the five-step loop into a **host** with **12 Loop Extension Points** and made every feature a **Capability** — a folder `capabilities/<id>/capability.json` declaring owned skills/agents, lifecycle hooks, a federated config slice, and loop-extension registrations (`step` / `contribution` / `gate`). 32 capabilities ship today (20 `role:feature`, 12 `role:runtime`). The architecture is in place; the **ecosystem is not**.

View File

@@ -1,11 +1,27 @@
# Cross-AI Plan Convergence via Existing Orchestration Commands
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-05-24); see "Ratification" below
- **Date:** 2026-05-24
- **Issue:** #15
Current orchestration commands (`/gsd-autonomous` and `/gsd-progress --next --auto`) route planning through `gsd-plan-phase` and only use local/Claude subagent review paths. The cross-AI convergence path already exists (`/gsd-plan-review-convergence`, `/gsd-review`, `review.default_reviewers`, `review.models.*`) but is not wired into these orchestrators. This creates a gap: users can configure cross-AI reviewers yet still get local-only planning in autonomous/auto-chain execution.
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive after the shipped implementation was independently re-verified; the Status field had read "Proposed" for roughly 8 weeks after the underlying decision had already landed.
**Evidence the decision shipped:**
- Primary, parity, and alias surfaces are present verbatim: `commands/gsd/progress.md:4,28` (`--next --converge`, `--cross-ai` alias, reviewer flags, `--max-cycles N`) and `commands/gsd/autonomous.md:4,40-41` (`--converge`, `--cross-ai` alias).
- The `plan_strategy=local|converge` seam is implemented in `gsd-core/workflows/next.md:260-313` (`PLAN_STRATEGY` parsing, `CONVERGENCE_ARGS` build, feature-gate check, Route-3 override) and mirrored in `gsd-core/workflows/autonomous.md:19-90,378-419`.
- Fail-fast-on-disabled-gate behavior matches the ADR's Failure Policy exactly: `next.md:279-292` and the equivalent block in `autonomous.md` check `workflow.plan_review_convergence` via `config-get` and abort with the exact `gsd config-set workflow.plan_review_convergence true` instruction — no silent downgrade to `local`.
- The config contract is shipped: `gsd-core/bin/shared/config-schema.manifest.json:36` (`workflow.plan_review_convergence`), `:54` (`review.default_reviewers`), `:123,141` (`review.models.*`); documented identically in `docs/CONFIGURATION.md:225,316` and `docs/COMMANDS.md:620-622,850-852`.
- Dedicated regression tests exist: `tests/adr-15-progress-converge.test.cjs` (179 lines, describe block titled `'ADR-15: /gsd:progress --next --auto --converge (#1190)'`) and `tests/autonomous-converge.test.cjs` (225 lines, covering the parity surface under `'autonomous --converge flag (#711)'` — this file does not itself reference ADR-15 by name).
- Landing commits: `092340d18` (`fix(#711): wire autonomous convergence flag`, 2026-06-10, parity surface) and `0b3a2e5f9` (`feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)`, 2026-06-14) — the latter's commit body states "ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface" and confirms the wiring gap the ADR called out is closed.
- No later ADR references or supersedes ADR-15: `grep -rl 'ADR-15' docs/adr/*.md` returns only `docs/adr/README.md`'s own index row (line 158), which still lists it as "Proposed" — the stale bookkeeping entry this ratification corrects.
**Governance state:** Issue #15 CLOSED — stateReason COMPLETED (closed 2026-05-25T03:12:26Z). Follow-up test-coverage issue #1190 ("test(coverage): fill Proposed-ADR test gaps") also CLOSED — stateReason COMPLETED (closed 2026-06-14T19:52:24Z).
## Decision
Do not add a new command. Add convergence as an orchestration policy in existing commands, with `/gsd-progress` as the primary operator surface.

View File

@@ -1,9 +1,26 @@
# ADR-1577: Untrusted-input boundary + opt-in injection blocking
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-06-25); see "Ratification" below
- **Issue:** [#1577](https://github.com/open-gsd/gsd-core/issues/1577)
- **Part of:** [#1573](https://github.com/open-gsd/gsd-core/issues/1573) (harden the agent layer against documented LLM failure modes)
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive; the Proposed status had gone unconfirmed for 22 days since the ADR landed on 2026-06-25.
**Evidence the decision shipped:**
- Issue #1577 is closed (`state=CLOSED`, `stateReason=COMPLETED`, closed 2026-06-24T21:07:24Z) as split A of the umbrella #1573, scoped exactly to this ADR's decision.
- `hooks/gsd-read-injection-scanner.js:118` extends the scanner to `SCANNED_TOOLS = new Set(['Read', 'WebFetch', 'WebSearch'])`, wired via `hooks/hooks.json:34`'s `"Read|WebFetch|WebSearch"` matcher — closing the WebFetch/WebSearch gap named in Context.
- `hooks/gsd-read-injection-scanner.js:212` gates blocking on `cfg.security?.injection_blocking === true`, read directly via `fs.readFileSync`/`JSON.parse` (independent of `src/configuration.cts`'s key whitelist, so no drop risk).
- `security.injection_blocking` is a registered config key end-to-end: `gsd-core/bin/shared/config-schema.manifest.json:109` lists it and `gsd-core/bin/shared/config-defaults.manifest.json:103` defaults it `false`; `src/configuration.cts:47` builds `VALID_CONFIG_KEYS` from that manifest and `src/config-schema.cts:61` (`isValidConfigKey`) consults it.
- `gsd-core/references/untrusted-input-boundary.md` exists and is `@`-included by exactly the 10 ingest agents named in the Decision: `gsd-advisor-researcher`, `gsd-ai-researcher`, `gsd-assumptions-analyzer`, `gsd-doc-classifier`, `gsd-doc-synthesizer`, `gsd-domain-researcher`, `gsd-phase-researcher`, `gsd-project-researcher`, `gsd-research-synthesizer`, `gsd-ui-researcher`.
- `docs/explanation/security-model.md:150-165` documents the PostToolUse pre-filter framing and names all 10 agents; `docs/CONFIGURATION.md:910` documents `security.injection_blocking` with a direct link to ADR-1577.
- `tests/read-injection-scanner.security.test.cjs` runs `SCAN-WF-01`, `SCAN-WF-02`, `SCAN-WF-03`, and `SCAN-WS-01` against the real hook subprocess for WebFetch/WebSearch payloads and asserts real detections.
- `tests/injection-blocking-config.test.cjs` asserts `isValidConfigKey('security.injection_blocking')` is true, `isValidConfigKey('security')` is false, and `CONFIG_DEFAULTS.security.injection_blocking === false`.
**Governance state:** owning issue #1577 — CLOSED, stateReason COMPLETED, closed 2026-06-24T21:07:24Z.
## Context
The research/doc-ingest agents concatenate text returned by WebFetch / WebSearch / Read into their context with no data/instruction separation, and the `gsd-read-injection-scanner` hook only scanned the `Read` tool — leaving WebFetch/WebSearch (the largest untrusted channel) unscanned. Prompt injection via fetched content is a documented LLM failure mode (arXiv [2506.05739](https://arxiv.org/abs/2506.05739), [2507.15219](https://arxiv.org/abs/2507.15219), [2504.20472](https://arxiv.org/abs/2504.20472)).

View File

@@ -37,6 +37,41 @@ addenda. The boundary:
this ADR, leaving 550 to own the contract and this ADR to own the mechanism. Until that is
agreed, 550's addenda remain authoritative and this ADR is non-binding.
## Why this is still `Proposed` (audited 2026-07-17)
**What shipped.** The audit confirmed the enforcement mechanism this ADR describes is real
and in place, not aspirational. All seven Decision points are present in
`src/prohibition-enforcement.cts` and `src/probe-core.cts`: `runProhibitionEnforcement`
(`src/prohibition-enforcement.cts:654-732`), `dispositionForProhibition`
(`src/probe-core.cts:470-516`), the vacuity guards `isNonVacuousNodeTestRed`
(`src/prohibition-enforcement.cts:351-355`) and `isNonVacuousNodeTestPass`, and
`defaultProveFailFirst`'s node-test branch (`src/prohibition-enforcement.cts:591-627`) —
which does implement the #1906 mandatory-`cleanFixture` causation control exactly as
Decision 4 / the 2026-07-03 addendum describe, not merely as a documented intent. Test
coverage is substantial (`tests/prohibition-enforcement.test.cjs`, 1336 lines), and all six
contributing issues (#644, #1259, #1278, #1279, #1346, #1906) are closed as COMPLETED on
GitHub.
**The blocker.** This ADR names its own precondition for becoming binding, in its own words:
"on accepting this ADR, replace ADR-550's 2026-06-12 / #1259 / #1279 / #1346 / #1278
enforcement addenda with a one-line pointer to this ADR ... Until that is agreed, 550's
addenda remain authoritative and this ADR is non-binding." That dedup has not happened.
Direct read of `docs/adr/550-spec-phase-probe-contract.md` confirms all four named addenda —
"Addendum (2026-06-12; updated 2026-06-15)", "Addendum (2026-06-15, #1279)", "Addendum
(2026-06-21, #1346)", and "Addendum (2026-06-15): optional `check` descriptor ... (#1278)" —
remain in ADR-550 in full, verbatim; none has been collapsed to a pointer. A later, separate
addendum in ADR-550 (2026-06-22, from #1607) does cross-reference ADR-1606 for the
*recall/representation-side* "Alternatives considered," but that is additive scaffolding, not
the enforcement-addenda dedup this ADR names as its own condition — the four target addenda
are untouched by it. No commit, PR, or tracked issue was found executing the dedup.
**Unblock condition.** Edit `docs/adr/550-spec-phase-probe-contract.md` to collapse the four
named addenda (2026-06-12/2026-06-15 update, #1279, #1346, #1278) into the one-line pointer
this ADR calls for, then flip both ADR-550's cross-reference and this ADR's Status in the
same PR. To check in minutes: grep `docs/adr/550-spec-phase-probe-contract.md` for
`## Addendum (2026-06-12`, `#1279`, `#1346`, and `#1278` — if those headings still carry the
full addendum text rather than a one-line pointer, the precondition remains unmet.
## Context
ADR-550 D4 originally specified the `test` tier as a "hard gate in both interactive and

View File

@@ -1,6 +1,6 @@
# ADR 1610: workflow & agent size-budget ratchet (per-file byte baseline + tier hard caps) [Proposed]
# ADR 1610: workflow & agent size-budget ratchet (per-file byte baseline + tier hard caps) [Accepted]
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-06-22); see "Ratification" below
- **Date:** 2026-06-22
> **Provenance.** Drafted 2026-06-22 to give an already-shipped architectural governance
@@ -13,6 +13,22 @@
> `scripts/lib/allowlist-ratchet.cjs` on `next`. The rationale here is lifted from those
> tests' own doc comments (the decision was documented in-code but never as an ADR).
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive after independent re-verification of the evidence below; the ADR had sat in `Proposed` for 25 days after the decision it documents had already shipped.
**Evidence the decision shipped:**
- Owning issue #1074 ("replace tier-max workflow size-budget ratchet with a per-file baseline + loose hard caps") is CLOSED, `stateReason: COMPLETED`, closed 2026-06-12T14:00:56Z.
- All three landing PRs are MERGED: #1089 (`test(#1074): add additive per-file workflow size baseline guard`, 2026-06-12T03:59:57Z), #1096 (`test(#1074): swap workflow size enforcement to baseline + loose hard caps`, 2026-06-12T13:30:44Z), #1097 (`test(#1074): agent-size-budget per-file baseline + line→byte rebase`, 2026-06-12T13:58:44Z).
- `scripts/workflow-size.cjs:32-35` — `lfByteCount()` implements the CRLF→LF-normalized byte count described in Decision point 2 (#683).
- `scripts/workflow-size.cjs:64-72,80-82` — `measureMdFiles`/`measureWorkflows` is the single shared measurement path cited in Decision point 5, re-exported for both the guard and `scripts/update-size-baseline.cjs`.
- `scripts/lib/allowlist-ratchet.cjs:180` exports `assertFileBaseline` — the per-file baseline assertion named in Decision point 3 and Cross-references.
- `tests/workflow-size-budget.test.cjs:95-97,102` defines `XL_CAP = 98304` (96 KiB), `LARGE_CAP = 61440` (60 KiB), `DEFAULT_CAP = 40960` (40 KiB), `NEW_FILE_CAP = 32768` (32 KiB) — the exact numbers quoted in Decision point 3.
**Governance:** owning issue #1074, `stateReason: COMPLETED`, closed 2026-06-12T14:00:56Z.
## Context
`gsd-core/workflows/*.md` and `agents/*.md` are loaded **verbatim into agent context** every

View File

@@ -1,10 +1,25 @@
# Existing Code Onboarding Module owns deterministic repo-state detection and onboarding route selection
- **Status:** Proposed
- **Status:** Accepted — ratified 2026-07-17 (originally Proposed 2026-07-06); see "Ratification" below
- **Date:** 2026-07-06
- **Issue:** #1990
- **Implementation:** PR #1994
## Ratification (2026-07-17): Proposed → Accepted
Ratified by explicit maintainer directive after independent re-verification of the evidence below; the Status field sat at Proposed for 11 days after the decision shipped.
**Evidence the decision shipped:**
- Issue #1990 ("Add /gsd:onboard for existing-codebase setup") is CLOSED, stateReason COMPLETED, closed 2026-07-07T04:23:56Z; PR #1994 ("feat(#1990): add brownfield onboarding workflow") is MERGED into `next` at 2026-07-07T04:23:55Z with body `Closes #1990`.
- `src/onboard-projection.cts` (15,248 bytes) and its compiled `gsd-core/bin/lib/onboard-projection.cjs` are both present on disk, implementing the projection this ADR describes.
- `src/init.cts:56` imports the projection as `onboardProjection`, and `src/init-command-router.cts:75` wires the `onboard:` route that consumes it — confirming the Init Command Module integration (the ADR's literal handler name "initOnboard" is not itself a grep-matched symbol; the consumption is the router entry plus the destructured import).
- `gsd-core/workflows/onboard.md`, `commands/gsd/onboard.md`, and `skills/gsd-onboard/SKILL.md` all exist on disk, matching the "What stays OUTSIDE this Module" boundary.
- `tests/onboard-command.test.cjs` (25,129 bytes) contains named tests covering the load-bearing gate order — `routes planning artifacts without PROJECT.md to partial planning`, `fast mode routes incomplete planning to partial-planning before the complete-map gate (regression #1990: fast map gate misroute)` — vendor exclusion (`ignores generated and vendor directories when detecting existing code`), and package-manifest brownfield detection (`treats package manifests as brownfield even without source files`).
- Six commits tagged `#1990` landed the ADR, the projection module, and doc/index updates: `3c7d722ed`, `e8fb05e96`, `d0b8eacd3`, `1171499f3`, `bc751a64e`, `192764f0c`.
**Governance state:** Owning issue #1990 — CLOSED, stateReason COMPLETED, closed 2026-07-07T04:23:56Z.
## Context
GSD already ships strong individual primitives for adopting an existing codebase: `/gsd:map-codebase` (parallel codebase analysis), `/gsd:ingest-docs` (classify and consolidate existing ADR/PRD/SPEC/RFC docs), and `/gsd:new-project` (planning initialization). What it lacked was a single guided entry point that inspects a brownfield repository and tells the user *which primitive runs first*.

View File

@@ -1,4 +1,4 @@
# ADR-0175: Harden release-workflow version validation — reject leading zeros and pre-check npm
# ADR-218: Harden release-workflow version validation — reject leading zeros and pre-check npm
- **Status:** Accepted (2026-05-24)
- **Date:** 2026-05-24

Some files were not shown because too many files have changed in this diff Show More