* fix(3678): executor must respect commit_docs:false; teach SDK skip envelope
Closes#3678
When `commit_docs: false` in `.planning/config.json`, the SDK's
`cmdCommit` correctly short-circuits and returns
`{committed: false, hash: null, reason: 'skipped_commit_docs_false'}`
without staging or committing anything. The agent prompt at
`agents/gsd-executor.md:710-720` (final_commit block) tells the executor
to call `gsd-sdk query commit "docs(...)" --files .planning/...` but says
NOTHING about how to interpret a skipped return. With no explicit
instruction, the LLM improvises raw `git add` / `git add -f` / `git commit`
to "fulfill" the per-plan commit step it was told to make, which leaks
gitignored `.planning/` artifacts into the user's git history (exactly
what the reporter observed).
Three coordinated fixes:
1. **agents/gsd-executor.md final_commit block** — adds explicit handling
text for all three SDK return envelopes (`committed:true`, `skipped:true
commit_docs`, `skipped:true gitignored`, `committed:false other reasons`).
States plainly: "Do not fall back to raw `git add` / `git commit` /
`git add -f` when the SDK returns `skipped: true`."
2. **get-shit-done/bin/lib/commands.cjs cmdCommit** — adds `skipped: true`
to both skip-path envelopes so agents see "skipped" as a first-class
success signal rather than inferring "no commit happened, I must
improvise" from absent `hash` / `committed:false`. Backward-compatible:
existing callers reading `committed` / `hash` / `reason` are unaffected.
3. **tests/bug-3678-executor-commit-docs-respect.test.cjs** — 7-test
regression covering:
- A1/A2: agent prompt mentions the skip envelope AND explicitly forbids
raw-git fallback (`source-text-is-the-product` exception)
- B1: SDK envelope carries `committed:false`, `skipped:true`, canonical
`reason: 'skipped_commit_docs_false'` (frozen enum)
- B2: git index empty after commit_docs:false skip (no `.planning/` staged)
- B3: HEAD unchanged after commit_docs:false skip
- C1/C2: structural ban on `git add -f` / `git add --force` in any agent
or workflow body (prohibition-sentence exception preserves audit prose)
Verification:
- node --test tests/bug-3678-*: 7/7 pass
- Targeted regression (10 commit/executor-adjacent files): 135/135 pass
- Full docker suite (gsd-test-summary): 11751/11740 pass / 0 fail
(the 11 added are this test plus a few collateral pickups)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(changeset): add fragment for #3678 fix (Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>)
* chore(changeset): set PR number 3679 (Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>)
* fix(3678): preserve skip-aware carve-out in executor completion checklist
The new `final_commit` prose at lines 717-741 teaches the executor to
treat `skipped:true` as success and forbids raw-git fallback, but the
downstream completion checklist still contained an unconditional
"Final metadata commit made" checkbox. An LLM executor reading an
unchecked mandatory box may attempt to satisfy it via raw `git add`,
re-introducing the exact regression this PR is meant to prevent.
Update the checklist line to carve out the intentional-skip case and
add a regression test asserting the carve-out remains present.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(3683): normalize /gsd:<cmd> → /gsd-<cmd> in command, workflow, and reference bodies
Extends #3677's agent-body normalizer to all body text staged through
copyWithPathReplacement (commands, workflows, references). The initial
isCommand guard was structurally redundant — normalizeAgentBodyForRuntime
already self-gates on shouldNormalizeHyphenNamespaceInAgentBody(runtime),
so dropping it covers all hyphen-name runtimes (Claude / Qwen / Hermes)
without affecting colon-canonical runtimes (Gemini).
Addresses the user-visible symptom in #3683: workflows like
get-shit-done/workflows/discuss-phase.md (7 colon refs) leaked /gsd:<cmd>
markers to the model context, which the model echoed at the end of
/gsd-discuss-phase runs.
Source-prose drift caught by the new cross-reference invariant test:
- commands/gsd/plan-phase.md: removed a slash-form mention of the
deleted /gsd-research-phase command (#3042)
- commands/gsd/profile-user.md: replaced a slash-form artifact
reference with a backticked bare name (the referenced item is a
skill config, not a user-callable slash command)
Tests:
- tests/bug-3683-command-colon-namespace-leak.test.cjs — runtime-form
regression for commands/gsd/*.md staging
- tests/bug-3683-command-cross-reference-invariant.test.cjs — locks
cross-reference coherence: every /gsd-X / /gsd:X reference in a
command body must resolve to commands/gsd/X.md (so a future rename
forces every cross-reference to update)
- tests/bug-3683-workflow-colon-namespace-leak.test.cjs — runtime-form
regression for workflows + references; negative test for gemini
asserting the colon form is preserved
Fixes#3683
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(changeset): remove undefined cycle reference
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(3677): normalize /gsd:<cmd> → /gsd-<cmd> in agent bodies for hyphen-name runtimes
Closes#3677
The executor agent bodies installed to `~/.claude/agents/gsd-*.md` (and
the Qwen / Hermes equivalents) still contained retired `/gsd:<cmd>`
colon-form references in their prose. Every GSD skill / agent has
registered under the canonical hyphen `name:` form since #2808, so the
colon form is unroutable — Claude Code rejects it with `Unknown command:
/gsd:execute-phase. Did you mean /gsd-execute-phase?`. Reporter measured
~28 agent files / ~96 leaked refs on a full Claude global install.
This is the agent-body surface of the same class of bug as the two
already-fixed sibling surfaces:
- #3583 (SKILL.md skill bodies) — fixed via #3629
- #3584 (user-facing runtime "Next step: /gsd:…" emissions) — fixed via #3606
The agent-body surface in `bin/install.js`'s agent install loop was
never covered: the Claude-default / Qwen / Hermes branches register
hyphen `name:` but copy bodies verbatim (Qwen/Hermes do branding-only
swaps; Claude-default falls through with no body conversion at all), so
the colon refs leak.
Fix:
1. Add a pure predicate `shouldNormalizeHyphenNamespaceInAgentBody(runtime)`
backed by an explicit allow-list `HYPHEN_NAME_AGENT_RUNTIMES =
{claude, qwen, hermes}`. Unknown / future runtimes default to false
(better to leak than to mangle).
2. Add `normalizeAgentBodyForRuntime(content, runtime, cmdNames)` that
conditionally applies the shared `transformContentToHyphen` from
`scripts/fix-slash-commands.cjs` (same transform #3629 used for
SKILL.md bodies).
3. Call `normalizeAgentBodyForRuntime(content, runtime, readGsdCommandNames())`
in the agent install loop right before `fs.writeFileSync`, so it
composes with all the existing runtime branches. For Gemini and
self-converting runtimes the predicate short-circuits, so their
convertClaudeAgentToXAgent output is not re-rewritten.
4. Export both functions from `bin/install.js` for the regression test.
Regression test (`tests/bug-3677-agent-colon-namespace-leak.test.cjs`):
24 tests across 4 groups — A (exports exist), B (predicate matrix
covering all 15 runtimes in the layout table + an unknown-runtime case),
C (normalize helper applies/skips correctly for claude/qwen/hermes/
gemini/copilot), D (sanity check of the underlying transform).
Verification:
- node --test tests/bug-3677-*: 24/24 pass
- Sibling-regression (6 slash-namespace test files): 76/76 pass
- All install-minimal-all-runtimes suites: 54/54 pass after `npm run
build:sdk` (the prior 27 fails were pre-existing — missing local
sdk/dist build, not introduced by this change)
- Full docker suite (gsd-test-summary): 11769/0 fail
(11751 baseline + 18 new = my 24 tests with some collateral pickups)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(changeset): set PR number 3680 (Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>)
* test(3677): port real-source efficacy + idempotence tests from #3681
Adds describe group E with 5 behavioral tests credited to John Turner
(johnzilla, PR #3681 — closed in favor of this PR by its author):
E0: command roster is populated and includes symptom commands
E1: every agents/gsd-*.md transforms clean — real-source efficacy
E2: idempotent — repeat transform on hyphenated input is a no-op
E3: word boundary — /gsd:plan-phase-extra is not a roster match
E4: rewrites bare gsd:<cmd> shorthand (no leading slash)
E1 is the test that would have caught the original bug — pure-function
tests can pass while the install.js wiring silently bypasses the
transform. E2 guards against double-rewrite mangling during reinstall.
29/29 tests pass (24 original + 5 ported).
Co-Authored-By: John Turner <johnzilla@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: John Turner <johnzilla@users.noreply.github.com>
Phase 2 of #3660 / ADR-3660. Routes both lifecycle verbs through the
Runtime Artifact Layout Module landed in Phase 1 (#3663):
- Add installRuntimeArtifacts(runtime, configDir, scope, resolvedProfile)
and uninstallRuntimeArtifacts(runtime, configDir, scope) as the public
orchestrators. Both pre-prune stale gsd-* entries before staged copy;
installRuntimeArtifacts brackets the prune+copy with preserveUserArtifacts /
restoreUserArtifacts so user-owned content (e.g. gsd-dev-preferences) survives
wipe-and-replace for claude/qwen/hermes runtimes.
- Add applyRuntimeContentRewritesInPlace as the per-runtime path/branding
post-stage step (preserves byte-output equivalence with the legacy
copyCommandsAs* pipeline, including Qwen/Hermes branding rewrites).
- Add _copyStaged, _removeGsdEntries kind-aware filesystem helpers.
- Add _runLegacyInstallMigrations, _runLegacyUninstallCleanup as thin
dispatchers over existing ADR-0008 legacy migrations (Hermes flat->nested
per #2841, dev-preferences-as-skill per #2973). For Hermes, also clean up
the intermediate skills/gsd/gsd-*/ layout that pre-Phase-2 installs left
on disk.
- Delete the 9 copyCommandsAs*Skills functions (Codex / Cursor / Windsurf /
Trae / CodeBuddy / Copilot / Claude / Antigravity / Augment) and the
_copyCommandsAsSkillsViaConverter helper. All test entry points migrated
to call installRuntimeArtifacts directly through the unified seam.
- Collapse the 9-branch uninstall ladder to one uninstallRuntimeArtifacts
call plus preserved non-layout side-effects (Codex TOML, Copilot
instructions, hooks).
- Unify install dispatcher: a single _isSkillsRuntime gate routes all 11
skills runtimes through installRuntimeArtifacts for both full and core/
minimal profiles. Removes 11 per-runtime if-else branches (3 minimal-mode
shim branches + 8 dead after-the-gate branches).
Net delta on bin/install.js: 11,495 -> 11,174 (-321 LOC).
New tests:
- tests/install-uninstall-layout-loop.test.cjs (34 tests) - per-runtime
fixture assertions on install/uninstall/legacy-migration ordering.
- tests/install-hermes-regressions.test.cjs (6 tests) - covers the six
defects surfaced by iterative review: Hermes upgrade leaves stale dirs,
--hermes --profile=core fall-through, --qwen --profile=core fall-through,
minimal-mode dev-preferences migration skipped (Hermes/Qwen/Claude-global),
and ordering bug in _runLegacyInstallMigrations.
Existing tests (10,038 prior + 40 new) all green: 10,078/10,078 pass.
Refs #3664
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After PR #3649 (merge 40a442b2), test (windows-latest, 24) failed at npm ci with zero stdout/stderr (16s opaque exit). Identical code passed on the PR. Root cause: the npm ci and npm run build:sdk steps in .github/workflows/test.yml (jobs test and coverage) lacked an effective shell — neither step-level nor via defaults.run.shell. On windows-latest, Actions defaults to pwsh; npm.cmd → node.exe → npm-cli.js child-process chain under pwsh can swallow stderr.
Pin shell: bash on the 4 outlier steps. Add tests/workflow-shell-pinning.test.cjs that scans every workflow file referencing windows-latest, computes effective shell with proper workflow- and job-level defaults.run.shell inheritance, and fails CI if any run: npm|npx … step lacks an effective shell.
Closes#3672
Replace the blunt kindPrefix !== '' guard in _syncGsdDir with a
manifest-membership discriminator. For non-empty prefix runtimes, behavior
is unchanged (prefix match). For Hermes (empty prefix), a directory is
GSD-owned iff its stem appears in the canonical manifest; only those dirs
are removal candidates when absent from the staged set. Dirs not in the
manifest (user-owned) are preserved unconditionally. When no manifest is
provided (legacy callers), the removal pass is skipped (conservative fallback).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remove the fs.existsSync(dest) guard in applySurface so _syncGsdDir is
always called. _syncGsdDir already does mkdirSync(..., { recursive: true })
so the destination is created when absent — recovering partially-initialized
or user-deleted runtime config dirs. Also threads manifest through to
_syncGsdDir as optional 4th arg (used by P1-2 Hermes fix).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remove the 2 empty try/finally wrappers (finally bodies contained only
comments, no cleanup actions). Inline the assertions directly; add a
comment noting stagedDir lifecycle ownership. No behavior change.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace all 13 try/finally blocks with t.after() per-test cleanup hooks.
Import createTempDir/cleanup from tests/helpers.cjs; use createTempDir
inside createFixtureSkillsDir and createFixtureAgentsDir. Remove os import
(now unused at top level).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
convertClaudeCommandToClaudeSkill(content, skillName, runtime, cmdNames) uses
the runtime arg to gate Hermes/Qwen branding and version: frontmatter emission
(#2808, #3583). Previously the layout module called it with only 2 args so the
runtime-specific formatting was never applied.
Changes:
- skillsKind() gains a runtime param (5th arg after converterName).
- stage() computes cmdNames = readGsdCommandNames() once per call (perf: avoids
repeated fs.readdirSync in the converter) and wraps the real converter so all
4 args are forwarded.
- readGsdCommandNames added to bin/install.js GSD_TEST_MODE exports block so the
stage closure can call it without requiring the script separately.
- All switch arms updated to pass the canonical runtime string.
Converters that do not inspect runtime/cmdNames (Cursor, Codex, Copilot, etc.)
accept and ignore the extra arguments — no behaviour change for those runtimes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously runtime-artifact-layout.cjs set process.env.GSD_TEST_MODE at the
top level before requiring bin/install.js. Any caller that required this module
would silently inherit the GSD_TEST_MODE='1' side-effect for the lifetime of
the process.
Replace with a lazy loader (loadInstallExports / getInstallExports) that saves
the current GSD_TEST_MODE, sets it to '1' only if it was undefined, calls
require(), and restores the original value in a finally block. The exported
module.exports cache (_installExports) ensures the require is run at most once.
Converter names are now passed as strings to skillsKind; the stage closure
resolves them via getInstallExports()[converterName] at call time, so the
installer is not loaded until a stage() function is actually invoked.
Verification: node -e "require('./...runtime-artifact-layout.cjs'); console.log(process.env.GSD_TEST_MODE)"
prints undefined.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Both findInstallSourceRoot and findAgentsSourceRoot now accept an optional
runtimeConfigDir. When provided, they first check <runtimeConfigDir>/.gsd-source:
read the stored path, verify it exists, and return it. Fall through to the
path.dirname walk-up only when the marker is absent or points to a missing path.
The factory functions (commandsKind, agentsKind, skillsKind) now receive configDir
from the switch arms and pass it through to the finders. listSurface in surface.cjs
passes its runtimeConfigDir argument to findInstallSourceRoot so installed-from-source
layouts use the marker rather than the walk-up.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Both findInstallSourceRoot and findAgentsSourceRoot previously returned a
path.join(__dirname, '..', '..', '..', ...) fallback when the walk-up loop
found nothing. Replace with throw per CLAUDE.md "No Relative Path Traversal"
policy: .. chains silently break on CWD changes; a throw with the failing
__dirname in the message is an unambiguous diagnostic.
The walk-up loop already uses path.dirname iteratively (no literal ..); only
the dead-end fallback used the banned pattern.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When kindPrefix === '' (Hermes: destSubpath=skills/gsd, no per-skill prefix),
startsWith('') always returns true so the prior removal loop would delete any
dir not in the staged set — including user-owned skill dirs. Guard the entire
removal block behind kindPrefix !== '' so non-staged dirs are never pruned when
there is no prefix to distinguish GSD-owned from user-owned entries.
TDD: failing test added first asserting user-custom-skill is preserved through
a _syncGsdDir call with kindPrefix=''.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
User-visible: applySurface now prunes skills/gsd-*/ dirs on profile
switch, closing the structural gap behind #3659.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Flip status from Proposed → Accepted and record the branch,
implementation reference, and Phase 2 dependency.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Tests verify that kind.stage(resolvedProfile) produces correct directory
structure: .md files for commands, agent .md files for agents, and
gsd-<stem>/SKILL.md dirs for skills kinds.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
RED-2/GREEN-2: wildcard skills='*' stages all *.md as <prefix><stem>/SKILL.md
with converter applied. Includes existence guard and error cleanup.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(3596): adversarial security/prompt-injection abuse suite
Adds `tests/security-prompt-injection.test.cjs` and a fixtures
directory at `tests/fixtures/adversarial/security/` covering the
attack classes enumerated in #3596:
- Command substitution / backticks / heredoc payloads in workstream
names — sentinel-file probes prove no shell is spawned, slugifier
neutralises the input.
- Path traversal through `--ws` and slash-bearing workstream names —
rejected with structured `--json-errors` payload, no stack trace,
no filesystem mutation outside the project root.
- Fake `<system>` / `[SYSTEM]` / `<<SYS>>` / `[INST]` boundary tags —
sanitizeForPrompt neutralises every form; structural negative
property locked across all six styles in one place.
- Zero-width / bidi-override codepoints — stripped per the documented
codepoint set; asserted via codePoint inspection, not regex
literals.
- Hostile read of CONTEXT.md / PLAN.md / ROADMAP.md fixtures —
`gsd-read-injection-scanner.js` surfaces the advisory; excluded
paths and non-Read tools stay silent; malformed JSON does not
crash the hook.
- Hostile write of `.planning/` files — `gsd-prompt-guard.js` emits
a `PreToolUse` advisory; non-Write/Edit tools stay silent.
- Fake `ghp_*` / `sk-*` env tokens — never echoed in CLI stdout or
stderr under hostile inputs; covered under
`// allow-test-rule: structural-regression-guard` because the only
way to assert byte-level absence is `.includes(token)` against the
captured streams.
- `validatePath`, `validateShellArg`, `validatePhaseNumber`,
`validateFieldName` — focused negative-input contract pins.
Pinned behavior gaps (documented, NOT fixed in this PR):
- `<instructions>` is intentionally whitelisted by both the scanner
and the sanitiser (GSD's own prompt scaffolding). Two REGRESSION
GUARD tests lock that contract.
- The current `scanForInjection` does NOT flag malicious markdown
links (javascript:/data:/embedded-credentials URLs). PINNED with
negative-proof so any future scope extension fails the assertion
and forces a deliberate update to the acceptance map.
- `prompt-builder.ts` does not yet wrap plan/context markdown in an
"untrusted data" envelope. That seam lives on the TS side and is
covered by `sdk/src/prompt-builder.test.ts`; out of scope for a
CJS test file. Mentioned in the file header.
Verification:
- `node --test tests/security-prompt-injection.test.cjs` → 73 tests
pass.
- `node scripts/lint-no-source-grep.cjs` → 0 violations across
546 test files (one `allow-test-rule: structural-regression-guard`
annotation on this file for the token-absence assertions).
- `node scripts/run-tests.cjs` → 9730 tests pass, 0 fail.
Refs #3596
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(3596): allow adversarial fixtures in scan + harden graphify status parse
* fix(3596): skip adversarial security fixtures in secret scan
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds tests/feat-3598-generator-correctness.test.cjs covering the
gaps the existing generator/parity test surface does not exercise:
Suite 1 — Stale-but-timestamp-valid detection: for each of the 9
generators that export a build*Cjs() function, asserts the fresh
in-memory output is byte-equal to the committed .generated.cjs.
Catches manual edits and partial-write drift that timestamp-only
freshness checks miss.
Suite 2 — Determinism: calls each build*Cjs() twice and asserts the
outputs are identical. Catches time/random/iteration-order regressions.
Suite 3 — Runtime/SDK alias parity: asserts canonical command names
and alias sets are identical between
get-shit-done/bin/lib/command-aliases.generated.cjs (runtime) and
sdk/src/query/command-aliases.generated.ts (SDK). Beyond timestamp
freshness.
Suite 4 — No duplicate aliases in the live registry: behavioral
equivalent of the issue's example test. The generator has no
fixture/--source seam (it reads in-memory COMMAND_DEFINITIONS_BY_FAMILY),
so the structural invariant is asserted on the deployed surface; a
collision fails with both colliding canonicals named.
Suite 5 — build-hooks.js atomicity: runs the build twice, snapshots
hooks/dist/ each time, asserts byte-identical output. Also asserts
no orphaned .dist-staging-* sibling directories remain and that every
.js file shipped to dist parses (positive proof of the vm.Script
syntax guard).
26 tests pass on macOS / Node 24.
Closes#3598
* docs(3660): add Runtime Artifact Layout Module ADR + CONTEXT.md entry
Records the architectural decision to add a Runtime Artifact Layout Module
that owns per-runtime artifact placement (commands / agents / skills) as a
typed seam. Closes the bug class behind #3659 where the Runtime Surface
Module forgets artifact kinds the install/uninstall pipelines track.
CONTEXT.md gains the new module entry directly after the Skill Surface
Budget Module since the two seams compose.
Pure docs — no code changes. Phase 1 implementation lands in a separate PR.
Closes#3660
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(3660): fix ADR parity and clarify phase boundaries
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Resolves conflict in .github/workflows/test.yml: keep the 6 new drift-check
steps from main (plan-scan, secrets, schema-detect, decisions,
workstream-name-policy, Shared Module hand-sync) before the split-lane test
runs from this PR. PR's dedicated `coverage` job replaces main's per-matrix
`Run tests with coverage` step.
Other conflicting files (CONTRIBUTING.md, get-shit-done/bin/lib/init.cjs,
package.json) auto-merged cleanly. Changeset files (.changeset/*) brought in
from main as adds.