* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
12 KiB
GSD Core security model
Explanation — This document describes why GSD Core has the security posture it does and how the layers fit together. It is not a reference for every hook parameter. For the
/gsd-secure-phasecommand and its options, see Commands. For the implementation-level hook architecture, see Architecture § Hook System. For the org-wide security baseline (scanner controls, incident checklists, ownership model), see SECURITY.md.
Why AI-driven development needs a dedicated security posture
A conventional code editor does not execute arbitrary packages on your behalf.
GSD Core does. The research → plan → execute pipeline automates the full path
from "name a package" to "run npm install <package>", from "write a
planning artifact" to "use that artifact as an LLM system prompt". Each
automation step removes a human from the loop — and each removal is a
potential attack surface.
GSD Core's security model is built around one organising principle: defence in depth. No single control is assumed to be perfect. Several overlapping layers each reduce a distinct class of risk, and together they make the attack surface substantially harder to exploit without eliminating it entirely. The honest summary at the end of this document explains what the system cannot protect against.
Layer 1 — Supply-chain protection: the Package Legitimacy Gate
The threat
AI models hallucinate package names. This is not a fringe failure mode: 2025 research documents roughly 20 % of AI-generated package references as hallucinated names that do not correspond to legitimate packages. A subset of those hallucinated names — approximately 43 % in the same research — recur consistently across prompts, meaning an attacker can observe which names AI tools commonly produce and pre-register those names on npm, PyPI, or crates.io with malicious post-install scripts. The technique is called slopsquatting.
The insidious quality of slopsquatting is that a hallucinated name that passes
npm view looks legitimate. The registry entry proves only that someone
registered the name — not that the package does what the AI said it does, not
that it has any legitimate users, and not that its install scripts are safe.
Without a gate, a hallucinated name would flow undetected through GSD's
researcher → planner → executor pipeline and eventually run as
npm install <attacker-package> on your machine.
How the gate works
The gate operates across three pipeline stages:
Research stage. When gsd-phase-researcher recommends external packages,
it runs slopcheck install <pkgs> --json against each one. The results are
written to a ## Package Legitimacy Audit table in RESEARCH.md. Packages
tagged [SLOP] (high-confidence hallucination or attacker-registered) are
stripped from RESEARCH.md entirely before the file is saved. They never
reach the planner.
Planning stage. gsd-planner reads the Audit table. For any package
tagged [SUS] (suspicious: newly registered, low download count, no source
repository, or naming pattern close to a popular package) or [ASSUMED]
(sourced from WebSearch rather than direct registry verification), the planner
inserts a checkpoint:human-verify task before the install step. The
checkpoint includes a direct link to the registry page and specific things to
look for: maintainer history, issue-tracker activity, absence of suspicious
install scripts.
Execution stage. If an install fails, gsd-executor surfaces a
checkpoint and stops. It does not silently try an alternative package name —
which could itself be malicious. This is an explicit rule in the executor's
behaviour (RULE 3 in the executor agent definition).
Why WebSearch packages are always [ASSUMED]
Package names discovered through WebSearch are tagged [ASSUMED] regardless
of whether npm view succeeds. A package that exists on the registry is not
the same as a package that is safe to install. npm view proves registration,
not legitimacy. The [ASSUMED] tag triggers the same human-verify checkpoint
as [SUS], ensuring that any unverified web-discovered recommendation always
gets a human review before installation.
Ecosystem coverage
The researcher uses registry-specific verification commands rather than a single generic check:
- Node.js:
npm view - Python:
pip index versions - Rust:
cargo search
This covers cross-ecosystem hallucination, which occurs at roughly 9 % according to 2025 USENIX research — cases where an AI recommends a package that exists in one ecosystem but not the one actually in use.
Graceful degradation
If slopcheck is unavailable (not installed, or the pip install fails at
research time), GSD applies the strictest possible fallback: every
recommended package is tagged [ASSUMED], and the planner gates every
install with a checkpoint:human-verify task. Research and planning proceed
normally — the system never hard-fails on a missing tool dependency. This
is intentionally stricter than the normal flow: slopcheck unavailability means
every package install gets a human checkpoint.
The slopcheck tool is MIT-licensed and pip-installable. If it is ever
abandoned, the [ASSUMED]-gate fallback ensures human-checkpoint coverage is
maintained regardless.
Layer 2 — Prompt injection defences
The threat
GSD Core generates Markdown files that become LLM system prompts. The
research pipeline reads external web content; the planning pipeline
incorporates user-supplied text (--text-file, --prd); the execution
pipeline writes planning artifacts that are later re-read as agent context.
Any user-controlled text flowing into these artifacts is a potential
indirect prompt injection vector — an attacker-controlled string that,
once inside a system prompt, attempts to override the agent's instructions or
exfiltrate information.
How the defences work
GSD Core addresses prompt injection at three levels.
Input validation (security.cjs). The gsd-core/bin/lib/security.cjs
module is the central security utility. It provides:
- Path traversal prevention: user-supplied file paths (
--text-file,--prd) are validated to resolve within the project directory, with macOS/var→/private/varsymlink resolution handled explicitly - Prompt injection detection: known injection patterns (role overrides, instruction bypasses, system tag injections) are scanned in user-supplied text before it enters any planning artifact
- Safe JSON parsing: a wrapper that prevents prototype-pollution attacks via crafted JSON payloads
- Shell argument validation: arguments passed to subshell commands are validated before use
Runtime hook: gsd-prompt-guard.js. This hook fires on every Write or
Edit call that targets .planning/ files. It scans the content being written
for the same injection patterns as security.cjs (a subset inlined directly
into the hook for independence — the hook does not require() the module, so
it runs even if the module path changes). Detection is advisory-only: the
hook logs the finding but does not block the write. The rationale is that a
false-positive block on a legitimate planning write would be more disruptive
than a missed injection in a secondary scan layer.
Runtime hook: gsd-read-injection-scanner.js. This hook fires on the
output of every Read tool call. It scans the content that was just read for
injected instructions in untrusted content — catching cases where an attacker
has embedded instructions in a file that GSD is about to incorporate into an
agent's context.
CI scanner. prompt-injection-scan.test.cjs scans all agent, workflow,
and command files for embedded injection vectors as part of the test suite.
This catches injection attempts in the GSD source itself — for example, a
supply-chain attack that modified a workflow file to add a role-override
instruction.
Read Injection Scanner vs Prompt Guard
The two hooks cover complementary surfaces. gsd-prompt-guard.js watches
writes to planning artifacts — it catches injection being planted.
gsd-read-injection-scanner.js watches reads of any file — it catches
injection being ingested from external content (a dependency's README, a
third-party config file, a user-provided document). Together they bracket
the ingest → store → re-read lifecycle.
Layer 3 — Repository and dependency integrity
Upstream of GSD's runtime behaviour, the open-gsd organisation enforces
controls at the repository and package level. These are documented in full in
docs/security/baseline.md and are summarised
here for completeness.
Dependency integrity. All third-party dependencies are pinned via
package-lock.json and verified against published checksums before install.
A scripts/check-npm-integrity.cjs gate detects invalid versions, missing
packages, and extraneous packages at CI time. This mitigates dependency
confusion and typosquatting attacks against GSD's own dependencies.
Secret scanning. Every commit and PR is scanned for hardcoded secrets.
Intentional test fixtures must be annotated with the project-standard
exclusion grammar (see SECURITY.md for the annotation format). Un-annotated
suppressions fail CI.
Locale-safe text scanning. Output and user-facing strings are scanned for Unicode homoglyphs, bidirectional override characters, and invisible Unicode — the class of attacks documented in CVE-2021-42574 ("Trojan Source") that can hide malicious content in diffs.
Trade-offs and limits
The security model described here meaningfully reduces the attack surface for AI-driven development. It does not eliminate supply-chain risk.
What the Package Legitimacy Gate reduces: The probability that a
hallucinated or attacker-registered package reaches npm install without
a human checkpoint. The [SLOP] gate removes high-confidence bad packages
entirely; the [SUS] / [ASSUMED] gates require human review before
execution. This substantially raises the cost of a successful slopsquatting
attack.
What the Package Legitimacy Gate does not eliminate: A legitimate package
that is later compromised (account takeover, dependency confusion in its own
tree) is not caught by slopcheck, which checks registration signals at
research time. Lock files and npm audit at the dependency-integrity layer
are the controls for that class of attack.
What the prompt injection defences reduce: The probability that user-controlled text in planning artifacts successfully overrides agent instructions. Pattern-matching on known injection forms catches the common cases; novel jailbreaks or low-signal injections may pass undetected. The advisory-only posture means detection is logged but not blocked — a deliberate choice that preserves workflow continuity at the cost of not hard-stopping on a detection.
What the prompt injection defences do not eliminate: A sufficiently creative injection that does not match known patterns, or an injection that arrives through a channel the hooks do not cover (for example, content injected into a dependency's published README that is read by a subagent browsing documentation). Defence in depth means each layer makes the attack harder, not that any single layer makes it impossible.
Reporting vulnerabilities. Report via private GitHub security advisory at
https://github.com/open-gsd/gsd-core/security/advisories/new. Do not open
public issues. See SECURITY.md for the response timeline
and disclosure policy.
Related
- Commands — includes
/gsd-secure-phaseand/gsd-code-reviewwith security-relevant flags - Architecture § Hook System — implementation detail on every hook, its event trigger, and safety properties
- SECURITY.md — vulnerability reporting, org-wide security baseline, secret-scan exclusion governance, and dependency integrity verification
- Docs index