* test(3596): adversarial security/prompt-injection abuse suite
Adds `tests/security-prompt-injection.test.cjs` and a fixtures
directory at `tests/fixtures/adversarial/security/` covering the
attack classes enumerated in #3596:
- Command substitution / backticks / heredoc payloads in workstream
names — sentinel-file probes prove no shell is spawned, slugifier
neutralises the input.
- Path traversal through `--ws` and slash-bearing workstream names —
rejected with structured `--json-errors` payload, no stack trace,
no filesystem mutation outside the project root.
- Fake `<system>` / `[SYSTEM]` / `<<SYS>>` / `[INST]` boundary tags —
sanitizeForPrompt neutralises every form; structural negative
property locked across all six styles in one place.
- Zero-width / bidi-override codepoints — stripped per the documented
codepoint set; asserted via codePoint inspection, not regex
literals.
- Hostile read of CONTEXT.md / PLAN.md / ROADMAP.md fixtures —
`gsd-read-injection-scanner.js` surfaces the advisory; excluded
paths and non-Read tools stay silent; malformed JSON does not
crash the hook.
- Hostile write of `.planning/` files — `gsd-prompt-guard.js` emits
a `PreToolUse` advisory; non-Write/Edit tools stay silent.
- Fake `ghp_*` / `sk-*` env tokens — never echoed in CLI stdout or
stderr under hostile inputs; covered under
`// allow-test-rule: structural-regression-guard` because the only
way to assert byte-level absence is `.includes(token)` against the
captured streams.
- `validatePath`, `validateShellArg`, `validatePhaseNumber`,
`validateFieldName` — focused negative-input contract pins.
Pinned behavior gaps (documented, NOT fixed in this PR):
- `<instructions>` is intentionally whitelisted by both the scanner
and the sanitiser (GSD's own prompt scaffolding). Two REGRESSION
GUARD tests lock that contract.
- The current `scanForInjection` does NOT flag malicious markdown
links (javascript:/data:/embedded-credentials URLs). PINNED with
negative-proof so any future scope extension fails the assertion
and forces a deliberate update to the acceptance map.
- `prompt-builder.ts` does not yet wrap plan/context markdown in an
"untrusted data" envelope. That seam lives on the TS side and is
covered by `sdk/src/prompt-builder.test.ts`; out of scope for a
CJS test file. Mentioned in the file header.
Verification:
- `node --test tests/security-prompt-injection.test.cjs` → 73 tests
pass.
- `node scripts/lint-no-source-grep.cjs` → 0 violations across
546 test files (one `allow-test-rule: structural-regression-guard`
annotation on this file for the token-absence assertions).
- `node scripts/run-tests.cjs` → 9730 tests pass, 0 fail.
Refs #3596
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(3596): allow adversarial fixtures in scan + harden graphify status parse
* fix(3596): skip adversarial security fixtures in secret scan
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: add /gsd-spec-phase — Socratic spec refinement with ambiguity scoring (#2213)
Introduces `/gsd-spec-phase <phase>` as an optional pre-step before discuss-phase.
Clarifies WHAT a phase delivers (requirements, boundaries, acceptance criteria) with
quantitative ambiguity scoring before discuss-phase handles HOW to implement.
- `commands/gsd/spec-phase.md` — slash command routing to workflow
- `get-shit-done/workflows/spec-phase.md` — full Socratic interview loop (up to 6
rounds, 5 rotating perspectives: Researcher, Simplifier, Boundary Keeper, Failure
Analyst, Seed Closer) with weighted 4-dimension ambiguity gate (≤ 0.20 to write SPEC.md)
- `get-shit-done/templates/spec.md` — SPEC.md template with falsifiable requirements
(Current/Target/Acceptance per requirement), Boundaries, Acceptance Criteria,
Ambiguity Report, and Interview Log; includes two full worked examples
- `get-shit-done/workflows/discuss-phase.md` — new `check_spec` step detects
`{padded_phase}-SPEC.md` at startup; displays "Found SPEC.md — N requirements
locked. Focusing on implementation decisions."; `analyze_phase` respects `spec_loaded`
flag to skip "what/why" gray areas; `write_context` emits `<spec_lock>` section
with boundary summary and canonical ref to SPEC.md
- `docs/ARCHITECTURE.md` — update command/workflow counts (74→75, 71→72)
Closes#2213
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(hooks): add gsd-read-injection-scanner PostToolUse hook (#2201)
Adds a new PostToolUse hook that scans content returned by the Read tool
for prompt injection patterns, including four summarisation-specific patterns
(retention-directive, permanence-claim, etc.) that survive context compression.
Defense-in-depth for long GSD sessions where the context summariser cannot
distinguish user instructions from content read from external files.
- Advisory-only (warns without blocking), consistent with gsd-prompt-guard.js
- LOW severity for 1-2 patterns, HIGH for 3+
- Inlined pattern library (hook independence)
- Exclusion list: .planning/, REVIEW.md, CHECKPOINT, security docs, hook sources
- Wired in install.js as PostToolUse matcher: Read, timeout: 5s
- Added to MANAGED_HOOKS for staleness detection
- 19 tests covering all 13 acceptance criteria (SCAN-01–07, EXCL-01–06, EDGE-01–06)
Closes#2201
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): add read-injection-scanner files to prompt-injection-scan allowlist
Test payloads in tests/read-injection-scanner.test.cjs and inlined patterns
in hooks/gsd-read-injection-scanner.js legitimately contain injection strings.
Add both to the CI script allowlist to prevent false-positive failures.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): assert exitCode, stdout, and signal explicitly in EDGE-05
Addresses CodeRabbit feedback: the success path discarded the return
value so a malformed-JSON input that produced stdout would still pass.
Now captures and asserts all three observable properties.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md
- Replace require('node:assert') with require('node:assert/strict') across
all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
as that is appropriate for per-function resource guards, not lifecycle hooks
Closes#1674
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* perf(tests): add --test-concurrency=4 to test runner for parallel file execution
Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): allowlist verify.test.cjs in prompt-injection scanner
tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add base64-scan.sh and secret-scan.sh to prompt injection scanner
allowlist (scanner was flagging its own pattern strings)
- Skip executable bit check on Windows (no Unix permissions)
- Skip bash script execution tests on Windows (requires Git Bash)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add CI security pipeline to catch prompt injection attacks, base64-obfuscated
payloads, leaked secrets, and .planning/ directory commits in PRs.
This is critical for get-shit-done because the entire codebase is markdown
prompts — a prompt injection in a workflow file IS the attack surface.
New files:
- scripts/prompt-injection-scan.sh: scans for instruction override, role
manipulation, system boundary injection, DAN/jailbreak, and tool call
injection patterns in changed files
- scripts/base64-scan.sh: extracts base64 blobs >= 40 chars, decodes them,
and checks decoded content against injection patterns (skips data URIs
and binary content)
- scripts/secret-scan.sh: detects AWS keys, OpenAI/Anthropic keys, GitHub
PATs, Stripe keys, private key headers, and generic credential patterns
- .github/workflows/security-scan.yml: runs all three scans plus a
.planning/ directory check on every PR
- .base64scanignore / .secretscanignore: per-repo false positive allowlists
- tests/security-scan.test.cjs: 51 tests covering script existence,
pattern matching, false positive avoidance, and workflow structure
All scripts support --diff (CI), --file, and --dir modes. Cross-platform
(macOS + Linux). SHA-pinned actions. Environment variables used for
github context in run blocks (no direct interpolation).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>