* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
218 lines
7.4 KiB
Bash
Executable File
218 lines
7.4 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# prompt-injection-scan.sh — Scan files for prompt injection patterns
|
|
#
|
|
# Usage:
|
|
# scripts/prompt-injection-scan.sh --diff origin/main # CI mode: scan changed .md files
|
|
# scripts/prompt-injection-scan.sh --file path/to/file # Scan a single file
|
|
# scripts/prompt-injection-scan.sh --dir agents/ # Scan all files in a directory
|
|
#
|
|
# Exit codes:
|
|
# 0 = clean
|
|
# 1 = findings detected
|
|
# 2 = usage error
|
|
set -euo pipefail
|
|
|
|
# ─── Patterns ────────────────────────────────────────────────────────────────
|
|
# Each pattern is a POSIX extended regex. Keep alphabetized by category.
|
|
|
|
PATTERNS=(
|
|
# Instruction override
|
|
'ignore[[:space:]]+(all[[:space:]]+)?(previous|prior|above|earlier|preceding)[[:space:]]+(instructions|prompts|rules|directives|context)'
|
|
'disregard[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules)'
|
|
'forget[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules|context)'
|
|
'override[[:space:]]+(all[[:space:]]+)?(system|previous|safety)[[:space:]]+(instructions|prompts|rules|checks|filters|guards)'
|
|
'override[[:space:]]+(system|safety|security)[[:space:]]'
|
|
|
|
# Role manipulation
|
|
'you[[:space:]]+are[[:space:]]+now[[:space:]]+(a|an|my)[[:space:]]'
|
|
'from[[:space:]]+now[[:space:]]+on[[:space:]]+(you|pretend|act|behave)'
|
|
'pretend[[:space:]]+(you[[:space:]]+are|to[[:space:]]+be)[[:space:]]'
|
|
'act[[:space:]]+as[[:space:]]+(a|an|if|my)[[:space:]]'
|
|
'roleplay[[:space:]]+as[[:space:]]'
|
|
'assume[[:space:]]+the[[:space:]]+role[[:space:]]+of[[:space:]]'
|
|
|
|
# System prompt extraction
|
|
'output[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'reveal[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'show[[:space:]]+me[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'print[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'what[[:space:]]+(is|are)[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'repeat[[:space:]]+(your|the|all)[[:space:]]+(system[[:space:]]+)?(prompt|instructions|rules)'
|
|
|
|
# Fake message boundaries
|
|
'</?system>'
|
|
'</?assistant>'
|
|
'</?human>'
|
|
'\[SYSTEM\]'
|
|
'\[/SYSTEM\]'
|
|
'\[INST\]'
|
|
'\[/INST\]'
|
|
'<<SYS>>'
|
|
'<</SYS>>'
|
|
|
|
# Tool call injection / code execution in markdown
|
|
'eval[[:space:]]*\([[:space:]]*["\x27]'
|
|
'exec[[:space:]]*\([[:space:]]*["\x27]'
|
|
'Function[[:space:]]*\([[:space:]]*["\x27].*return'
|
|
|
|
# Jailbreak / DAN patterns
|
|
'do[[:space:]]+anything[[:space:]]+now'
|
|
'DAN[[:space:]]+mode'
|
|
'developer[[:space:]]+mode[[:space:]]+(enabled|output|activated)'
|
|
'jailbreak'
|
|
'bypass[[:space:]]+(safety|content|security)[[:space:]]+(filter|check|rule|guard)'
|
|
)
|
|
|
|
# ─── Allowlist ───────────────────────────────────────────────────────────────
|
|
# Files that legitimately discuss injection patterns (security docs, tests, this script)
|
|
ALLOWLIST=(
|
|
'scripts/prompt-injection-scan.sh'
|
|
'scripts/base64-scan.sh'
|
|
'scripts/secret-scan.sh'
|
|
'tests/security-scan.security.test.cjs'
|
|
'tests/security.test.cjs'
|
|
'tests/prompt-injection-scan.security.test.cjs'
|
|
'tests/verify.test.cjs'
|
|
'gsd-core/bin/lib/security.cjs'
|
|
'hooks/gsd-prompt-guard.js'
|
|
'hooks/gsd-read-injection-scanner.js'
|
|
'tests/read-injection-scanner.security.test.cjs'
|
|
'tests/read-injection-scanner.property.test.cjs'
|
|
'tests/security-prompt-injection.security.test.cjs'
|
|
'tests/list-seeds.test.cjs'
|
|
'tests/fixtures/adversarial/security/'
|
|
'SECURITY.md'
|
|
# These files contain intentional injection examples / security-model prose
|
|
# and are not attack vectors — they explain/demonstrate injection patterns.
|
|
'TEST-EXAMPLES.md'
|
|
'explanation/security-model.md'
|
|
# The untrusted-input boundary reference quotes injection phrases
|
|
# ("ignore previous instructions", "you are now…") as examples agents must
|
|
# NOT comply with — it is the defense, not an attack vector.
|
|
'references/untrusted-input-boundary.md'
|
|
# Security regression tests for input validators — fixtures must contain
|
|
# real injection payloads to prove the validator rejects them. See
|
|
# DEFECT.PROMPT-INJECTION-SCAN-COLLISION in CONTEXT.md.
|
|
'tests/windsurf-conversion.test.cjs'
|
|
)
|
|
|
|
is_allowlisted() {
|
|
local file="$1"
|
|
for allowed in "${ALLOWLIST[@]}"; do
|
|
if [[ "$file" == *"$allowed"* ]]; then
|
|
return 0
|
|
fi
|
|
done
|
|
return 1
|
|
}
|
|
|
|
# ─── File Collection ─────────────────────────────────────────────────────────
|
|
|
|
collect_files() {
|
|
local mode="$1"
|
|
shift
|
|
|
|
case "$mode" in
|
|
--diff)
|
|
local base="${1:-origin/main}"
|
|
# Get changed files in the diff, filter to scannable extensions
|
|
git diff --name-only --diff-filter=ACMR "$base"...HEAD 2>/dev/null \
|
|
| grep -E '\.(md|cjs|js|json|yml|yaml|sh)$' || true
|
|
;;
|
|
--file)
|
|
if [[ -f "$1" ]]; then
|
|
echo "$1"
|
|
else
|
|
echo "Error: file not found: $1" >&2
|
|
exit 2
|
|
fi
|
|
;;
|
|
--dir)
|
|
local dir="$1"
|
|
if [[ ! -d "$dir" ]]; then
|
|
echo "Error: directory not found: $dir" >&2
|
|
exit 2
|
|
fi
|
|
find "$dir" -type f \( -name '*.md' -o -name '*.cjs' -o -name '*.js' -o -name '*.json' -o -name '*.yml' -o -name '*.yaml' -o -name '*.sh' \) \
|
|
! -path '*/node_modules/*' ! -path '*/.git/*' ! -path '*/dist/*' 2>/dev/null || true
|
|
;;
|
|
--stdin)
|
|
cat
|
|
;;
|
|
*)
|
|
echo "Usage: $0 --diff [base] | --file <path> | --dir <path> | --stdin" >&2
|
|
exit 2
|
|
;;
|
|
esac
|
|
}
|
|
|
|
# ─── Scanner ─────────────────────────────────────────────────────────────────
|
|
|
|
scan_file() {
|
|
local file="$1"
|
|
local found=0
|
|
|
|
if is_allowlisted "$file"; then
|
|
return 0
|
|
fi
|
|
|
|
for pattern in "${PATTERNS[@]}"; do
|
|
# Use grep -iE for case-insensitive extended regex
|
|
# -n for line numbers, -c for count mode first to check
|
|
local matches
|
|
matches=$(grep -inE -e "$pattern" "$file" 2>/dev/null || true)
|
|
if [[ -n "$matches" ]]; then
|
|
if [[ $found -eq 0 ]]; then
|
|
echo "FAIL: $file"
|
|
found=1
|
|
fi
|
|
echo "$matches" | while IFS= read -r line; do
|
|
echo " $line"
|
|
done
|
|
fi
|
|
done
|
|
|
|
return $found
|
|
}
|
|
|
|
# ─── Main ────────────────────────────────────────────────────────────────────
|
|
|
|
main() {
|
|
if [[ $# -eq 0 ]]; then
|
|
echo "Usage: $0 --diff [base] | --file <path> | --dir <path>" >&2
|
|
exit 2
|
|
fi
|
|
|
|
local mode="$1"
|
|
shift
|
|
|
|
local files
|
|
files=$(collect_files "$mode" "$@")
|
|
|
|
if [[ -z "$files" ]]; then
|
|
echo "prompt-injection-scan: no files to scan"
|
|
exit 0
|
|
fi
|
|
|
|
local total=0
|
|
local failed=0
|
|
|
|
while IFS= read -r file; do
|
|
[[ -z "$file" ]] && continue
|
|
total=$((total + 1))
|
|
if ! scan_file "$file"; then
|
|
failed=$((failed + 1))
|
|
fi
|
|
done <<< "$files"
|
|
|
|
echo ""
|
|
echo "prompt-injection-scan: scanned $total files, $failed with findings"
|
|
|
|
if [[ $failed -gt 0 ]]; then
|
|
exit 1
|
|
fi
|
|
exit 0
|
|
}
|
|
|
|
main "$@"
|