Files
msd-core/scripts/prompt-injection-scan.sh
Rezolv e87fb409ee enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint

Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age
freshness proxy (state_commits_behind / state_commit_stale) through
state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as
health W024. The proxy is advisory: classify() deliberately does NOT consume
it (ADR-1787 locks the classification/routing boundary — a signal, not a route).

Composes with #3099 and #1882 (both merged to next after this branch): the
commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE
diagnostic reads `last_activity` — two different fields, not "two staleness
signals on one field." A new regression test asserts a STATE.md carrying both
an unparseable last_activity AND a valid state_head resolves each independently
(diagnostic fires once; freshness reads state_head, commits_behind 0).

Rebased onto next (flattened): resolved the add/add conflicts in
src/smart-entry.cts (kept both the #2573 freshness import/derivation and the
#3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both
describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B).
Tests: smart-entry 62, state/state-transition/health/verify 639, all pass.

* chore(#2573): allowlist health-validation test in the prompt-injection scan

The scanner's `exec('` code-execution pattern matches the benign
`re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests
(pre-existing: 16 such calls on next, this PR adds none). The file entered the
diff-mode scan's changed-file set only because #2573's W024 state_head
assertions touch it. Allowlist it alongside the other test files that carry
pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION
class). Scanner self-test 38/0; diff scan 14 files, 0 findings.
2026-08-11 17:10:23 -04:00

261 lines
10 KiB
Bash
Executable File

#!/usr/bin/env bash
# prompt-injection-scan.sh — Scan files for prompt injection patterns
#
# Usage:
# scripts/prompt-injection-scan.sh --diff origin/main # CI mode: scan changed .md files
# scripts/prompt-injection-scan.sh --file path/to/file # Scan a single file
# scripts/prompt-injection-scan.sh --dir agents/ # Scan all files in a directory
#
# Exit codes:
# 0 = clean
# 1 = findings detected
# 2 = usage error
set -euo pipefail
# ─── Patterns ────────────────────────────────────────────────────────────────
# Each pattern is a POSIX extended regex. Keep alphabetized by category.
#
# Left-boundary prefix `(^|[^[:alnum:]])`: several trigger words are also
# suffixes of ordinary English words or camelCase identifiers (fact/impact/
# contract/artifact/interact all end in "act"; retrieval/medieval end in
# "eval"; blueprint/reprint/fingerprint end in "print"; describeFunction/
# wrapFunction end in "Function"; Jordan/Sudan end in "dan"), so an
# unanchored keyword matches as a false-positive substring. `\b` is a GNU
# grep extension and this script must also run under BSD/macOS grep, so the
# boundary is spelled out as `(^|[^[:alnum:]])` instead. This never narrows
# real detections: a genuine attack phrase is always preceded by start-of-
# line, whitespace, or punctuation, never by another alnum character glued
# directly onto the keyword. Only patterns whose leading keyword is provably
# not a real-word suffix are left unanchored (#3175 audit).
PATTERNS=(
# Instruction override
'ignore[[:space:]]+(all[[:space:]]+)?(previous|prior|above|earlier|preceding)[[:space:]]+(instructions|prompts|rules|directives|context)'
'disregard[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules)'
'forget[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules|context)'
'override[[:space:]]+(all[[:space:]]+)?(system|previous|safety)[[:space:]]+(instructions|prompts|rules|checks|filters|guards)'
'override[[:space:]]+(system|safety|security)[[:space:]]'
# Role manipulation
'you[[:space:]]+are[[:space:]]+now[[:space:]]+(a|an|my)[[:space:]]'
'from[[:space:]]+now[[:space:]]+on[[:space:]]+(you|pretend|act|behave)'
'pretend[[:space:]]+(you[[:space:]]+are|to[[:space:]]+be)[[:space:]]'
'(^|[^[:alnum:]])act[[:space:]]+as[[:space:]]+(a|an|if|my)[[:space:]]'
'roleplay[[:space:]]+as[[:space:]]'
'assume[[:space:]]+the[[:space:]]+role[[:space:]]+of[[:space:]]'
# System prompt extraction
'output[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
'reveal[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
'show[[:space:]]+me[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
'(^|[^[:alnum:]])print[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
'what[[:space:]]+(is|are)[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
'repeat[[:space:]]+(your|the|all)[[:space:]]+(system[[:space:]]+)?(prompt|instructions|rules)'
# Fake message boundaries
'</?system>'
'</?assistant>'
'</?human>'
'\[SYSTEM\]'
'\[/SYSTEM\]'
'\[INST\]'
'\[/INST\]'
'<<SYS>>'
'<</SYS>>'
# Tool call injection / code execution in markdown
#
# The quote-or-apostrophe class below is spelled ["'"'"'"] (a literal `'`
# via bash's close-quote/escape/reopen idiom), not `["\x27]` — `\x27` is a
# GNU-grep-only hex escape; BSD/macOS grep treats it as four literal
# characters (", \, x, 2, 7) and never matches an actual apostrophe, so
# `eval('...')` (single-quoted) silently went undetected on macOS while
# passing on GNU-grep CI runners. Found auditing #3175; fixed here since it
# is the same unanchored/portability defect class as the boundary fix.
'(^|[^[:alnum:]])eval[[:space:]]*\([[:space:]]*["'"'"']'
'exec[[:space:]]*\([[:space:]]*["'"'"']'
'(^|[^[:alnum:]])Function[[:space:]]*\([[:space:]]*["'"'"'].*return'
# Jailbreak / DAN patterns
'do[[:space:]]+anything[[:space:]]+now'
'(^|[^[:alnum:]])DAN[[:space:]]+mode'
'developer[[:space:]]+mode[[:space:]]+(enabled|output|activated)'
'jailbreak'
'bypass[[:space:]]+(safety|content|security)[[:space:]]+(filter|check|rule|guard)'
)
# ─── Allowlist ───────────────────────────────────────────────────────────────
# Files that legitimately discuss injection patterns (security docs, tests, this script)
ALLOWLIST=(
'scripts/prompt-injection-scan.sh'
'scripts/base64-scan.sh'
'scripts/secret-scan.sh'
'tests/security-scan.security.test.cjs'
'tests/security.test.cjs'
'tests/prompt-injection-scan.security.test.cjs'
'tests/verify.test.cjs'
'gsd-core/bin/lib/security.cjs'
'hooks/gsd-prompt-guard.js'
'hooks/gsd-read-injection-scanner.js'
'tests/read-injection-scanner.security.test.cjs'
'tests/read-injection-scanner.property.test.cjs'
'tests/security-prompt-injection.security.test.cjs'
'tests/list-seeds.test.cjs'
'tests/fixtures/adversarial/security/'
'SECURITY.md'
# These files contain intentional injection examples / security-model prose
# and are not attack vectors — they explain/demonstrate injection patterns.
'TEST-EXAMPLES.md'
'explanation/security-model.md'
# The untrusted-input boundary reference quotes injection phrases
# ("ignore previous instructions", "you are now…") as examples agents must
# NOT comply with — it is the defense, not an attack vector.
'references/untrusted-input-boundary.md'
# Security regression tests for input validators — fixtures must contain
# real injection payloads to prove the validator rejects them. See
# DEFECT.PROMPT-INJECTION-SCAN-COLLISION in CONTEXT.md.
'tests/windsurf-conversion.test.cjs'
# RuleTester fixtures for the local/no-unguarded-nonportable-exec ESLint rule
# contain shell-exec command strings (exec("sh -c …"), execFileSync('bash',['-c',…]))
# as test DATA the rule must lint — not attack vectors. ADR-1703 Phase 3 (#1720).
'tests/no-unguarded-nonportable-exec.rule.test.cjs'
# RuleTester fixtures for the local/no-bare-npm-exec ESLint rule contain npm
# exec command strings (execFileSync('npm', ['install'])) as test DATA the rule
# must lint — not attack vectors. ADR-1703 Phase 4 (#1726).
'tests/no-bare-npm-exec.rule.test.cjs'
# #2547 — the Kimi field-shadowing regression proves gsd-prompt-guard still
# SCANS the reconstructed edit[].new content when a model-supplied new_string
# tries to shadow it. The fixture must be a real injection phrase or the test
# asserts nothing: it is the payload the guard is required to catch, carried
# as test DATA. Same class as the read-injection-scanner suites above.
'tests/kimi-payload-field-shadowing.security.test.cjs'
# Phase-ID grammar regression tests exercise `RegExp.prototype.exec` via
# `re.exec('<phase-id>')` against fixtures like 'MANIFOLD-64-auth' / 'CK-64-auth'.
# The scanner's `exec('` code-execution pattern matches that benign method call,
# not an attack vector — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the
# test fixtures above. Pre-existing content (16 such calls on `next`); it surfaces
# here only because #2573's W024 `state_head` assertions make the file appear in
# the changed-file set the diff-mode scan walks.
'tests/health-validation.test.cjs'
)
is_allowlisted() {
local file="$1"
for allowed in "${ALLOWLIST[@]}"; do
if [[ "$file" == *"$allowed"* ]]; then
return 0
fi
done
return 1
}
# ─── File Collection ─────────────────────────────────────────────────────────
collect_files() {
local mode="$1"
shift
case "$mode" in
--diff)
local base="${1:-origin/main}"
# Get changed files in the diff, filter to scannable extensions
git diff --name-only --diff-filter=ACMR "$base"...HEAD 2>/dev/null \
| grep -E '\.(md|cjs|js|json|yml|yaml|sh)$' || true
;;
--file)
if [[ -f "$1" ]]; then
echo "$1"
else
echo "Error: file not found: $1" >&2
exit 2
fi
;;
--dir)
local dir="$1"
if [[ ! -d "$dir" ]]; then
echo "Error: directory not found: $dir" >&2
exit 2
fi
find "$dir" -type f \( -name '*.md' -o -name '*.cjs' -o -name '*.js' -o -name '*.json' -o -name '*.yml' -o -name '*.yaml' -o -name '*.sh' \) \
! -path '*/node_modules/*' ! -path '*/.git/*' ! -path '*/dist/*' 2>/dev/null || true
;;
--stdin)
cat
;;
*)
echo "Usage: $0 --diff [base] | --file <path> | --dir <path> | --stdin" >&2
exit 2
;;
esac
}
# ─── Scanner ─────────────────────────────────────────────────────────────────
scan_file() {
local file="$1"
local found=0
if is_allowlisted "$file"; then
return 0
fi
for pattern in "${PATTERNS[@]}"; do
# Use grep -iE for case-insensitive extended regex
# -n for line numbers, -c for count mode first to check
local matches
matches=$(grep -inE -e "$pattern" "$file" 2>/dev/null || true)
if [[ -n "$matches" ]]; then
if [[ $found -eq 0 ]]; then
echo "FAIL: $file"
found=1
fi
echo "$matches" | while IFS= read -r line; do
echo " $line"
done
fi
done
return $found
}
# ─── Main ────────────────────────────────────────────────────────────────────
main() {
if [[ $# -eq 0 ]]; then
echo "Usage: $0 --diff [base] | --file <path> | --dir <path>" >&2
exit 2
fi
local mode="$1"
shift
local files
files=$(collect_files "$mode" "$@")
if [[ -z "$files" ]]; then
echo "prompt-injection-scan: no files to scan"
exit 0
fi
local total=0
local failed=0
while IFS= read -r file; do
[[ -z "$file" ]] && continue
total=$((total + 1))
if ! scan_file "$file"; then
failed=$((failed + 1))
fi
done <<< "$files"
echo ""
echo "prompt-injection-scan: scanned $total files, $failed with findings"
if [[ $failed -gt 0 ]]; then
exit 1
fi
exit 0
}
main "$@"