Files
msd-core/docs/features/improved-prompt-injection-scanner.md
Tom Boucher 1e67ec9737 enhance(#3908): the scanners distinguish an empty diff from one they could not compute (#3937)
* feat(#3908): the scanners distinguish an empty diff from one they could not compute

collect_files ended 2>/dev/null || true, which destroyed the evidence three ways: the redirect discarded git's diagnostic, the pipe replaced git's status with grep's, and || true forced success regardless. Four distinct conditions - an established-empty diff, a bad ref, no repository, and a repository with no commits - all reported clean, and a secret scanner reporting clean because git failed is indistinguishable from an all-clear to any gate consuming it.

git now runs separately from the filter so its status and diagnostic both survive. An established-empty diff exits NO_INPUT; a scope that could not be established exits UNAVAILABLE; the usage sites move off 2 to USAGE. || true is retained on the filter alone, where it is correct: a diff of only images is empty, not failed.

Codes are sourced from a generated shell fragment rather than written into three scripts, so a re-allocation cannot desync them, and a missing fragment fails loudly instead of falling back to literals. The security workflow is updated in the same change: without it, a docs-only PR would newly fail the job.

* fix(#3908): keep scanner stderr out of the file list, and drop try/finally from test bodies

Capturing git and find output with 2>&1 was right for the failure path but wrong for the success path: a warning emitted alongside a successful diff flowed into the file list and was treated as a filename. stderr is now captured separately, forwarded as a warning on success and as the diagnostic on failure, and never folded into the list.

Also converts the control tests' try/finally blocks to t.after(), which CONTRIBUTING bans inside a test body because it masks failures.

* chore(#3908): backfill changeset pr number

* docs(#3908): record the scanners' four-outcome exit contract

SECURITY.md is root-level, so the docs gate correctly held: a Changed fragment owes a file under docs/. The contract also belongs where the feature is described, as REQ-SCAN-INJ-05.

docs/FEATURES.md is GENERATED from per-feature fragments (#3840) - the first edit went into the generated file and gen-features --check caught it, which is the same edit-the-output drift this epic exists to close. The fragment is the source; FEATURES.md is regenerated.

---------

Co-authored-by: sim <sim@local>
2026-08-27 13:11:13 -04:00

2.5 KiB
Raw Blame History

id, title, group
id title group
99 Improved Prompt Injection Scanner v1.34.0 Features

Hook: gsd-prompt-guard.js, gsd-read-injection-scanner.js Script: scripts/prompt-injection-scan.sh, scripts/base64-scan.sh

Purpose: Defense-in-depth detection of prompt injection attempts in planning artifacts and ingested content. Live hooks inline their own pattern subsets for hook independence (they do not import from security.cts). The CI scanner (scanForInjection in security.cts) provides a centralized engine for codebase-wide scanning in tests.

Requirements:

  • REQ-SCAN-INJ-01: Live hooks MUST detect invisible Unicode characters (zero-width spaces, soft hyphens, Unicode tag block U+E0000–E007F)

  • REQ-SCAN-INJ-02: Live hooks MUST detect known injection patterns (instruction override, role manipulation, system-prompt extraction, fake message boundaries). Base64-decode scanning is a CI-time control (scripts/base64-scan.sh), not a live hook — live hooks match a base64-exfiltration phrase regex only, they do not decode.

  • REQ-SCAN-INJ-03: Scanner MUST apply entropy analysis — Entropy analysis (scanEntropyAnomalies) was removed in #2198 as dead code (zero production callers; live hooks do not perform entropy analysis). This requirement is deferred pending a maintainable live implementation.

  • REQ-SCAN-INJ-04: Scanner MUST remain advisory-only — detection is logged, not blocking

  • REQ-SCAN-INJ-05: A scanner that could not establish its file list MUST NOT report clean (#3908). The CI scanners (prompt-injection-scan.sh, base64-scan.sh, secret-scan.sh) distinguish four outcomes rather than collapsing them into exit 0:

    Outcome Exit Meaning
    scanned, no findings 0 files were in scope and none matched
    findings 1 the scan's own verdict
    nothing in scope NO_INPUT the diff resolved and was genuinely empty — e.g. a docs-only PR
    could not scan UNAVAILABLE the file list was never established: a bad ref, no repository, or a repository with no commits

    Codes come from the exit-code registry (ADR-3889), sourced from gsd-core/bin/shared/exit-codes.sh, never written into the scripts. Every one is non-zero, so a caller written if ! scanner; then behaves identically for a clean scan and trips for everything else — this can turn a false green red, never a red green. .github/workflows/security-scan.yml treats nothing in scope as a pass and could not scan as a failure; previously the latter passed silently, having scanned nothing.