* fix(#4016): imperative-override patterns tolerate filler words The narrow imperative-override family tolerates no filler between the verb and the noun, so a planted "Forget all of your instructions" (measured in a real public transcript) matched none of the 14 patterns and both consuming hooks stayed silent. One combined filler-tolerant pattern is appended; the narrow four stay untouched to keep the change merge-friendly. Known trade-offs, disclosed in #4016: linter-doc prose like "ignore rules on a single line" now trips a LOW advisory, and the overlap with the narrow patterns means one sentence can count twice toward severity thresholds. Regression tests assert the previously-missed phrasings fire in BOTH consuming hooks (gsd-prompt-guard and gsd-read-injection-scanner), not just in the raw pattern list, per the agent brief in #4016. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb * chore(#4016): changeset fragment for PR #4061 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb * test(#4016): pin the disclosed linter-doc FP as single-pattern LOW, never blocking Review follow-up on PR #4061: the combined filler-tolerant pattern's disclosed false-positive class (linter-doc prose such as "use eslint-disable-next-line to ignore rules on a single line") was documented in prose only. Two tests now pin it: - the prose matches exactly ONE shared pattern (the #4016 combined pattern, not a narrow one), so it cannot silently start double-counting toward the 3+ HIGH threshold; - through the real gsd-read-injection-scanner subprocess with security.injection_blocking=true, the prose yields a single-finding LOW advisory and no block decision — with an in-test positive control proving a 3+-pattern payload DOES block in the same directory, so the non-blocking assertion cannot pass vacuously. Samples are fragment-built like the existing SAMPLES rows so this file's own diff does not trip the CI injection scanner. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#4016): replace the five narrow imperative-override patterns with one superset The first cut appended a filler-tolerant combined pattern next to the five narrow verb patterns. Both consumers count one finding per matching pattern toward the severity threshold, so the overlap made one sentence count twice: "Ignore previous instructions. Forget your instructions." scored 2 (LOW) on next and 3 (HIGH, blockable) on the branch. It also left `override` out of the combined pattern. Replace the narrow family (ignore x2, disregard, forget, override) with ONE superset pattern over ignore|disregard|forget|discard|override. At least one filler (all|of|the|your|my|system|previous|prior|above|earlier) must sit between verb and noun, enforced by a lookahead with no repetition; the two noun-less/bare forms the old list accepted (`disregard (all) previous`, `forget instructions`) are kept as explicit tails so the new pattern is a strict superset. Bare "override rules" / "ignore instructions" are ordinary repo prose (6 measured hits across docs and source) and stay unmatched. Corpus measurement over 3019 .md/.js/.cjs/.mjs files (injection-sample tests excluded): the old family hit 2 lines, the new pattern hits 3, the only new one being a documented injection example in planner-reversibility.md that the old family missed (the issue's own class). Tests: SAMPLES reshaped to the 10-entry list; superset proof table (17 legacy phrasings, each matching exactly one pattern); five issue phrasings including `override all of your previous instructions` counted exactly once through both hook subprocesses; double-count regression (1 finding, LOW); design pin that bare verb+noun matches nothing; linter-doc FP pin split into bare (silent) and determined (single LOW, never blocks). All fragment-built; the CI prompt-injection scanner reports 0 findings on every touched file. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1 * chore(#4016): changeset body in the canonical bold-lead format .changeset/README.md Format: a leading bold change sentence, then an em-dash explanation. Also drops the verbatim planted phrase from the body so the rendered CHANGELOG line does not trip the pattern it describes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1 * fix(#4016): render a bounded pattern label in the prompt-guard advisory, pin plural prompts Review round 4 of PR #4061 left two nits open. 1. gsd-prompt-guard.js pushed `pattern.source` verbatim into its typed finding and, through renderFinding, into the user-facing advisory. With the #4016 superset pattern that source is 300 characters, so a genuine hit surfaced an advisory dominated by a raw regex dump. The read scanner has trimmed its equivalent since #3523 (`\s+` -> `-`, strip `()\`, cut at 50). That transform is hoisted into hooks/lib/injection-patterns.js as `describePattern` and used by BOTH hooks, so one finding renders the same label everywhere. Byte-identical to the scanner's old inline output for all 10 patterns (measured). No new staging dependency: both hooks already require this module. 2. The noun alternation `prompts?` had no positive coverage for the plural branch. One filler-regression row now exercises `... previous prompts ...` and runs through the existing once-per-hook, exactly-one-pattern loops. The parity test's prompt-guard count assertion moves off substring-matching the advisory prose onto the typed `findings` surface added in #3546, per CONTRIBUTING's raw-text-matching prohibition. New test: the superset source exceeds the bound (positive control), the prompt guard never embeds it, and both hooks carry the identical label in `findings[0].match`. Tests: parity, read-scanner, kimi field-shadowing, prompt-injection-scan, hooks-crash-policy, dead-exports: 206 run, 196 pass, 0 fail, 10 pre-existing platform skips. eslint clean; changeset lint ok; hooks runtime-build-seam lint ok; the CI prompt-injection scanner reports 0 findings on the PR diff. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016GGp8kEB5zCDmJ6TYHP1Nj --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
76 lines
4.4 KiB
JavaScript
76 lines
4.4 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* injection-patterns.js — the shared prompt-injection pattern list (#3504, epic #1900).
|
|
*
|
|
* Single source of truth for the standard injection signatures used by BOTH
|
|
* gsd-prompt-guard.js (PreToolUse scan of writes into .planning/) and
|
|
* gsd-read-injection-scanner.js (PostToolUse scan of Read/WebFetch/WebSearch
|
|
* content). Previously each hook carried a byte-identical copy ("inlined for
|
|
* hook independence") that could silently drift — a pattern tightened in one
|
|
* would stop protecting the other surface.
|
|
*
|
|
* Why a shared lib require is safe here (the old inlining rationale, retired):
|
|
* the installer stages hooks/lib/ from the GSD_HOOK_LIB_FILES allowlist for the
|
|
* shared-bundle surfaces (Claude-family settings.json runtimes and Kimi), the
|
|
* Cursor stager auto-discovers require('./lib/...') in staged scripts and fails
|
|
* the install loudly when a helper is missing (#2587), and the plugin path
|
|
* ships this directory wholesale.
|
|
*
|
|
* Deliberately NOT unified with src/security.cts's scanForInjection set: that
|
|
* set runs inside the compiled lib tree; hooks must stay loadable standalone
|
|
* without it. The two lists are different surfaces by design, not drift.
|
|
*
|
|
* Keep this file free of literal 'gsd:' text — the stager rewrites that marker
|
|
* in staged hook content.
|
|
*/
|
|
|
|
const INJECTION_PATTERNS = Object.freeze([
|
|
// #4016: ONE filler-tolerant imperative-override pattern. It REPLACES the
|
|
// five narrow verb patterns (ignore x2 / disregard / forget / override) that
|
|
// used to sit here: they tolerated no filler between verb and noun, so a
|
|
// planted "forget all of your ..." phrasing (measured in the wild) matched
|
|
// nothing. A single pattern rather than narrow-plus-combined because both
|
|
// consumers count one finding PER PATTERN toward the severity threshold
|
|
// (3 = HIGH, blockable under security.injection_blocking): overlapping
|
|
// patterns made one sentence count twice and pushed a two-phrasing LOW
|
|
// payload to HIGH (PR #4061 review).
|
|
// verbs ignore|disregard|forget|discard|override. `override` carries
|
|
// the old `override system|previous prompt|instructions` family;
|
|
// `discard` is the synonym seen in the same planted phrasings.
|
|
// fillers at least ONE of all|of|the|your|my|<qualifier> must sit between
|
|
// verb and noun (a lookahead, no repetition). Bare "override
|
|
// rules" / "ignore instructions" are ordinary repo prose
|
|
// (measured: 6 hits across docs and source) and stay unmatched.
|
|
// tails `disregard (all) previous` with no noun, and bare `forget` +
|
|
// `instructions`, are kept verbatim because the old narrow
|
|
// patterns accepted them; the superset is pinned by
|
|
// tests/injection-patterns-parity.security.test.cjs.
|
|
// known FP determined linter-doc prose such as "ignore (the) rules" trips
|
|
// a single-pattern LOW advisory; one pattern can never block.
|
|
/(?:ignore|disregard|forget|discard|override)\s+(?=(?:all|of|the|your|my|system|previous|prior|above|earlier)\s)(?:all\s+)?(?:of\s+)?(?:the\s+|your\s+|my\s+)?(?:(?:system|previous|prior|above|earlier)\s+)?(?:instructions|directives|prompts?|rules)|disregard\s+(?:all\s+)?previous|forget\s+instructions/i,
|
|
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
|
|
/act\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
|
|
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
|
|
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
|
|
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
|
|
/<\/?(?:system|assistant|human)>/i,
|
|
/\[SYSTEM\]/i,
|
|
/\[INST\]/i,
|
|
/<<\s*SYS\s*>>/i,
|
|
]);
|
|
|
|
// Short single-line label for a matched pattern, used for both the rendered
|
|
// advisory prose and the typed `findings[].match` field in BOTH consumer
|
|
// hooks. The raw regex source is not user-facing: the #4016 superset pattern
|
|
// is ~280 characters, and gsd-prompt-guard.js echoed it verbatim into its
|
|
// advisory (PR #4061 review nit). Byte-identical to the transform
|
|
// gsd-read-injection-scanner.js carried inline since #3523, hoisted here so
|
|
// one finding renders the same way in both hooks. Both hooks already require
|
|
// this module, so this adds no new staging dependency.
|
|
function describePattern(pattern) {
|
|
return pattern.source.replace(/\\s\+/g, '-').replace(/[()\\]/g, '').substring(0, 50);
|
|
}
|
|
|
|
module.exports = { INJECTION_PATTERNS, describePattern };
|