Files
msd-core/hooks/lib/injection-patterns.js
allcounter 723ea08dc2 fix(#4016): imperative-override injection patterns tolerate filler words (#4061)
* fix(#4016): imperative-override patterns tolerate filler words

The narrow imperative-override family tolerates no filler between the
verb and the noun, so a planted "Forget all of your instructions"
(measured in a real public transcript) matched none of the 14 patterns
and both consuming hooks stayed silent.

One combined filler-tolerant pattern is appended; the narrow four stay
untouched to keep the change merge-friendly. Known trade-offs, disclosed
in #4016: linter-doc prose like "ignore rules on a single line" now
trips a LOW advisory, and the overlap with the narrow patterns means one
sentence can count twice toward severity thresholds.

Regression tests assert the previously-missed phrasings fire in BOTH
consuming hooks (gsd-prompt-guard and gsd-read-injection-scanner), not
just in the raw pattern list, per the agent brief in #4016.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* chore(#4016): changeset fragment for PR #4061

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* test(#4016): pin the disclosed linter-doc FP as single-pattern LOW, never blocking

Review follow-up on PR #4061: the combined filler-tolerant pattern's
disclosed false-positive class (linter-doc prose such as "use
eslint-disable-next-line to ignore rules on a single line") was
documented in prose only. Two tests now pin it:

- the prose matches exactly ONE shared pattern (the #4016 combined
  pattern, not a narrow one), so it cannot silently start double-counting
  toward the 3+ HIGH threshold;
- through the real gsd-read-injection-scanner subprocess with
  security.injection_blocking=true, the prose yields a single-finding
  LOW advisory and no block decision — with an in-test positive control
  proving a 3+-pattern payload DOES block in the same directory, so the
  non-blocking assertion cannot pass vacuously.

Samples are fragment-built like the existing SAMPLES rows so this file's
own diff does not trip the CI injection scanner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#4016): replace the five narrow imperative-override patterns with one superset

The first cut appended a filler-tolerant combined pattern next to the five
narrow verb patterns. Both consumers count one finding per matching pattern
toward the severity threshold, so the overlap made one sentence count twice:
"Ignore previous instructions. Forget your instructions." scored 2 (LOW) on
next and 3 (HIGH, blockable) on the branch. It also left `override` out of
the combined pattern.

Replace the narrow family (ignore x2, disregard, forget, override) with ONE
superset pattern over ignore|disregard|forget|discard|override. At least one
filler (all|of|the|your|my|system|previous|prior|above|earlier) must sit
between verb and noun, enforced by a lookahead with no repetition; the two
noun-less/bare forms the old list accepted (`disregard (all) previous`,
`forget instructions`) are kept as explicit tails so the new pattern is a
strict superset. Bare "override rules" / "ignore instructions" are ordinary
repo prose (6 measured hits across docs and source) and stay unmatched.

Corpus measurement over 3019 .md/.js/.cjs/.mjs files (injection-sample tests
excluded): the old family hit 2 lines, the new pattern hits 3, the only new
one being a documented injection example in planner-reversibility.md that
the old family missed (the issue's own class).

Tests: SAMPLES reshaped to the 10-entry list; superset proof table (17 legacy
phrasings, each matching exactly one pattern); five issue phrasings including
`override all of your previous instructions` counted exactly once through
both hook subprocesses; double-count regression (1 finding, LOW); design pin
that bare verb+noun matches nothing; linter-doc FP pin split into bare
(silent) and determined (single LOW, never blocks). All fragment-built; the
CI prompt-injection scanner reports 0 findings on every touched file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* chore(#4016): changeset body in the canonical bold-lead format

.changeset/README.md Format: a leading bold change sentence, then an em-dash
explanation. Also drops the verbatim planted phrase from the body so the
rendered CHANGELOG line does not trip the pattern it describes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* fix(#4016): render a bounded pattern label in the prompt-guard advisory, pin plural prompts

Review round 4 of PR #4061 left two nits open.

1. gsd-prompt-guard.js pushed `pattern.source` verbatim into its typed
   finding and, through renderFinding, into the user-facing advisory. With
   the #4016 superset pattern that source is 300 characters, so a genuine hit
   surfaced an advisory dominated by a raw regex dump. The read scanner has
   trimmed its equivalent since #3523 (`\s+` -> `-`, strip `()\`, cut at 50).
   That transform is hoisted into hooks/lib/injection-patterns.js as
   `describePattern` and used by BOTH hooks, so one finding renders the same
   label everywhere. Byte-identical to the scanner's old inline output for
   all 10 patterns (measured). No new staging dependency: both hooks already
   require this module.

2. The noun alternation `prompts?` had no positive coverage for the plural
   branch. One filler-regression row now exercises `... previous prompts ...`
   and runs through the existing once-per-hook, exactly-one-pattern loops.

The parity test's prompt-guard count assertion moves off substring-matching
the advisory prose onto the typed `findings` surface added in #3546, per
CONTRIBUTING's raw-text-matching prohibition. New test: the superset source
exceeds the bound (positive control), the prompt guard never embeds it, and
both hooks carry the identical label in `findings[0].match`.

Tests: parity, read-scanner, kimi field-shadowing, prompt-injection-scan,
hooks-crash-policy, dead-exports: 206 run, 196 pass, 0 fail, 10 pre-existing
platform skips. eslint clean; changeset lint ok; hooks runtime-build-seam lint
ok; the CI prompt-injection scanner reports 0 findings on the PR diff.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016GGp8kEB5zCDmJ6TYHP1Nj

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 06:58:45 -04:00

76 lines
4.4 KiB
JavaScript

'use strict';
/**
* injection-patterns.js — the shared prompt-injection pattern list (#3504, epic #1900).
*
* Single source of truth for the standard injection signatures used by BOTH
* gsd-prompt-guard.js (PreToolUse scan of writes into .planning/) and
* gsd-read-injection-scanner.js (PostToolUse scan of Read/WebFetch/WebSearch
* content). Previously each hook carried a byte-identical copy ("inlined for
* hook independence") that could silently drift — a pattern tightened in one
* would stop protecting the other surface.
*
* Why a shared lib require is safe here (the old inlining rationale, retired):
* the installer stages hooks/lib/ from the GSD_HOOK_LIB_FILES allowlist for the
* shared-bundle surfaces (Claude-family settings.json runtimes and Kimi), the
* Cursor stager auto-discovers require('./lib/...') in staged scripts and fails
* the install loudly when a helper is missing (#2587), and the plugin path
* ships this directory wholesale.
*
* Deliberately NOT unified with src/security.cts's scanForInjection set: that
* set runs inside the compiled lib tree; hooks must stay loadable standalone
* without it. The two lists are different surfaces by design, not drift.
*
* Keep this file free of literal 'gsd:' text — the stager rewrites that marker
* in staged hook content.
*/
const INJECTION_PATTERNS = Object.freeze([
// #4016: ONE filler-tolerant imperative-override pattern. It REPLACES the
// five narrow verb patterns (ignore x2 / disregard / forget / override) that
// used to sit here: they tolerated no filler between verb and noun, so a
// planted "forget all of your ..." phrasing (measured in the wild) matched
// nothing. A single pattern rather than narrow-plus-combined because both
// consumers count one finding PER PATTERN toward the severity threshold
// (3 = HIGH, blockable under security.injection_blocking): overlapping
// patterns made one sentence count twice and pushed a two-phrasing LOW
// payload to HIGH (PR #4061 review).
// verbs ignore|disregard|forget|discard|override. `override` carries
// the old `override system|previous prompt|instructions` family;
// `discard` is the synonym seen in the same planted phrasings.
// fillers at least ONE of all|of|the|your|my|<qualifier> must sit between
// verb and noun (a lookahead, no repetition). Bare "override
// rules" / "ignore instructions" are ordinary repo prose
// (measured: 6 hits across docs and source) and stay unmatched.
// tails `disregard (all) previous` with no noun, and bare `forget` +
// `instructions`, are kept verbatim because the old narrow
// patterns accepted them; the superset is pinned by
// tests/injection-patterns-parity.security.test.cjs.
// known FP determined linter-doc prose such as "ignore (the) rules" trips
// a single-pattern LOW advisory; one pattern can never block.
/(?:ignore|disregard|forget|discard|override)\s+(?=(?:all|of|the|your|my|system|previous|prior|above|earlier)\s)(?:all\s+)?(?:of\s+)?(?:the\s+|your\s+|my\s+)?(?:(?:system|previous|prior|above|earlier)\s+)?(?:instructions|directives|prompts?|rules)|disregard\s+(?:all\s+)?previous|forget\s+instructions/i,
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
/act\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
/<\/?(?:system|assistant|human)>/i,
/\[SYSTEM\]/i,
/\[INST\]/i,
/<<\s*SYS\s*>>/i,
]);
// Short single-line label for a matched pattern, used for both the rendered
// advisory prose and the typed `findings[].match` field in BOTH consumer
// hooks. The raw regex source is not user-facing: the #4016 superset pattern
// is ~280 characters, and gsd-prompt-guard.js echoed it verbatim into its
// advisory (PR #4061 review nit). Byte-identical to the transform
// gsd-read-injection-scanner.js carried inline since #3523, hoisted here so
// one finding renders the same way in both hooks. Both hooks already require
// this module, so this adds no new staging dependency.
function describePattern(pattern) {
return pattern.source.replace(/\\s\+/g, '-').replace(/[()\\]/g, '').substring(0, 50);
}
module.exports = { INJECTION_PATTERNS, describePattern };