Files
msd-core/docs/adr/1577-untrusted-input-boundary-and-injection-blocking.md
Alex V. a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00

3.0 KiB

ADR-1577: Untrusted-input boundary + opt-in injection blocking

  • Status: Proposed
  • Issue: #1577
  • Part of: #1573 (harden the agent layer against documented LLM failure modes)

Context

The research/doc-ingest agents concatenate text returned by WebFetch / WebSearch / Read into their context with no data/instruction separation, and the gsd-read-injection-scanner hook only scanned the Read tool — leaving WebFetch/WebSearch (the largest untrusted channel) unscanned. Prompt injection via fetched content is a documented LLM failure mode (arXiv 2506.05739, 2507.15219, 2504.20472).

Two mechanisms were considered for the hook-level control:

  1. Redaction — strip the detected content before it reaches the model. This requires hookSpecificOutput.updatedToolOutput, which is unused anywhere in this repo and not verifiable in CI for a PostToolUse hook. Claiming redaction the code can't reliably perform would re-introduce exactly the overclaim this work set out to remove.
  2. Circuit-breaker — a PostToolUse hook that, after the fetch has executed and the content is already in the transcript, emits decision: "block" to halt the agent's next step. It does not redact content already in context.

Decision

  • Extend the scanner to match Read | WebFetch | WebSearch, documented honestly as a pattern-based pre-filter, not a model-level guard.
  • Make the prompt-level boundary the primary control: a shared gsd-core/references/untrusted-input-boundary.md, @-included by the 10 ingest agents, instructs treat-fetched-text-as-data, self-scan before use, task-anchoring, and a fresh random delimiter per quoted wrap. This is the layer that keeps an injection from being followed even while it sits in context.
  • Ship hook-level blocking as an opt-in circuit-breaker: security.injection_blocking (a registered config key; default advisory). Documentation states plainly that enabling it halts further processing on a HIGH detection — it does not retroactively redact the already-fetched content. Redaction via updatedToolOutput is deferred until that field's behavior is verifiable in this runtime.

Consequences

  • Non-breaking. The default posture is advisory; no existing default changes. Blocking is reached only by an explicit opt-in.
  • The strongest guarantee is prompt-level (data/instruction separation), which is unenforced at runtime — this is defense-in-depth (arXiv 2503.00061), not a hard sandbox. A determined adaptive attacker or a weaker model may still be influenced.
  • Localized docs are managed separately; only the canonical English docs/explanation/security-model.md is updated here.
  • Follow-up: if/when updatedToolOutput redaction is confirmed supported, the circuit-breaker can be upgraded to an actual redactor without changing the opt-in surface.