* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
3.0 KiB
3.0 KiB
ADR-1577: Untrusted-input boundary + opt-in injection blocking
- Status: Proposed
- Issue: #1577
- Part of: #1573 (harden the agent layer against documented LLM failure modes)
Context
The research/doc-ingest agents concatenate text returned by WebFetch / WebSearch / Read into their context with no data/instruction separation, and the gsd-read-injection-scanner hook only scanned the Read tool — leaving WebFetch/WebSearch (the largest untrusted channel) unscanned. Prompt injection via fetched content is a documented LLM failure mode (arXiv 2506.05739, 2507.15219, 2504.20472).
Two mechanisms were considered for the hook-level control:
- Redaction — strip the detected content before it reaches the model. This requires
hookSpecificOutput.updatedToolOutput, which is unused anywhere in this repo and not verifiable in CI for a PostToolUse hook. Claiming redaction the code can't reliably perform would re-introduce exactly the overclaim this work set out to remove. - Circuit-breaker — a PostToolUse hook that, after the fetch has executed and the content is already in the transcript, emits
decision: "block"to halt the agent's next step. It does not redact content already in context.
Decision
- Extend the scanner to match
Read | WebFetch | WebSearch, documented honestly as a pattern-based pre-filter, not a model-level guard. - Make the prompt-level boundary the primary control: a shared
gsd-core/references/untrusted-input-boundary.md,@-included by the 10 ingest agents, instructs treat-fetched-text-as-data, self-scan before use, task-anchoring, and a fresh random delimiter per quoted wrap. This is the layer that keeps an injection from being followed even while it sits in context. - Ship hook-level blocking as an opt-in circuit-breaker:
security.injection_blocking(a registered config key; default advisory). Documentation states plainly that enabling it halts further processing on a HIGH detection — it does not retroactively redact the already-fetched content. Redaction viaupdatedToolOutputis deferred until that field's behavior is verifiable in this runtime.
Consequences
- Non-breaking. The default posture is advisory; no existing default changes. Blocking is reached only by an explicit opt-in.
- The strongest guarantee is prompt-level (data/instruction separation), which is unenforced at runtime — this is defense-in-depth (arXiv 2503.00061), not a hard sandbox. A determined adaptive attacker or a weaker model may still be influenced.
- Localized docs are managed separately; only the canonical English
docs/explanation/security-model.mdis updated here. - Follow-up: if/when
updatedToolOutputredaction is confirmed supported, the circuit-breaker can be upgraded to an actual redactor without changing the opt-in surface.