Files
msd-core/tests/untrusted-input-isolation.test.cjs
Alex V. a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00

58 lines
2.8 KiB
JavaScript

'use strict';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const ROOT = path.join(__dirname, '..');
const REF = path.join(ROOT, 'gsd-core', 'references', 'untrusted-input-boundary.md');
const INGEST_AGENTS = [
'gsd-phase-researcher', 'gsd-project-researcher', 'gsd-domain-researcher',
'gsd-ai-researcher', 'gsd-advisor-researcher', 'gsd-research-synthesizer',
'gsd-doc-classifier', 'gsd-doc-synthesizer',
// AC #2 named agents: gsd-ui-researcher carries the full WebSearch/WebFetch
// toolset (web ingress); gsd-assumptions-analyzer reads 5-15 codebase source
// files (external/source-document ingress per the boundary).
'gsd-ui-researcher', 'gsd-assumptions-analyzer',
];
describe('untrusted-input isolation (#12)', () => {
test('shared reference exists with the data/instruction directive', () => {
assert.ok(fs.existsSync(REF), 'untrusted-input-boundary.md must exist');
const src = fs.readFileSync(REF, 'utf8');
assert.match(src, /<security_context>/);
assert.match(src, /treated as data/i);
assert.match(src, /never as instructions/i);
});
test('reference contains randomized-marker instruction (honest PPA 2506.05739)', () => {
const src = fs.readFileSync(REF, 'utf8');
// Must mention randomness near a DATA marker — fixed/predictable markers are spoofable
assert.match(src, /random|fresh|unique|nonce/i,
'reference must instruct agents to generate a fresh/random delimiter per wrap');
assert.match(src, /DATA_/,
'reference must still reference DATA_ marker pattern');
});
test('reference contains self-guard/self-scan instruction (honest PromptArmor 2507.15219)', () => {
const src = fs.readFileSync(REF, 'utf8');
// Must instruct agent to scan/inspect content itself before using it
assert.match(src, /inspect|scan.{0,30}before|act as.{0,30}guard|self.{0,10}guard|self.{0,10}scan/i,
'reference must instruct agents to self-inspect content for embedded instructions before use');
});
test('reference contains task-anchor instruction (honest Referencing 2504.20472)', () => {
const src = fs.readFileSync(REF, 'utf8');
// Must instruct agent to act only on its assigned task and ignore off-task instructions in data
assert.match(src, /only.{0,40}(?:your|the).{0,20}(?:task|assignment)|assigned task|not tied to/i,
'reference must instruct agents to act only on their assigned task and ignore instructions in data not tied to that task');
});
for (const name of INGEST_AGENTS) {
test(`${name} @-includes the untrusted-input-boundary reference`, () => {
const src = fs.readFileSync(path.join(ROOT, 'agents', `${name}.md`), 'utf8');
assert.match(src, /references\/untrusted-input-boundary\.md/, `${name} missing the @-include`);
});
}
});