* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
58 lines
2.8 KiB
JavaScript
58 lines
2.8 KiB
JavaScript
'use strict';
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const ROOT = path.join(__dirname, '..');
|
|
const REF = path.join(ROOT, 'gsd-core', 'references', 'untrusted-input-boundary.md');
|
|
const INGEST_AGENTS = [
|
|
'gsd-phase-researcher', 'gsd-project-researcher', 'gsd-domain-researcher',
|
|
'gsd-ai-researcher', 'gsd-advisor-researcher', 'gsd-research-synthesizer',
|
|
'gsd-doc-classifier', 'gsd-doc-synthesizer',
|
|
// AC #2 named agents: gsd-ui-researcher carries the full WebSearch/WebFetch
|
|
// toolset (web ingress); gsd-assumptions-analyzer reads 5-15 codebase source
|
|
// files (external/source-document ingress per the boundary).
|
|
'gsd-ui-researcher', 'gsd-assumptions-analyzer',
|
|
];
|
|
|
|
describe('untrusted-input isolation (#12)', () => {
|
|
test('shared reference exists with the data/instruction directive', () => {
|
|
assert.ok(fs.existsSync(REF), 'untrusted-input-boundary.md must exist');
|
|
const src = fs.readFileSync(REF, 'utf8');
|
|
assert.match(src, /<security_context>/);
|
|
assert.match(src, /treated as data/i);
|
|
assert.match(src, /never as instructions/i);
|
|
});
|
|
|
|
test('reference contains randomized-marker instruction (honest PPA 2506.05739)', () => {
|
|
const src = fs.readFileSync(REF, 'utf8');
|
|
// Must mention randomness near a DATA marker — fixed/predictable markers are spoofable
|
|
assert.match(src, /random|fresh|unique|nonce/i,
|
|
'reference must instruct agents to generate a fresh/random delimiter per wrap');
|
|
assert.match(src, /DATA_/,
|
|
'reference must still reference DATA_ marker pattern');
|
|
});
|
|
|
|
test('reference contains self-guard/self-scan instruction (honest PromptArmor 2507.15219)', () => {
|
|
const src = fs.readFileSync(REF, 'utf8');
|
|
// Must instruct agent to scan/inspect content itself before using it
|
|
assert.match(src, /inspect|scan.{0,30}before|act as.{0,30}guard|self.{0,10}guard|self.{0,10}scan/i,
|
|
'reference must instruct agents to self-inspect content for embedded instructions before use');
|
|
});
|
|
|
|
test('reference contains task-anchor instruction (honest Referencing 2504.20472)', () => {
|
|
const src = fs.readFileSync(REF, 'utf8');
|
|
// Must instruct agent to act only on its assigned task and ignore off-task instructions in data
|
|
assert.match(src, /only.{0,40}(?:your|the).{0,20}(?:task|assignment)|assigned task|not tied to/i,
|
|
'reference must instruct agents to act only on their assigned task and ignore instructions in data not tied to that task');
|
|
});
|
|
|
|
for (const name of INGEST_AGENTS) {
|
|
test(`${name} @-includes the untrusted-input-boundary reference`, () => {
|
|
const src = fs.readFileSync(path.join(ROOT, 'agents', `${name}.md`), 'utf8');
|
|
assert.match(src, /references\/untrusted-input-boundary\.md/, `${name} missing the @-include`);
|
|
});
|
|
}
|
|
});
|