Files
msd-core/docs/explanation/security-model.md
Alex V. a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00

14 KiB

GSD Core security model

Explanation — This document describes why GSD Core has the security posture it does and how the layers fit together. It is not a reference for every hook parameter. For the /gsd-secure-phase command and its options, see Commands. For the implementation-level hook architecture, see Architecture § Hook System. For the org-wide security baseline (scanner controls, incident checklists, ownership model), see SECURITY.md.


Why AI-driven development needs a dedicated security posture

A conventional code editor does not execute arbitrary packages on your behalf. GSD Core does. The research → plan → execute pipeline automates the full path from "name a package" to "run npm install <package>", from "write a planning artifact" to "use that artifact as an LLM system prompt". Each automation step removes a human from the loop — and each removal is a potential attack surface.

GSD Core's security model is built around one organising principle: defence in depth. No single control is assumed to be perfect. Several overlapping layers each reduce a distinct class of risk, and together they make the attack surface substantially harder to exploit without eliminating it entirely. The honest summary at the end of this document explains what the system cannot protect against.


Layer 1 — Supply-chain protection: the Package Legitimacy Gate

The threat

AI models hallucinate package names. This is not a fringe failure mode: 2025 research documents roughly 20 % of AI-generated package references as hallucinated names that do not correspond to legitimate packages. A subset of those hallucinated names — approximately 43 % in the same research — recur consistently across prompts, meaning an attacker can observe which names AI tools commonly produce and pre-register those names on npm, PyPI, or crates.io with malicious post-install scripts. The technique is called slopsquatting.

The insidious quality of slopsquatting is that a hallucinated name that passes npm view looks legitimate. The registry entry proves only that someone registered the name — not that the package does what the AI said it does, not that it has any legitimate users, and not that its install scripts are safe. Without a gate, a hallucinated name would flow undetected through GSD's researcher → planner → executor pipeline and eventually run as npm install <attacker-package> on your machine.

How the gate works

The gate operates across three pipeline stages:

Research stage. When gsd-phase-researcher recommends external packages, it runs slopcheck install <pkgs> --json against each one. The results are written to a ## Package Legitimacy Audit table in RESEARCH.md. Packages tagged [SLOP] (high-confidence hallucination or attacker-registered) are stripped from RESEARCH.md entirely before the file is saved. They never reach the planner.

Planning stage. gsd-planner reads the Audit table. For any package tagged [SUS] (suspicious: newly registered, low download count, no source repository, or naming pattern close to a popular package) or [ASSUMED] (sourced from WebSearch rather than direct registry verification), the planner inserts a checkpoint:human-verify task before the install step. The checkpoint includes a direct link to the registry page and specific things to look for: maintainer history, issue-tracker activity, absence of suspicious install scripts.

Execution stage. If an install fails, gsd-executor surfaces a checkpoint and stops. It does not silently try an alternative package name — which could itself be malicious. This is an explicit rule in the executor's behaviour (RULE 3 in the executor agent definition).

Why WebSearch packages are always [ASSUMED]

Package names discovered through WebSearch are tagged [ASSUMED] regardless of whether npm view succeeds. A package that exists on the registry is not the same as a package that is safe to install. npm view proves registration, not legitimacy. The [ASSUMED] tag triggers the same human-verify checkpoint as [SUS], ensuring that any unverified web-discovered recommendation always gets a human review before installation.

Ecosystem coverage

The researcher uses registry-specific verification commands rather than a single generic check:

  • Node.js: npm view
  • Python: pip index versions
  • Rust: cargo search

This covers cross-ecosystem hallucination, which occurs at roughly 9 % according to 2025 USENIX research — cases where an AI recommends a package that exists in one ecosystem but not the one actually in use.

Graceful degradation

If slopcheck is unavailable (not installed, or the pip install fails at research time), GSD applies the strictest possible fallback: every recommended package is tagged [ASSUMED], and the planner gates every install with a checkpoint:human-verify task. Research and planning proceed normally — the system never hard-fails on a missing tool dependency. This is intentionally stricter than the normal flow: slopcheck unavailability means every package install gets a human checkpoint.

The slopcheck tool is MIT-licensed and pip-installable. If it is ever abandoned, the [ASSUMED]-gate fallback ensures human-checkpoint coverage is maintained regardless.


Layer 2 — Prompt injection defences

The threat

GSD Core generates Markdown files that become LLM system prompts. The research pipeline reads external web content; the planning pipeline incorporates user-supplied text (--text-file, --prd); the execution pipeline writes planning artifacts that are later re-read as agent context. Any user-controlled text flowing into these artifacts is a potential indirect prompt injection vector — an attacker-controlled string that, once inside a system prompt, attempts to override the agent's instructions or exfiltrate information.

How the defences work

GSD Core addresses prompt injection at three levels.

Input validation (security.cjs). The gsd-core/bin/lib/security.cjs module is the central security utility. It provides:

  • Path traversal prevention: user-supplied file paths (--text-file, --prd) are validated to resolve within the project directory, with macOS /var → /private/var symlink resolution handled explicitly
  • Prompt injection detection: known injection patterns (role overrides, instruction bypasses, system tag injections) are scanned in user-supplied text before it enters any planning artifact
  • Safe JSON parsing: a wrapper that prevents prototype-pollution attacks via crafted JSON payloads
  • Shell argument validation: arguments passed to subshell commands are validated before use

Runtime hook: gsd-prompt-guard.js. This hook fires on every Write or Edit call that targets .planning/ files. It scans the content being written for the same injection patterns as security.cjs (a subset inlined directly into the hook for independence — the hook does not require() the module, so it runs even if the module path changes). Detection is advisory-only: the hook logs the finding but does not block the write. The rationale is that a false-positive block on a legitimate planning write would be more disruptive than a missed injection in a secondary scan layer.

Runtime hook: gsd-read-injection-scanner.js. This hook fires on the output of every Read, WebFetch, and WebSearch tool call. It scans the content that was just read or fetched for injected instructions in untrusted content — catching cases where an attacker has embedded instructions in a file or remote resource that GSD is about to incorporate into an agent's context. The 10 research and doc-ingest agents additionally carry a shared <security_context> data/instruction boundary (defined in gsd-core/references/untrusted-input-boundary.md): gsd-project-researcher, gsd-phase-researcher, gsd-ui-researcher, gsd-assumptions-analyzer, gsd-advisor-researcher, gsd-doc-classifier, gsd-doc-synthesizer, gsd-research-synthesizer, gsd-ai-researcher, and gsd-domain-researcher. Any content fetched or read by those agents is treated as data, never as instructions, regardless of what the content claims to be.

Opt-in blocking (security.injection_blocking). By default all injection detections are advisory-only (logged, not blocked). Setting security.injection_blocking = true in .planning/config.json (a registered config key — gsd config-set security.injection_blocking true) upgrades HIGH-confidence detections to blocking. Be precise about what this does: the scanner is a PostToolUse hook, so it runs after the Read/WebFetch/WebSearch has already executed and the fetched content is already in the model's transcript. Blocking does not retroactively redact that content — it emits decision: "block", which halts the agent's next step and feeds the detection back as the reason, so the agent is stopped from acting further on the flagged result instead of silently continuing. LOW detections remain advisory under this setting. This flag is opt-in; the default (advisory-only) is preserved to avoid breaking existing workflows. The prompt-level boundary above (treat fetched text as data, never instructions) is the layer that keeps an injection from being followed even while it sits in context; the hook is a coarse pattern pre-filter and circuit-breaker, not a redactor.

CI scanner. prompt-injection-scan.security.test.cjs scans all agent, workflow, and command files for embedded injection vectors as part of the test suite. This catches injection attempts in the GSD source itself — for example, a supply-chain attack that modified a workflow file to add a role-override instruction.

Read Injection Scanner vs Prompt Guard

The two hooks cover complementary surfaces. gsd-prompt-guard.js watches writes to planning artifacts — it catches injection being planted. gsd-read-injection-scanner.js watches reads and remote fetches — it catches injection being ingested from external content (a dependency's README, a third-party config file, a user-provided document, or any URL fetched via WebFetch or WebSearch). The in-prompt <security_context> boundary in research agents provides an additional containment layer: even if an injected string reaches an agent, it is structurally separated from the instruction region. Together these controls bracket the ingest → store → re-read lifecycle.


Layer 3 — Repository and dependency integrity

Upstream of GSD's runtime behaviour, the open-gsd organisation enforces controls at the repository and package level. These are documented in full in docs/security/baseline.md and are summarised here for completeness.

Dependency integrity. All third-party dependencies are pinned via package-lock.json and verified against published checksums before install. A scripts/check-npm-integrity.cjs gate detects invalid versions, missing packages, and extraneous packages at CI time. This mitigates dependency confusion and typosquatting attacks against GSD's own dependencies.

Secret scanning. Every commit and PR is scanned for hardcoded secrets. Intentional test fixtures must be annotated with the project-standard exclusion grammar (see SECURITY.md for the annotation format). Un-annotated suppressions fail CI.

Locale-safe text scanning. Output and user-facing strings are scanned for Unicode homoglyphs, bidirectional override characters, and invisible Unicode — the class of attacks documented in CVE-2021-42574 ("Trojan Source") that can hide malicious content in diffs.


Trade-offs and limits

The security model described here meaningfully reduces the attack surface for AI-driven development. It does not eliminate supply-chain risk.

What the Package Legitimacy Gate reduces: The probability that a hallucinated or attacker-registered package reaches npm install without a human checkpoint. The [SLOP] gate removes high-confidence bad packages entirely; the [SUS] / [ASSUMED] gates require human review before execution. This substantially raises the cost of a successful slopsquatting attack.

What the Package Legitimacy Gate does not eliminate: A legitimate package that is later compromised (account takeover, dependency confusion in its own tree) is not caught by slopcheck, which checks registration signals at research time. Lock files and npm audit at the dependency-integrity layer are the controls for that class of attack.

What the prompt injection defences reduce: The probability that user-controlled text in planning artifacts successfully overrides agent instructions. Pattern-matching on known injection forms catches the common cases; novel jailbreaks or low-signal injections may pass undetected. The advisory-only posture means detection is logged but not blocked — a deliberate choice that preserves workflow continuity at the cost of not hard-stopping on a detection.

What the prompt injection defences do not eliminate: A sufficiently creative injection that does not match known patterns, or an injection that arrives through a channel the hooks do not cover. The previously uncovered channel of content injected into a dependency's published README and read by a subagent browsing documentation is now scanned at ingress by gsd-read-injection-scanner.js (which covers WebFetch and WebSearch output) and structurally isolated in-prompt by the <security_context> boundary in research agents — but novel jailbreaks and low-signal injections may still pass undetected. Defence in depth means each layer makes the attack harder, not that any single layer makes it impossible.

Reporting vulnerabilities. Report via private GitHub security advisory at https://github.com/open-gsd/gsd-core/security/advisories/new. Do not open public issues. See SECURITY.md for the response timeline and disclosure policy.


  • Commands — includes /gsd-secure-phase and /gsd-code-review with security-relevant flags
  • Architecture § Hook System — implementation detail on every hook, its event trigger, and safety properties
  • SECURITY.md — vulnerability reporting, org-wide security baseline, secret-scan exclusion governance, and dependency integrity verification
  • Docs index