feat(security): package legitimacy gate against slopsquatting (#3215)

* feat(security): package legitimacy gate against slopsquatting (#2827)

GSD's research → plan → execute pipeline had no install-time legitimacy
gate: a hallucinated package name that passes `npm view` could flow all
the way to `gsd-executor` running `npm install <malicious-pkg>` with no
human checkpoint. This PR closes that gap.

Changes:
- gsd-phase-researcher: runs slopcheck on every recommended package;
  emits `## Package Legitimacy Audit` table; strips [SLOP] packages;
  ecosystem-specific verification (pip/npm/cargo); WebSearch-sourced
  packages tagged [ASSUMED]; ctx7 fallback uses `command -v` guard
  instead of `npx --yes`
- gsd-planner: injects `checkpoint:human-verify` before [ASSUMED]/[SUS]
  installs; adds T-{phase}-SC STRIDE row to <threat_model> template;
  ctx7 fallback also uses `command -v` guard
- gsd-executor: RULE 3 excludes package installs from auto-fix; failed
  installs surface as checkpoints, never silent substitutions
- tests/package-legitimacy-gate.test.cjs: 24 structural assertions
  covering the full gate (node:test + node:assert, no raw .includes())
- docs: USER-GUIDE, COMMANDS, ARCHITECTURE updated with gate description
- .changeset: Security fragment for v1.51 release notes

Closes #2827

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs: expand Package Legitimacy Gate documentation

Add full user-facing depth to the gate docs across USER-GUIDE,
COMMANDS, and ARCHITECTURE:

- USER-GUIDE: rewrite gate section with concrete RESEARCH.md/PLAN.md
  examples, slopcheck verdict table, [ASSUMED] WebSearch tagging
  explanation, slopcheck-unavailable troubleshooting, and graceful
  degradation behavior
- COMMANDS.md: expand /gsd-plan-phase gate note with verdict bullets;
  add install-failure checkpoint behavior to /gsd-execute-phase
- ARCHITECTURE.md: expand gate section with threat model rationale,
  layer table, claim provenance integration, ecosystem coverage, and
  graceful degradation semantics

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): harden package legitimacy checkpoint semantics

* fix(planner): satisfy size gates and tighten package gate wording

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-05-08 09:08:06 -04:00
committed by GitHub
parent 397c34142a
commit b37c487325
8 changed files with 740 additions and 34 deletions

View File

@@ -0,0 +1,5 @@
---
type: Security
pr: 3215
---
**Package legitimacy gate added** — GSD now runs slopcheck against every researcher-recommended package before it enters RESEARCH.md; slopsquatted ([SLOP]) packages are removed at the source and suspicious ([SUS]) or assumed ([ASSUMED]) packages force a `checkpoint:human-verify` task before the executor installs them. The `npx --yes` auto-download pattern is replaced with a `command -v` guard across all three agent files, and executor RULE 3 explicitly excludes package-manager installs from auto-fix scope.

View File

@@ -33,19 +33,26 @@ When you need library or framework documentation, check in this order:
Step 1 — Resolve library ID:
```bash
npx --yes ctx7@latest library <name> "<query>"
if command -v ctx7 &>/dev/null; then
ctx7 library <name> "<query>"
else
echo "ctx7 not found — install with: npm install -g ctx7 (verify at npmjs.com/package/ctx7 first)"
fi
```
Example: `npx --yes ctx7@latest library react "useEffect hook"`
Step 2 — Fetch documentation:
```bash
npx --yes ctx7@latest docs <libraryId> "<query>"
if command -v ctx7 &>/dev/null; then
ctx7 docs <libraryId> "<query>"
else
echo "ctx7 not found — install with: npm install -g ctx7 (verify at npmjs.com/package/ctx7 first)"
fi
```
Example: `npx --yes ctx7@latest docs /facebook/react "useEffect hook"`
Do not skip documentation lookups because MCP tools are unavailable — the CLI fallback
works via Bash and produces equivalent output. Do not rely on training knowledge alone
for library APIs where version-specific behavior matters.
for library APIs where version-specific behavior matters. Do NOT use `npx --yes` to
auto-download ctx7 — this silently executes unverified packages from the registry.
</documentation_lookup>
<project_context>
@@ -168,7 +175,30 @@ No user permission needed for Rules 1-3.
**Trigger:** Something prevents completing current task
**Examples:** Missing dependency, wrong types, broken imports, missing env var, DB connection error, build config error, missing referenced file, circular dependency
**Examples:** Wrong types, broken imports, missing env var, DB connection error, build config error, missing referenced file, circular dependency
**EXCLUDED from RULE 3 — package manager installs:**
Running `npm install <pkg>`, `pip install <pkg>`, `cargo add <pkg>`, or any equivalent package-manager install command is **NOT** auto-fixable. If a referenced package fails to install or cannot be found:
1. Do NOT attempt to install a similarly-named alternative.
2. Do NOT retry with a different package name.
3. Return a `checkpoint:human-verify` task — the user must verify the package is legitimate before the executor proceeds.
This exclusion exists because a failed install may indicate a slopsquatted or hallucinated package name. Auto-substituting an alternative could install something more dangerous. If a package install fails, emit:
```xml
<task type="checkpoint:human-verify" gate="blocking-human">
<what-built>Package install failed — human verification required</what-built>
<how-to-verify>
`[package-name]` could not be installed. Before proceeding:
1. Verify the package exists and is legitimate: https://npmjs.com/package/[package-name]
2. Confirm the package name is spelled correctly in PLAN.md
3. If the package does not exist, return to /gsd-research-phase to find the correct package
</how-to-verify>
<resume-signal>Type "verified" with the correct package name, or "abort" to stop the phase</resume-signal>
</task>
```
Use `gate="blocking-human"` for package-legitimacy checkpoints so they are unambiguously excluded from auto-approval behavior.
---
@@ -265,7 +295,7 @@ For full automation-first patterns, server lifecycle, CLI handling:
**Auto-mode checkpoint behavior** (when `AUTO_CFG` is `"true"`):
- **checkpoint:human-verify** → Auto-approve. Log `⚡ Auto-approved: [what-built]`. Continue to next task.
- **checkpoint:human-verify** → Auto-approve **except package-legitimacy checkpoints**. If checkpoint has `gate="blocking-human"` OR its purpose indicates package legitimacy verification (`what-built` mentions `Package verification required before install` or `Package install failed — human verification required`), do **not** auto-approve. STOP and return checkpoint_return_format for explicit human confirmation.
- **checkpoint:decision** → Auto-select first option (planners front-load the recommended choice). Log `⚡ Auto-selected: [option name]`. Continue to next task.
- **checkpoint:human-action** → STOP normally. Auth gates cannot be automated — return structured checkpoint message using checkpoint_return_format.

View File

@@ -26,10 +26,12 @@ Spawned by `/gsd-plan-phase` (integrated) or `/gsd-research-phase` (standalone).
- Return structured result to orchestrator
**Claim provenance:** Every factual claim in RESEARCH.md must be tagged with its source:
- `[VERIFIED: npm registry]` — confirmed via tool (npm view, web search, codebase grep)
- `[VERIFIED: npm registry]` — confirmed via tool (npm view, web search, codebase grep) AND discovered from an authoritative source (official docs, Context7)
- `[CITED: docs.example.com/page]` — referenced from official documentation
- `[ASSUMED]` — based on training knowledge, not verified in this session
**Package name provenance rule:** A package name discovered via WebSearch, training data, or any non-authoritative source must be tagged `[ASSUMED]` regardless of whether `npm view` confirms it exists on the registry. Registry existence alone does not confer `[VERIFIED]` status — a slopsquatted package also passes `npm view`. Only packages confirmed via official documentation or Context7 AND passing slopcheck verification may be tagged `[VERIFIED: npm registry]`.
Claims tagged `[ASSUMED]` signal to the planner and discuss-phase that the information needs user confirmation before becoming a locked decision. Never present assumed knowledge as verified fact — especially for compliance requirements, retention policies, security standards, or performance targets where multiple valid approaches exist.
</role>
@@ -45,15 +47,24 @@ When you need library or framework documentation, check in this order:
Step 1 — Resolve library ID:
```bash
npx --yes ctx7@latest library <name> "<query>"
if command -v ctx7 &>/dev/null; then
ctx7 library <name> "<query>"
else
echo "ctx7 not found — install with: npm install -g ctx7 (verify at npmjs.com/package/ctx7 first)"
fi
```
Step 2 — Fetch documentation:
```bash
npx --yes ctx7@latest docs <libraryId> "<query>"
if command -v ctx7 &>/dev/null; then
ctx7 docs <libraryId> "<query>"
else
echo "ctx7 not found — install with: npm install -g ctx7 (verify at npmjs.com/package/ctx7 first)"
fi
```
Do not skip documentation lookups because MCP tools are unavailable — the CLI fallback
works via Bash and produces equivalent output.
works via Bash and produces equivalent output. Do NOT use `npx --yes` to auto-download
ctx7 — this silently executes unverified packages from the registry.
</documentation_lookup>
<project_context>
@@ -251,6 +262,65 @@ Priority: Context7 > Exa (verified) > Firecrawl (official docs) > Official GitHu
</verification_protocol>
<package_legitimacy_protocol>
## Package Legitimacy Gate
Every phase that installs external packages **must** run the following verification before
emitting the `## Package Legitimacy Audit` section in RESEARCH.md.
### Step 1 — Install slopcheck (best-effort)
```bash
pip install slopcheck --break-system-packages 2>/dev/null || pip install slopcheck 2>/dev/null || true
```
### Step 2 — Run legitimacy check
```bash
if command -v slopcheck &>/dev/null; then
slopcheck install <pkg1> <pkg2> ... --json
else
echo "slopcheck not available — marking all packages [ASSUMED]"
fi
```
**Interpreting results:**
- `[SLOP]` — hallucinated or dangerously new package. **Remove entirely** from all RESEARCH.md recommendations. List in audit table under `Disposition: REMOVED`.
- `[SUS]` — suspicious (new, low-downloads, or no source repo). **Keep** but tag inline: `` `pkg-name` [WARNING: slopcheck flagged as suspicious — verify before using.] ``
- `[OK]` — clean. Proceed normally.
**Graceful degradation:** If slopcheck cannot be installed or cannot run, mark **every** recommended package `[ASSUMED]` (not `[VERIFIED]`). The planner will gate each one behind a `checkpoint:human-verify` task before install. This is strictly safer than the current baseline — never a hard failure.
### Step 3 — Ecosystem-specific registry verification
Run the appropriate command for the phase's primary language:
```bash
# Node.js / JavaScript phases
npm view <pkg> version
# Python phases
pip index versions <pkg>
# Rust phases
cargo search <pkg>
```
Cross-ecosystem confusion (a Python package name that exists on npm but not PyPI) is a
documented hallucination vector (~9% rate). Always verify on the correct ecosystem registry.
### Step 4 — Check for suspicious postinstall scripts (Node.js phases)
```bash
npm view <pkg> scripts.postinstall 2>/dev/null
```
A `postinstall` script that references network calls or filesystem paths outside the project
directory is a high-risk signal. Flag such packages `[SUS]` even if slopcheck rates them `[OK]`.
</package_legitimacy_protocol>
<output_format>
## RESEARCH.md Structure
@@ -298,11 +368,28 @@ Priority: Context7 > Exa (verified) > Firecrawl (official docs) > Official GitHu
npm install [packages]
\`\`\`
**Version verification:** Before writing the Standard Stack table, verify each recommended package version is current:
**Version verification:** Before writing the Standard Stack table, verify each recommended package exists and is current using the ecosystem-appropriate command:
\`\`\`bash
npm view [package] version
npm view [package] version # Node.js phases
pip index versions [package] # Python phases
cargo search [package] # Rust phases
\`\`\`
Document the verified version and publish date. Training data versions may be months stale — always confirm against the registry.
Document the verified version and publish date. Training data versions may be months stale — always confirm against the correct ecosystem registry.
## Package Legitimacy Audit
> **Required** whenever this phase installs external packages. Run the Package Legitimacy Gate protocol before completing this section.
| Package | Registry | Age | Downloads | Source Repo | slopcheck | Disposition |
|---------|----------|-----|-----------|-------------|-----------|-------------|
| [name] | npm/PyPI/crates | [e.g., 8 yrs] | [e.g., 50M/wk] | [github.com/org/repo or "none"] | [OK] | Approved |
| [name] | npm | [e.g., 3 days] | [e.g., 0] | none | [SLOP] | REMOVED |
| [name] | npm | [e.g., 2 mo] | [e.g., 800/wk] | [github.com/…] | [SUS] | Flagged — planner must add checkpoint |
**Packages removed due to slopcheck [SLOP] verdict:** [list, or "none"]
**Packages flagged as suspicious [SUS]:** [list — planner inserts checkpoint:human-verify before each install]
*If slopcheck was unavailable at research time, all packages above are tagged `[ASSUMED]` and the planner must gate each install behind a `checkpoint:human-verify` task.*
## Architecture Patterns

View File

@@ -35,7 +35,7 @@ Your job: Produce PLAN.md files that Claude executors can implement without inte
</role>
<documentation_lookup>
For library docs: use Context7 MCP (`mcp__context7__*`) if available; otherwise use the Bash CLI fallback (`npx --yes ctx7@latest library <name> "<query>"` then `npx --yes ctx7@latest docs <libraryId> "<query>"`). The CLI fallback works via Bash when MCP is unavailable.
For library docs: prefer Context7 MCP. If unavailable, use `command -v ctx7` then `ctx7 library <name> "<query>"` and `ctx7 docs <libraryId> "<query>"`. Never use `npx --yes ctx7@latest`.
</documentation_lookup>
<project_context>
@@ -480,6 +480,7 @@ Output: [Artifacts created]
|-----------|----------|-----------|-------------|-----------------|
| T-{phase}-01 | {S/T/R/I/D/E} | {function/endpoint/file} | mitigate | {specific: e.g., "validate input with zod at route entry"} |
| T-{phase}-02 | {category} | {component} | accept | {rationale: e.g., "no PII, low-value target"} |
| T-{phase}-SC | Tampering | npm/pip/cargo installs | mitigate | slopcheck + blocking human checkpoint for [ASSUMED]/[SUS] |
</threat_model>
<verification>
@@ -615,6 +616,14 @@ Read ROADMAP.md `**Requirements:**` line for this phase. Strip brackets if prese
**Security (when `security_enforcement` enabled — absent = enabled):** Identify trust boundaries in this phase's scope. Map STRIDE categories to applicable tech stack from RESEARCH.md security domain. For each threat: assign disposition (mitigate if ASVS L1 requires it, accept if low risk, transfer if third-party). Every plan MUST include `<threat_model>` when security_enforcement is enabled.
**Package legitimacy gate (npm/pip/cargo only):**
- Require RESEARCH.md `## Package Legitimacy Audit` before package-manager install tasks.
- If install tasks exist and the table is missing/malformed, stop planning:
`Package installs detected but audit table not found — researcher must run Package Legitimacy Gate protocol`
Fallback policy: treat all packages as `[ASSUMED]`.
- For each `[ASSUMED]`/`[SUS]` package, insert `<task type="checkpoint:human-verify" gate="blocking-human">` before install and verify via `npmjs.com/package`, `pypi.org/project`, or `crates.io/crates`.
- `[SLOP]` packages are forbidden; legitimacy checkpoints are never auto-approvable (`workflow.auto_advance` ignored). Keep `T-{phase}-SC` in `<threat_model>`.
**Step 1: State the Goal**
Take phase goal from ROADMAP.md. Must be outcome-shaped, not task-shaped.
- Good: "Working chat interface" (outcome)
@@ -655,11 +664,6 @@ Message list component wiring:
**Step 5: Identify Key Links**
"Where is this most likely to break?" Key links = critical connections where breakage causes cascading failures.
For chat interface:
- Input onSubmit -> API call (if broken: typing works but sending doesn't)
- API save -> database (if broken: appears to send but doesn't persist)
- Component -> real data (if broken: shows placeholder, not messages)
## Must-Haves Output Format
```yaml
@@ -689,20 +693,6 @@ must_haves:
pattern: "prisma\\.message\\.(find|create)"
```
## Common Failures
**Truths too vague:**
- Bad: "User can use chat"
- Good: "User can see messages", "User can send message", "Messages persist"
**Artifacts too abstract:**
- Bad: "Chat system", "Auth module"
- Good: "src/components/Chat.tsx", "src/app/api/auth/login/route.ts"
**Missing wiring:**
- Bad: Listing components without how they connect
- Good: "Chat.tsx fetches from /api/chat via useEffect on mount"
</goal_backward>
<checkpoints>

View File

@@ -433,7 +433,11 @@ ui-phase → UI-SPEC.md (design contract, optional)
plan-phase
├── Research gate (blocks if RESEARCH.md has unresolved open questions)
├── Phase Researcher → RESEARCH.md
│ └── Package Legitimacy Gate: slopcheck on every package; [SLOP] removed,
│ [SUS]/[ASSUMED] flagged; Audit table written to RESEARCH.md
├── Planner (with reachability check) → PLAN.md files
│ └── checkpoint:human-verify injected before [ASSUMED]/[SUS] installs;
│ T-{phase}-SC STRIDE row added for install-bearing plans
├── Plan Checker → Verify loop (max 3x)
├── Requirements coverage gate (REQ-IDs → plans)
└── Decision coverage gate (CONTEXT.md `<decisions>` → plans, BLOCKING — #2492)
@@ -653,6 +657,30 @@ Debounce: 5 tool uses between repeated warnings. Severity escalation (WARNING→
- Missing bridge files handled gracefully (subagents, fresh sessions)
- Context monitor is advisory — never issues imperative commands that override user preferences
### Package Legitimacy Gate (v1.51)
The researcher → planner → executor pipeline includes a supply-chain gate against slopsquatting (AI-hallucinated package names pre-registered with malicious post-install scripts).
**Threat model:** GSD automates the full path from "researcher names a package" to "executor runs `npm install`". A hallucinated name that passes `npm view` (proving only registration, not legitimacy) would previously flow through undetected. ~20% of AI-generated package references are hallucinated; ~43% of those names recur consistently across prompts, making pre-registration economically viable for attackers.
**Gate layers:**
| Layer | Component | Action |
|-------|-----------|--------|
| Research | `gsd-phase-researcher` | Runs `slopcheck install <pkgs> --json`; writes `## Package Legitimacy Audit` table to RESEARCH.md; strips `[SLOP]` packages before RESEARCH.md is written |
| Planning | `gsd-planner` | Reads Audit table; inserts `checkpoint:human-verify` before any `[ASSUMED]` or `[SUS]` install task; adds `T-{phase}-SC` STRIDE supply-chain row to `<threat_model>` |
| Execution | `gsd-executor` | RULE 3 excludes package installation from auto-fix scope; failed installs surface as checkpoints, never silent substitutions |
**Claim provenance integration:** Package names discovered via WebSearch are tagged `[ASSUMED]` (not `[VERIFIED]`) regardless of `npm view` result. This extends the existing `[ASSUMED]` / `[VERIFIED]` / `[CITED]` provenance system by enforcing the provenance tag as a hard gate at the install boundary — `[ASSUMED]` always generates a `checkpoint:human-verify` in PLAN.md.
**Ecosystem coverage:** The researcher uses registry-specific verification commands — `npm view` (Node), `pip index versions` (Python), `cargo search` (Rust) — rather than a single generic check. This catches cross-ecosystem hallucination (~9% rate documented in 2025 USENIX research).
**Graceful degradation:** If `slopcheck` is unavailable, every recommended package is tagged `[ASSUMED]` and gated with a checkpoint. Research and planning proceed; the system never hard-fails on a missing tool dependency.
**External dependency:** `slopcheck` (MIT, pip-installable). If abandoned, the `[ASSUMED]`-gate fallback maintains human-checkpoint coverage.
---
### Security Hooks (v1.27)
**Prompt Guard** (`gsd-prompt-guard.js`):

View File

@@ -162,6 +162,17 @@ Research, plan, and verify a phase.
- With `--research`: force-refresh — re-spawn researcher unconditionally, no prompt.
- With `--view`: print existing RESEARCH.md to stdout, no spawn. Errors if RESEARCH.md missing.
**Package Legitimacy Gate (v1.51):**
When the researcher recommends external packages, it runs `slopcheck install <pkg> --json` on each one and writes a `## Package Legitimacy Audit` table to RESEARCH.md recording Registry, Age, Downloads, Source Repo, and slopcheck verdict. Verdicts:
- `[SLOP]` — package removed from RESEARCH.md entirely; never reaches the planner
- `[SUS]` — package flagged; planner inserts `checkpoint:human-verify` before the install task
- `[OK]` — package approved; no checkpoint added
Packages sourced from WebSearch are tagged `[ASSUMED]` (not `[VERIFIED]`) and treated the same as `[SUS]` — they get a human checkpoint before install. If `slopcheck` cannot be installed, every recommended package is tagged `[ASSUMED]` and gated.
See [Package Legitimacy Gate in the User Guide](USER-GUIDE.md#package-legitimacy-gate-v151) for the full checkpoint format, verdict table, and troubleshooting.
```bash
/gsd-plan-phase 1 # Research + plan + verify phase 1
/gsd-plan-phase 3 --skip-research # Plan without research (familiar domain)
@@ -227,6 +238,8 @@ Execute all plans in a phase with wave-based parallelization, or run a specific
**Prerequisites:** Phase has PLAN.md files
**Produces:** per-plan `{phase}-{N}-SUMMARY.md`, git commits, and `{phase}-VERIFICATION.md` when the phase is fully complete
**Package install failures (v1.51):** If a plan's install step fails, the executor surfaces a `checkpoint:human-verify` and stops. It does not auto-install a similarly-named alternative. This is intentional — silently substituting package names is how slopsquatting spreads. Respond to the checkpoint after verifying the package on its registry page.
```bash
/gsd-execute-phase 1 # Execute phase 1
/gsd-execute-phase 1 --wave 2 # Execute only Wave 2

View File

@@ -721,6 +721,87 @@ The `security.cjs` module scans for known injection patterns (role overrides, in
---
### Package Legitimacy Gate (v1.51)
AI coding tools hallucinate package names. Attackers pre-register those names on npm, PyPI, and crates.io with malicious post-install scripts — a technique called *slopsquatting*. A hallucinated name that passes `npm view` looks legitimate, so it would flow undetected through GSD's research → plan → execute pipeline all the way to `npm install <malicious-pkg>` running on your machine.
v1.51 adds a three-layer gate that stops this before it reaches your shell.
#### What you'll see
**In RESEARCH.md** — every phase that recommends external packages now includes a `## Package Legitimacy Audit` table:
```markdown
## Package Legitimacy Audit
| Package | Registry | Age | Downloads | Source Repo | slopcheck | Disposition |
|---------|----------|-----|-----------|-------------|-----------|-------------|
| express | npm | 13 yrs | 100M+/wk | github.com/expressjs/express | [OK] | Approved |
| some-new-util | npm | 3 days | 47 | none | [SLOP] | REMOVED |
| api-bridge | npm | 6 mo | 1.2k/wk | github.com/user/api-bridge | [SUS] | Flagged |
**Packages removed due to slopcheck:** some-new-util
**Packages flagged as suspicious:** api-bridge — planner will require human verification before install
```
`[SLOP]` packages are removed from RESEARCH.md entirely. They never reach the planner.
**In PLAN.md** — if a package is tagged `[ASSUMED]` (sourced from WebSearch, not registry-verified) or `[SUS]` (slopcheck suspicious), the plan includes a verification checkpoint *before* the install task:
```xml
<task type="checkpoint:human-verify">
<what-built>Package verification required before install</what-built>
<how-to-verify>
Verify these packages before proceeding:
- `api-bridge` [SUS — 6 months old, 1.2k downloads/week, GitHub repo present]
Check: https://npmjs.com/package/api-bridge
Look for: maintainer history, issue tracker activity, no suspicious install scripts
</how-to-verify>
<resume-signal>Type "verified" once you've confirmed all packages are legitimate</resume-signal>
</task>
```
**During execution** — if an install fails, the executor surfaces a checkpoint and stops. It does not silently try a similarly-named alternative (which could be even more dangerous).
#### Slopcheck verdicts
| Verdict | Meaning | GSD action |
|---------|---------|------------|
| `[OK]` | Package passes all legitimacy checks | Proceeds — no checkpoint added |
| `[SUS]` | Suspicious signals (new, low downloads, no source repo, etc.) | Flagged in Audit table; planner adds `checkpoint:human-verify` before install |
| `[SLOP]` | High-confidence hallucination or attacker-registered package | Removed from RESEARCH.md; never reaches planner |
#### Claim provenance and WebSearch packages
Package names discovered through WebSearch are always tagged `[ASSUMED]` in RESEARCH.md, regardless of whether `npm view` succeeds. A package that exists on the registry is not the same as a package that's safe to install — `npm view` only proves registration, not legitimacy.
`[ASSUMED]` packages trigger the same `checkpoint:human-verify` gate as `[SUS]` packages. You'll see the checkpoint with a link to the registry page and guidance on what to look for.
#### If slopcheck isn't installed
GSD attempts `pip install slopcheck` at research time. If that fails:
- Every recommended package is tagged `[ASSUMED]`
- The planner gates every install with a `checkpoint:human-verify` task
- Research and planning complete normally — nothing hard-fails
This is intentionally stricter than the normal flow: slopcheck unavailability means every package install gets a human checkpoint, which is the safest fallback.
To install slopcheck manually:
```bash
pip install slopcheck
# verify: slopcheck install express --json
```
#### slopcheck dependency
`slopcheck` is a MIT-licensed Python tool maintained by ToxSec (the researcher who documented the slopsquatting attack surface). It checks packages across npm, PyPI, crates.io, RubyGems, Go modules, Maven, and Packagist using multi-signal heuristics: registry age, download count, source-repo linkage, naming distance to popular packages, and registry-specific suspicion patterns.
If `slopcheck` is ever unavailable or abandoned, GSD's `[ASSUMED]`-gate fallback ensures you always get a human checkpoint before any install — the system never silently degrades to the pre-v1.51 behavior.
---
### Execution Wave Coordination
```

View File

@@ -0,0 +1,472 @@
'use strict';
/**
* Package Legitimacy Gate — structural contract tests (#2827)
*
* Verifies that the three agents (researcher, planner, executor) contain the
* interlocking instruction text that forms the slopsquatting defence gate.
*/
const { describe, test, before } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const AGENTS = path.join(__dirname, '..', 'agents');
const RESEARCHER = path.join(AGENTS, 'gsd-phase-researcher.md');
const PLANNER = path.join(AGENTS, 'gsd-planner.md');
const EXECUTOR = path.join(AGENTS, 'gsd-executor.md');
function parseSections(md) {
const lines = md.split('\n');
const sections = [];
let current = { heading: '__preamble__', body: [] };
let inFence = false;
for (const line of lines) {
if (line.trimStart().startsWith('```')) inFence = !inFence;
if (!inFence && /^#{1,3} /.test(line)) {
sections.push(current);
current = { heading: line.replace(/^#+\s*/, '').trim(), body: [] };
continue;
}
current.body.push(line);
}
sections.push(current);
return sections;
}
function extractCodeBlocks(text) {
const blocks = [];
const lines = text.split('\n');
let inside = false;
let buf = [];
for (const line of lines) {
if (line.trimStart().startsWith('```')) {
if (inside) {
blocks.push(buf.join('\n'));
buf = [];
}
inside = !inside;
continue;
}
if (inside) buf.push(line);
}
return blocks;
}
function extractResearchTemplate(content) {
const lines = content.split('\n');
let inside = false;
let isMarkdownFence = false;
let buf = [];
for (const line of lines) {
if (!inside && line.startsWith('```markdown')) {
inside = true;
isMarkdownFence = true;
buf = [];
continue;
}
if (inside && line.startsWith('```') && isMarkdownFence) {
const candidate = buf.join('\n');
if (/^\s*#\s+Phase\b/m.test(candidate)) return candidate;
inside = false;
isMarkdownFence = false;
continue;
}
if (inside) buf.push(line);
}
return '';
}
function extractPlanTemplate(content) {
const blocks = extractCodeBlocks(content);
for (const block of blocks) {
if (/^\s*<threat_model>/m.test(block)) return block;
}
return '';
}
function extractXmlElement(text, tag) {
const start = text.indexOf(`<${tag}>`);
const end = text.indexOf(`</${tag}>`);
if (start === -1 || end === -1) return '';
return text.slice(start, end + tag.length + 3);
}
function normalizeTokens(text) {
return text
.toLowerCase()
.replace(/https?:\/\//g, ' ')
.replace(/[\[\]]/g, '')
.replace(/[^a-z0-9{}:_-]+/g, ' ')
.trim()
.split(/\s+/)
.filter(Boolean);
}
function hasAllTokens(text, required) {
const tokenSet = new Set(normalizeTokens(text));
return required.every((token) => tokenSet.has(token.toLowerCase()));
}
function anyLineHasAll(lines, required) {
return lines.some((line) => hasAllTokens(line, required));
}
function parseMarkdownTable(lines) {
const tableLines = lines.filter((line) => /^\s*\|/.test(line));
if (tableLines.length < 2) return null;
const toCells = (line) => line
.trim()
.replace(/^\|/, '')
.replace(/\|$/, '')
.split('|')
.map((cell) => cell.trim());
const headers = toCells(tableLines[0]);
const rows = tableLines
.slice(2)
.map(toCells)
.filter((cells) => cells.length === headers.length)
.map((cells) => ({
cells,
fields: Object.fromEntries(headers.map((h, i) => [h, cells[i]])),
}));
return { headers, rows };
}
function parseMarkdownTables(lines) {
const groups = [];
let current = [];
for (const line of lines) {
if (/^\s*\|/.test(line)) {
current.push(line);
continue;
}
if (current.length > 0) {
groups.push(current);
current = [];
}
}
if (current.length > 0) groups.push(current);
return groups
.map((group) => parseMarkdownTable(group))
.filter(Boolean);
}
function lineIndexes(lines, predicate) {
const indexes = [];
for (let i = 0; i < lines.length; i += 1) {
if (predicate(lines[i], i)) indexes.push(i);
}
return indexes;
}
function inNearbyWindow(sourceIndexes, targetIndexes, distance) {
return sourceIndexes.some((src) => targetIndexes.some((dst) => Math.abs(src - dst) <= distance));
}
function readModel(filePath) {
const text = fs.readFileSync(filePath, 'utf-8');
return {
text,
lines: text.split('\n'),
sections: parseSections(text),
codeBlocks: extractCodeBlocks(text),
};
}
describe('gsd-phase-researcher.md — slopcheck invocation', () => {
let model;
before(() => {
model = readModel(RESEARCHER);
});
test('contains slopcheck install command in a fenced code block', () => {
const found = model.codeBlocks.some((block) => hasAllTokens(block, ['slopcheck', 'install']));
assert.ok(found, 'researcher must invoke slopcheck install inside a fenced code block');
});
test('slopcheck invocation includes --json flag', () => {
const found = model.codeBlocks.some((block) =>
hasAllTokens(block, ['slopcheck', 'install']) && hasAllTokens(block, ['json'])
);
assert.ok(found, 'slopcheck invocation must pass --json');
});
test('guards slopcheck invocation with command availability check', () => {
const hasCommandV = model.codeBlocks.some((block) => hasAllTokens(block, ['command', '-v', 'slopcheck']));
const hasWhich = model.codeBlocks.some((block) => hasAllTokens(block, ['which', 'slopcheck']));
assert.ok(hasCommandV || hasWhich, 'researcher must guard slopcheck invocation with command -v or which');
});
test('documents graceful degradation when slopcheck is unavailable', () => {
const hasAssumedLine = anyLineHasAll(model.lines, ['assumed']);
const hasSlopcheckUnavailableLine = model.lines.some((line) => {
const slopcheckMention = hasAllTokens(line, ['slopcheck']);
const unavailableMention =
hasAllTokens(line, ['not', 'available']) ||
hasAllTokens(line, ['not', 'found']) ||
hasAllTokens(line, ['unavailable']) ||
hasAllTokens(line, ['cannot', 'installed']);
return slopcheckMention && unavailableMention;
});
assert.ok(
hasAssumedLine && hasSlopcheckUnavailableLine,
'researcher must document [ASSUMED] fallback when slopcheck cannot run'
);
});
});
describe('gsd-phase-researcher.md — Package Legitimacy Audit section in template', () => {
let templateSections;
before(() => {
const model = readModel(RESEARCHER);
const template = extractResearchTemplate(model.text);
templateSections = parseSections(template);
});
test('RESEARCH.md template contains Package Legitimacy Audit section', () => {
const section = templateSections.find((s) => s.heading === 'Package Legitimacy Audit');
assert.ok(section, 'RESEARCH.md template must include a Package Legitimacy Audit section');
});
test('Package Legitimacy Audit table has required columns', () => {
const section = templateSections.find((s) => s.heading === 'Package Legitimacy Audit');
assert.ok(section, 'Package Legitimacy Audit section must exist');
const table = parseMarkdownTable(section.body);
assert.ok(table, 'Package Legitimacy Audit section must include a markdown table');
const expected = ['Package', 'Registry', 'Age', 'Downloads', 'slopcheck', 'Disposition'];
for (const column of expected) {
assert.ok(table.headers.includes(column), `audit table must have "${column}" column`);
}
});
test('audit section documents [SLOP], [SUS], and [OK] dispositions', () => {
const section = templateSections.find((s) => s.heading === 'Package Legitimacy Audit');
assert.ok(section, 'Package Legitimacy Audit section must exist');
const table = parseMarkdownTable(section.body);
assert.ok(table, 'Package Legitimacy Audit section must include a markdown table');
const rowTexts = table.rows.map((row) => row.cells.join(' '));
const slop = rowTexts.some((value) => hasAllTokens(value, ['slop']));
const sus = rowTexts.some((value) => hasAllTokens(value, ['sus']));
const ok = rowTexts.some((value) => hasAllTokens(value, ['ok']));
assert.ok(slop, 'audit section must document [SLOP] disposition');
assert.ok(sus, 'audit section must document [SUS] disposition');
assert.ok(ok, 'audit section must document [OK] disposition');
});
});
describe('gsd-phase-researcher.md — ecosystem-specific package verification', () => {
let model;
before(() => {
model = readModel(RESEARCHER);
});
test('documents pip index versions for Python phases', () => {
assert.ok(anyLineHasAll(model.lines, ['pip', 'index', 'versions']), 'researcher must document pip index versions');
});
test('documents cargo search for Rust phases', () => {
assert.ok(anyLineHasAll(model.lines, ['cargo', 'search']), 'researcher must document cargo search');
});
});
describe('gsd-phase-researcher.md — no npx --yes auto-download', () => {
let model;
before(() => {
model = readModel(RESEARCHER);
});
test('does not invoke npx --yes inside a code block', () => {
const found = model.codeBlocks.some((block) => hasAllTokens(block, ['npx', '--yes']));
assert.equal(found, false, 'researcher must not invoke npx --yes in any code block');
});
test('ctx7 CLI fallback uses command -v guard', () => {
const found = model.codeBlocks.some((block) => hasAllTokens(block, ['command', '-v', 'ctx7']));
assert.ok(found, 'ctx7 CLI fallback must guard with command -v ctx7 before invocation');
});
});
describe('gsd-phase-researcher.md — WebSearch-origin package tagging', () => {
let model;
before(() => {
model = readModel(RESEARCHER);
});
test('packages discovered via WebSearch are tagged [ASSUMED]', () => {
const webSearchLines = lineIndexes(model.lines, (line) => hasAllTokens(line, ['websearch']));
const assumedLines = lineIndexes(model.lines, (line) => hasAllTokens(line, ['assumed']));
assert.ok(webSearchLines.length > 0, 'researcher file must mention WebSearch');
assert.ok(assumedLines.length > 0, 'researcher file must mention [ASSUMED]');
assert.ok(
inNearbyWindow(webSearchLines, assumedLines, 25),
'researcher must instruct WebSearch-discovered packages are tagged [ASSUMED] in nearby guidance'
);
});
});
describe('gsd-planner.md — checkpoint gate for [ASSUMED]/[SUS] packages', () => {
let model;
before(() => {
model = readModel(PLANNER);
});
test('checkpoint:human-verify guidance references [ASSUMED] and [SUS]', () => {
const hasCheckpoint = anyLineHasAll(model.lines, ['checkpoint:human-verify']);
const hasAssumed = anyLineHasAll(model.lines, ['assumed']);
const hasSus = anyLineHasAll(model.lines, ['sus']);
assert.ok(hasCheckpoint && hasAssumed, 'planner must gate [ASSUMED] packages behind checkpoint:human-verify');
assert.ok(hasCheckpoint && hasSus, 'planner must gate [SUS] packages behind checkpoint:human-verify');
});
test('package-legitimacy checkpoint uses blocking-human gate and non-auto-approvable language', () => {
const hasBlockingHumanGate = anyLineHasAll(model.lines, ['checkpoint:human-verify', 'blocking-human']);
const hasNeverAutoApproveRule = model.lines.some((line) =>
hasAllTokens(line, ['never', 'auto-approvable']) ||
hasAllTokens(line, ['never', 'auto', 'approvable'])
);
assert.ok(hasBlockingHumanGate, 'planner legitimacy checkpoint must use gate="blocking-human"');
assert.ok(hasNeverAutoApproveRule, 'planner must state legitimacy checkpoints are never auto-approvable');
});
test('package verification checkpoint includes registry URL guidance', () => {
const hasRegistryGuidance = model.lines.some((line) =>
hasAllTokens(line, ['npmjs', 'package']) ||
hasAllTokens(line, ['pypi', 'project']) ||
hasAllTokens(line, ['crates', 'crates'])
);
assert.ok(hasRegistryGuidance, 'planner package-verify checkpoint must include registry URL examples');
});
});
describe('gsd-planner.md — supply-chain row in threat_model template', () => {
let planTemplate;
let threatModelBlock;
before(() => {
const model = readModel(PLANNER);
planTemplate = extractPlanTemplate(model.text);
threatModelBlock = extractXmlElement(planTemplate, 'threat_model');
});
test('PLAN.md template contains threat_model element', () => {
assert.ok(/^\s*<threat_model>/m.test(planTemplate), 'PLAN.md template must include <threat_model>');
});
test('threat_model template includes supply-chain row with mitigate disposition', () => {
const tables = parseMarkdownTables(threatModelBlock.split('\n'));
const strideTable = tables.find((table) => table.headers.includes('Threat ID'));
assert.ok(strideTable, 'threat_model must include STRIDE threat register table');
const supplyChainRow = strideTable.rows.find((row) => hasAllTokens(row.cells[0] || '', ['t-{phase}-sc']));
assert.ok(supplyChainRow, 'threat_model must include T-{phase}-SC supply-chain row');
const disposition = supplyChainRow.cells[3] || '';
assert.ok(hasAllTokens(disposition, ['mitigate']), 'supply-chain threat disposition must be mitigate');
});
});
describe('gsd-planner.md — no npx --yes auto-download', () => {
test('does not invoke npx --yes inside a code block', () => {
const model = readModel(PLANNER);
const found = model.codeBlocks.some((block) => hasAllTokens(block, ['npx', '--yes']));
assert.equal(found, false, 'planner must not invoke npx --yes in any code block');
});
});
describe('gsd-executor.md — package installs excluded from RULE 3 auto-fix', () => {
let model;
before(() => {
model = readModel(EXECUTOR);
});
test('does not invoke npx --yes inside a code block', () => {
const found = model.codeBlocks.some((block) => hasAllTokens(block, ['npx', '--yes']));
assert.equal(found, false, 'executor must not invoke npx --yes in any code block');
});
test('RULE 3 section explicitly excludes package-manager installs', () => {
const rule3Line = lineIndexes(model.lines, (line) => hasAllTokens(line, ['rule', '3']))[0];
assert.notEqual(rule3Line, undefined, 'executor must contain RULE 3 section');
const window = model.lines.slice(rule3Line, rule3Line + 35);
const hasInstallCommands =
anyLineHasAll(window, ['npm', 'install']) ||
anyLineHasAll(window, ['pip', 'install']) ||
anyLineHasAll(window, ['cargo', 'add']);
const hasExclusionLanguage =
anyLineHasAll(window, ['excluded']) ||
anyLineHasAll(window, ['not', 'auto-fixable']) ||
anyLineHasAll(window, ['do', 'not']);
assert.ok(hasInstallCommands && hasExclusionLanguage, 'RULE 3 must explicitly exclude package-manager installs');
});
test('failed package installs surface checkpoint:human-verify', () => {
const rule3Line = lineIndexes(model.lines, (line) => hasAllTokens(line, ['rule', '3']))[0];
assert.notEqual(rule3Line, undefined, 'executor must contain RULE 3 section');
const window = model.lines.slice(rule3Line, rule3Line + 50);
const hasFailureLanguage =
anyLineHasAll(window, ['failed', 'install']) ||
anyLineHasAll(window, ['install', 'fails']) ||
anyLineHasAll(window, ['install', 'failed']);
const hasCheckpoint = anyLineHasAll(window, ['checkpoint:human-verify']);
assert.ok(
hasFailureLanguage && hasCheckpoint,
'executor must emit checkpoint:human-verify when package install fails'
);
});
test('auto mode does not auto-approve package-legitimacy checkpoints', () => {
const autoModeLine = lineIndexes(model.lines, (line) => hasAllTokens(line, ['auto-mode', 'checkpoint', 'behavior']))[0];
assert.notEqual(autoModeLine, undefined, 'executor must define auto-mode checkpoint behavior');
const window = model.lines.slice(autoModeLine, autoModeLine + 25);
const hasExceptionRule =
anyLineHasAll(window, ['except', 'package-legitimacy', 'checkpoints']) ||
anyLineHasAll(window, ['do', 'not', 'auto-approve']) ||
anyLineHasAll(window, ['blocking-human']);
assert.ok(
hasExceptionRule,
'executor auto mode must explicitly block auto-approval for package-legitimacy checkpoints'
);
});
});