Files
msd-core/hooks/gsd-read-injection-scanner.js
Tom Boucher e32a53b974 feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (#133)
* test(113): add per-rule failing tests + hostile fixture for markdown link payloads

RED phase for issue #113 — scanForInjection() currently returns { clean: true }
for markdown links containing javascript:, data:text/html, userinfo credentials,
and token-in-query payloads.

Changes:
- tests/fixtures/adversarial/security/context-malicious-markdown-link.md:
  Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME,
  MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign
  negative controls (data:image/png, mailto:, https://github.com, port-only URL).
- tests/security-prompt-injection.test.cjs:
  - Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion
    to "malicious-markdown-link fixture is flagged by scanner" (forward-looking).
  - Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings
    with ruleId, file, line, match fields.
  - Added parity guard: every MARKDOWN_LINK_PATTERNS source string from
    security.cjs must appear in gsd-read-injection-scanner.js hook source.

D3 false-positive grep: 0 legitimate matches — no allowlist entries needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook)

GREEN phase for issue #113.

Rule details (all with primary source citations):

  MD-LINK-JS-SCHEME
    Flags ](javascript:...) regardless of case.
    Source: OWASP XSS Prevention Cheat Sheet
    https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html

  MD-LINK-DATA-SCHEME
    Flags data: URIs NOT in the explicit safe-list.
    Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf).
    data:image/svg+xml is intentionally BLOCKED — SVG can host <script>.
    Source: OWASP File Upload Cheat Sheet — SVG Files
    https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files

  MD-LINK-USERINFO
    Flags https?://user:pass@host in markdown link targets.
    Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo).
    Source: RFC 3986 §3.2.1 (userinfo syntax)
    https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
    RFC 9110 §4.2.4 (HTTP deprecates userinfo)
    https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4

  MD-LINK-TOKEN-IN-QUERY
    Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey,
    secret, password, client_secret, code — regardless of value.
    Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1
    https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
    D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed.

Architecture:
- scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection()
  extended with structuredFindings (ruleId, file, line, match) via opts.file.
- hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence
  (same pattern sources, verified by parity test).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(113): flip PINNED malicious-markdown-link assertion and add parity guard

REFACTOR phase — tightening test rigor after test-rigor skill review:

1. Fixture assertion now enumerates all 4 expected ruleIds explicitly:
   [MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY].
   Previously findings.length > 0 would pass even if 3 of 4 rules were broken.

2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for
   1-based line numbers) instead of `typeof f.line === 'number'` (vacuous).

3. match field assertions tightened to check the hostile content is present:
   - MD-LINK-JS-SCHEME: /javascript:/i in match
   - MD-LINK-DATA-SCHEME: /data:/i in match
   - MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker)
   - MD-LINK-TOKEN-IN-QUERY: /token=/i in match

4. Parity test checks actual RegExp .source strings (not just lengths), verifying
   the hook contains the exact canonical pattern sources character-for-character.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility

1. .changeset/113-malicious-markdown-links.md — required Security fragment
   for the user-facing markdown-link scanner changes in this PR (changeset-lint
   was failing with FAIL_MISSING_FRAGMENT).

2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR
   set for mutation state subcommands whose CJS contract is always exit-0.
   On Windows/Node 24 the SDK bridge returns result.ok===false for validation
   failures (e.g. state record-metric --phase 1 with no --plan/--duration),
   causing dispatchViaSdk() to call error() (exit 1) instead of output({error})
   (exit 0). The fix maps SDK non-ok results to JSON output for the affected
   mutation commands (record-metric, advance-plan, record-session, add-decision,
   add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS
   contract on all platforms.

tests/state.test.cjs:1161 "returns error when required fields missing" passes
locally (104/104 pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 10:15:03 -04:00

204 lines
7.3 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
// gsd-hook-version: {{GSD_VERSION}}
// GSD Read Injection Scanner — PostToolUse hook (#2201)
// Scans file content returned by the Read tool for prompt injection patterns.
// Catches poisoned content at ingestion before it enters conversation context.
//
// Defense-in-depth: long GSD sessions hit context compression, and the
// summariser does not distinguish user instructions from content read from
// external files. Poisoned instructions that survive compression become
// indistinguishable from trusted context. This hook warns at ingestion time.
//
// Triggers on: Read tool PostToolUse events
// Action: Advisory warning (does not block) — logs detection for awareness
// Severity: LOW (1–2 patterns), HIGH (3+ patterns)
//
// False-positive exclusion: .planning/, REVIEW.md, CHECKPOINT, security docs,
// hook source files — these legitimately contain injection-like strings.
const path = require('path');
// Summarisation-specific patterns (novel — not in gsd-prompt-guard.js).
// These target instructions specifically designed to survive context compression.
const SUMMARISATION_PATTERNS = [
/when\s+(?:summari[sz]ing|compressing|compacting),?\s+(?:retain|preserve|keep)\s+(?:this|these)/i,
/this\s+(?:instruction|directive|rule)\s+is\s+(?:permanent|persistent|immutable)/i,
/preserve\s+(?:these|this)\s+(?:rules?|instructions?|directives?)\s+(?:in|through|after|during)/i,
/(?:retain|keep)\s+(?:this|these)\s+(?:in|through|after)\s+(?:summar|compress|compact)/i,
];
// Markdown link patterns — mirrors scripts/security.cjs MARKDOWN_LINK_PATTERNS, inlined for hook independence.
// Issue #113: detect javascript:, data: (non-safe-list), userinfo credentials, and token-in-query.
//
// Sources:
// MD-LINK-JS-SCHEME: OWASP XSS Prevention
// https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html
// MD-LINK-DATA-SCHEME: OWASP File Upload (SVG unsafe)
// https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files
// MD-LINK-USERINFO: RFC 3986 §3.2.1, RFC 9110 §4.2.4
// https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
// https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4
// MD-LINK-TOKEN-IN-QUERY: RFC 9700 §4.3.1
// https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
const DATA_URI_SAFE_MIME_RE = /^data:(image\/(png|jpe?g|gif|webp|bmp|ico|avif|heic)|font\/(woff2?|otf|ttf))(;[^,]*)?,/i;
const MARKDOWN_LINK_PATTERNS = [
{
pattern: /\]\(\s*javascript:/i,
ruleId: 'MD-LINK-JS-SCHEME',
},
{
pattern: /\]\(\s*data:/i,
ruleId: 'MD-LINK-DATA-SCHEME',
safePredicate: (line) => {
const m = line.match(/\]\(\s*(data:[^)]*)/i);
if (!m) return false;
return DATA_URI_SAFE_MIME_RE.test(m[1]);
},
},
{
pattern: /\]\(\s*https?:\/\/[^/\s]+:[^/@\s]+@/i,
ruleId: 'MD-LINK-USERINFO',
},
{
pattern: /[?&](token|access_token|id_token|refresh_token|api_key|apikey|secret|password|client_secret|code)=/i,
ruleId: 'MD-LINK-TOKEN-IN-QUERY',
},
];
// Standard injection patterns — mirrors gsd-prompt-guard.js, inlined for hook independence.
const INJECTION_PATTERNS = [
/ignore\s+(all\s+)?previous\s+instructions/i,
/ignore\s+(all\s+)?above\s+instructions/i,
/disregard\s+(all\s+)?previous/i,
/forget\s+(all\s+)?(your\s+)?instructions/i,
/override\s+(system|previous)\s+(prompt|instructions)/i,
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
/act\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
/<\/?(?:system|assistant|human)>/i,
/\[SYSTEM\]/i,
/\[INST\]/i,
/<<\s*SYS\s*>>/i,
];
const ALL_PATTERNS = [...INJECTION_PATTERNS, ...SUMMARISATION_PATTERNS];
function isExcludedPath(filePath) {
const p = filePath.replace(/\\/g, '/');
return (
p.includes('/.planning/') ||
p.includes('.planning/') ||
/(?:^|\/)REVIEW\.md$/i.test(p) ||
/CHECKPOINT/i.test(path.basename(p)) ||
/[/\\](?:security|techsec|injection)[/\\.]/i.test(p) ||
/security\.cjs$/.test(p) ||
p.includes('/.claude/hooks/')
);
}
let inputBuf = '';
const stdinTimeout = setTimeout(() => process.exit(0), 5000);
process.stdin.setEncoding('utf8');
process.stdin.on('data', chunk => { inputBuf += chunk; });
process.stdin.on('end', () => {
clearTimeout(stdinTimeout);
try {
const data = JSON.parse(inputBuf);
if (data.tool_name !== 'Read') {
process.exit(0);
}
const filePath = data.tool_input?.file_path || '';
if (!filePath) {
process.exit(0);
}
if (isExcludedPath(filePath)) {
process.exit(0);
}
// Extract content from tool_response — string (cat -n output) or object form
let content = '';
const resp = data.tool_response;
if (typeof resp === 'string') {
content = resp;
} else if (resp && typeof resp === 'object') {
const c = resp.content;
if (Array.isArray(c)) {
content = c.map(b => (typeof b === 'string' ? b : b.text || '')).join('\n');
} else if (c != null) {
content = String(c);
}
}
if (!content || content.length < 20) {
process.exit(0);
}
const findings = [];
for (const pattern of ALL_PATTERNS) {
if (pattern.test(content)) {
// Trim pattern source for readable output
findings.push(pattern.source.replace(/\\s\+/g, '-').replace(/[()\\]/g, '').substring(0, 50));
}
}
// Markdown link patterns (issue #113)
const lines = content.split('\n');
for (const entry of MARKDOWN_LINK_PATTERNS) {
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
const m = line.match(entry.pattern);
if (!m) continue;
if (entry.safePredicate && entry.safePredicate(line)) continue;
findings.push(`${entry.ruleId}:${m[0].substring(0, 40)}`);
}
}
// Invisible Unicode (zero-width, RTL override, soft hyphen, BOM)
if (/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD\u2060-\u2069]/.test(content)) {
findings.push('invisible-unicode');
}
// Unicode tag block U+E0000–E007F (invisible instruction injection vector)
try {
if (/[\u{E0000}-\u{E007F}]/u.test(content)) {
findings.push('unicode-tag-block');
}
} catch {
// Engine does not support Unicode property escapes — skip this check
}
if (findings.length === 0) {
process.exit(0);
}
const severity = findings.length >= 3 ? 'HIGH' : 'LOW';
const fileName = path.basename(filePath);
const detail = severity === 'HIGH'
? 'Multiple patterns — strong injection signal. Review the file for embedded instructions before proceeding.'
: 'Single pattern match may be a false positive (e.g., documentation). Proceed with awareness.';
const output = {
hookSpecificOutput: {
hookEventName: 'PostToolUse',
additionalContext:
`\u26a0\ufe0f READ INJECTION SCAN [${severity}]: File "${fileName}" triggered ` +
`${findings.length} pattern(s): ${findings.join(', ')}. ` +
`This content is now in your conversation context. ${detail} ` +
`Source: ${filePath}`,
},
};
process.stdout.write(JSON.stringify(output));
} catch {
// Silent fail — never block tool execution
process.exit(0);
}
});