Commit Graph

3 Commits

Author SHA1 Message Date
Tom Boucher
e32a53b974 feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (#133)
* test(113): add per-rule failing tests + hostile fixture for markdown link payloads

RED phase for issue #113 — scanForInjection() currently returns { clean: true }
for markdown links containing javascript:, data:text/html, userinfo credentials,
and token-in-query payloads.

Changes:
- tests/fixtures/adversarial/security/context-malicious-markdown-link.md:
  Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME,
  MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign
  negative controls (data:image/png, mailto:, https://github.com, port-only URL).
- tests/security-prompt-injection.test.cjs:
  - Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion
    to "malicious-markdown-link fixture is flagged by scanner" (forward-looking).
  - Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings
    with ruleId, file, line, match fields.
  - Added parity guard: every MARKDOWN_LINK_PATTERNS source string from
    security.cjs must appear in gsd-read-injection-scanner.js hook source.

D3 false-positive grep: 0 legitimate matches — no allowlist entries needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook)

GREEN phase for issue #113.

Rule details (all with primary source citations):

  MD-LINK-JS-SCHEME
    Flags ](javascript:...) regardless of case.
    Source: OWASP XSS Prevention Cheat Sheet
    https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html

  MD-LINK-DATA-SCHEME
    Flags data: URIs NOT in the explicit safe-list.
    Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf).
    data:image/svg+xml is intentionally BLOCKED — SVG can host <script>.
    Source: OWASP File Upload Cheat Sheet — SVG Files
    https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files

  MD-LINK-USERINFO
    Flags https?://user:pass@host in markdown link targets.
    Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo).
    Source: RFC 3986 §3.2.1 (userinfo syntax)
    https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
    RFC 9110 §4.2.4 (HTTP deprecates userinfo)
    https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4

  MD-LINK-TOKEN-IN-QUERY
    Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey,
    secret, password, client_secret, code — regardless of value.
    Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1
    https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
    D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed.

Architecture:
- scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection()
  extended with structuredFindings (ruleId, file, line, match) via opts.file.
- hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence
  (same pattern sources, verified by parity test).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(113): flip PINNED malicious-markdown-link assertion and add parity guard

REFACTOR phase — tightening test rigor after test-rigor skill review:

1. Fixture assertion now enumerates all 4 expected ruleIds explicitly:
   [MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY].
   Previously findings.length > 0 would pass even if 3 of 4 rules were broken.

2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for
   1-based line numbers) instead of `typeof f.line === 'number'` (vacuous).

3. match field assertions tightened to check the hostile content is present:
   - MD-LINK-JS-SCHEME: /javascript:/i in match
   - MD-LINK-DATA-SCHEME: /data:/i in match
   - MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker)
   - MD-LINK-TOKEN-IN-QUERY: /token=/i in match

4. Parity test checks actual RegExp .source strings (not just lengths), verifying
   the hook contains the exact canonical pattern sources character-for-character.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility

1. .changeset/113-malicious-markdown-links.md — required Security fragment
   for the user-facing markdown-link scanner changes in this PR (changeset-lint
   was failing with FAIL_MISSING_FRAGMENT).

2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR
   set for mutation state subcommands whose CJS contract is always exit-0.
   On Windows/Node 24 the SDK bridge returns result.ok===false for validation
   failures (e.g. state record-metric --phase 1 with no --plan/--duration),
   causing dispatchViaSdk() to call error() (exit 1) instead of output({error})
   (exit 0). The fix maps SDK non-ok results to JSON output for the affected
   mutation commands (record-metric, advance-plan, record-session, add-decision,
   add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS
   contract on all platforms.

tests/state.test.cjs:1161 "returns error when required fields missing" passes
locally (104/104 pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 10:15:03 -04:00
Jeremy McSpadden
3856b53098 Merge remote-tracking branch 'origin/main' into fix/2406-ship-read-injection-scanner
# Conflicts:
#	CHANGELOG.md
2026-04-18 11:37:47 -05:00
Tom Boucher
c35997fb0b feat(hooks): add gsd-read-injection-scanner PostToolUse hook (#2201) (#2328)
* feat: add /gsd-spec-phase — Socratic spec refinement with ambiguity scoring (#2213)

Introduces `/gsd-spec-phase <phase>` as an optional pre-step before discuss-phase.
Clarifies WHAT a phase delivers (requirements, boundaries, acceptance criteria) with
quantitative ambiguity scoring before discuss-phase handles HOW to implement.

- `commands/gsd/spec-phase.md` — slash command routing to workflow
- `get-shit-done/workflows/spec-phase.md` — full Socratic interview loop (up to 6
  rounds, 5 rotating perspectives: Researcher, Simplifier, Boundary Keeper, Failure
  Analyst, Seed Closer) with weighted 4-dimension ambiguity gate (≤ 0.20 to write SPEC.md)
- `get-shit-done/templates/spec.md` — SPEC.md template with falsifiable requirements
  (Current/Target/Acceptance per requirement), Boundaries, Acceptance Criteria,
  Ambiguity Report, and Interview Log; includes two full worked examples
- `get-shit-done/workflows/discuss-phase.md` — new `check_spec` step detects
  `{padded_phase}-SPEC.md` at startup; displays "Found SPEC.md — N requirements
  locked. Focusing on implementation decisions."; `analyze_phase` respects `spec_loaded`
  flag to skip "what/why" gray areas; `write_context` emits `<spec_lock>` section
  with boundary summary and canonical ref to SPEC.md
- `docs/ARCHITECTURE.md` — update command/workflow counts (74→75, 71→72)

Closes #2213

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(hooks): add gsd-read-injection-scanner PostToolUse hook (#2201)

Adds a new PostToolUse hook that scans content returned by the Read tool
for prompt injection patterns, including four summarisation-specific patterns
(retention-directive, permanence-claim, etc.) that survive context compression.

Defense-in-depth for long GSD sessions where the context summariser cannot
distinguish user instructions from content read from external files.

- Advisory-only (warns without blocking), consistent with gsd-prompt-guard.js
- LOW severity for 1-2 patterns, HIGH for 3+
- Inlined pattern library (hook independence)
- Exclusion list: .planning/, REVIEW.md, CHECKPOINT, security docs, hook sources
- Wired in install.js as PostToolUse matcher: Read, timeout: 5s
- Added to MANAGED_HOOKS for staleness detection
- 19 tests covering all 13 acceptance criteria (SCAN-01–07, EXCL-01–06, EDGE-01–06)

Closes #2201

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): add read-injection-scanner files to prompt-injection-scan allowlist

Test payloads in tests/read-injection-scanner.test.cjs and inlined patterns
in hooks/gsd-read-injection-scanner.js legitimately contain injection strings.
Add both to the CI script allowlist to prevent false-positive failures.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(test): assert exitCode, stdout, and signal explicitly in EDGE-05

Addresses CodeRabbit feedback: the success path discarded the return
value so a malformed-JSON input that produced stdout would still pass.
Now captures and asserts all three observable properties.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 17:22:31 -04:00