* test(113): add per-rule failing tests + hostile fixture for markdown link payloads
RED phase for issue #113 — scanForInjection() currently returns { clean: true }
for markdown links containing javascript:, data:text/html, userinfo credentials,
and token-in-query payloads.
Changes:
- tests/fixtures/adversarial/security/context-malicious-markdown-link.md:
Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME,
MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign
negative controls (data:image/png, mailto:, https://github.com, port-only URL).
- tests/security-prompt-injection.test.cjs:
- Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion
to "malicious-markdown-link fixture is flagged by scanner" (forward-looking).
- Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings
with ruleId, file, line, match fields.
- Added parity guard: every MARKDOWN_LINK_PATTERNS source string from
security.cjs must appear in gsd-read-injection-scanner.js hook source.
D3 false-positive grep: 0 legitimate matches — no allowlist entries needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook)
GREEN phase for issue #113.
Rule details (all with primary source citations):
MD-LINK-JS-SCHEME
Flags ](javascript:...) regardless of case.
Source: OWASP XSS Prevention Cheat Sheet
https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html
MD-LINK-DATA-SCHEME
Flags data: URIs NOT in the explicit safe-list.
Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf).
data:image/svg+xml is intentionally BLOCKED — SVG can host <script>.
Source: OWASP File Upload Cheat Sheet — SVG Files
https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files
MD-LINK-USERINFO
Flags https?://user:pass@host in markdown link targets.
Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo).
Source: RFC 3986 §3.2.1 (userinfo syntax)
https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
RFC 9110 §4.2.4 (HTTP deprecates userinfo)
https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4
MD-LINK-TOKEN-IN-QUERY
Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey,
secret, password, client_secret, code — regardless of value.
Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1
https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed.
Architecture:
- scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection()
extended with structuredFindings (ruleId, file, line, match) via opts.file.
- hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence
(same pattern sources, verified by parity test).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(113): flip PINNED malicious-markdown-link assertion and add parity guard
REFACTOR phase — tightening test rigor after test-rigor skill review:
1. Fixture assertion now enumerates all 4 expected ruleIds explicitly:
[MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY].
Previously findings.length > 0 would pass even if 3 of 4 rules were broken.
2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for
1-based line numbers) instead of `typeof f.line === 'number'` (vacuous).
3. match field assertions tightened to check the hostile content is present:
- MD-LINK-JS-SCHEME: /javascript:/i in match
- MD-LINK-DATA-SCHEME: /data:/i in match
- MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker)
- MD-LINK-TOKEN-IN-QUERY: /token=/i in match
4. Parity test checks actual RegExp .source strings (not just lengths), verifying
the hook contains the exact canonical pattern sources character-for-character.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility
1. .changeset/113-malicious-markdown-links.md — required Security fragment
for the user-facing markdown-link scanner changes in this PR (changeset-lint
was failing with FAIL_MISSING_FRAGMENT).
2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR
set for mutation state subcommands whose CJS contract is always exit-0.
On Windows/Node 24 the SDK bridge returns result.ok===false for validation
failures (e.g. state record-metric --phase 1 with no --plan/--duration),
causing dispatchViaSdk() to call error() (exit 1) instead of output({error})
(exit 0). The fix maps SDK non-ok results to JSON output for the affected
mutation commands (record-metric, advance-plan, record-session, add-decision,
add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS
contract on all platforms.
tests/state.test.cjs:1161 "returns error when required fields missing" passes
locally (104/104 pass).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): neutralize spaced+closing injection markers; fix audit-uat resolved status
scanForInjection recognizes — adds <user> tags, whitespace-padded tags
(e.g. <user >), closing [/SYSTEM]/[/INST] markers, and closing <</SYS>>
markers. Five new regression tests confirm each gap is closed.
whose result column reads PASS or resolved, so items that were already
confirmed do not appear as outstanding in audit-uat --raw. Two new
regression tests cover item-level PASS and file-level status: passed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: add closing-tag assertion for spaced <user > sanitization
The test for 'neutralizes spaced tags like <user >' only asserted that the
opening token '<user' was removed. A spaced closing tag '</user >' could
survive sanitization undetected. Added assert.ok(!result.includes('</user'))
to the same test block so both sides of the tag are verified.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: clarify capture_thought is an optional convention (#1873)
Issue #1873 merged /gsd:extract-learnings with an optional
capture_thought hook, but the docs never explained what the tool is
or where it comes from — readers couldn't tell whether it was a
bundled GSD tool, a required dependency, or something they had to
install. This surfaced in a user question on that issue's thread.
Clarify in docs/FEATURES.md §112 and the workflow file that
capture_thought is a convention — any MCP server exposing a tool
with that name will be used; if none is present, LEARNINGS.md
remains the primary output and the step is a silent no-op.
No behavioral change. All 23 extract-learnings tests still pass.
* fix(security): add human to detection message; test [/INST] closing form neutralization
- Detection message now lists <human> alongside <system>/<assistant>/<user>
- Sanitizer regex extended to cover [/INST] closing form (was only [INST])
- Detection pattern extended to cover [/INST] closing form
- New sanitizeForPrompt test asserts [/INST] is neutralized
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(config): add workflow.security_* keys to VALID_CONFIG_KEYS
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add language tag to fenced code block in FEATURES.md
Fixes MD040 lint finding in PR #2379 — the capture_thought tool
signature example was missing a javascript language identifier.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tests): allowlist execute-phase.md in prompt-injection scan
execute-phase.md grew to ~51K chars after the code-review gate step
was added in #1630, tripping the 50K size heuristic in the injection
scanner. The limit is calibrated for user-supplied input — trusted
workflow source files that legitimately exceed it are allowlisted
individually, following the same pattern as discuss-phase.md.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(security): improve prompt injection scanner with 4 detection layers (#1838)
- Layer 1: Unicode tag block U+E0000–U+E007F detection in strict mode (2025 supply-chain attack vector)
- Layer 2: Character-spacing obfuscation, delimiter injection (<system>/<assistant>/<user>/<human>), and long hex sequence patterns
- Layer 3: validatePromptStructure() — validates XML tag structure of agent/workflow files against known-valid tag set
- Layer 4: scanEntropyAnomalies() — Shannon entropy analysis flagging high-entropy paragraphs (>5.5 bits/char)
All layers implemented TDD (RED→GREEN): 31 new tests written first, verified failing, then implemented.
Full suite: 2559 tests, 0 failures. security.cjs: 99.6% stmt coverage.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(planner): add reachability_check step to prevent unreachable code
Closes#1495
* fix: trim gsd-planner.md below 50000-char limit after rebase
The reachability_check addition pushed the file to 50,275 chars when
merged with the assign_waves additions from #1600. Condense both sections
while preserving all logic; file is now 49,859 chars.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): normalize CRLF before prompt-stuffing length check
On Windows, git checks out files with CRLF line endings. JavaScript's
String.length counts \r characters, so a 49,859-byte file measures as
51,126 chars on Windows — falsely tripping the 50,000-char security
scanner. Normalize CRLF → LF before measuring in security.cjs and in
the reachability-check test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: trim gsd-planner.md to stay under 50000-char limit after merge with main
After rebasing onto main (which added mcp_tool_usage block), combined content
reached 50031 chars. Remove suggested log format from assign_waves rule to
bring file to 49972 chars, well under the 50000-char security scanner limit.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
The original PR (#1337) used \Z in a JavaScript regex, which is a
Perl/Python/Ruby anchor — JavaScript interprets it as a literal match
for the character 'Z', silently truncating expected text containing
that letter. Replace with a two-pass approach: try next-key lookahead
first, fall back to greedy match to end-of-string.
Also remove the redundant `to=all:` pattern in sanitizeForDisplay()
since it is a subset of the existing `to=[^:\s]+:` pattern.
Add regression tests proving the Z-truncation bug and verifying
expected blocks at end-of-section parse correctly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>