Commit Graph

7 Commits

Author SHA1 Message Date
Tom Boucher
325fc25c01 fix(#3569): require a digit-bearing phase id in the stats heading scan (#3591)
* test(#3569): pin stats phase-id shape — inline-code mentions produce no phantom row

Failing-first regression for #3569: cmdStats' heading scan accepted any
word as a phase id, so prose mentioning ### Phase N: inside inline code
inflated phases_total and disagreed with roadmap analyze. New adversarial
fixture phase-heading-inside-inline-code.md (blockquote + bare mention),
parity assertion against roadmap analyze, and over-narrowing guards for
decimal / milestone-prefixed / letter-prefixed ids.

* fix(#3569): require a digit-bearing phase id in the stats heading scan

cmdStats' hand-rolled heading pattern captured any word as a phase id,
so a ### Phase N: token inside an inline code span (the issue's
blockquote) produced a phantom Not-Started row that could never
complete, inflating phases_total and deflating percent forever. The id
capture is now the canonical #3036 shape roadmap.cts uses (digit
required; letter-prefixed, decimal, and milestone-prefixed ids keep
counting), so stats and roadmap analyze agree.

* fix(#3569): sanction the stats id-shape literal; correct zero-padded expectation

Review findings: the phase-id drift guard requires the // phase-id-owner:
comment directly above the regex (same form as roadmap.cts); the
milestone-prefixed over-narrowing guard must expect normalizePhaseName's
zero-padded 02-01 form, not the raw 2-01 token.

* chore(#3569): add changeset fragment

* chore(#3569): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-17 11:39:01 -04:00
Tom Boucher
4a1ed2531f enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check

Specifies ADR-2313 D6 before the implementation exists, so the tests
bind to the contract rather than to whatever the code happens to do.

RED is established by construction, not by a remote run:
checkCodexModelPosture and POSTURE_REASON are absent from the compiled
lib today, so every row fails on the missing export. A remote checkpoint
here would prove only that the function is missing, which is already
known — so the run is deliberately deferred to the combined green
checkpoint rather than spent proving a tautology.

That makes the NEGATIVE PROOFS the rows that carry real signal. Every
positive row passes even for a naive implementation that greps
/model\s*=/ over the whole file. Six rows fail it: light-tier
service_tier/model_verbosity decoupling (#774), hand-added keys, a
commented pin, the model_verbosity key-prefix collision, the runtime
no-op ordering, and the headline case — a literal `model = "sonnet"`
inside the developer_instructions ''' block, which the emitter fills
with agent prompts that discuss models constantly.

Row 14's fixture was verified to discriminate before being written: a
whole-file scan matches it and a header-slice scan does not. Without
that check the test would pass trivially and prove nothing, which is the
vacuous-test failure this epic has already hit repeatedly.

Adversarial TOML fixtures are hand-authored against the real Codex shape
rather than generated by generateCodexAgentToml, per #2371 — a fixture
from the writer can only confirm what the writer already believed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3242): validate codex .toml model posture, not just presence

Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export,
not a branch inside checkAgentsInstalled — that function carries 33
upstream dependents, cyclomatic 25, and sits in two traced process
flows, so it is deliberately left untouched.

It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf.
That is what Phase 1's constant move bought: agent-install-check is
documented as pure read/verify and imports only leaves, so reaching the
rule through model-resolver would have dragged config-loader into it.

Reads liberally, judges strictly, and never guesses. Tolerates comments,
key order, whitespace, CRLF, and a BOM; anchors on full key names so
model_verbosity does not satisfy a `model` probe; treats extra
hand-added keys as none of its business, since the check is a predicate
on the two fields the posture owns rather than a whitelist over the
document. An unreadable file becomes a named violation and the loop
keeps going.

The scan covers only the header slice — the lines before the
developer_instructions ''' marker. The emitter writes agent prompts into
that block and GSD's prompts discuss models constantly, so a whole-file
scan reports violations for prose. This is the highest-risk defect in
the phase and the reason its fixture was verified to discriminate before
being written.

The non-codex short-circuit runs before any filesystem call, so a stray
.toml under another runtime is never inspected.

Wired through cmdValidateAgents as an additive codex_posture key, so a
violating install is visible from a command a user actually runs rather
than only from a library nothing calls.

Also fixes a test defect found while implementing: .gitattributes forces
`* text=auto eol=lf` repo-wide, so the committed CRLF fixture was
normalized to LF in the index — `git ls-files --eol` reported `i/lf
w/crlf`, the working copy being stale pre-normalization bytes. The CRLF
row was asserting against a file that could not survive a fresh clone.
CRLF is now derived at runtime, which puts it under the test's control
rather than git's, instead of adding a .gitattributes exception that
fights a deliberate repo-wide policy and that anyone could re-normalize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3242): document the codex posture check where a user will look

Three quadrants, filed by where the reader actually arrives.

How-to (recover-and-troubleshoot.md, under Install and update problems)
is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400
about an unsupported model" — and opens with the verbatim error string.
Someone hitting this does not know the words "posture" or "ADR-2313";
they have a 400 in their terminal and will search for that.

Reference (COMMANDS.md) had no `validate agents` entry at all, though
sibling gsd-tools subcommands are documented. Adding user-visible output
to an undocumented command and then linking to it from the new how-to
would have left a dangling reference. The entry carries the
violation-reason table, since the frozen POSTURE_REASON enum is the
machine-readable contract a reader needs rather than the prose.

Both surfaces state that presence and posture are separate verdicts — a
missing agent lands in `missing`, never as a posture violation. That is
a deliberate design decision and would otherwise be invisible to someone
watching one command emit both.

Explanation stays in ADR-2313, which already covers D6 and the
liberal-parse/strict-judge boundary. Pointing at it beats duplicating it
into COMMANDS.md and creating two copies to drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): close two false negatives in the posture scan

Both found by an isolated reviewer and reproduced before fixing. Both
made the check report clean when it was not — the worst direction for
this function, since the how-to tells users an empty violations list
means the install is posture-clean.

A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML,
but the key pattern required a bare identifier, so the pin was silently
invisible. Bare, "double" and 'single' quoted forms now normalize to the
same key name.

The block marker was found by unanchored whole-content search and used
to truncate the header. A `description` value merely containing the
literal text `developer_instructions = '''` truncated the scan before a
real pin, and a user who hand-reordered `model` to sit after the block —
still legal TOML — was never scanned at all.

Fixed by changing the strategy rather than the regex: find the block's
range, anchored at line start, and scan every line OUTSIDE it. That
covers both failures and is strictly more correct than truncation, while
still never reading prompt prose. An unterminated block excludes the
rest of the file, which fails toward a false positive — the safe
direction, since misreading prose as a pin wastes a user's time while
the alternative hides a real one.

Also corrects two overclaims of mine. The how-to named "v1.11", a
version that does not exist — package.json is 1.10.0 and unreleased — so
it now describes the boundary by behavior and links the ADR. And the
test matrix asserted that a naive whole-file scan "fails exactly rows
12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16
produce the correct result against that baseline too. They guard real
but *different* mistakes, and the matrix now says which one each catches
instead of attributing them all to the header-slice defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3242): skip symlinked agent files instead of following them

Security review finding. The scan listed entries with readdirSync and
read them with readFileSync, which follows symlinks — so a symlink in
the agents directory pointing anywhere would have its contents read, and
any line matching the model pattern echoed into cmdValidateAgents'
output through the `value` field. A read-and-echo primitive on an
arbitrary path.

It needs write access to the agents directory, so it crosses no new
trust boundary today. Fixed anyway, for two reasons.

This repo already does it correctly next door: cmdEffortSync filters
with lstatSync().isFile() and the comment "Skip symlinks — only write
regular files to avoid clobbering symlink targets." Being inconsistent
with a sibling in the same subsystem IS the defect.

And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these
files. Establishing symlink-following as the house pattern for Codex
.toml handling here would hand Phase 3 a worse starting point while it
writes rather than reads.

Skipped silently rather than reported, matching the sibling: a symlinked
agent file is a structural install choice, which checkAgentsInstalled
owns, not a model-content posture defect. An lstat that itself throws
excludes the file rather than crashing the scan.

That does narrow the guarantee slightly, so the how-to now says an empty
list means every REGULAR .toml is clean, and tells anyone symlinking
their configs to check the targets by hand. Claiming a clean bill of
health over files the check declined to open would be the same kind of
false confidence the two false negatives above produced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3242): backfill changeset pr number (#3290)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 21:50:49 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
e32a53b974 feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (#133)
* test(113): add per-rule failing tests + hostile fixture for markdown link payloads

RED phase for issue #113 — scanForInjection() currently returns { clean: true }
for markdown links containing javascript:, data:text/html, userinfo credentials,
and token-in-query payloads.

Changes:
- tests/fixtures/adversarial/security/context-malicious-markdown-link.md:
  Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME,
  MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign
  negative controls (data:image/png, mailto:, https://github.com, port-only URL).
- tests/security-prompt-injection.test.cjs:
  - Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion
    to "malicious-markdown-link fixture is flagged by scanner" (forward-looking).
  - Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings
    with ruleId, file, line, match fields.
  - Added parity guard: every MARKDOWN_LINK_PATTERNS source string from
    security.cjs must appear in gsd-read-injection-scanner.js hook source.

D3 false-positive grep: 0 legitimate matches — no allowlist entries needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook)

GREEN phase for issue #113.

Rule details (all with primary source citations):

  MD-LINK-JS-SCHEME
    Flags ](javascript:...) regardless of case.
    Source: OWASP XSS Prevention Cheat Sheet
    https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html

  MD-LINK-DATA-SCHEME
    Flags data: URIs NOT in the explicit safe-list.
    Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf).
    data:image/svg+xml is intentionally BLOCKED — SVG can host <script>.
    Source: OWASP File Upload Cheat Sheet — SVG Files
    https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files

  MD-LINK-USERINFO
    Flags https?://user:pass@host in markdown link targets.
    Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo).
    Source: RFC 3986 §3.2.1 (userinfo syntax)
    https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
    RFC 9110 §4.2.4 (HTTP deprecates userinfo)
    https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4

  MD-LINK-TOKEN-IN-QUERY
    Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey,
    secret, password, client_secret, code — regardless of value.
    Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1
    https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
    D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed.

Architecture:
- scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection()
  extended with structuredFindings (ruleId, file, line, match) via opts.file.
- hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence
  (same pattern sources, verified by parity test).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(113): flip PINNED malicious-markdown-link assertion and add parity guard

REFACTOR phase — tightening test rigor after test-rigor skill review:

1. Fixture assertion now enumerates all 4 expected ruleIds explicitly:
   [MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY].
   Previously findings.length > 0 would pass even if 3 of 4 rules were broken.

2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for
   1-based line numbers) instead of `typeof f.line === 'number'` (vacuous).

3. match field assertions tightened to check the hostile content is present:
   - MD-LINK-JS-SCHEME: /javascript:/i in match
   - MD-LINK-DATA-SCHEME: /data:/i in match
   - MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker)
   - MD-LINK-TOKEN-IN-QUERY: /token=/i in match

4. Parity test checks actual RegExp .source strings (not just lengths), verifying
   the hook contains the exact canonical pattern sources character-for-character.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility

1. .changeset/113-malicious-markdown-links.md — required Security fragment
   for the user-facing markdown-link scanner changes in this PR (changeset-lint
   was failing with FAIL_MISSING_FRAGMENT).

2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR
   set for mutation state subcommands whose CJS contract is always exit-0.
   On Windows/Node 24 the SDK bridge returns result.ok===false for validation
   failures (e.g. state record-metric --phase 1 with no --plan/--duration),
   causing dispatchViaSdk() to call error() (exit 1) instead of output({error})
   (exit 0). The fix maps SDK non-ok results to JSON output for the affected
   mutation commands (record-metric, advance-plan, record-session, add-decision,
   add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS
   contract on all platforms.

tests/state.test.cjs:1161 "returns error when required fields missing" passes
locally (104/104 pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 10:15:03 -04:00
Tom Boucher
835dd6ab44 test(3596): adversarial security/prompt-injection abuse suite (#3654)
* test(3596): adversarial security/prompt-injection abuse suite

Adds `tests/security-prompt-injection.test.cjs` and a fixtures
directory at `tests/fixtures/adversarial/security/` covering the
attack classes enumerated in #3596:

  - Command substitution / backticks / heredoc payloads in workstream
    names — sentinel-file probes prove no shell is spawned, slugifier
    neutralises the input.
  - Path traversal through `--ws` and slash-bearing workstream names —
    rejected with structured `--json-errors` payload, no stack trace,
    no filesystem mutation outside the project root.
  - Fake `<system>` / `[SYSTEM]` / `<<SYS>>` / `[INST]` boundary tags —
    sanitizeForPrompt neutralises every form; structural negative
    property locked across all six styles in one place.
  - Zero-width / bidi-override codepoints — stripped per the documented
    codepoint set; asserted via codePoint inspection, not regex
    literals.
  - Hostile read of CONTEXT.md / PLAN.md / ROADMAP.md fixtures —
    `gsd-read-injection-scanner.js` surfaces the advisory; excluded
    paths and non-Read tools stay silent; malformed JSON does not
    crash the hook.
  - Hostile write of `.planning/` files — `gsd-prompt-guard.js` emits
    a `PreToolUse` advisory; non-Write/Edit tools stay silent.
  - Fake `ghp_*` / `sk-*` env tokens — never echoed in CLI stdout or
    stderr under hostile inputs; covered under
    `// allow-test-rule: structural-regression-guard` because the only
    way to assert byte-level absence is `.includes(token)` against the
    captured streams.
  - `validatePath`, `validateShellArg`, `validatePhaseNumber`,
    `validateFieldName` — focused negative-input contract pins.

Pinned behavior gaps (documented, NOT fixed in this PR):

  - `<instructions>` is intentionally whitelisted by both the scanner
    and the sanitiser (GSD's own prompt scaffolding). Two REGRESSION
    GUARD tests lock that contract.
  - The current `scanForInjection` does NOT flag malicious markdown
    links (javascript:/data:/embedded-credentials URLs). PINNED with
    negative-proof so any future scope extension fails the assertion
    and forces a deliberate update to the acceptance map.
  - `prompt-builder.ts` does not yet wrap plan/context markdown in an
    "untrusted data" envelope. That seam lives on the TS side and is
    covered by `sdk/src/prompt-builder.test.ts`; out of scope for a
    CJS test file. Mentioned in the file header.

Verification:

  - `node --test tests/security-prompt-injection.test.cjs` → 73 tests
    pass.
  - `node scripts/lint-no-source-grep.cjs` → 0 violations across
    546 test files (one `allow-test-rule: structural-regression-guard`
    annotation on this file for the token-absence assertions).
  - `node scripts/run-tests.cjs` → 9730 tests pass, 0 fail.

Refs #3596

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(3596): allow adversarial fixtures in scan + harden graphify status parse

* fix(3596): skip adversarial security fixtures in secret scan

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 00:34:53 -04:00
Tom Boucher
b1317633db test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite (#3633)
* test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite

Lands the adversarial parser-input corpus that CONTRIBUTING.md
§"QA Matrix Requirements / Parser and project-file inputs" and
TEST-EXAMPLES.md §"Parser Adversarial Fixtures" describe.

New tests/fixtures/adversarial/ layout:

  frontmatter/
    duplicate-keys.md            — same key twice (collapses last-wins)
    crlf-mixed.md                — CRLF endings throughout
    unclosed-block.md            — `---` open with no close
    unicode-keys-and-values.md   — non-ASCII + emoji + Greek
    null-byte-value.md           — U+0000 in a value
    huge-bounded.md              — 2000-item array, ~30KB

  roadmap/
    phase-heading-inside-fenced-code.md   — #2787 fence shadowing
    nested-fenced-code.md                 — outer + inner ``` blocks
    unicode-phase-titles.md               — JP / Greek / emoji titles
    repeated-phase-ids.md                 — phase 1 declared twice
    decimal-phase-mixed.md                — 2 vs 2.1 vs 2.10 vs 21
    markdown-headings-inside-html-comment.md — comment shadowing

Test files (all node:test, no try/finally in test bodies, no source-grep,
no raw-text matching on stdout/file content):

  tests/feat-3594-parser-adversarial-frontmatter.test.cjs (12 tests)
    Loads each fixture, pins parser invariants on extractFrontmatter()
    return shape. Cross-corpus "does not throw on any fixture" sweep.

  tests/feat-3594-parser-adversarial-roadmap.test.cjs (18 tests)
    Loads each fixture into a temp project's .planning/ROADMAP.md and
    drives `gsd-tools roadmap get-phase <N>` via the runCli harness
    introduced by #3593. Asserts on the typed JSON payload.

  tests/feat-3594-parser-property-style.test.cjs (2 tests)
    Deterministic mulberry32 PRNG generates 500 malformed-ish
    frontmatter inputs per test. Pins (a) extractFrontmatter is total
    over the corpus (no null-deref TypeError, always returns a plain
    object on success), (b) the suite completes well under 2 seconds
    (quadratic-regression guard).

Known-open bugs surfaced and pinned (intentionally NOT fixed in this
PR — separate issues warranted):

  - CJS roadmap parser matches `## Phase N:` headings inside fenced
    code blocks (the SDK parser tracks fences per the #2787 comment in
    sdk/src/query/roadmap.ts but the CJS path has not caught up).
  - CJS roadmap parser matches `## Phase N:` headings inside HTML
    comments.

Both are documented in-test with the "currently STILL matches it
(open: needs <fix>)" naming pattern so the day the production fix
lands, flipping the assertion from `found: true` to `found: false` is
the regression guard.

Test totals:
  - 32 new feat-3594-* tests (12 frontmatter + 18 roadmap + 2 property)
  - 108/108 pass when running together with the pre-existing
    frontmatter.test.cjs + roadmap.test.cjs suites (76 of theirs).

Closes #3594

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(3594): use Fisher-Yates shuffle for deterministic seeded inputs

Replaces `arr.sort(() => rng() - 0.5)` with a Fisher-Yates shuffle
driven by the supplied PRNG. The sort-based shuffle is non-transitive:
V8's TimSort behavior on non-transitive comparators is engine-defined,
so the same seed produced different orderings across Node versions —
undermining the test's stated reproducibility guarantee.

Fisher-Yates is O(n), transitive (no comparator at all), and consumes
exactly n-1 RNG values in a fixed order. The mulberry32 seed now
determines the input sequence end-to-end.

Codex review on PR #3633.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 00:23:27 -04:00