* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local>
347 lines
15 KiB
Bash
Executable File
347 lines
15 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# prompt-injection-scan.sh — Scan files for prompt injection patterns
|
|
#
|
|
# Usage:
|
|
# scripts/prompt-injection-scan.sh --diff origin/main # CI mode: scan changed .md files
|
|
# scripts/prompt-injection-scan.sh --file path/to/file # Scan a single file
|
|
# scripts/prompt-injection-scan.sh --dir agents/ # Scan all files in a directory
|
|
#
|
|
# Exit codes (ADR-3889, #3908 — registered in gsd-core/bin/shared/exit-codes.json):
|
|
# 0 = clean
|
|
# 1 = findings detected
|
|
# $EXIT_USAGE (64) = usage error (bad argv, missing --file/--dir target)
|
|
# $EXIT_NO_INPUT (66) = ran; scope established; zero files in scope (genuinely empty)
|
|
# $EXIT_UNAVAILABLE (69) = could not establish scope (bad ref, not a repo, unreadable dir)
|
|
set -euo pipefail
|
|
|
|
# ─── Exit-code registry (ADR-3889, #3908) ────────────────────────────────────
|
|
# Resolved relative to THIS script's location, not the caller's cwd. Loud,
|
|
# non-zero failure if the fragment is missing — never fall back to a guessed
|
|
# literal integer, and never let a missing registry silently degrade to
|
|
# exit 0.
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
EXIT_CODES_SH="$SCRIPT_DIR/../gsd-core/bin/shared/exit-codes.sh"
|
|
if [[ ! -f "$EXIT_CODES_SH" ]]; then
|
|
echo "prompt-injection-scan: FATAL: exit-code registry not found at $EXIT_CODES_SH" >&2
|
|
echo " Regenerate with: node scripts/gen-exit-code-registry.cjs --write" >&2
|
|
exit 1
|
|
fi
|
|
# shellcheck disable=SC1090
|
|
. "$EXIT_CODES_SH"
|
|
|
|
# ─── Patterns ────────────────────────────────────────────────────────────────
|
|
# Each pattern is a POSIX extended regex. Keep alphabetized by category.
|
|
#
|
|
# Left-boundary prefix `(^|[^[:alnum:]])`: several trigger words are also
|
|
# suffixes of ordinary English words or camelCase identifiers (fact/impact/
|
|
# contract/artifact/interact all end in "act"; retrieval/medieval end in
|
|
# "eval"; blueprint/reprint/fingerprint end in "print"; describeFunction/
|
|
# wrapFunction end in "Function"; Jordan/Sudan end in "dan"), so an
|
|
# unanchored keyword matches as a false-positive substring. `\b` is a GNU
|
|
# grep extension and this script must also run under BSD/macOS grep, so the
|
|
# boundary is spelled out as `(^|[^[:alnum:]])` instead. This never narrows
|
|
# real detections: a genuine attack phrase is always preceded by start-of-
|
|
# line, whitespace, or punctuation, never by another alnum character glued
|
|
# directly onto the keyword. Only patterns whose leading keyword is provably
|
|
# not a real-word suffix are left unanchored (#3175 audit).
|
|
|
|
PATTERNS=(
|
|
# Instruction override
|
|
'ignore[[:space:]]+(all[[:space:]]+)?(previous|prior|above|earlier|preceding)[[:space:]]+(instructions|prompts|rules|directives|context)'
|
|
'disregard[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules)'
|
|
'forget[[:space:]]+(all[[:space:]]+)?(previous|prior|above)[[:space:]]+(instructions|prompts|rules|context)'
|
|
'override[[:space:]]+(all[[:space:]]+)?(system|previous|safety)[[:space:]]+(instructions|prompts|rules|checks|filters|guards)'
|
|
'override[[:space:]]+(system|safety|security)[[:space:]]'
|
|
|
|
# Role manipulation
|
|
'you[[:space:]]+are[[:space:]]+now[[:space:]]+(a|an|my)[[:space:]]'
|
|
'from[[:space:]]+now[[:space:]]+on[[:space:]]+(you|pretend|act|behave)'
|
|
'pretend[[:space:]]+(you[[:space:]]+are|to[[:space:]]+be)[[:space:]]'
|
|
'(^|[^[:alnum:]])act[[:space:]]+as[[:space:]]+(a|an|if|my)[[:space:]]'
|
|
'roleplay[[:space:]]+as[[:space:]]'
|
|
'assume[[:space:]]+the[[:space:]]+role[[:space:]]+of[[:space:]]'
|
|
|
|
# System prompt extraction
|
|
'output[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'reveal[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'show[[:space:]]+me[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'(^|[^[:alnum:]])print[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'what[[:space:]]+(is|are)[[:space:]]+(your|the)[[:space:]]+(system[[:space:]]+)?(prompt|instructions)'
|
|
'repeat[[:space:]]+(your|the|all)[[:space:]]+(system[[:space:]]+)?(prompt|instructions|rules)'
|
|
|
|
# Fake message boundaries
|
|
'</?system>'
|
|
'</?assistant>'
|
|
'</?human>'
|
|
'\[SYSTEM\]'
|
|
'\[/SYSTEM\]'
|
|
'\[INST\]'
|
|
'\[/INST\]'
|
|
'<<SYS>>'
|
|
'<</SYS>>'
|
|
|
|
# Tool call injection / code execution in markdown
|
|
#
|
|
# The quote-or-apostrophe class below is spelled ["'"'"'"] (a literal `'`
|
|
# via bash's close-quote/escape/reopen idiom), not `["\x27]` — `\x27` is a
|
|
# GNU-grep-only hex escape; BSD/macOS grep treats it as four literal
|
|
# characters (", \, x, 2, 7) and never matches an actual apostrophe, so
|
|
# `eval('...')` (single-quoted) silently went undetected on macOS while
|
|
# passing on GNU-grep CI runners. Found auditing #3175; fixed here since it
|
|
# is the same unanchored/portability defect class as the boundary fix.
|
|
#
|
|
# `exec` stays receiver-blind on purpose. A left boundary that excludes a
|
|
# preceding `.` would drop every member-position `.exec('…')` — including
|
|
# `require('child_process').exec('…')`, the single most common Node spelling
|
|
# of the vector this pattern exists to catch — and a receiver allowlist
|
|
# cannot restore it, because the literal `child_process` is not adjacent to
|
|
# `.exec`. The cost is that `RegExp.prototype.exec`, which takes a subject
|
|
# string rather than code, also matches; files that legitimately call it are
|
|
# handled by ALLOWLIST below, never by narrowing the pattern.
|
|
'(^|[^[:alnum:]])eval[[:space:]]*\([[:space:]]*["'"'"']'
|
|
'exec[[:space:]]*\([[:space:]]*["'"'"']'
|
|
'(^|[^[:alnum:]])Function[[:space:]]*\([[:space:]]*["'"'"'].*return'
|
|
|
|
# Jailbreak / DAN patterns
|
|
'do[[:space:]]+anything[[:space:]]+now'
|
|
'(^|[^[:alnum:]])DAN[[:space:]]+mode'
|
|
'developer[[:space:]]+mode[[:space:]]+(enabled|output|activated)'
|
|
'jailbreak'
|
|
'bypass[[:space:]]+(safety|content|security)[[:space:]]+(filter|check|rule|guard)'
|
|
)
|
|
|
|
# ─── Allowlist ───────────────────────────────────────────────────────────────
|
|
# Files that legitimately discuss injection patterns (security docs, tests, this script)
|
|
ALLOWLIST=(
|
|
'scripts/prompt-injection-scan.sh'
|
|
'scripts/base64-scan.sh'
|
|
'scripts/secret-scan.sh'
|
|
'tests/security-scan.security.test.cjs'
|
|
'tests/security.test.cjs'
|
|
'tests/prompt-injection-scan.security.test.cjs'
|
|
'tests/verify.test.cjs'
|
|
'gsd-core/bin/lib/security.cjs'
|
|
'hooks/gsd-prompt-guard.js'
|
|
'hooks/gsd-read-injection-scanner.js'
|
|
'tests/read-injection-scanner.security.test.cjs'
|
|
'tests/read-injection-scanner.property.test.cjs'
|
|
'tests/security-prompt-injection.security.test.cjs'
|
|
'tests/list-seeds.test.cjs'
|
|
'tests/fixtures/adversarial/security/'
|
|
'SECURITY.md'
|
|
# These files contain intentional injection examples / security-model prose
|
|
# and are not attack vectors — they explain/demonstrate injection patterns.
|
|
'TEST-EXAMPLES.md'
|
|
'explanation/security-model.md'
|
|
# The untrusted-input boundary reference quotes injection phrases
|
|
# ("ignore previous instructions", "you are now…") as examples agents must
|
|
# NOT comply with — it is the defense, not an attack vector.
|
|
'references/untrusted-input-boundary.md'
|
|
# Security regression tests for input validators — fixtures must contain
|
|
# real injection payloads to prove the validator rejects them. See
|
|
# DEFECT.PROMPT-INJECTION-SCAN-COLLISION in CONTEXT.md.
|
|
'tests/windsurf-conversion.test.cjs'
|
|
# RuleTester fixtures for the local/no-unguarded-nonportable-exec ESLint rule
|
|
# contain shell-exec command strings (exec("sh -c …"), execFileSync('bash',['-c',…]))
|
|
# as test DATA the rule must lint — not attack vectors. ADR-1703 Phase 3 (#1720).
|
|
'tests/no-unguarded-nonportable-exec.rule.test.cjs'
|
|
# RuleTester fixtures for the local/no-bare-npm-exec ESLint rule contain npm
|
|
# exec command strings (execFileSync('npm', ['install'])) as test DATA the rule
|
|
# must lint — not attack vectors. ADR-1703 Phase 4 (#1726).
|
|
'tests/no-bare-npm-exec.rule.test.cjs'
|
|
# #2547 — the Kimi field-shadowing regression proves gsd-prompt-guard still
|
|
# SCANS the reconstructed edit[].new content when a model-supplied new_string
|
|
# tries to shadow it. The fixture must be a real injection phrase or the test
|
|
# asserts nothing: it is the payload the guard is required to catch, carried
|
|
# as test DATA. Same class as the read-injection-scanner suites above.
|
|
'tests/kimi-payload-field-shadowing.security.test.cjs'
|
|
# Phase-ID grammar regression tests exercise `RegExp.prototype.exec` via
|
|
# `re.exec('<phase-id>')` against fixtures like 'MANIFOLD-64-auth' / 'CK-64-auth'.
|
|
# The scanner's `exec('` code-execution pattern matches that benign method call,
|
|
# not an attack vector — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the
|
|
# test fixtures above. Pre-existing content (16 such calls on `next`); it surfaces
|
|
# here only because #2573's W024 `state_head` assertions make the file appear in
|
|
# the changed-file set the diff-mode scan walks.
|
|
'tests/health-validation.test.cjs'
|
|
# #2528 — same collision, same disposition: the continuation-grammar suite
|
|
# drives the tokenizer regexes directly via `re.exec('05-80-20')`, so the
|
|
# argument is the subject string, not a command. Exempted per file rather
|
|
# than by narrowing the `exec(` pattern: a left boundary excluding a preceding
|
|
# `.` would drop `require('child_process').exec('…')`, and a receiver
|
|
# allowlist cannot reach it either, because the literal `child_process` is
|
|
# not adjacent to `.exec`. See the note at the pattern itself.
|
|
'tests/continuation-grammar-parity.test.cjs'
|
|
# #3676 row 11b — quick-batch's task-list parser must treat a
|
|
# prompt-injection-shaped task description as inert data, never
|
|
# interpreted. The fixture has to be a real "ignore all previous
|
|
# instructions…" phrase or the test asserts nothing: it is the payload the
|
|
# parser is required to carry byte-for-byte through createBatch and STATE
|
|
# rendering, never a command. Same DEFECT.PROMPT-INJECTION-SCAN-COLLISION
|
|
# class as the input-validator fixtures above.
|
|
'tests/quick-batch.test.cjs'
|
|
)
|
|
|
|
is_allowlisted() {
|
|
local file="$1"
|
|
for allowed in "${ALLOWLIST[@]}"; do
|
|
if [[ "$file" == *"$allowed"* ]]; then
|
|
return 0
|
|
fi
|
|
done
|
|
return 1
|
|
}
|
|
|
|
# ─── File Collection ─────────────────────────────────────────────────────────
|
|
|
|
collect_files() {
|
|
local mode="$1"
|
|
shift
|
|
|
|
case "$mode" in
|
|
--diff)
|
|
local base="${1:-origin/main}"
|
|
# Run git separately from the filter pipe so its OWN exit status (not
|
|
# grep's) decides whether the diff could be established. stdout and
|
|
# stderr are captured SEPARATELY (never merged with `2>&1`) so that a
|
|
# warning git writes to stderr on an otherwise successful diff can
|
|
# never be mistaken for a filename in the file list. On failure the
|
|
# captured stderr is emitted as the diagnostic; on success it is
|
|
# forwarded as a warning, never folded into the file list.
|
|
# `|| true` on the filter below is CORRECT (not gratuitous): `grep -E`
|
|
# exits 1 when nothing matches the scannable extensions (e.g. a diff
|
|
# touching only non-scannable file types), which is a legitimate empty
|
|
# result, not a failure to run.
|
|
local raw status err_file
|
|
err_file=$(mktemp)
|
|
raw=$(git diff --name-only --diff-filter=ACMR "$base"...HEAD 2>"$err_file")
|
|
status=$?
|
|
if (( status != 0 )); then
|
|
cat "$err_file" >&2
|
|
rm -f "$err_file"
|
|
exit "$EXIT_UNAVAILABLE"
|
|
fi
|
|
if [[ -s "$err_file" ]]; then
|
|
echo "Warning: git diff emitted stderr output:" >&2
|
|
cat "$err_file" >&2
|
|
fi
|
|
rm -f "$err_file"
|
|
printf '%s\n' "$raw" | grep -E '\.(md|cjs|js|json|yml|yaml|sh)$' || true
|
|
;;
|
|
--file)
|
|
if [[ -f "$1" ]]; then
|
|
echo "$1"
|
|
else
|
|
echo "Error: file not found: $1" >&2
|
|
exit "$EXIT_USAGE"
|
|
fi
|
|
;;
|
|
--dir)
|
|
local dir="$1"
|
|
if [[ ! -d "$dir" ]]; then
|
|
echo "Error: directory not found: $dir" >&2
|
|
exit "$EXIT_USAGE"
|
|
fi
|
|
# Same treatment as --diff: a `find` that fails (e.g. permission
|
|
# denied) must not be reported as an empty directory, and stdout/stderr
|
|
# are captured separately so a stderr warning never enters the file list.
|
|
local raw status err_file
|
|
err_file=$(mktemp)
|
|
raw=$(find "$dir" -type f \( -name '*.md' -o -name '*.cjs' -o -name '*.js' -o -name '*.json' -o -name '*.yml' -o -name '*.yaml' -o -name '*.sh' \) \
|
|
! -path '*/node_modules/*' ! -path '*/.git/*' ! -path '*/dist/*' 2>"$err_file")
|
|
status=$?
|
|
if (( status != 0 )); then
|
|
cat "$err_file" >&2
|
|
rm -f "$err_file"
|
|
exit "$EXIT_UNAVAILABLE"
|
|
fi
|
|
if [[ -s "$err_file" ]]; then
|
|
echo "Warning: find emitted stderr output:" >&2
|
|
cat "$err_file" >&2
|
|
fi
|
|
rm -f "$err_file"
|
|
printf '%s\n' "$raw"
|
|
;;
|
|
--stdin)
|
|
cat
|
|
;;
|
|
*)
|
|
echo "Usage: $0 --diff [base] | --file <path> | --dir <path> | --stdin" >&2
|
|
exit "$EXIT_USAGE"
|
|
;;
|
|
esac
|
|
}
|
|
|
|
# ─── Scanner ─────────────────────────────────────────────────────────────────
|
|
|
|
scan_file() {
|
|
local file="$1"
|
|
local found=0
|
|
|
|
if is_allowlisted "$file"; then
|
|
return 0
|
|
fi
|
|
|
|
for pattern in "${PATTERNS[@]}"; do
|
|
# Use grep -iE for case-insensitive extended regex
|
|
# -n for line numbers, -c for count mode first to check
|
|
local matches
|
|
matches=$(grep -inE -e "$pattern" "$file" 2>/dev/null || true)
|
|
if [[ -n "$matches" ]]; then
|
|
if [[ $found -eq 0 ]]; then
|
|
echo "FAIL: $file"
|
|
found=1
|
|
fi
|
|
echo "$matches" | while IFS= read -r line; do
|
|
echo " $line"
|
|
done
|
|
fi
|
|
done
|
|
|
|
return $found
|
|
}
|
|
|
|
# ─── Main ────────────────────────────────────────────────────────────────────
|
|
|
|
main() {
|
|
if [[ $# -eq 0 ]]; then
|
|
echo "Usage: $0 --diff [base] | --file <path> | --dir <path>" >&2
|
|
exit "$EXIT_USAGE"
|
|
fi
|
|
|
|
local mode="$1"
|
|
shift
|
|
|
|
local files
|
|
files=$(collect_files "$mode" "$@")
|
|
|
|
if [[ -z "$files" ]]; then
|
|
# collect_files already exited (UNAVAILABLE/USAGE) for anything that
|
|
# could not establish scope. Reaching here with an empty result means
|
|
# scope WAS established and is genuinely empty — that is NO_INPUT, not a
|
|
# silent clean pass.
|
|
echo "prompt-injection-scan: no files to scan"
|
|
exit "$EXIT_NO_INPUT"
|
|
fi
|
|
|
|
local total=0
|
|
local failed=0
|
|
|
|
while IFS= read -r file; do
|
|
[[ -z "$file" ]] && continue
|
|
total=$((total + 1))
|
|
if ! scan_file "$file"; then
|
|
failed=$((failed + 1))
|
|
fi
|
|
done <<< "$files"
|
|
|
|
echo ""
|
|
echo "prompt-injection-scan: scanned $total files, $failed with findings"
|
|
|
|
if [[ $failed -gt 0 ]]; then
|
|
exit 1
|
|
fi
|
|
exit 0
|
|
}
|
|
|
|
main "$@"
|