* enhancement(#537): migrate code-review-flags to TS source of truth Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish). Second module after the semver-compare pilot (#541). Behaviour is preserved byte-for-behaviour (characterization test added in tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union. The require() path is unchanged, so code-review.md and the bug-3727 test keep working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps): 001-legacy-orphan-files, context-utilization, redaction, artifacts, command-arg-projection, clock, ui-safety-gate, review-reviewer-selection, clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a gitignored .cjs at the same require() path; behaviour preserved byte-for- behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): add @types/node, drop hand-rolled node-globals shim ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks migrating the ~49 remaining bin/lib modules that use node:fs/path/os/ child_process. Build + full suite (3030 pass) + lint all green; no .cts type changes were needed (real Node types matched the shim). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2) ADR-457 build-at-publish. Clean leaves: installer-migration-report, prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now strict-typed and removed from that exclude list): secrets, phase-lifecycle, workstream-name-policy, decisions, validate, schema-detect. Plus runtime-name-policy. Strict type fixes narrow unknown->concrete domain types (no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate runtime-slash to TS (cross-import proof) ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports ./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict (no declaration files; NodeNext .cjs->.cts mapping), emitting a correct require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules, which must be migrated in dependency order (leaves-up). Suite green, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3) ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder, plan-scan, fallow-runner, project-root, installer-migration-authoring, update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict typing fixed real issues (narrowing unknown, qualified fs/path calls, removed unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4) ADR-457 build-at-publish: configuration, state-document, shell-command- projection (42 dependents), security, command-aliases. shell-command- projection keeps a namespace child_process import for mock-intercept testability. loadConfig/migrateOnDisk emit synchronously (every caller uses them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) — full suite (3030 pass) confirms behaviour preserved. configuration/ state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5) ADR-457 build-at-publish: config-schema, model-profiles, 002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser. First batch importing already-migrated siblings (configuration, model-catalog, shell-command-projection, redaction, security) via ./sibling.cjs specifiers. Strict type narrowing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6) ADR-457 build-at-publish: graphify, install-profiles, intel, installer-migrations, worktree-safety. installer-migrations preserves its dynamic require() loader for numbered migration modules (scoped lint suppressions). Strict typing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Wave 2 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate Wave-3 modules to TS (batch 7) ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout, command-routing-hub, drift. Uses `import x = require()` for export= siblings; drift's lazy require of runtime-slash hoisted to a top-level import (verified non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate small Wave-4 modules to TS (batch 8) ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router, surface, roadmap-upgrade. Typed the hub router handler results as the HubResult discriminated union; surface drops 4 genuinely-unused imports. Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9) ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts, preserving all 63 exports via export=. All sibling deps already migrated (shell-command-projection, model-profiles, model-catalog, worktree-safety, planning-workspace, project-root, configuration, config-schema). Strict types, no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour preserved (independently verified: core's shard 3030 pass / 0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): make ESLint-coverage + test-sprawl checks migration-aware #551 test hardcoded 12 now-migrated modules as "hand-written, must be linted"; that invariant is obsoleted by the ADR-457 migration. Rewrite it to a filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted (covers package-identity, which has no TS source). Also eslint-ignore config-types.cjs (has a src counterpart) and drop the redundant tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests), which tripped the lint-test-file-count ratchet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10) ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state command routers + workstream-inventory. Router handler results typed against core's exported shapes; behaviour preserved (caught+fixed a --verify boolean flag regression mid-migration). Full suite green across all shards (only the 4 local gpg-env changeset-notes failures remain; CI passes them). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11) ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter, learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green across all shards (only the 4 local gpg-env failures remain). Also broadens atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form while still asserting platformWriteSync is called (safety guard intact). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate config + profile-output to TS (batch 12) ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited caller tolerates it). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 5 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13) ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour preserved (dead toPosixPath import dropped from audit; inline requires hoisted). Suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate commands + state hubs to TS (batch 14) ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents). All exports preserved; inner requires kept non-hoisted where load-order matters (install.js, per-call security); acquireStateLock cast inlined to preserve the err.code source token a structural test inspects. Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 6 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate milestone to TS (batch 15a, hand-authored) ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly (subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to optional, matching its real always-optional call contract (unblocks remaining 2-arg output callers). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules) ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify (1615), init (2113). Adds src/package-identity.d.cts so verify can import the permanently value-baked package-identity.cjs under strict TS. Fixes two regressions the migration introduced in verify: restore cmdValidateHealth's `return result` (callers/tests read result.warnings — it is NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour intact). Full suite green across all shards (only the 4 local gpg-env failures); lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is deleted per ADR-457's final step. Also gitignore the tsc-generated config-types.cjs (was still committed) for consistency with every other emitted artifact. package-identity.cjs stays value-baked (declared via src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g <dir>` and git installs run the `prepare` lifecycle — which was missing — so the unpacked install shipped without the compiled .cjs and failed at startup with "Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job). Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does NOT run for registry consumers (they get the pre-built tarball), only for source/local/pack installs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are always present on disk: - check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the `prepare` build (tsc) — but it runs before deps are installed, so tsc is absent and it misreported the lockfile as out of sync. Add --ignore-scripts (a lockfile check must not build). - the lint-tests job installs with --ignore-scripts (no prepare build), but lint:skill-deps require()s the built install-profiles.cjs. Add an explicit `npm run build:lib` step after install. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step) prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`, which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT — breaking it with "Invalid format". build:lib (tsc) is silent on success and is all the unpacked/source install needs (the smoke-unpacked assertions exercise gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches prepack. build:hooks still runs on prepublishOnly for real publishes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): wire Stryker mutation gate to build-at-publish layout The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were generated artifacts and (b) included modules with no coverage in the command's test set. Rework: mutation.yml now derives changed COVERED modules from src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc was ~3x over the 30-min CI budget). NOTE: with the gate now correctly measuring the covered modules, their actual mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap (adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver — a maintainer decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): raise mutation coverage of covered modules above the 50 gate Adds focused example-based unit tests that kill surviving mutants in the two lowest-scoring covered modules: - tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9% - tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4% Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered modules now scores 82.25% (>= break 50); every covered module is >= 68%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix The serial Stryker run timed out at 30 min once the migration's added tests made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the gate completes well under budget — folded into this PR (was tracked as #609) because it's a prerequisite for this PR's mutation gate to pass. - scripts/mutation-matrix.cjs: single source of truth (covered-module -> test files) computing changed covered modules from git diff -> {has_work, matrix}. - mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name "Stryker mutation score (changed files only)" so branch protection is unchanged. Per-shard jobs report as "Stryker (<module>)". - stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back to the full command locally). Closes #609. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note Per-module sharding revealed that active-workstream-store (46.5%) and frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from the bloated 300-test command; on their own tests they were below the gate. Add focused unit tests: - tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9% - tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4% Both wired into scripts/mutation-matrix.cjs (per-module test map) and stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50 with only their own tests (config-schema/context-utilization/prompt-budget/ adr-parser already did). Also removes the leftover blacksmith TODO comment — GitHub-hosted runners only; speed comes from parallel per-module shards. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the way the per-module CI shard runs it) — an earlier ~98% reading was inflated by accidentally running the full multi-module command. Add 96 targeted tests to tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans): scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear the gate on their own tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
448 lines
17 KiB
TypeScript
448 lines
17 KiB
TypeScript
/**
|
||
* Security — Input validation, path traversal prevention, and prompt injection guards
|
||
*
|
||
* This module centralizes security checks for GSD tooling. Because GSD generates
|
||
* markdown files that become LLM system prompts (agent instructions, workflow state,
|
||
* phase plans), any user-controlled text that flows into these files is a potential
|
||
* indirect prompt injection vector.
|
||
*
|
||
* Threat model:
|
||
* 1. Path traversal: user-supplied file paths escape the project directory
|
||
* 2. Prompt injection: malicious text in arguments/PRDs embeds LLM instructions
|
||
* 3. Shell metacharacter injection: user text interpreted by shell
|
||
* 4. JSON injection: malformed JSON crashes or corrupts state
|
||
* 5. Regex DoS: crafted input causes catastrophic backtracking
|
||
*
|
||
* ADR-457 build-at-publish: the hand-written bin/lib/security.cjs collapsed
|
||
* to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour
|
||
* from the prior hand-written .cjs; only types are added.
|
||
*/
|
||
|
||
import fs from 'node:fs';
|
||
import path from 'node:path';
|
||
|
||
// ─── Path Traversal Prevention ──────────────────────────────────────────────
|
||
|
||
/**
|
||
* Validate that a file path resolves within an allowed base directory.
|
||
* Prevents path traversal attacks via ../ sequences, symlinks, or absolute paths.
|
||
*/
|
||
export function validatePath(filePath: unknown, baseDir: unknown, opts: { allowAbsolute?: boolean } = {}): { safe: boolean; resolved: string; error?: string } {
|
||
if (!filePath || typeof filePath !== 'string') {
|
||
return { safe: false, resolved: '', error: 'Empty or invalid file path' };
|
||
}
|
||
if (!baseDir || typeof baseDir !== 'string') {
|
||
return { safe: false, resolved: '', error: 'Empty or invalid base directory' };
|
||
}
|
||
if (filePath.includes('\0')) {
|
||
return { safe: false, resolved: '', error: 'Path contains null bytes' };
|
||
}
|
||
let resolvedBase: string;
|
||
try {
|
||
resolvedBase = fs.realpathSync(path.resolve(baseDir));
|
||
} catch {
|
||
resolvedBase = path.resolve(baseDir);
|
||
}
|
||
let resolvedPath: string;
|
||
if (path.isAbsolute(filePath)) {
|
||
if (!opts.allowAbsolute) {
|
||
return { safe: false, resolved: '', error: 'Absolute paths not allowed' };
|
||
}
|
||
resolvedPath = path.resolve(filePath);
|
||
} else {
|
||
resolvedPath = path.resolve(baseDir, filePath);
|
||
}
|
||
try {
|
||
resolvedPath = fs.realpathSync(resolvedPath);
|
||
} catch {
|
||
const parentDir = path.dirname(resolvedPath);
|
||
try {
|
||
const realParent = fs.realpathSync(parentDir);
|
||
resolvedPath = path.join(realParent, path.basename(resolvedPath));
|
||
} catch {
|
||
// Parent doesn't exist either — keep the resolved path as-is
|
||
}
|
||
}
|
||
const normalizedBase = resolvedBase + path.sep;
|
||
const normalizedPath = resolvedPath + path.sep;
|
||
if (resolvedPath !== resolvedBase && !normalizedPath.startsWith(normalizedBase)) {
|
||
return {
|
||
safe: false,
|
||
resolved: resolvedPath,
|
||
error: `Path escapes allowed directory: ${resolvedPath} is outside ${resolvedBase}`,
|
||
};
|
||
}
|
||
return { safe: true, resolved: resolvedPath };
|
||
}
|
||
|
||
/**
|
||
* Validate a file path and throw on traversal attempt.
|
||
* Convenience wrapper around validatePath for use in CLI commands.
|
||
*/
|
||
export function requireSafePath(filePath: unknown, baseDir: unknown, label: string | null | undefined, opts: { allowAbsolute?: boolean } = {}): string {
|
||
const result = validatePath(filePath, baseDir, opts);
|
||
if (!result.safe) {
|
||
throw new Error(`${label || 'Path'} validation failed: ${result.error}`);
|
||
}
|
||
return result.resolved;
|
||
}
|
||
|
||
// ─── Prompt Injection Detection ────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Patterns that indicate prompt injection attempts in user-supplied text.
|
||
* These patterns catch common indirect prompt injection techniques where
|
||
* an attacker embeds LLM instructions in text that will be read by an agent.
|
||
*
|
||
* Note: This is defense-in-depth — not a complete solution. The primary defense
|
||
* is proper input/output boundaries in agent prompts.
|
||
*/
|
||
export const INJECTION_PATTERNS: RegExp[] = [
|
||
// Direct instruction override attempts
|
||
/ignore\s+(all\s+)?previous\s+instructions/i,
|
||
/ignore\s+(all\s+)?above\s+instructions/i,
|
||
/disregard\s+(all\s+)?previous/i,
|
||
/forget\s+(all\s+)?(your\s+)?instructions/i,
|
||
/override\s+(system|previous)\s+(prompt|instructions)/i,
|
||
|
||
// Role/identity manipulation
|
||
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
|
||
/act\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
|
||
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
|
||
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
|
||
|
||
// System prompt extraction
|
||
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
|
||
/what\s+(?:are|is)\s+your\s+(?:system\s+)?(?:prompt|instructions)/i,
|
||
|
||
// Hidden instruction markers (XML/HTML tags that mimic system messages)
|
||
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
|
||
// Requires > to close the tag (not just whitespace) to avoid matching generic types like Promise<User | null>
|
||
/<\/?(?:system|assistant|human)>/i,
|
||
/\[SYSTEM\]/i,
|
||
/\[\/?(INST)\]/i,
|
||
/<<\s*SYS\s*>>/i,
|
||
|
||
// Exfiltration attempts
|
||
/(?:send|post|fetch|curl|wget)\s+(?:to|from)\s+https?:\/\//i,
|
||
/(?:base64|btoa|encode)\s+(?:and\s+)?(?:send|exfiltrate|output)/i,
|
||
|
||
// Tool manipulation
|
||
/(?:run|execute|call|invoke)\s+(?:the\s+)?(?:bash|shell|exec|spawn)\s+(?:tool|command)/i,
|
||
];
|
||
|
||
// Explicit safe-list for data: MIME types that are benign in link targets.
|
||
// Note: image/svg+xml is intentionally NOT in this list (SVG can host <script>).
|
||
const DATA_URI_SAFE_MIME_RE = /^data:(image\/(png|jpe?g|gif|webp|bmp|ico|avif|heic)|font\/(woff2?|otf|ttf))(;[^,]*)?,/i;
|
||
|
||
interface MarkdownLinkPattern {
|
||
pattern: RegExp;
|
||
ruleId: string;
|
||
safePredicate?: (line: string) => boolean;
|
||
}
|
||
|
||
export const MARKDOWN_LINK_PATTERNS: MarkdownLinkPattern[] = [
|
||
{
|
||
pattern: /\]\(\s*javascript:/i,
|
||
ruleId: 'MD-LINK-JS-SCHEME',
|
||
},
|
||
{
|
||
pattern: /\]\(\s*data:/i,
|
||
ruleId: 'MD-LINK-DATA-SCHEME',
|
||
safePredicate: (line: string) => {
|
||
const m = line.match(/\]\(\s*(data:[^)]*)/i);
|
||
if (!m) return false;
|
||
return DATA_URI_SAFE_MIME_RE.test(m[1]);
|
||
},
|
||
},
|
||
{
|
||
pattern: /\]\(\s*https?:\/\/[^/\s]+:[^/@\s]+@/i,
|
||
ruleId: 'MD-LINK-USERINFO',
|
||
},
|
||
{
|
||
pattern: /[?&](token|access_token|id_token|refresh_token|api_key|apikey|secret|password|client_secret|code)=/i,
|
||
ruleId: 'MD-LINK-TOKEN-IN-QUERY',
|
||
},
|
||
];
|
||
|
||
interface ObfuscationPatternEntry {
|
||
pattern: RegExp;
|
||
message: string;
|
||
}
|
||
|
||
const OBFUSCATION_PATTERN_ENTRIES: ObfuscationPatternEntry[] = [
|
||
{
|
||
pattern: /\b(\w\s){4,}\w\b/,
|
||
message: 'Character-spacing obfuscation pattern detected (e.g. "i g n o r e")',
|
||
},
|
||
{
|
||
pattern: /<\/?(system|human|assistant|user)\s*>/i,
|
||
message: 'Delimiter injection pattern: <system>/<human>/<assistant>/<user> tag detected',
|
||
},
|
||
{
|
||
pattern: /0x[0-9a-fA-F]{16,}/,
|
||
message: 'Long hex sequence detected — possible encoded payload',
|
||
},
|
||
];
|
||
|
||
interface StructuredFinding {
|
||
ruleId: string;
|
||
file: string | undefined;
|
||
line: number;
|
||
match: string;
|
||
}
|
||
|
||
/**
|
||
* Scan text for potential prompt injection patterns.
|
||
* Returns an array of findings (empty = clean).
|
||
*/
|
||
export function scanForInjection(text: unknown, opts: { strict?: boolean; file?: string } = {}): { clean: boolean; findings: string[]; structuredFindings: StructuredFinding[] } {
|
||
if (!text || typeof text !== 'string') {
|
||
return { clean: true, findings: [], structuredFindings: [] };
|
||
}
|
||
|
||
const findings: string[] = [];
|
||
const structuredFindings: StructuredFinding[] = [];
|
||
|
||
for (const pattern of INJECTION_PATTERNS) {
|
||
if (pattern.test(text)) {
|
||
findings.push(`Matched injection pattern: ${pattern.source}`);
|
||
}
|
||
}
|
||
|
||
for (const entry of OBFUSCATION_PATTERN_ENTRIES) {
|
||
if (entry.pattern.test(text)) {
|
||
findings.push(entry.message);
|
||
}
|
||
}
|
||
|
||
const lines = text.split('\n');
|
||
for (const entry of MARKDOWN_LINK_PATTERNS) {
|
||
for (let i = 0; i < lines.length; i++) {
|
||
const line = lines[i];
|
||
const m = line.match(entry.pattern);
|
||
if (!m) continue;
|
||
if (entry.safePredicate && entry.safePredicate(line)) continue;
|
||
const matchText = m[0];
|
||
findings.push(`Matched markdown link pattern [${entry.ruleId}]: ${matchText}`);
|
||
structuredFindings.push({
|
||
ruleId: entry.ruleId,
|
||
file: opts.file,
|
||
line: i + 1,
|
||
match: matchText,
|
||
});
|
||
}
|
||
}
|
||
|
||
if (opts.strict) {
|
||
// Check for suspicious Unicode that could hide instructions
|
||
// (zero-width chars, RTL override, homoglyph attacks)
|
||
if (/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/.test(text)) {
|
||
findings.push('Contains suspicious zero-width or invisible Unicode characters');
|
||
}
|
||
|
||
// Layer 1: Unicode tag block U+E0000–E007F (2025 supply-chain attack vector)
|
||
// These characters are invisible and can embed hidden instructions
|
||
if (/[\uDB40\uDC00-\uDB40\uDC7F]/u.test(text) || /[\u{E0000}-\u{E007F}]/u.test(text)) {
|
||
findings.push('Contains Unicode tag block characters (U+E0000–E007F) — invisible instruction injection vector');
|
||
}
|
||
|
||
// Check for extremely long strings that could be prompt stuffing.
|
||
// Normalize CRLF → LF before measuring so Windows checkouts don't inflate the count.
|
||
const normalizedLength = text.replace(/\r\n/g, '\n').replace(/\r/g, '\n').length;
|
||
if (normalizedLength > 50000) {
|
||
findings.push(`Suspicious text length: ${normalizedLength} chars (potential prompt stuffing)`);
|
||
}
|
||
}
|
||
|
||
return { clean: findings.length === 0, findings, structuredFindings };
|
||
}
|
||
|
||
/**
|
||
* Sanitize text that will be embedded in agent prompts or planning documents.
|
||
* Strips known injection markers while preserving legitimate content.
|
||
*/
|
||
export function sanitizeForPrompt(text: unknown): string {
|
||
if (!text || typeof text !== 'string') return text as string;
|
||
|
||
let sanitized = text;
|
||
|
||
// Strip zero-width characters that could hide instructions
|
||
sanitized = sanitized.replace(/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/g, '');
|
||
|
||
// Neutralize XML/HTML tags that mimic system boundaries
|
||
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
|
||
sanitized = sanitized.replace(/<(\/?)\s*(?:system|assistant|human|user)\s*>/gi,
|
||
(_, slash: string) => `<${slash || ''}system-text>`);
|
||
|
||
// Neutralize [SYSTEM] / [INST] / [/INST] markers
|
||
sanitized = sanitized.replace(/\[(\/?)(SYSTEM|INST)\]/gi, (_, slash: string, tag: string) => `[${slash}${tag.toUpperCase()}-TEXT]`);
|
||
|
||
// Neutralize <<SYS>> and <</SYS>> markers (Llama-style delimiters)
|
||
sanitized = sanitized.replace(/<<\/?\s*SYS\s*>>/gi, '«SYS-TEXT»');
|
||
|
||
return sanitized;
|
||
}
|
||
|
||
/**
|
||
* Sanitize text that will be displayed back to the user.
|
||
* Removes protocol-like leak markers that should never surface in checkpoints.
|
||
*/
|
||
export function sanitizeForDisplay(text: unknown): string {
|
||
if (!text || typeof text !== 'string') return text as string;
|
||
|
||
let sanitized = sanitizeForPrompt(text);
|
||
|
||
const protocolLeakPatterns = [
|
||
/^\s*(?:assistant|user|system)\s+to=[^:\s]+:[^\n]+$/i,
|
||
/^\s*<\|(?:assistant|user|system)[^|]*\|>\s*$/i,
|
||
];
|
||
|
||
sanitized = sanitized
|
||
.split('\n')
|
||
.filter(line => !protocolLeakPatterns.some(pattern => pattern.test(line)))
|
||
.join('\n');
|
||
|
||
return sanitized;
|
||
}
|
||
|
||
// ─── Shell Safety ───────────────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Validate that a string is safe to use as a shell argument when quoted.
|
||
*/
|
||
export function validateShellArg(value: unknown, label: string | null | undefined): string {
|
||
if (!value || typeof value !== 'string') {
|
||
throw new Error(`${label || 'Argument'}: empty or invalid value`);
|
||
}
|
||
if (value.includes('\0')) {
|
||
throw new Error(`${label || 'Argument'}: contains null bytes`);
|
||
}
|
||
if (/[$`]/.test(value) && /\$\(|`/.test(value)) {
|
||
throw new Error(`${label || 'Argument'}: contains potential command substitution`);
|
||
}
|
||
return value;
|
||
}
|
||
|
||
// ─── JSON Safety ──────────────────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Safely parse JSON with error handling and optional size limits.
|
||
*/
|
||
export function safeJsonParse(text: unknown, opts: { maxLength?: number; label?: string } = {}): { ok: boolean; value?: unknown; error?: string } {
|
||
const maxLength = opts.maxLength || 1048576;
|
||
const label = opts.label || 'JSON';
|
||
if (!text || typeof text !== 'string') {
|
||
return { ok: false, error: `${label}: empty or invalid input` };
|
||
}
|
||
if (text.length > maxLength) {
|
||
return { ok: false, error: `${label}: input exceeds ${maxLength} byte limit (got ${text.length})` };
|
||
}
|
||
try {
|
||
const value = JSON.parse(text) as unknown;
|
||
return { ok: true, value };
|
||
} catch (err) {
|
||
const msg = err instanceof Error ? err.message : String(err);
|
||
return { ok: false, error: `${label}: parse error — ${msg}` };
|
||
}
|
||
}
|
||
|
||
// ─── Phase/Argument Validation ─────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Validate a phase number argument.
|
||
*/
|
||
export function validatePhaseNumber(phase: unknown): { valid: boolean; normalized?: string; error?: string } {
|
||
if (!phase || typeof phase !== 'string') {
|
||
return { valid: false, error: 'Phase number is required' };
|
||
}
|
||
const trimmed = phase.trim();
|
||
if (/^\d{1,4}[A-Z]?(?:\.\d{1,3})*$/i.test(trimmed)) {
|
||
return { valid: true, normalized: trimmed };
|
||
}
|
||
if (/^[A-Z][A-Z0-9]*(?:-[A-Z0-9]+){1,4}$/i.test(trimmed) && trimmed.length <= 30) {
|
||
return { valid: true, normalized: trimmed };
|
||
}
|
||
return { valid: false, error: `Invalid phase number format: "${trimmed}"` };
|
||
}
|
||
|
||
/**
|
||
* Validate a STATE.md field name to prevent injection into regex patterns.
|
||
*/
|
||
export function validateFieldName(field: unknown): { valid: boolean; error?: string } {
|
||
if (!field || typeof field !== 'string') {
|
||
return { valid: false, error: 'Field name is required' };
|
||
}
|
||
if (/^[A-Za-z][A-Za-z0-9 _.\-/]{0,60}$/.test(field)) {
|
||
return { valid: true };
|
||
}
|
||
return { valid: false, error: `Invalid field name: "${field}"` };
|
||
}
|
||
|
||
// ─── Layer 3: Structural Schema Validation ──────────────────────────────────────────────────────────────────────────
|
||
|
||
const KNOWN_VALID_TAGS = new Set([
|
||
'objective', 'process', 'step', 'success_criteria', 'critical_rules',
|
||
'available_agent_types', 'purpose', 'required_reading',
|
||
]);
|
||
|
||
/**
|
||
* Validate the XML structure of a prompt file.
|
||
*/
|
||
export function validatePromptStructure(text: unknown, fileType: string): { valid: boolean; violations: string[] } {
|
||
if (!text || typeof text !== 'string') {
|
||
return { valid: true, violations: [] };
|
||
}
|
||
if (fileType !== 'agent' && fileType !== 'workflow') {
|
||
return { valid: true, violations: [] };
|
||
}
|
||
const violations: string[] = [];
|
||
const tagRegex = /<([A-Za-z][A-Za-z0-9_-]*)/g;
|
||
let match: RegExpExecArray | null;
|
||
while ((match = tagRegex.exec(text)) !== null) {
|
||
const tag = match[1].toLowerCase();
|
||
if (!KNOWN_VALID_TAGS.has(tag)) {
|
||
violations.push(`Unknown XML tag in ${fileType} file: <${tag}>`);
|
||
}
|
||
}
|
||
return { valid: violations.length === 0, violations };
|
||
}
|
||
|
||
// ─── Layer 4: Paragraph-Level Entropy Anomaly Detection ─────────────────────────────────────────────────────────────────────
|
||
|
||
function shannonEntropy(text: string): number {
|
||
if (!text || text.length === 0) return 0;
|
||
const freq: Record<string, number> = {};
|
||
for (const ch of text) {
|
||
freq[ch] = (freq[ch] || 0) + 1;
|
||
}
|
||
const len = text.length;
|
||
let entropy = 0;
|
||
for (const count of Object.values(freq)) {
|
||
const p = count / len;
|
||
entropy -= p * Math.log2(p);
|
||
}
|
||
return entropy;
|
||
}
|
||
|
||
/**
|
||
* Scan text for paragraphs with anomalously high Shannon entropy.
|
||
*/
|
||
export function scanEntropyAnomalies(text: unknown): { clean: boolean; findings: string[] } {
|
||
if (!text || typeof text !== 'string') {
|
||
return { clean: true, findings: [] };
|
||
}
|
||
const findings: string[] = [];
|
||
const paragraphs = text.split(/\n\n+/);
|
||
for (const para of paragraphs) {
|
||
if (para.length <= 50) continue;
|
||
const entropy = shannonEntropy(para);
|
||
if (entropy > 5.5) {
|
||
findings.push(
|
||
`High-entropy paragraph detected (${entropy.toFixed(2)} bits/char) — possible encoded payload`
|
||
);
|
||
}
|
||
}
|
||
return { clean: findings.length === 0, findings };
|
||
}
|