Files
msd-core/src/security.cts
Tom Boucher c1885df9e5 chore(#2143): prohibition-with-teeth + migrate remaining ad-hoc table sites — Phase 4 (final) (#2253)
* chore(#2143): prohibition-with-teeth + migrate remaining table sites — Phase 4

Phase 4 of epic #2143 (ADR-2143 §7). Completes the markdown table/mutation
consolidation by (a) giving the ad-hoc-parsing prohibition teeth and (b)
migrating the last ad-hoc table sites onto the shared seam.

- src/markdown-table.cts: new formatting-preserving `updateTableCell` primitive
  (self-contained, ragged-row-tolerant header/delimiter/cell-range scan; splices
  only the target cell's raw span, preserving all other bytes incl. padding/CRLF;
  no-op-preserves-padding when a transformer returns the current value). Exports
  splitTableRow/isDelimiterRow/findTableStartOffset for tolerant reuse.
- eslint-rules/no-adhoc-markdown-parsing.cjs: TABLE-REGEX detector extended to
  `new RegExp(<literal|static-template>)`; new `.replace()`-mutation detector for
  roadmap/state/content receivers with a table/section-shaped pattern.
- scripts/lint-table-schema-drift.cjs (wired into lint:ci): fails if a TABLE_SCHEMA
  header drifts from its authored table; tests import its logic (single source).
- Migrated onto the seam (behaviour-preserving vs pre-Phase-4 HEAD, verified
  byte-diff old-vs-new): roadmap.cts cmdRoadmapUpdatePlanProgress, phase.cts
  cmdPhaseComplete + traceability, milestone.cts cmdRequirementsMarkComplete,
  uat.cts read path, state.cts metrics/decisions/By-Phase.
- Incidental correctness gains from the migration: a decoy table can no longer
  swallow a phase-progress update (## Progress scoping); a ragged neighbouring
  row no longer silently aborts an edit; completing integer phase N no longer
  touches a decimal sub-phase N.x row; record-metric no longer drops trailing
  section content or duplicates the ## Performance Metrics section.
- Kept justified allow-adhoc-markdown markers only where genuinely not a table
  (security.cts <|role|> token) or a loose non-GFM section (uat human-verify).

Two orthogonal isolated reviews (correctness/adversarial + security) passed;
correctness found 4 behaviour regressions in the first migration pass, all fixed
and re-verified byte-identical-or-better vs OLD.

Surfaced for maintainer (pre-existing, ambiguous domain logic, NOT changed here):
templates/state.md places a By-Phase table under ## Performance Metrics while
cmdStateRecordMetric assumes a Plan|Duration|Tasks|Files table.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): match traceability row by first-cell value, not Requirement header

Phase 4's migration matched the REQUIREMENTS.md traceability row by a column
literally named `Requirement` (`row['Requirement']`), but real tables head that
column `REQ-ID`. The by-name lookup found nothing, so `phase complete` and
`requirements mark-complete` left the Status cell `Pending` (regressed #2769 /
#2203, caught by gsd-test — 8 failures, both node 22/24).

- src/phase.cts, src/milestone.cts: match the row by its FIRST cell's value
  (the requirement-ID column) regardless of that column's HEADER name, via
  `Object.values(row)[0]` (updateTableCell builds the record in header order).
  This mirrors OLD's first-cell `\|\s*<id>\s*\|` anchor, restoring header-name
  independence while keeping the seam.
- src/milestone.cts hasTable: broadened from `Requirement`-only to also
  recognize `Requirement ID` / `REQ-ID` / `REQ ID` headers, kept in sync with
  the now-positional rowMatch/hasRow so a REQ-ID-headed table participates in
  the ADR-2143 §6 write-set and the #2140 table_unmatched drift check (it was
  silently omitted before — a checkbox-only partial reconcile against a REQ-ID
  table could report as fully reconciled). The `Requirement`-headed path is
  byte-identical to OLD.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2143): replace stale structural milestone guards with behavioural suite

The `milestone.cjs regex global state fix` block was a source-structure guard
(allow-test-rule: structural-regression-guard) — it readFileSync'd the compiled
milestone.cjs and asserted removed regex idioms (`tablePattern.test`,
`afterTable !== reqContent`, `doneTable = new RegExp(...)`). Phase 4's migration
deleted those regexes (table update is now updateTableCell), making the
assertions obsolete. Per the Test Cleanup rule, replace them in-PR with a
behavioural suite driving the compiled CLI:

- multi-ID mark-complete flips all IDs (guards the lastIndex/global-state class),
- Pending->Complete flip under both `REQ-ID` and `Requirement` headers (#2769),
- idempotent already_complete detection with no corruption,
- REQ-ID-headed table participates in write_set (traceability entry, applied),
- REQ-ID-headed table trips #2140 table_unmatched drift on a missing row.

Pruned the now-nonexistent structural-regression-guard entry from the
lint-allow-test-rule-refs allowlist (the source-text-is-the-product entry for
the same file remains valid).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): backfill PR number 2253

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): record-metric targets its own metrics table, not By-Phase velocity

`state record-metric` appended its per-plan row (`| Phase 1 P1 | 5min | 3 tasks |
4 files |`) into the FIRST table under `## Performance Metrics` — which on a real
template-derived STATE.md is the By-Phase velocity table `| Phase | Plans | Total
| Avg/Plan |`, polluting it on EVERY plan completion (execute-plan.md:414 is a
per-plan call). The command's own metrics table is `| Plan | Duration | Tasks |
Files |`, which the template does not ship, so the row never reached it; the
scaffold branch also emitted a wrong `| Phase | Plan | Duration | Notes |` header
matching neither the row nor the canonical table.

Pre-existing (predates Phase 4); surfaced while migrating this site and fixed here
per no-defer, on the user's explicit go-ahead.

- src/state.cts cmdStateRecordMetric: locate the metrics table by its own header
  shape (`Plan|Duration|Tasks|Files`, via splitTableRow/isDelimiterRow) rather
  than "first table in the section". When the section exists but has no metrics
  table (only the By-Phase table), self-heal by appending a fresh **Per-Plan
  Metrics:** table to the END of the section body — By-Phase table, Recent Trend
  and footer preserved verbatim, no duplicate `## Performance Metrics` heading,
  created stays false. Absent-section scaffold header corrected to the canonical
  `| Plan | Duration | Tasks | Files |`. Ragged-tolerance + None-yet preserved.
- Not touching templates/state.md (golden-install-parity hashed) — record-metric
  self-creates the table on first use instead.

Failing-first regression test (tests/state.test.cjs) demonstrates the By-Phase
pollution on the pre-fix build, then green after. Verified: no pollution, self-
heal idempotency, both-tables isolation, content/heading preservation, flags,
None-yet, corrected scaffold header (23-check adversarial harness + all existing
record-metric scenarios).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2143): deleteSection seam primitive (level-bounded whole-section removal)

ADR-2143 §4 shipped withSection/collectSection (replace a section BODY) but no
way to DELETE a section (heading + body). Phase 4 suppressed the phase-remove
section delete instead of building it. deleteSection(content, predicate, opts)
locates the section via the collectSection machinery and splices out from the
heading's start offset to the next same-or-higher-level heading — so a level-3
`### Phase N` delete stops at a following level-2 `## Progress`, never past it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): phase remove no longer deletes ## Progress on last-phase removal

updateRoadmapAfterPhaseRemoval deleted a `### Phase N` detail section with a
greedy raw regex whose lazy scan, on the LAST phase, ran to EOF and destroyed
the following `## Progress` heading and its entire tracking table — silent data
loss, uncovered by tests (removal tests only exercised a middle phase). Migrated
onto the new deleteSection seam (level-bounded, stops at `## Progress`); dropped
the allow-adhoc-markdown SECTION-DELETION suppression. Failing-first regression
(tests/phase.test.cjs) removes the LAST phase and asserts the ## Progress heading
+ table survive; middle-phase removal is byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2143): deleteTableRow seam primitive (row removal, ragged-tolerant)

Sibling of updateTableCell: locates the first GFM table, matches a DATA row by
predicate (ragged-tolerant record build, header order), and splices out that
row's whole line preserving every other byte. Returns {ok:false,reason} on no
table / no match. Enables migrating the phase-remove Progress-table row delete
off its ad-hoc regex (ADR-2143 §7 — the "future row-delete seam" Phase 4 punted).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): phase remove deletes the Progress row via deleteTableRow

The Progress-table row delete used a whole-document regex with two defects:
(a) `\.?\s` required whitespace after the phase number, so a COMPACT row
`|2|Beta|` was never deleted (stale row left behind); (b) unscoped — it could
strike a row in a different table (e.g. an earlier `| Phase | Requirements |`
table). Migrated onto deleteTableRow, scoped to the `## Progress` section
(mirrors deriveProgressFromRoadmap), matching the row by first-cell phase number
(integer zero-pad-insensitive; decimal exact; removing `2` never touches `2.5`).
Both allow-adhoc-markdown suppressions removed. New behavioural tests: compact
unpadded row deleted; padded byte-parity on the surviving rows (their ordinal
correctly renumbers via the pre-existing renumber block).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): deleteTableRow leaves no dangling newline on last EOL-less row

Deleting the final row of a table with no trailing EOL sliced from the row's
start to end-of-string, stranding the newline that terminated the previous line.
Back rowStart over the preceding \r?\n in that branch so the table ends cleanly.
(Caught by the primitive's own unit test on gsd-test; local scenario checks
missed the no-trailing-EOL edge.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): migrate read-only section-collects onto collectSection

Six hand-rolled `## Section` read-extract regexes replaced by the collectSection
seam (behaviour-preserving; extracted bodies feed the same downstream parsers):
state.cts matchSessionSection (## Session / ## Session Continuity) + ## Blockers,
smart-entry.cts ## Blockers, audit.cts ## Current Focus + ## Open Questions.
Removes 6 allow-adhoc-markdown "pending #1372" suppressions. Incidental fix: the
old Session regex `## Session[ \t]*\n` silently failed on a CRLF `## Session\r\n`
heading (Windows STATE.md), nulling all session fields; collectSection is
CRLF-safe, so session state now resolves on Windows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): fence-safe state-transition section writes + dedup stripFrontmatter

- milestoneCompleteCore's `## Current Position` and `## Operator Next Steps`
  section resets used fence-blind raw regexes that a fenced `##` inside the body
  could truncate/mis-target (#2130/#2067/#2080 class). Migrated onto a
  fence-aware tokenizeHeadings-based helper (resetSectionVerbatim) that is
  byte-identical to the old output on the canonical path (9/9 fixtures) and
  correctly ignores a fenced fake heading (proven robustness gain).
- mutateCurrentPositionFirstTime: hand-rolled locate+splice → collectSection +
  replaceSection (byte-parity).
- stripFrontmatter was inlined byte-identically in state.cts AND
  state-transition.cts; hoisted the single canonical copy into frontmatter.cts
  (both call sites now import it) + unit tests — eliminates the divergence risk
  per CLAUDE.md "Generative Fix Divergence". Removes 3 allow-adhoc-markdown /
  #1372 markers.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): name-address By-Phase sum + uat parse, eslint recall hole, catches

- state.cts By-Phase "Total plans completed" sum: positional 2nd-cell regex →
  name-addressed splitTableRow read (correct on a reordered header, where the
  old code silently summed the wrong column). Marker removed.
- uat.cts parseVerificationItems: loose pipe regex → splitTableRow within the
  existing table/numbered/bullet union scan (item list byte-identical; does NOT
  reintroduce the reverted strict-parseMarkdownTable item-drop). Marker removed.
- eslint no-adhoc-markdown-parsing: close the `new RegExp(identifier)` recall
  hole — resolve a const-declared table-shaped regex identifier (mirrors the
  .replace() detector) + RuleTester cases; param/call args stay out (boundary).
- commands.cts: delete a lying comment that claimed the scaffold date "stays on
  raw UTC / deferred" — #2136 already moved it to realClock.localToday().
- Empty catches (classified, not blind-swept): removed 4 dead try/catch;
  fixed 3 error-hiding (phase-insert decimal-dir I/O collision now fails loud;
  phase-remove rename partial-failure surfaced; milestone-archive true count via
  finally); left best-effort swallows with justification comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): extractFencedBlock seam + migrate api-coverage named fence

parseCoverageMatrix extracted its ```coverage fenced block with an ad-hoc regex
(the last real allow-adhoc-markdown suppression). Added extractFencedBlock to the
markdown-sectionizer seam (reuses stripFencedCode's CommonMark fence engine —
info-string match, ~~~/backtick, nesting, indent) and migrated onto it; byte-
parity on the parsed CoverageMatrix across 8 fixtures. Only security.cts:367
(a genuine `<|role|>` protocol-token false-positive, not a GFM table) remains
marked in src/ — the "prohibition with teeth" goal (nothing grandfathered but a
true FP) is met.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): By-Phase row insert is name-addressed (insertTableRow seam)

updatePerformanceMetricsSection's INSERT-new-row branch located the By-Phase
table with a canonical-column-order-only regex + a hardcoded positional row
literal, so on a reordered header it silently inserted nothing — inconsistent
with the now name-addressed UPDATE and SUM halves of the same function. Added
insertTableRow (markdown-table seam sibling of updateTableCell/deleteTableRow:
name-addressed, header-order-agnostic, EOL-preserving) and migrated the branch
onto it, mapping By-Phase values by column NAME. Canonical-order output is
byte-identical; a reordered header now inserts a correctly-mapped row; a
pre-existing CRLF mixed-EOL splice glitch is incidentally fixed. Retired the
now-dead byPhaseTablePattern const.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): phase-list checkbox flip via updateBullet seam

Added updateBullet (markdown-sectionizer): a fence-aware, offset-tracked
single-bullet write primitive (GFM 1–4-space marker tolerance) — the write
counterpart to read-only iterateBullets. Migrated mutateMilestonePhase's
phase-list checkbox flip (`- [ ] Phase N …` → `- [x] … (completed <date>)`)
off its whole-slice regex onto it, same milestone-slice scope + clock seam.
Byte-identical across simple / idempotent / metachar-title / double-space /
CRLF scenarios.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): scope the Progress-ordinal renumber to ## Progress via seam

phase remove's integer-renumber decremented Progress-table phase ordinals with a
whole-document `content.replace(/(\|\s*)(\d+)(\.\s)/g, …)` — unscoped, so it also
rewrote any `| N. …` cell in an unrelated/decoy table (same class as the batch-2
row-delete scoping bug). Migrated onto updateTableCell, scoped to the ## Progress
section, decrementing each affected row's leading phase ordinal by column name.
Byte-identical on canonical Progress tables + multi-row + decimal-sibling cases;
a decoy `| 3. … |` row before ## Progress is now correctly left untouched. The
sibling heading / checkbox-bullet / PLAN.md-filename / Depends-on-prose renumbers
are not GFM-table mutations (outside ADR-2143's table/section mandate) — left as-is.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2143): review fixes — scope traceability write, restore Current Position H3-stop

Adversarial review of the remediation (BLOCK verdict) — all 9 findings fixed:
- F1 (BLOCKER): requirements mark-complete / phase complete flipped the checkbox
  but NOT the traceability row on the shipped template, because updateTableCell
  bound to the FIRST table (## Out of Scope, no Status column) instead of the
  ## Traceability table — the #2140 silent-divergence class, re-introduced by the
  seam migration and missed by tests (fixtures had Traceability first). Scoped
  the write + hasRow probe to the ## Traceability section slice (updateTraceability
  Cell helper) in milestone.cts + phase.cts. Failing-first tests on the
  Out-of-Scope-before-Traceability layout; the #2769 first-cell match preserved.
- F2 (MAJOR): mutateCurrentPositionFirstTime restored to locateCurrentPosition
  (STOP_H2_PLUS) — collectSection's default H2-stop swallowed a level-3 subsection
  and the field regexes clobbered it (#2130 class).
- F3/F8: Progress-ordinal renumber re-escapes via escapeCell + keys padding
  recovery by row index (was de-escaping `\|` and losing padding on dup values).
- F4: insertTableRow escapes cell values internally.
- F5: updateBullet accepts a tab after the marker (`[ \t]{1,4}`).
- F7: resetSectionVerbatim consumes CRLF blank lines (byte-parity on CRLF).
- F6/F9: corrected two misleading comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): data-loss + CRLF-session user-facing fixes (#2253)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2143): de-flake the G10 windsurf ReDoS-guard wall-clock assertion

The G10 test asserted `elapsedMs < 1000` for a 200k-char payload — a wall-clock
assertion (CLAUDE.md: never assert on wall-clock time) that flaked on a loaded
node24 bench at ~1.1s. It was redundant: runHook's spawnSync `timeout: 10000`
already SIGKILLs a catastrophic-backtracking hook, so the exit-0 assertion is the
real ReDoS guard. Removed the timing assertion; kept exit-0 + documented the
subprocess-timeout mechanism. Surfaced (not caused) by this branch's gsd-test
runs loading the bench; unrelated to the markdown-parsing changes but fixed in
place per the no-flaky-tests rule.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 14:25:44 -04:00

486 lines
19 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* Security — Input validation, path traversal prevention, and prompt injection guards
*
* This module centralizes security checks for GSD tooling. Because GSD generates
* markdown files that become LLM system prompts (agent instructions, workflow state,
* phase plans), any user-controlled text that flows into these files is a potential
* indirect prompt injection vector.
*
* Threat model:
* 1. Path traversal: user-supplied file paths escape the project directory
* 2. Prompt injection: malicious text in arguments/PRDs embeds LLM instructions
* 3. Shell metacharacter injection: user text interpreted by shell
* 4. JSON injection: malformed JSON crashes or corrupts state
* 5. Regex DoS: crafted input causes catastrophic backtracking
*
* ADR-457 build-at-publish: the hand-written bin/lib/security.cjs collapsed
* to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour
* from the prior hand-written .cjs; only types are added.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
// ─── Path Traversal Prevention ──────────────────────────────────────────────
/**
* Validate that a file path resolves within an allowed base directory.
* Prevents path traversal attacks via ../ sequences, symlinks, or absolute paths.
*/
export function validatePath(filePath: unknown, baseDir: unknown, opts: { allowAbsolute?: boolean } = {}): { safe: boolean; resolved: string; error?: string } {
if (!filePath || typeof filePath !== 'string') {
return { safe: false, resolved: '', error: 'Empty or invalid file path' };
}
if (!baseDir || typeof baseDir !== 'string') {
return { safe: false, resolved: '', error: 'Empty or invalid base directory' };
}
if (filePath.includes('\0')) {
return { safe: false, resolved: '', error: 'Path contains null bytes' };
}
let resolvedBase: string;
try {
resolvedBase = fs.realpathSync(path.resolve(baseDir));
} catch {
resolvedBase = path.resolve(baseDir);
}
let resolvedPath: string;
if (path.isAbsolute(filePath)) {
if (!opts.allowAbsolute) {
return { safe: false, resolved: '', error: 'Absolute paths not allowed' };
}
resolvedPath = path.resolve(filePath);
} else {
resolvedPath = path.resolve(baseDir, filePath);
}
try {
resolvedPath = fs.realpathSync(resolvedPath);
} catch {
const parentDir = path.dirname(resolvedPath);
try {
const realParent = fs.realpathSync(parentDir);
resolvedPath = path.join(realParent, path.basename(resolvedPath));
} catch {
// Parent doesn't exist either — keep the resolved path as-is
}
}
const normalizedBase = resolvedBase + path.sep;
const normalizedPath = resolvedPath + path.sep;
if (resolvedPath !== resolvedBase && !normalizedPath.startsWith(normalizedBase)) {
return {
safe: false,
resolved: resolvedPath,
error: `Path escapes allowed directory: ${resolvedPath} is outside ${resolvedBase}`,
};
}
return { safe: true, resolved: resolvedPath };
}
/**
* Load the opt-in trusted global roots allowlist from config.
*
* Reads `config.agent_skills_security.trusted_global_roots` (an array of
* path strings). Each entry is canonicalized via realpathSync: non-strings
* are dropped, leading `~/` is expanded to `os.homedir()`, entries that are
* not absolute after expansion are dropped (project-relative paths are
* rejected as a security boundary), and entries that do not exist on disk are
* dropped (a non-existent root is not trustworthy). The canonical realpath is
* used for all subsequent checks and as the stored value — this closes the
* case-insensitive bypass on macOS APFS (`/users/alice` vs `/Users/alice`)
* and ensures trust doesn't drift across re-invocations if a root is
* re-created at a different target. Results are de-duplicated by canonical path.
*/
export function loadTrustedGlobalRoots(config: unknown): string[] {
const roots = (config as Record<string, unknown> | null | undefined)
?.['agent_skills_security'] as Record<string, unknown> | undefined;
const raw = roots?.['trusted_global_roots'];
if (!Array.isArray(raw)) return [];
// Compute canonical homedir once for case-insensitive-safe comparison.
let realHome: string;
try {
realHome = fs.realpathSync(os.homedir());
} catch {
realHome = os.homedir();
}
const seen = new Set<string>();
const result: string[] = [];
for (const entry of raw) {
if (typeof entry !== 'string') continue;
let expanded: string;
if (entry === '~') {
expanded = os.homedir();
} else if (entry.startsWith('~/')) {
expanded = path.join(os.homedir(), entry.slice(2));
} else {
expanded = entry;
}
if (!path.isAbsolute(expanded)) continue; // reject project-relative
// Canonicalize: resolve symlinks and normalise case. If the path doesn't
// exist or can't be read, skip it — a non-existent root is not trustworthy.
let real: string;
try {
real = fs.realpathSync(expanded);
} catch {
continue; // non-existent or unreadable — skip
}
// Reject dangerously broad roots: filesystem root (e.g. '/' or 'C:\' or UNC '\\server\share').
// Normalize both sides by stripping trailing path separators before comparing so that
// Windows UNC shares (where path.parse().root includes a trailing separator) are caught.
const stripTrailingSep = (p: string): string => p.replace(/[\\/]+$/, '');
if (stripTrailingSep(path.parse(real).root) === stripTrailingSep(real)) continue;
// Reject homedir itself (canonical compare closes case-insensitive bypass).
// Apply stripTrailingSep for robustness on platforms where realpathSync may
// or may not include a trailing separator on the homedir path.
if (stripTrailingSep(real) === stripTrailingSep(realHome)) continue;
if (seen.has(real)) continue;
seen.add(real);
result.push(real);
}
return result;
}
/**
* Validate a file path and throw on traversal attempt.
* Convenience wrapper around validatePath for use in CLI commands.
*/
export function requireSafePath(filePath: unknown, baseDir: unknown, label: string | null | undefined, opts: { allowAbsolute?: boolean } = {}): string {
const result = validatePath(filePath, baseDir, opts);
if (!result.safe) {
throw new Error(`${label || 'Path'} validation failed: ${result.error}`);
}
return result.resolved;
}
// ─── Prompt Injection Detection ────────────────────────────────────────────────────
/**
* Patterns that indicate prompt injection attempts in user-supplied text.
* These patterns catch common indirect prompt injection techniques where
* an attacker embeds LLM instructions in text that will be read by an agent.
*
* Note: This is defense-in-depth — not a complete solution. The primary defense
* is proper input/output boundaries in agent prompts.
*/
export const INJECTION_PATTERNS: RegExp[] = [
// Direct instruction override attempts
/ignore\s+(all\s+)?previous\s+instructions/i,
/ignore\s+(all\s+)?above\s+instructions/i,
/disregard\s+(all\s+)?previous/i,
/forget\s+(all\s+)?(your\s+)?instructions/i,
/override\s+(system|previous)\s+(prompt|instructions)/i,
// Role/identity manipulation
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
/act\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
// System prompt extraction
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
/what\s+(?:are|is)\s+your\s+(?:system\s+)?(?:prompt|instructions)/i,
// Hidden instruction markers (XML/HTML tags that mimic system messages)
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
// Requires > to close the tag (not just whitespace) to avoid matching generic types like Promise<User | null>
/<\/?(?:system|assistant|human)>/i,
/\[SYSTEM\]/i,
/\[\/?(INST)\]/i,
/<<\s*SYS\s*>>/i,
// Exfiltration attempts
/(?:send|post|fetch|curl|wget)\s+(?:to|from)\s+https?:\/\//i,
/(?:base64|btoa|encode)\s+(?:and\s+)?(?:send|exfiltrate|output)/i,
// Tool manipulation
/(?:run|execute|call|invoke)\s+(?:the\s+)?(?:bash|shell|exec|spawn)\s+(?:tool|command)/i,
];
// Explicit safe-list for data: MIME types that are benign in link targets.
// Note: image/svg+xml is intentionally NOT in this list (SVG can host <script>).
const DATA_URI_SAFE_MIME_RE = /^data:(image\/(png|jpe?g|gif|webp|bmp|ico|avif|heic)|font\/(woff2?|otf|ttf))(;[^,]*)?,/i;
interface MarkdownLinkPattern {
pattern: RegExp;
ruleId: string;
safePredicate?: (line: string) => boolean;
}
export const MARKDOWN_LINK_PATTERNS: MarkdownLinkPattern[] = [
{
pattern: /\]\(\s*javascript:/i,
ruleId: 'MD-LINK-JS-SCHEME',
},
{
pattern: /\]\(\s*data:/i,
ruleId: 'MD-LINK-DATA-SCHEME',
safePredicate: (line: string) => {
const m = line.match(/\]\(\s*(data:[^)]*)/i);
if (!m) return false;
return DATA_URI_SAFE_MIME_RE.test(m[1]);
},
},
{
pattern: /\]\(\s*https?:\/\/[^/\s]+:[^/@\s]+@/i,
ruleId: 'MD-LINK-USERINFO',
},
{
pattern: /[?&](token|access_token|id_token|refresh_token|api_key|apikey|secret|password|client_secret|code)=/i,
ruleId: 'MD-LINK-TOKEN-IN-QUERY',
},
];
interface ObfuscationPatternEntry {
pattern: RegExp;
message: string;
}
const OBFUSCATION_PATTERN_ENTRIES: ObfuscationPatternEntry[] = [
{
pattern: /\b(\w\s){4,}\w\b/,
message: 'Character-spacing obfuscation pattern detected (e.g. "i g n o r e")',
},
{
pattern: /<\/?(system|human|assistant|user)\s*>/i,
message: 'Delimiter injection pattern: <system>/<human>/<assistant>/<user> tag detected',
},
{
pattern: /0x[0-9a-fA-F]{16,}/,
message: 'Long hex sequence detected — possible encoded payload',
},
];
interface StructuredFinding {
ruleId: string;
file: string | undefined;
line: number;
match: string;
}
/**
* Scan text for potential prompt injection patterns.
* Returns an array of findings (empty = clean).
*/
export function scanForInjection(text: unknown, opts: { strict?: boolean; file?: string } = {}): { clean: boolean; findings: string[]; structuredFindings: StructuredFinding[] } {
if (!text || typeof text !== 'string') {
return { clean: true, findings: [], structuredFindings: [] };
}
const findings: string[] = [];
const structuredFindings: StructuredFinding[] = [];
for (const pattern of INJECTION_PATTERNS) {
if (pattern.test(text)) {
findings.push(`Matched injection pattern: ${pattern.source}`);
}
}
for (const entry of OBFUSCATION_PATTERN_ENTRIES) {
if (entry.pattern.test(text)) {
findings.push(entry.message);
}
}
const lines = text.split('\n');
for (const entry of MARKDOWN_LINK_PATTERNS) {
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
const m = line.match(entry.pattern);
if (!m) continue;
if (entry.safePredicate && entry.safePredicate(line)) continue;
const matchText = m[0];
findings.push(`Matched markdown link pattern [${entry.ruleId}]: ${matchText}`);
structuredFindings.push({
ruleId: entry.ruleId,
file: opts.file,
line: i + 1,
match: matchText,
});
}
}
if (opts.strict) {
// Check for suspicious Unicode that could hide instructions
// (zero-width chars, RTL override, homoglyph attacks)
if (/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/.test(text)) {
findings.push('Contains suspicious zero-width or invisible Unicode characters');
}
// Layer 1: Unicode tag block U+E0000–E007F (2025 supply-chain attack vector)
// These characters are invisible and can embed hidden instructions
if (/[\uDB40\uDC00-\uDB40\uDC7F]/u.test(text) || /[\u{E0000}-\u{E007F}]/u.test(text)) {
findings.push('Contains Unicode tag block characters (U+E0000–E007F) — invisible instruction injection vector');
}
// Check for extremely long strings that could be prompt stuffing.
// Normalize CRLF → LF before measuring so Windows checkouts don't inflate the count.
const normalizedLength = text.replace(/\r\n/g, '\n').replace(/\r/g, '\n').length;
if (normalizedLength > 50000) {
findings.push(`Suspicious text length: ${normalizedLength} chars (potential prompt stuffing)`);
}
}
return { clean: findings.length === 0, findings, structuredFindings };
}
/**
* Sanitize text that will be embedded in agent prompts or planning documents.
* Strips known injection markers while preserving legitimate content.
*/
export function sanitizeForPrompt(text: unknown): string {
if (!text || typeof text !== 'string') return text as string;
let sanitized = text;
// Strip zero-width characters that could hide instructions
sanitized = sanitized.replace(/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/g, '');
// Neutralize XML/HTML tags that mimic system boundaries
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
sanitized = sanitized.replace(/<(\/?)\s*(?:system|assistant|human|user)\s*>/gi,
(_, slash: string) => `<${slash || ''}system-text>`);
// Neutralize [SYSTEM] / [INST] / [/INST] markers
sanitized = sanitized.replace(/\[(\/?)(SYSTEM|INST)\]/gi, (_, slash: string, tag: string) => `[${slash}${tag.toUpperCase()}-TEXT]`);
// Neutralize <<SYS>> and <</SYS>> markers (Llama-style delimiters)
sanitized = sanitized.replace(/<<\/?\s*SYS\s*>>/gi, '«SYS-TEXT»');
return sanitized;
}
/**
* Sanitize text that will be displayed back to the user.
* Removes protocol-like leak markers that should never surface in checkpoints.
*/
export function sanitizeForDisplay(text: unknown): string {
if (!text || typeof text !== 'string') return text as string;
let sanitized = sanitizeForPrompt(text);
const protocolLeakPatterns = [
/^\s*(?:assistant|user|system)\s+to=[^:\s]+:[^\n]+$/i,
/^\s*<\|(?:assistant|user|system)[^|]*\|>\s*$/i, // allow-adhoc-markdown: not a GFM table-cell scan — matches `<|role|>` protocol-leak marker tokens (prompt-injection sanitization), a false-positive on the table-regex pipe+cell-class fingerprint
];
sanitized = sanitized
.split('\n')
.filter(line => !protocolLeakPatterns.some(pattern => pattern.test(line)))
.join('\n');
return sanitized;
}
// ─── Shell Safety ───────────────────────────────────────────────────────────────────────
/**
* Validate that a string is safe to use as a shell argument when quoted.
*/
export function validateShellArg(value: unknown, label: string | null | undefined): string {
if (!value || typeof value !== 'string') {
throw new Error(`${label || 'Argument'}: empty or invalid value`);
}
if (value.includes('\0')) {
throw new Error(`${label || 'Argument'}: contains null bytes`);
}
if (/[$`]/.test(value) && /\$\(|`/.test(value)) {
throw new Error(`${label || 'Argument'}: contains potential command substitution`);
}
return value;
}
// ─── JSON Safety ──────────────────────────────────────────────────────────────────────────
/**
* Safely parse JSON with error handling and optional size limits.
*/
export function safeJsonParse(text: unknown, opts: { maxLength?: number; label?: string } = {}): { ok: boolean; value?: unknown; error?: string } {
const maxLength = opts.maxLength || 1048576;
const label = opts.label || 'JSON';
if (!text || typeof text !== 'string') {
return { ok: false, error: `${label}: empty or invalid input` };
}
if (text.length > maxLength) {
return { ok: false, error: `${label}: input exceeds ${maxLength} byte limit (got ${text.length})` };
}
try {
const value = JSON.parse(text) as unknown;
return { ok: true, value };
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
return { ok: false, error: `${label}: parse error — ${msg}` };
}
}
// ─── Phase/Argument Validation ─────────────────────────────────────────────────────────
/**
* Validate a phase number argument.
*/
export function validatePhaseNumber(phase: unknown): { valid: boolean; normalized?: string; error?: string } {
if (!phase || typeof phase !== 'string') {
return { valid: false, error: 'Phase number is required' };
}
const trimmed = phase.trim();
if (/^\d{1,4}[A-Z]?(?:\.\d{1,3})*$/i.test(trimmed)) {
return { valid: true, normalized: trimmed };
}
if (/^[A-Z][A-Z0-9]*(?:-[A-Z0-9]+){1,4}$/i.test(trimmed) && trimmed.length <= 30) {
return { valid: true, normalized: trimmed };
}
return { valid: false, error: `Invalid phase number format: "${trimmed}"` };
}
/**
* Validate a STATE.md field name to prevent injection into regex patterns.
*/
export function validateFieldName(field: unknown): { valid: boolean; error?: string } {
if (!field || typeof field !== 'string') {
return { valid: false, error: 'Field name is required' };
}
if (/^[A-Za-z][A-Za-z0-9 _.\-/]{0,60}$/.test(field)) {
return { valid: true };
}
return { valid: false, error: `Invalid field name: "${field}"` };
}
// ─── Layer 3: Structural Schema Validation ──────────────────────────────────────────────────────────────────────────
const KNOWN_VALID_TAGS = new Set([
'objective', 'process', 'step', 'success_criteria', 'critical_rules',
'available_agent_types', 'purpose', 'required_reading',
]);
/**
* Validate the XML structure of a prompt file.
*/
export function validatePromptStructure(text: unknown, fileType: string): { valid: boolean; violations: string[] } {
if (!text || typeof text !== 'string') {
return { valid: true, violations: [] };
}
if (fileType !== 'agent' && fileType !== 'workflow') {
return { valid: true, violations: [] };
}
const violations: string[] = [];
const tagRegex = /<([A-Za-z][A-Za-z0-9_-]*)/g;
let match: RegExpExecArray | null;
while ((match = tagRegex.exec(text)) !== null) {
const tag = match[1].toLowerCase();
if (!KNOWN_VALID_TAGS.has(tag)) {
violations.push(`Unknown XML tag in ${fileType} file: <${tag}>`);
}
}
return { valid: violations.length === 0, violations };
}
// NOTE (#2198): scanEntropyAnomalies + shannonEntropy were removed as dead exports.
// They had zero production callers — the live hooks (gsd-prompt-guard.js,
// gsd-read-injection-scanner.js) inline their own pattern subsets for hook
// independence and never called these functions. scanForInjection is retained
// below: it serves as the CI codebase-scanner engine
// (tests/prompt-injection-scan.security.test.cjs), not as a live hook.