Files
msd-core/src/security.cts
Tom Boucher bbdf7e8e84 chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero

Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold.

THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser,
not grep) found what the epic never enumerated: ADR-4650 named seven containment
implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13
files, several guarding a write or an `fs.rmSync`. Two verified by reading rather
than pattern-matching — `research-store.cts` comments its own as "ensure the
resolved file path stays inside the store dir" immediately before a write, and
`capability-lifecycle.cts` gates `fs.rmSync` with one.

So the epic's Done-when "one containment predicate, used at every site" was FALSE
when Phase 3 reported it satisfied. It is true now: the rule is clean across
src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist.

WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose
first argument is a managed root and whose later arguments derive from argv. That
is a taint analysis over 2046 call sites, in ESLint, without type information;
"derives from argv" is not locally decidable. Any approximation either floods or
is trivially evaded, and a rule that fires on hundreds of correct sites earns an
allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is
the COMPARISON, not the join, and that has one recognizable shape.

  Arm 1  X.startsWith(Y + sep)            the hand-rolled containment idiom
  Arm 2  a containment predicate called as a bare statement, answer discarded

Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper
was called". Its example `validatePath(x, root).resolved` is already
structurally impossible — Phase 3 un-exported `validatePath` — so the remaining
expressible failure is ignoring the answer, which is the defect that recurred
five times in this epic. The census found exactly one live instance
(`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers
stop re-deriving the path the comment above it was extracted to stop them
re-deriving.

The rule deliberately does NOT try to catch validate-one-path-use-another where
the answer is used but a different variable flows onward. That needs flow
analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the
two are complementary.

PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a
lexical site onto the realpath family broke four tests and was caught only by the
matrix. Every migrated site was triaged individually. The six
installer-migrations tree-walks and the six capability-lifecycle gates take the
LEXICAL family because their operands are already realpath-resolved and they
deliberately treat the final component as a link; boundary sites take realpath.

TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken.
`installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT
`target === root` by contract, while the canonical comparison ACCEPTS it. Swapped
naively, a migration could `rmdir` the user's config root and a third-party
descriptor could write at configHome itself. Both keep `=== root` as an explicit
additional arm alongside the predicate call — the predicate decides containment,
the call site keeps its own extra condition (ADR-4650 decision 6).

ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was
byte-identical to `isContainedIn` and said so in its own docstring.
`isContainedIn` is now exported for callers that have already resolved both
operands and need only the comparison, with a doc note that a caller which has
NOT resolved them must use a full predicate instead.

THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and
carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable
reason. Two justifications: (a) not a containment decision — an ancestor-walk loop
condition, sub-repo grouping, worktree identity matching, declared-path coverage;
(b) it IS containment but the canonical predicate is unreachable —
`capability-validator.cjs` is a committed pre-build `.cjs` and the compiled
`security.cjs` is untracked build output, so requiring it would break a fresh
clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure.
The marker was renamed from `allow-lexical-prefix-match` mid-phase because that
name asserted only (a) and would have stated something false at the (b) sites.

A marker suppresses BEFORE the violation counter increments, so a file whose
every occurrence is marked still reports `staleAllowlistEntry` — otherwise a
drained entry lingers and silently re-permits the site later.

DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/`
file made `npm run lint` fail with the rule's full guidance message; removing it
returned the tree to clean. Both halves recorded — red alone proves nothing,
since a rule red for an unrelated reason looks identical.

DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two
distinct rejection messages and two manual realpath calls; routing it through
`tryWithinRoot` collapses them to one message, and a missing module now surfaces
as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The
security property is preserved and slightly strengthened — the candidate is
realpathed and containment re-checked, and the dangling-symlink oracle closure
comes along with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4654): record the containment ratchet in CONTEXT.md and the security model

Both entries previously described the seam without the thing that keeps it a
seam. They now state what the rule bans, and — more usefully for whoever reads
this next — what it deliberately does NOT attempt: deciding per path.join call
whether an argument came from user input. That question is not locally
decidable, and an approximation across ~2000 join sites would earn an exemption
list of hundreds, which is the opposite of a ratchet.

Also records the marker's two legitimate justifications and that its reason is
mandatory, so the escape stays reviewable rather than becoming a mute button.

Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json
regenerated and confirmed byte-identical rather than assumed — which also
confirms eslint-rules/ is not a shipped path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): close review findings and the two matrix failures

MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my
evidence for collapsing it was wrong. I searched tests/ for the literal string
"resolves outside its install root", found nothing, and reported that no test
asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so
the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module
pointing OUTSIDE the install root is not loaded" — the test guarding the exact
property I claimed was preserved. defaultRequireFromInstallRoot now does both
checks again with both messages byte-identical, each routed through the
canonical predicate, which is better than the original since that hand-rolled
both comparisons.

MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot
serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments,
so a suppression marker inside a plan body drifts the baseline exactly as an
edit does. Measured: with markers in place, two of the four still differed from
their committed checksums. The four shipped bodies are now byte-identical to
next, and the rule's config excludes those four paths BY NAME rather than by a
directory wildcard, so a NEW migration is still covered. Six containment
comparisons stay un-ratcheted there; that gap is recorded in the rule's Known
gaps, in CONTEXT.md and in the security model rather than left implicit.
Justification (c) is removed from the marker's documented reasons, because a
marker was proven unable to express it.

ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT
shape while permitting the incorrect one: startsWith(root) with no separator is
the genuinely unsafe form, since it accepts a sibling such as root-evil, and my
own test blessed it as valid. Flagging every bare startsWith would swamp the
rule, so that stays a STATED gap rather than a silent one. Closed for real: the
template-literal spelling, which the census never saw because it only inspected
plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated.
A separator reached through a const alias is now resolved via scope analysis.
And isContainedIn, exported in Phase 3, was missing from the discarded-result
set, so a bare no-op call went unflagged on the one function the epic funnels
through.

SECURITY REVIEW — the marker could over-suppress two ways: a block comment
worked identically to a line comment, and one marker silently covered every
violation sharing its line. It now requires a Line comment positioned after the
flagged node ends, so it anchors to the node it trails. Four sites had dropped
an unreachable-but-deliberate equality rejection against the root; each is
restored as the call site's own arm. eslint.config.mjs still documented the OLD
marker token, which my rename missed — it would have sent the next author in
circles.

A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a
stale eslint cache while twelve real violations existed. Every lint check here
now clears the cache first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): anchor a suppression marker to the violation it actually trails

The matrix caught this; my own test caught it, on its first execution. The case
"two violations on one line: trailing marker suppresses only the one it trails"
expected 1 error and got 0 — both were suppressed.

ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose
range started at or after the node's end. A trailing marker at the END of a line
sits after EVERY node on that line, so that condition held for all of them.
"After the node" does not identify WHICH node the marker trails. The fix reads
as correct and is not.

FIX: deferred reporting. Violations accumulate during traversal instead of being
reported immediately; at Program:exit each marker claims exactly ONE pending
violation — the one on its line whose end is nearest before the marker begins —
and every unclaimed violation is then counted and reported. One marker, one
suppression. An earlier violation sharing the line is still reported, which is
the property the security review asked for and the previous attempt only
appeared to deliver.

The counter now increments at flush time rather than during traversal, so a
suppressed occurrence still does not keep an allowlist entry alive.

AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test`
is hard-blocked here, so this rule's test file could only ever be executed on
the remote matrix — which is why a broken anchoring shipped into a run. ESLint's
programmatic Linter API is not a test runner, and exercising the rule through it
verifies every case locally in seconds. All 24 now pass locally, including the
two-on-one-line case that failed remotely. That loop should have been built
before the rule was first sent to the matrix rather than after it failed twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json

The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs
now carries both, with the how-to quadrant skipped for a stated reason rather
than an empty field. The audience for this deliverable is a contributor who
trips the rule, and the task-oriented guidance reaches them in the ESLint
message itself — which names the correct predicate, says how to choose between
the realpath and lexical families, cites the Phase 3 regression caused by
choosing wrong, and gives the marker syntax. A docs/how-to page would be a
second, driftable copy read by nobody at the moment of failure.

enablementSequence is recorded as what it actually is: a VERIFICATION sequence,
not an enablement one. The rule is never off, so there is no off-to-on
transition to describe.

scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not
evaluate against the mandated pr:0 placeholder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:17:46 -04:00

731 lines
30 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* Security — Input validation, path traversal prevention, and prompt injection guards
*
* This module centralizes security checks for GSD tooling. Because GSD generates
* markdown files that become LLM system prompts (agent instructions, workflow state,
* phase plans), any user-controlled text that flows into these files is a potential
* indirect prompt injection vector.
*
* Threat model:
* 1. Path traversal: user-supplied file paths escape the project directory
* 2. Prompt injection: malicious text in arguments/PRDs embeds LLM instructions
* 3. Shell metacharacter injection: user text interpreted by shell
* 4. JSON injection: malformed JSON crashes or corrupts state
* 5. Regex DoS: crafted input causes catastrophic backtracking
*
* ADR-457 build-at-publish: the hand-written bin/lib/security.cjs collapsed
* to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour
* from the prior hand-written .cjs; only types are added.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
// ─── Path Traversal Prevention ──────────────────────────────────────────────
/**
* THE containment comparison — the single place this repo decides whether an
* already-resolved path lies inside an already-resolved root (ADR-4650).
*
* Separator-aware on purpose: comparing the bare strings would accept a
* sibling that merely shares a prefix (`<root>-evil` against `<root>`), so both
* sides get a trailing separator before the prefix test. `target === root` is
* contained.
*
* `pathImpl` lets a caller supply `path.win32` / `path.posix` instead of the
* ambient module, so win32 separator semantics are testable off Windows.
*
* Exported for callers that have ALREADY resolved both operands themselves
* and need only this comparison step (e.g. a caller that owns its own
* `fs.realpathSync` calls to preserve an exists-vs-escaped tri-state). A
* caller that has NOT resolved its operands must NOT reach for this function
* directly — the comparison alone is not a containment check — and should use
* `assertWithinRoot` / `tryWithinRoot` (or the `assertWithinRootLexical` /
* `tryWithinRootLexical` pair) instead.
*/
export function isContainedIn(
resolvedTarget: string,
resolvedRoot: string,
pathImpl: { sep: string } = path,
): boolean {
if (resolvedTarget === resolvedRoot) return true;
return (resolvedTarget + pathImpl.sep).startsWith(resolvedRoot + pathImpl.sep); // allow-handrolled-containment: this IS the canonical comparison every other site routes through
}
/**
* Validate that a file path resolves within an allowed base directory.
* Prevents path traversal attacks via ../ sequences, symlinks, or absolute paths.
*/
function validatePath(filePath: unknown, baseDir: unknown, opts: { allowAbsolute?: boolean } = {}): { safe: boolean; resolved: string; error?: string } {
if (!filePath || typeof filePath !== 'string') {
return { safe: false, resolved: '', error: 'Empty or invalid file path' };
}
if (!baseDir || typeof baseDir !== 'string') {
return { safe: false, resolved: '', error: 'Empty or invalid base directory' };
}
if (filePath.includes('\0')) {
return { safe: false, resolved: '', error: 'Path contains null bytes' };
}
let resolvedBase: string;
try {
resolvedBase = fs.realpathSync(path.resolve(baseDir));
} catch {
resolvedBase = path.resolve(baseDir);
}
let resolvedPath: string;
if (path.isAbsolute(filePath)) {
if (!opts.allowAbsolute) {
return { safe: false, resolved: '', error: 'Absolute paths not allowed' };
}
resolvedPath = path.resolve(filePath);
} else {
resolvedPath = path.resolve(baseDir, filePath);
}
try {
resolvedPath = fs.realpathSync(resolvedPath);
} catch {
// realpathSync failed — either resolvedPath doesn't exist at all, or it's
// a dangling symlink (the link itself exists but its target doesn't).
// lstat (unlike stat/realpath) stats the link itself and does NOT follow
// it, so it succeeds for a dangling symlink and throws ENOENT for a
// genuinely absent path. That's the discriminator: without it, a dangling
// symlink to a non-existent OUTSIDE path would fall through to the
// parent-resolution fallback below and be re-accepted as an in-project
// path, while a symlink to an EXISTING outside path is correctly
// rejected via the realpathSync success branch above — a state
// difference an attacker can use as an existence oracle for arbitrary
// absolute paths.
try {
if (fs.lstatSync(resolvedPath).isSymbolicLink()) {
return { safe: false, resolved: '', error: 'Path is an unresolvable symbolic link' };
}
} catch {
// lstat also threw — resolvedPath (and its would-be link) genuinely
// doesn't exist. Fall through to ancestor resolution below.
}
// Walk up to the nearest ancestor that exists and realpath THAT, then
// re-append the remaining (not-yet-created) segments. This canonicalizes
// resolvedPath the same way resolvedBase was canonicalized above,
// regardless of how many leading directories are missing — a single
// parent-only check would leave resolvedPath un-canonicalized whenever
// the parent is also missing, which breaks the startsWith comparison
// below on any non-canonical cwd (e.g. macOS /var/... vs
// /private/var/...).
let ancestor = path.dirname(resolvedPath);
const remainder: string[] = [path.basename(resolvedPath)];
for (;;) {
try {
const realAncestor = fs.realpathSync(ancestor);
resolvedPath = path.join(realAncestor, ...remainder);
break;
} catch {
const parent = path.dirname(ancestor);
if (parent === ancestor) {
// Reached filesystem root without finding an existing ancestor —
// keep resolvedPath as-is.
break;
}
remainder.unshift(path.basename(ancestor));
ancestor = parent;
}
}
}
if (!isContainedIn(resolvedPath, resolvedBase)) {
return {
safe: false,
resolved: resolvedPath,
error: `Path escapes allowed directory: ${resolvedPath} is outside ${resolvedBase}`,
};
}
return { safe: true, resolved: resolvedPath };
}
/**
* Load the opt-in trusted global roots allowlist from config.
*
* Reads `config.agent_skills_security.trusted_global_roots` (an array of
* path strings). Each entry is canonicalized via realpathSync: non-strings
* are dropped, leading `~/` is expanded to `os.homedir()`, entries that are
* not absolute after expansion are dropped (project-relative paths are
* rejected as a security boundary), and entries that do not exist on disk are
* dropped (a non-existent root is not trustworthy). The canonical realpath is
* used for all subsequent checks and as the stored value — this closes the
* case-insensitive bypass on macOS APFS (`/users/alice` vs `/Users/alice`)
* and ensures trust doesn't drift across re-invocations if a root is
* re-created at a different target. Results are de-duplicated by canonical path.
*/
export function loadTrustedGlobalRoots(config: unknown): string[] {
const roots = (config as Record<string, unknown> | null | undefined)
?.['agent_skills_security'] as Record<string, unknown> | undefined;
const raw = roots?.['trusted_global_roots'];
if (!Array.isArray(raw)) return [];
// Compute canonical homedir once for case-insensitive-safe comparison.
let realHome: string;
try {
realHome = fs.realpathSync(os.homedir());
} catch {
realHome = os.homedir();
}
const seen = new Set<string>();
const result: string[] = [];
for (const entry of raw) {
if (typeof entry !== 'string') continue;
let expanded: string;
if (entry === '~') {
expanded = os.homedir();
} else if (entry.startsWith('~/')) {
expanded = path.join(os.homedir(), entry.slice(2));
} else {
expanded = entry;
}
if (!path.isAbsolute(expanded)) continue; // reject project-relative
// Canonicalize: resolve symlinks and normalise case. If the path doesn't
// exist or can't be read, skip it — a non-existent root is not trustworthy.
let real: string;
try {
real = fs.realpathSync(expanded);
} catch {
continue; // non-existent or unreadable — skip
}
// Reject dangerously broad roots: filesystem root (e.g. '/' or 'C:\' or UNC '\\server\share').
// Normalize both sides by stripping trailing path separators before comparing so that
// Windows UNC shares (where path.parse().root includes a trailing separator) are caught.
const stripTrailingSep = (p: string): string => p.replace(/[\\/]+$/, '');
if (stripTrailingSep(path.parse(real).root) === stripTrailingSep(real)) continue;
// Reject homedir itself (canonical compare closes case-insensitive bypass).
// Apply stripTrailingSep for robustness on platforms where realpathSync may
// or may not include a trailing separator on the homedir path.
if (stripTrailingSep(real) === stripTrailingSep(realHome)) continue;
if (seen.has(real)) continue;
seen.add(real);
result.push(real);
}
return result;
}
/**
* A path proven to resolve inside a declared root.
*
* A plain `string` is NOT assignable to `ContainedPath` — that asymmetry is
* the entire point. The shape being replaced (`validatePath`'s
* `{ resolved: string }`) returns a usable-looking path even when the answer
* is unsafe (the traversal branch still populates `resolved` with the
* escaping path), so a plain string in hand proves nothing. A
* `ContainedPath` can only be produced by `assertWithinRoot` /
* `tryWithinRoot` on their success paths, so possessing one is proof the
* containment check already passed.
*/
export type ContainedPath = string & { readonly __containedIn: unique symbol };
/**
* Named acceptance policy for what kind of candidate path is even considered.
*
* This replaces the old per-call-site `{ allowAbsolute: true }` boolean flag.
* At a call site, `{ allowAbsolute: true }` reads as "containment is relaxed
* here" — which is FALSE. An absolute path that resolves OUTSIDE the root is
* still rejected; the flag only ever controlled whether an absolute candidate
* was considered at all. `AbsoluteInsideRoot` states the real contract: an
* absolute candidate is accepted for consideration, but containment is
* enforced exactly as it is for a relative one.
*/
export const PathAcceptance = {
/** Relative candidates only; an absolute candidate is rejected outright. */
RelativeOnly: 'relative-only',
/**
* An absolute candidate is accepted — but ONLY if it still resolves inside the
* root. Containment is NOT relaxed by this policy; an absolute path outside the
* root is rejected exactly as a traversal is. This is the distinction the old
* `{ allowAbsolute: true }` flag failed to make at its call sites.
*/
AbsoluteInsideRoot: 'absolute-inside-root',
} as const;
export type PathAcceptancePolicy = (typeof PathAcceptance)[keyof typeof PathAcceptance];
/**
* Validate a file path and throw on traversal attempt.
* Convenience wrapper around validatePath for use in CLI commands.
*/
export function assertWithinRoot(candidate: unknown, root: unknown, label?: string | null, policy: PathAcceptancePolicy = PathAcceptance.RelativeOnly): ContainedPath {
const result = validatePath(candidate, root, { allowAbsolute: policy === PathAcceptance.AbsoluteInsideRoot });
if (!result.safe) {
throw new Error(`${label || 'Path'} validation failed: ${result.error}`);
}
return result.resolved as ContainedPath;
}
/**
* Validate a file path and return null on traversal attempt (no throw).
*
* Returns exactly `null` when unsafe — never `''`, never `result.resolved`.
* `validatePath` populates `resolved` with the ESCAPING path on the
* traversal branch, so returning it here would reproduce the defect this
* narrowing exists to remove.
*/
export function tryWithinRoot(candidate: unknown, root: unknown, policy: PathAcceptancePolicy = PathAcceptance.RelativeOnly): ContainedPath | null {
const result = validatePath(candidate, root, { allowAbsolute: policy === PathAcceptance.AbsoluteInsideRoot });
if (!result.safe) {
return null;
}
return result.resolved as ContainedPath;
}
/**
* Validate a file path and throw on traversal attempt.
* Convenience wrapper around validatePath for use in CLI commands.
*
* Delegates to assertWithinRoot so there is one implementation beneath both
* names; its declared return type is ContainedPath (a branded string, still
* assignable to string) so existing callers keep compiling untouched.
*/
export function requireSafePath(filePath: unknown, baseDir: unknown, label: string | null | undefined, policy: PathAcceptancePolicy = PathAcceptance.RelativeOnly): ContainedPath {
return assertWithinRoot(filePath, baseDir, label, policy);
}
/**
* LEXICAL containment — `path.resolve` only, never any filesystem access.
*
* Shares `isContainedIn` with the realpath-based predicate, so there is ONE
* containment decision in this repo; these differ only in how a path is
* RESOLVED before that decision, never in the decision itself (ADR-4650
* decisions 1 and 6).
*
* Use this — and say why at the call site — only where a symlink must be
* PRESERVED rather than resolved, or where the target legitimately does not
* exist yet. Three such cases exist: a destination validated before the
* `mkdirSync` that creates it, a migration that snapshots and restores a
* symlinked path AS A LINK, and a restore gate that refuses links outright.
* Everywhere else the realpath-based `assertWithinRoot` / `tryWithinRoot` is
* the correct predicate, because a lexical check CANNOT SEE A SYMLINK: a
* caller relying on one for a write-confinement guarantee must pair it with
* its own symlink refusal.
*
* `candidate` is resolved RELATIVE TO `root` (so an absolute candidate is
* taken as-is, matching `path.resolve` semantics). `target === root` is
* contained.
*
* DELIBERATELY ABSENT: no NUL-byte rejection here. The existing lexical
* callers do not reject NUL at this layer (one of them checks NUL itself,
* separately), and adding it here would change their behavior. Callers that
* need it keep their own check.
*/
export function tryWithinRootLexical(
candidate: unknown,
root: unknown,
opts: { pathImpl?: { resolve(...segments: string[]): string; sep: string } } = {},
): ContainedPath | null {
const p = opts.pathImpl || path;
if (typeof candidate !== 'string' || candidate === '') return null;
if (typeof root !== 'string' || root === '') return null;
const rootResolved = p.resolve(root);
const targetResolved = p.resolve(root, candidate);
return isContainedIn(targetResolved, rootResolved, p) ? (targetResolved as ContainedPath) : null;
}
export function assertWithinRootLexical(
candidate: unknown,
root: unknown,
label?: string | null,
opts: { pathImpl?: { resolve(...segments: string[]): string; sep: string } } = {},
): ContainedPath {
const contained = tryWithinRootLexical(candidate, root, opts);
if (contained === null) {
throw new Error(`${label || 'Path'} validation failed: lexical containment check failed`);
}
return contained;
}
// ─── Prompt Injection Detection ────────────────────────────────────────────────────
/**
* Patterns that indicate prompt injection attempts in user-supplied text.
* These patterns catch common indirect prompt injection techniques where
* an attacker embeds LLM instructions in text that will be read by an agent.
*
* Note: This is defense-in-depth — not a complete solution. The primary defense
* is proper input/output boundaries in agent prompts.
*/
export const INJECTION_PATTERNS: RegExp[] = [
// Direct instruction override attempts
/ignore\s+(all\s+)?previous\s+instructions/i,
/ignore\s+(all\s+)?above\s+instructions/i,
/disregard\s+(all\s+)?previous/i,
/forget\s+(all\s+)?(your\s+)?instructions/i,
/override\s+(system|previous)\s+(prompt|instructions)/i,
// Role/identity manipulation
/you\s+are\s+now\s+(?:a|an|the)\s+/i,
/\bact\s+as\s+(?:a|an|the)\s+(?!plan|phase|wave)/i,
/pretend\s+(?:you(?:'re| are)\s+|to\s+be\s+)/i,
/from\s+now\s+on,?\s+you\s+(?:are|will|should|must)/i,
// System prompt extraction
/(?:print|output|reveal|show|display|repeat)\s+(?:your\s+)?(?:system\s+)?(?:prompt|instructions)/i,
/what\s+(?:are|is)\s+your\s+(?:system\s+)?(?:prompt|instructions)/i,
// Hidden instruction markers (XML/HTML tags that mimic system messages)
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
// Requires > to close the tag (not just whitespace) to avoid matching generic types like Promise<User | null>
/<\/?(?:system|assistant|human)>/i,
/\[SYSTEM\]/i,
/\[\/?(INST)\]/i,
/<<\s*SYS\s*>>/i,
// Exfiltration attempts
/(?:send|post|fetch|curl|wget)\s+(?:to|from)\s+https?:\/\//i,
/(?:base64|btoa|encode)\s+(?:and\s+)?(?:send|exfiltrate|output)/i,
// Tool manipulation
/(?:run|execute|call|invoke)\s+(?:the\s+)?(?:bash|shell|exec|spawn)\s+(?:tool|command)/i,
];
// Explicit safe-list for data: MIME types that are benign in link targets.
// Note: image/svg+xml is intentionally NOT in this list (SVG can host <script>).
const DATA_URI_SAFE_MIME_RE = /^data:(image\/(png|jpe?g|gif|webp|bmp|ico|avif|heic)|font\/(woff2?|otf|ttf))(;[^,]*)?,/i;
interface MarkdownLinkPattern {
pattern: RegExp;
ruleId: string;
safePredicate?: (line: string) => boolean;
}
export const MARKDOWN_LINK_PATTERNS: MarkdownLinkPattern[] = [
{
pattern: /\]\(\s*javascript:/i,
ruleId: 'MD-LINK-JS-SCHEME',
},
{
pattern: /\]\(\s*data:/i,
ruleId: 'MD-LINK-DATA-SCHEME',
safePredicate: (line: string) => {
const m = line.match(/\]\(\s*(data:[^)]*)/i);
if (!m) return false;
return DATA_URI_SAFE_MIME_RE.test(m[1]);
},
},
{
pattern: /\]\(\s*https?:\/\/[^/\s]+:[^/@\s]+@/i,
ruleId: 'MD-LINK-USERINFO',
},
{
pattern: /[?&](token|access_token|id_token|refresh_token|api_key|apikey|secret|password|client_secret|code)=/i,
ruleId: 'MD-LINK-TOKEN-IN-QUERY',
},
];
interface ObfuscationPatternEntry {
pattern: RegExp;
message: string;
}
const OBFUSCATION_PATTERN_ENTRIES: ObfuscationPatternEntry[] = [
{
pattern: /\b(\w\s){4,}\w\b/,
message: 'Character-spacing obfuscation pattern detected (e.g. "i g n o r e")',
},
{
pattern: /<\/?(system|human|assistant|user)\s*>/i,
message: 'Delimiter injection pattern: <system>/<human>/<assistant>/<user> tag detected',
},
{
pattern: /0x[0-9a-fA-F]{16,}/,
message: 'Long hex sequence detected — possible encoded payload',
},
];
interface StructuredFinding {
ruleId: string;
file: string | undefined;
line: number;
match: string;
}
/**
* Scan text for potential prompt injection patterns.
* Returns an array of findings (empty = clean).
*/
export function scanForInjection(text: unknown, opts: { strict?: boolean; file?: string } = {}): { clean: boolean; findings: string[]; structuredFindings: StructuredFinding[] } {
if (!text || typeof text !== 'string') {
return { clean: true, findings: [], structuredFindings: [] };
}
const findings: string[] = [];
const structuredFindings: StructuredFinding[] = [];
for (const pattern of INJECTION_PATTERNS) {
if (pattern.test(text)) {
findings.push(`Matched injection pattern: ${pattern.source}`);
}
}
for (const entry of OBFUSCATION_PATTERN_ENTRIES) {
if (entry.pattern.test(text)) {
findings.push(entry.message);
}
}
const lines = text.split('\n');
for (const entry of MARKDOWN_LINK_PATTERNS) {
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
const m = line.match(entry.pattern);
if (!m) continue;
if (entry.safePredicate && entry.safePredicate(line)) continue;
const matchText = m[0];
findings.push(`Matched markdown link pattern [${entry.ruleId}]: ${matchText}`);
structuredFindings.push({
ruleId: entry.ruleId,
file: opts.file,
line: i + 1,
match: matchText,
});
}
}
if (opts.strict) {
// Check for suspicious Unicode that could hide instructions
// (zero-width chars, RTL override, homoglyph attacks)
if (/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/.test(text)) {
findings.push('Contains suspicious zero-width or invisible Unicode characters');
}
// Layer 1: Unicode tag block U+E0000–E007F (2025 supply-chain attack vector)
// These characters are invisible and can embed hidden instructions
if (/[\uDB40\uDC00-\uDB40\uDC7F]/u.test(text) || /[\u{E0000}-\u{E007F}]/u.test(text)) {
findings.push('Contains Unicode tag block characters (U+E0000–E007F) — invisible instruction injection vector');
}
// Check for extremely long strings that could be prompt stuffing.
// Normalize CRLF → LF before measuring so Windows checkouts don't inflate the count.
const normalizedLength = text.replace(/\r\n/g, '\n').replace(/\r/g, '\n').length;
if (normalizedLength > 50000) {
findings.push(`Suspicious text length: ${normalizedLength} chars (potential prompt stuffing)`);
}
}
return { clean: findings.length === 0, findings, structuredFindings };
}
/**
* Sanitize text that will be embedded in agent prompts or planning documents.
* Strips known injection markers while preserving legitimate content.
*/
export function sanitizeForPrompt(text: unknown): string {
if (!text || typeof text !== 'string') return text as string;
let sanitized = text;
// Strip zero-width characters that could hide instructions
sanitized = sanitized.replace(/[\u200B-\u200F\u2028-\u202F\uFEFF\u00AD]/g, '');
// Neutralize XML/HTML tags that mimic system boundaries
// Note: <instructions> is excluded — GSD uses it as legitimate prompt structure
sanitized = sanitized.replace(/<(\/?)\s*(?:system|assistant|human|user)\s*>/gi,
(_, slash: string) => `<${slash || ''}system-text>`);
// Neutralize [SYSTEM] / [INST] / [/INST] markers
sanitized = sanitized.replace(/\[(\/?)(SYSTEM|INST)\]/gi, (_, slash: string, tag: string) => `[${slash}${tag.toUpperCase()}-TEXT]`);
// Neutralize <<SYS>> and <</SYS>> markers (Llama-style delimiters)
sanitized = sanitized.replace(/<<\/?\s*SYS\s*>>/gi, '«SYS-TEXT»');
return sanitized;
}
/**
* Sanitize text that will be displayed back to the user.
* Removes protocol-like leak markers that should never surface in checkpoints.
*/
export function sanitizeForDisplay(text: unknown): string {
if (!text || typeof text !== 'string') return text as string;
let sanitized = sanitizeForPrompt(text);
const protocolLeakPatterns = [
/^\s*(?:assistant|user|system)\s+to=[^:\s]+:[^\n]+$/i,
/^\s*<\|(?:assistant|user|system)[^|]*\|>\s*$/i, // allow-adhoc-markdown: not a GFM table-cell scan — matches `<|role|>` protocol-leak marker tokens (prompt-injection sanitization), a false-positive on the table-regex pipe+cell-class fingerprint
];
sanitized = sanitized
.split('\n')
.filter(line => !protocolLeakPatterns.some(pattern => pattern.test(line)))
.join('\n');
return sanitized;
}
/**
* Sanitize a value that must render as a SINGLE LINE and is derived from a
* filesystem name (a phase directory's number/name token, an archived
* milestone label, a bare filename) — not from file/frontmatter CONTENT.
*
* Why this is NOT `sanitizeForDisplay`: that helper's job is multi-line
* prose — it strips whole protocol-leak LINES while deliberately preserving
* `\n` between legitimate ones (see its docstring and
* `tests/security.test.cjs`'s neighbouring describe). A filesystem name is
* the opposite shape: it is supposed to be one line, so a `\n`/`\r` inside
* one is never legitimate content to preserve — it is an attacker (or a
* doctored checkout) using the directory NAME itself as the injection
* vector. #3458's reproduction: a phase directory literally named
* `zz\n0 open items require decisions.\n\x1b[2K\x1b[1G FORGED`
* flows verbatim into `audit-open`'s human report (the phase-number
* fallback taken when the name doesn't match `PHASE_NUMBER_TOKEN_SOURCE`).
* `sanitizeForDisplay` would pass every one of those bytes straight through
* — by design, since it never touches control characters — so the embedded
* `\n` becomes a real newline in the report, printing a forged
* "0 open items require decisions." as its own line, and the raw ESC bytes
* reach the terminal.
*
* This helper closes that hole by ESCAPING (never silently stripping) the
* C0 control range (0x00–0x1F, including ESC 0x1B, CR, LF), DEL (0x7F), and
* the C1 range (0x80–0x9F) into a visible representation (`\n`, `\x1b`,
* ...). Escaping rather than stripping is deliberate: a reviewer reading the
* report should be able to SEE that a name was doctored, not have it quietly
* normalized away as if nothing happened. Every other character — including
* all ordinary printable and non-ASCII text — passes through byte-identical.
*/
export function sanitizeLabel(text: unknown): string {
if (!text || typeof text !== 'string') return text as string;
const NAMED_ESCAPES: Record<number, string> = {
0x00: '\\0',
0x07: '\\a',
0x08: '\\b',
0x09: '\\t',
0x0a: '\\n',
0x0b: '\\v',
0x0c: '\\f',
0x0d: '\\r',
0x1b: '\\x1b',
};
let out = '';
for (const ch of text) {
const code = ch.codePointAt(0) as number;
const isC0 = code <= 0x1f;
const isDel = code === 0x7f;
const isC1 = code >= 0x80 && code <= 0x9f;
if (isC0 || isDel || isC1) {
out += NAMED_ESCAPES[code] ?? `\\x${code.toString(16).padStart(2, '0')}`;
} else {
out += ch;
}
}
return out;
}
// ─── Shell Safety ───────────────────────────────────────────────────────────────────────
/**
* Validate that a string is safe to use as a shell argument when quoted.
*/
export function validateShellArg(value: unknown, label: string | null | undefined): string {
if (!value || typeof value !== 'string') {
throw new Error(`${label || 'Argument'}: empty or invalid value`);
}
if (value.includes('\0')) {
throw new Error(`${label || 'Argument'}: contains null bytes`);
}
if (/[$`]/.test(value) && /\$\(|`/.test(value)) {
throw new Error(`${label || 'Argument'}: contains potential command substitution`);
}
return value;
}
// ─── JSON Safety ──────────────────────────────────────────────────────────────────────────
/**
* Safely parse JSON with error handling and optional size limits.
*/
export function safeJsonParse(text: unknown, opts: { maxLength?: number; label?: string } = {}): { ok: boolean; value?: unknown; error?: string } {
const maxLength = opts.maxLength || 1048576;
const label = opts.label || 'JSON';
if (!text || typeof text !== 'string') {
return { ok: false, error: `${label}: empty or invalid input` };
}
if (text.length > maxLength) {
return { ok: false, error: `${label}: input exceeds ${maxLength} byte limit (got ${text.length})` };
}
try {
const value = JSON.parse(text) as unknown;
return { ok: true, value };
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
return { ok: false, error: `${label}: parse error — ${msg}` };
}
}
// ─── Phase/Argument Validation ─────────────────────────────────────────────────────────
/**
* Validate a phase number argument.
*/
export function validatePhaseNumber(phase: unknown): { valid: boolean; normalized?: string; error?: string } {
if (!phase || typeof phase !== 'string') {
return { valid: false, error: 'Phase number is required' };
}
const trimmed = phase.trim();
if (/^\d{1,4}[A-Z]?(?:\.\d{1,3})*$/i.test(trimmed)) {
return { valid: true, normalized: trimmed };
}
if (/^[A-Z][A-Z0-9]*(?:-[A-Z0-9]+){1,4}$/i.test(trimmed) && trimmed.length <= 30) {
return { valid: true, normalized: trimmed };
}
return { valid: false, error: `Invalid phase number format: "${trimmed}"` };
}
/**
* Validate a STATE.md field name to prevent injection into regex patterns.
*/
export function validateFieldName(field: unknown): { valid: boolean; error?: string } {
if (!field || typeof field !== 'string') {
return { valid: false, error: 'Field name is required' };
}
if (/^[A-Za-z][A-Za-z0-9 _.\-/]{0,60}$/.test(field)) {
return { valid: true };
}
return { valid: false, error: `Invalid field name: "${field}"` };
}
// ─── Layer 3: Structural Schema Validation ──────────────────────────────────────────────────────────────────────────
const KNOWN_VALID_TAGS = new Set([
'objective', 'process', 'step', 'success_criteria', 'critical_rules',
'available_agent_types', 'purpose', 'required_reading',
]);
/**
* Validate the XML structure of a prompt file.
*/
export function validatePromptStructure(text: unknown, fileType: string): { valid: boolean; violations: string[] } {
if (!text || typeof text !== 'string') {
return { valid: true, violations: [] };
}
if (fileType !== 'agent' && fileType !== 'workflow') {
return { valid: true, violations: [] };
}
const violations: string[] = [];
const tagRegex = /<([A-Za-z][A-Za-z0-9_-]*)/g;
let match: RegExpExecArray | null;
while ((match = tagRegex.exec(text)) !== null) {
const tag = match[1].toLowerCase();
if (!KNOWN_VALID_TAGS.has(tag)) {
violations.push(`Unknown XML tag in ${fileType} file: <${tag}>`);
}
}
return { valid: violations.length === 0, violations };
}
// NOTE (#2198): scanEntropyAnomalies + shannonEntropy were removed as dead exports.
// They had zero production callers — the live hooks (gsd-prompt-guard.js,
// gsd-read-injection-scanner.js) inline their own pattern subsets for hook
// independence and never called these functions. scanForInjection is retained
// below: it serves as the CI codebase-scanner engine
// (tests/prompt-injection-scan.security.test.cjs), not as a live hook.