fix(#3477): run untrusted key_links patterns on a linear-time engine (#3496)

`cmdVerifyKeyLinks` compiled `must_haves.key_links[].pattern` from plan frontmatter with `new RegExp()` and tested it against whole file contents, so a nested-quantifier pattern such as `(a+)+$` hung `verify-phase` indefinitely (CWE-1333). JavaScript has no regex-execution timeout.

Untrusted patterns now run on RE2 (re2js), whose match time is linear in input length — the class is closed by the engine, not by a heuristic screen. The screen lost in the ADR-0174 consolidation was deliberately NOT restored: it never worked, since `(a|a)*$`, `((a+))+$`, `(a+){2,}$` and `(a{1,3})+$` all evade it. A refused pattern's matcher returns false for every input, so it cannot report a match no matter what the caller does.

The engine is vendored at gsd-core/bin/lib/vendor/re2js.cjs because gsd-core/bin/** is copied into installed trees with no node_modules; runtime dependencies are unchanged. New ESLint rule local/no-external-require-in-bin enforces that invariant, which had been documented in a comment since the #3024/#2071 bug class and enforced nowhere.

Backreferences and look-around are unsupported by RE2 by construction — disclosed in a Changed changeset.

Closes #3477

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-08-14 14:34:36 -04:00
committed by GitHub
parent b946051a46
commit 895d9df96d
19 changed files with 9067 additions and 17 deletions

View File

@@ -0,0 +1,5 @@
---
type: Security
pr: 3496
---
**`verify key-links` can no longer be hung by a plan's `key_links` pattern** — the pattern was compiled straight from plan frontmatter with a backtracking engine and tested against whole file contents, so a nested-quantifier pattern such as `(a+)+$` pinned a CPU core indefinitely and stalled any `verify-phase` run that reached it. Untrusted patterns now execute on the RE2 engine, whose match time is linear in the input length, and a pattern that cannot be compiled is refused outright rather than guessed at — a refused pattern can never report a match. (#3477)

View File

@@ -0,0 +1,5 @@
---
type: Changed
pr: 3496
---
**`must_haves.key_links[].pattern` now uses RE2 syntax** — backreferences and look-around are no longer supported in a key-links pattern, because they are the constructs that require a backtracking engine and cannot be evaluated in guaranteed linear time. A pattern using them is reported as `pattern_neutralized: "unsupported"` with the link marked unverified, rather than being silently matched as literal text. Ordinary patterns, including every example shipped in the docs, are unaffected. (#3477)

9
.gitignore vendored
View File

@@ -288,6 +288,15 @@ __pycache__/
venv/
target/
vendor/
# Vendored third-party artifacts under version control on purpose — see
# gsd-core/bin/lib/vendor/README.md. bin/** must have zero external requires
# (installed trees ship with no node_modules), so this one directory is a
# deliberate, tracked exception to the blanket `vendor/` ignore above.
!/gsd-core/bin/lib/vendor/
# Source-side twin (tsc resolves a .cts module's relative imports against
# src/, not the output dir) of the same vendored artifact — see
# scripts/lint-vendored-deps.cjs, which keeps both copies in sync.
!/src/vendor/
*.log
.cache/
tmp/

View File

@@ -510,6 +510,7 @@
"update-context.cjs",
"validate-command-router.cjs",
"validate.cjs",
"vendor/re2js.cjs",
"verification-command-router.cjs",
"verification.cjs",
"verify-command-router.cjs",

View File

@@ -622,6 +622,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `worktree-base-ref.cjs` | Worktree base-ref drift detection and degrade decision (`evaluateWorktreeBaseDegrade`) plus no-clobber `worktree.baseRef` settings management for the `base-check`/`set-baseref` subcommands (#683) |
| `health-diagnostic-rules/worktree-health.cjs` | Health-diagnostic rules: worktree health checks (W020, W017, W027 — the split-off stale-worktree subject), ported behavior-preserving from `cmdValidateHealth` (ADR-3180 §8.2/§8.3/§8.5, Phase 11, #3309) |
| `worktree-safety.cjs` | Worktree-root resolution and non-destructive prune policy decisions; owns W017 health-check logic |
| `vendor/re2js.cjs` | **Vendored third-party artifact, not a GSD module.** Verbatim copy of `re2js`' CJS build — the RE2 linear-time regex engine used by `pattern.cjs` to evaluate untrusted `key_links` patterns without catastrophic backtracking (#3477). Vendored because `gsd-core/bin/**` is copied into installed trees that have no `node_modules`, so it may contain no external requires (enforced by `local/no-external-require-in-bin`). Never hand-edit; `scripts/lint-vendored-deps.cjs` byte-compares it against the pinned `re2js` devDependency in `lint:ci`. See `gsd-core/bin/lib/vendor/README.md` |
| `write-set.cjs` | Shared fail-loud `Result<T>` (`{ok:true,value}\|{ok:false,reason}`) and per-surface write-set contracts (ADR-2143, epic #2143) — `WriteOutcome` (`{surface,applied}`), `WriteSet` (`WriteOutcome[]`), and `writeSetComplete(ws)` (true only when the set is non-empty AND every surface applied, never an OR-into-one-flag); `markdown-table.cjs` re-exports `Result` from here so existing importers are unaffected; consumed by `milestone.cts`'s `requirements mark-complete` handler to report a structured per-surface (`checkbox`/`traceability`) write-set alongside its existing fields (fixes the structural half of #2140) |
[`docs/CLI-TOOLS.md`](CLI-TOOLS.md) may describe a subset of these modules; when it disagrees with the filesystem, this table and the directory listing are authoritative.

View File

@@ -0,0 +1,129 @@
'use strict';
/**
* no-external-require-in-bin
*
* Flag any `require(...)` / `import ... from '...'` under `gsd-core/bin/**`
* whose specifier is neither relative (`./`, `../`) nor a Node builtin
* (including the `node:` prefix form).
*
* ## Why
*
* `gsd-core/bin/**` is copied by the installer into trees that have NO
* `node_modules` (e.g. `~/.claude/gsd-core/`). An external (npm-package)
* `require()`/`import` under this tree resolves fine in THIS repo (where
* `node_modules/` exists) but throws `Cannot find module '<pkg>'` for every
* installed user, because the package is never shipped there. #3477
* follow-up: `src/pattern.cts` (compiled to `gsd-core/bin/lib/pattern.cjs`)
* shipped `import { RE2JS } from 're2js'` and broke `verify` for every
* installed user until the dependency was vendored under
* `gsd-core/bin/lib/vendor/`.
*
* The fix for a genuine external-package dependency is never "add it back to
* `dependencies`" — vendor the compiled artifact under
* `gsd-core/bin/lib/vendor/` (a verbatim, third-party copy; see
* `gsd-core/bin/lib/vendor/README.md`) and import it via a relative path
* instead.
*
* ## Why this is ALSO registered on src/**\/*.cts
*
* Every `src/**\/*.cts` module compiles 1:1 into `gsd-core/bin/lib/*.cjs`
* (ADR-457; `tsconfig.build.json` `rootDir: "src"`, `outDir:
* "gsd-core/bin/lib"`), and the emitted `.cjs` mirrors are almost entirely
* `eslint.config.mjs` global-`ignores`d as generated artifacts (lint the
* source, not the tsc output) — so a rule registered ONLY on
* `gsd-core/bin/**\/*.cjs` would never see a bad import re-introduced into an
* already-migrated module. `src/pattern.cts`'s `import { RE2JS } from
* 're2js'` is exactly this case: the compiled `gsd-core/bin/lib/pattern.cjs`
* is on the ignore list, so only catching it at the `.cts` source closes the
* gap. `TSImportEqualsDeclaration` (the `import x = require('./y.cjs')` form
* used throughout `src/**\/*.cts` for CommonJS interop) is handled alongside
* plain `ImportDeclaration` for this reason.
*/
const { builtinModules } = require('node:module');
const BUILTIN_MODULES = new Set(builtinModules);
/**
* Is `specifier` a Node builtin module (with or without the `node:` prefix)?
* @param {string} specifier
* @returns {boolean}
*/
function isBuiltinModule(specifier) {
const bare = specifier.startsWith('node:') ? specifier.slice('node:'.length) : specifier;
return BUILTIN_MODULES.has(bare) || BUILTIN_MODULES.has(specifier);
}
/**
* Is `specifier` a relative import (`./` or `../`)?
* @param {string} specifier
* @returns {boolean}
*/
function isRelativeSpecifier(specifier) {
return specifier.startsWith('./') || specifier.startsWith('../');
}
/** @type {import('eslint').Rule.RuleModule} */
const rule = {
meta: {
type: 'problem',
docs: {
description:
'Disallow require()/import of an external (non-relative, non-builtin) module under gsd-core/bin/** (installed trees have no node_modules)',
category: 'Portability',
},
schema: [],
messages: {
externalRequireInBin:
'External module "{{specifier}}" required/imported under gsd-core/bin/**: installed ' +
'trees have no node_modules (gsd-core/bin/** is copied verbatim into e.g. ' +
'~/.claude/gsd-core/), so this resolves here but throws "Cannot find module' +
'" for every installed user. Vendor the artifact under gsd-core/bin/lib/vendor/ ' +
'(see gsd-core/bin/lib/vendor/README.md) and import it via a relative path instead.',
},
},
create(context) {
/**
* @param {import('eslint').Rule.Node} node — reported node
* @param {string} specifier
*/
function check(node, specifier) {
if (typeof specifier !== 'string') return;
if (isRelativeSpecifier(specifier)) return;
if (isBuiltinModule(specifier)) return;
context.report({ node, messageId: 'externalRequireInBin', data: { specifier } });
}
return {
CallExpression(node) {
if (node.callee.type !== 'Identifier' || node.callee.name !== 'require') return;
const arg = node.arguments[0];
if (!arg || arg.type !== 'Literal' || typeof arg.value !== 'string') return;
check(node, arg.value);
},
ImportDeclaration(node) {
check(node, node.source.value);
},
ImportExpression(node) {
const arg = node.source;
if (!arg || arg.type !== 'Literal' || typeof arg.value !== 'string') return;
check(node, arg.value);
},
// `import foo = require('...')` — the CommonJS-interop form used
// throughout src/**/*.cts (every src/*.cts compiles 1:1 into
// gsd-core/bin/lib/*.cjs, so it is exactly as much "gsd-core/bin/**"
// content as a hand-written .cjs file is).
TSImportEqualsDeclaration(node) {
const ref = node.moduleReference;
if (!ref || ref.type !== 'TSExternalModuleReference') return;
const arg = ref.expression;
if (!arg || arg.type !== 'Literal' || typeof arg.value !== 'string') return;
check(node, arg.value);
},
};
},
};
module.exports = rule;

View File

@@ -29,6 +29,7 @@ import requireFsOpFallback from './eslint-rules/require-fs-op-fallback.cjs';
import noUnboundedSpawn from './eslint-rules/no-unbounded-spawn.cjs';
import noDuplicateFoldMarker from './eslint-rules/no-duplicate-fold-marker.cjs';
import requireSubprocessTimeout from './eslint-rules/require-subprocess-timeout.cjs';
import noExternalRequireInBin from './eslint-rules/no-external-require-in-bin.cjs';
const localPlugin = {
rules: {
@@ -52,6 +53,7 @@ const localPlugin = {
'no-unbounded-spawn': noUnboundedSpawn,
'no-duplicate-fold-marker': noDuplicateFoldMarker,
'require-subprocess-timeout': requireSubprocessTimeout,
'no-external-require-in-bin': noExternalRequireInBin,
},
};
@@ -293,6 +295,15 @@ export default tseslint.config(
'gsd-core/bin/lib/workflow-fragments.cjs',
// ADR-1671 Phase 5 (#2932): tsc-generated runtime artifact — lint the src/section-manifest.cts source.
'gsd-core/bin/lib/section-manifest.cjs',
// #3477 follow-up: verbatim third-party artifact vendored so gsd-core/bin/**
// carries zero external requires (installed trees have no node_modules).
// See gsd-core/bin/lib/vendor/README.md; never lint/edit these by hand.
'gsd-core/bin/lib/vendor/**',
// Source-side twin of the same vendored .d.cts (needed so tsc resolves
// types for the relative './vendor/re2js.cjs' import from
// src/pattern.cts — module resolution for a .cts source is relative to
// src/, not the output dir). Same verbatim-third-party exemption.
'src/vendor/**',
],
},
@@ -342,6 +353,14 @@ export default tseslint.config(
// repo/missing network (DEFECT.UNBOUNDED-SUBPROCESS in CONTEXT.md).
// The 8 pre-existing call sites this surfaced were migrated in #2896.
'local/require-subprocess-timeout': 'error',
// #3477 follow-up: every src/**/*.cts module compiles 1:1 into
// gsd-core/bin/lib/*.cjs, which ships into installed trees with no
// node_modules — and the emitted mirror is almost always
// eslint-ignored as a generated artifact (see the src/pattern.cts note
// in eslint-rules/no-external-require-in-bin.cjs), so this is the ONLY
// place a bad external import in an already-migrated module is still
// visible to lint.
'local/no-external-require-in-bin': 'error',
},
},
@@ -428,6 +447,34 @@ export default tseslint.config(
},
},
// ── gsd-core/bin/**/*.cjs only — no-external-require-in-bin ────────────────
// A NARROWER block than the combined glob above on purpose: gsd-core/bin/**
// is the ONLY surface in that shared glob that is copied verbatim into
// installed trees with no node_modules (scripts/**, eslint-rules/**,
// bin/lib/**, pi/**, examples/**, vscode/*.js, .kilo/plugins/*.js, and
// .opencode/plugins/*.js all run inside THIS repo checkout, where
// node_modules exists, and legitimately require npm packages). Registering
// this rule on the shared block above would falsely flag every one of
// those. #3477 follow-up: re2js was the live instance of this defect —
// src/pattern.cts (compiled to gsd-core/bin/lib/pattern.cjs) shipped
// `import { RE2JS } from 're2js'` and broke `verify` for every installed
// user until the dependency was vendored under gsd-core/bin/lib/vendor/.
{
files: ['gsd-core/bin/**/*.cjs'],
plugins: {
local: localPlugin,
},
languageOptions: {
sourceType: 'commonjs',
globals: {
...globals.node,
},
},
rules: {
'local/no-external-require-in-bin': 'error',
},
},
// ── hooks/**/*.js — enforcement hooks (#3059) ──────────────────────────────
{
files: ['hooks/**/*.js', 'hooks/**/*.cjs'],

37
gsd-core/bin/lib/vendor/README.md vendored Normal file
View File

@@ -0,0 +1,37 @@
# vendor/
This directory holds **verbatim, unmodified** copies of third-party build
artifacts that `gsd-core/bin/**` needs at runtime.
## Why
`gsd-core/bin/**` is copied by the installer into trees that have **no
`node_modules`** (e.g. `~/.claude/gsd-core/`). Any external (non-relative,
non-builtin) `require()`/`import` under `gsd-core/bin/**` breaks `verify`
(and everything else) for every installed user, because the module simply
cannot be resolved there. The fix is to vendor the compiled artifact
in-tree instead of depending on it being installed as an npm package.
`eslint-rules/no-external-require-in-bin.cjs` enforces this at lint time.
## Contents
- `re2js.cjs` — verbatim copy of `node_modules/re2js/build/index.cjs`
(upstream package `re2js`, pinned version see `package.json`
`devDependencies.re2js`). Used by `src/pattern.cts` (compiled to
`gsd-core/bin/lib/pattern.cjs`) for linear-time RE2 pattern compilation.
- `re2js.d.cts` — verbatim copy of `node_modules/re2js/build/index.d.cts`,
so TypeScript resolves types for the relative import from `src/pattern.cts`.
## Do not hand-edit
These files are **verbatim** copies of the upstream build output. Never
edit them directly — refresh them from `node_modules` instead:
```
cp node_modules/re2js/build/index.cjs gsd-core/bin/lib/vendor/re2js.cjs
cp node_modules/re2js/build/index.d.cts gsd-core/bin/lib/vendor/re2js.d.cts
```
`node scripts/lint-vendored-deps.cjs` fails CI if the vendored copy drifts
byte-for-byte from `node_modules/re2js/build/` or from the `re2js` version
pinned in `package.json` `devDependencies`.

6480
gsd-core/bin/lib/vendor/re2js.cjs vendored Normal file

File diff suppressed because one or more lines are too long

938
gsd-core/bin/lib/vendor/re2js.d.cts vendored Normal file
View File

@@ -0,0 +1,938 @@
// Generated by dts-bundle-generator v9.5.1
declare class DFA {
static MAX_CACHE_CLEARS: number;
static STATE_MEMORY_ESTIMATE: number;
constructor(prog: any, maxMem?: number);
prog: any;
stateCache: Map<any, any>;
stateCount: number;
startState: any;
stateLimit: number;
cacheClears: number;
failed: boolean;
clock: number;
computeClosure(pcs: any): {
pcs: Int32Array<ArrayBuffer>;
isMatch: boolean;
matchIDs: any[];
};
getState(pcs: any): any;
evictCache(): void;
step(state: any, charCode: any, anchor: any): any;
match(input: any, pos: any, anchor: any): boolean;
matchSet(input: any, pos: any, anchor: any): any[];
}
declare class Prog {
inst: any[];
start: number;
numCap: number;
lbStarts: any[];
numLb: number;
getInst(pc: any): any;
numInst(): number;
addInst(op: any): void;
skipNop(pc: any): any;
prefix(): (string | boolean)[];
startCond(): number;
patch(l: any, val: any): void;
append(l1: any, l2: any): any;
/**
*
* @returns {string}
*/
toString(): string;
}
export class RE2Set {
/** @type {number} */
static UNANCHORED: number;
/** @type {number} */
static ANCHOR_START: number;
/** @type {number} */
static ANCHOR_BOTH: number;
/**
* Constructs a new RE2Set with the specified anchor mode and flags.
* @param {number} [anchor=RE2Set.UNANCHORED] - The anchoring mode (e.g., RE2Set.UNANCHORED).
* @param {number} [flags=0] - The public flags to apply to all patterns in the set.
* @param {number} [maxMem=8388608] - The maximum memory in bytes to use for the DFA (default 8MB).
*/
constructor(anchor?: number, flags?: number, maxMem?: number);
anchor: number;
jsFlags: number;
maxMem: number;
re2Flags: number;
regexps: any[];
prog: Prog;
dfa: DFA;
dummyRe2: {
prog: Prog;
cond: number;
prefix: string;
prefixRune: number;
longest: boolean;
};
/**
* Adds a new regular expression pattern to the set.
* Patterns cannot be added after the set has been compiled.
* @param {string} pattern - The regular expression pattern to add.
* @returns {number} The integer index assigned to the added pattern.
* @throws {RE2JSCompileException} If patterns are added after compilation.
*/
add(pattern: string): number;
/**
* Compiles the added patterns into a single state machine.
* This is automatically called on the first match if not called explicitly.
* @returns {void}
*/
compile(): void;
/**
* Matches the input against the compiled set of regular expressions.
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to match against.
* @returns {number[]} An array of indices representing the patterns that successfully matched the input.
*/
match(input: string | number[] | Uint8Array): number[];
}
export class MatcherInput {
/**
* Return the MatcherInput for UTF_16 encoding.
* @returns {Utf16MatcherInput}
*/
static utf16(charSequence: any): Utf16MatcherInput;
/**
* Return the MatcherInput for UTF_8 encoding.
* @returns {Utf8MatcherInput}
*/
static utf8(input: any): Utf8MatcherInput;
}
/**
* Abstract the representations of input text supplied to Matcher.
*/
export class MatcherInputBase {
static Encoding: any;
getEncoding(): void;
/** @returns {string} */
asCharSequence(): string;
/** @returns {Uint8Array|number[]} */
asBytes(): Uint8Array | number[];
/** @returns {number} */
length(): number;
/**
*
* @returns {boolean}
*/
isUTF8Encoding(): boolean;
/**
*
* @returns {boolean}
*/
isUTF16Encoding(): boolean;
}
declare class Utf16MatcherInput extends MatcherInputBase {
/** @param {string|null} charSequence */
constructor(charSequence?: string | null);
charSequence: string;
getEncoding(): any;
/**
*
* @returns {number[]}
*/
asBytes(): number[];
}
declare class Utf8MatcherInput extends MatcherInputBase {
/** @param {Uint8Array|number[]|null} bytes */
constructor(bytes?: Uint8Array | number[] | null);
bytes: number[] | Uint8Array<ArrayBufferLike>;
getEncoding(): any;
}
/**
* A stateful iterator that interprets a regex {@code RE2JS} on a specific input.
*
* Conceptually, a Matcher consists of four parts:
* <ol>
* <li>A compiled regular expression {@code RE2JS}, set at construction and fixed for the lifetime
* of the matcher.</li>
*
* <li>The remainder of the input string, set at construction or {@link #reset()} and advanced by
* each match operation such as {@link #find}, {@link #matches} or {@link #lookingAt}.</li>
*
* <li>The current match information, accessible via {@link #start}, {@link #end}, and
* {@link #group}, and updated by each match operation.</li>
*
* <li>The append position, used and advanced by {@link #appendReplacement} and {@link #appendTail}
* if performing a search and replace from the input to an external {@code StringBuffer}.
*
* </ol>
*
*
* @author rsc@google.com (Russ Cox)
*/
export class Matcher {
/**
* V8 and WebKit have historical hard limits on the number of arguments
* that can be passed to a function. We cap replacer arguments to prevent
* Call Stack Overflow (DoS) vulnerabilities on massive ASTs.
*/
static MAX_REPLACER_ARGS: number;
/**
* Quotes '\' and '$' in {@code s}, so that the returned string could be used in
* {@link #appendReplacement} as a literal replacement of {@code s}.
*
* @param {string} str the string to be quoted
* @param {boolean} [javaMode=false] whether the replacement will be used in javaMode
* @returns {string} the quoted string
*/
static quoteReplacement(str: string, javaMode?: boolean): string;
/**
*
* @param {import('./index.js').RE2JS} pattern
* @param {string|number[]|Uint8Array|MatcherInputBase} input
*/
constructor(pattern: RE2JS, input: string | number[] | Uint8Array | MatcherInputBase);
/**
* The pattern being matched.
* @type {import('./index.js').RE2JS}
*/
patternInput: RE2JS;
/** @type {number} */
patternGroupCount: number;
/** @type {number[]} */
groups: number[];
/** @type {Record<string, number>} */
namedGroups: Record<string, number>;
/** @type {number} */
numberOfInstructions: number;
/**
* Returns the {@code RE2JS} associated with this {@code Matcher}.
* @returns {import('./index.js').RE2JS}
*/
pattern(): RE2JS;
/**
* Resets the {@code Matcher}, rewinding input and discarding any match information.
*
* @returns {Matcher} the {@code Matcher} itself, for chained method calls
*/
reset(): Matcher;
/** @type {number} */
matcherInputLength: number;
/** @type {number} */
appendPos: number;
hasMatch: boolean;
hasGroups: boolean;
anchorFlag: number;
/**
* Resets the {@code Matcher} and changes the input.
* @param {string|number[]|Uint8Array|MatcherInputBase} input
* @returns {Matcher} the {@code Matcher} itself, for chained method calls
*/
resetMatcherInput(input: string | number[] | Uint8Array | MatcherInputBase): Matcher;
matcherInput: MatcherInputBase;
/**
* Returns the start of the named group of the most recent match, or -1 if the group was not
* matched.
* @param {string|number} [group=0]
* @returns {number}
*/
start(group?: string | number): number;
/**
* Returns the end of the named group of the most recent match, or -1 if the group was not
* matched.
* @param {string|number} [group=0]
* @returns {number}
*/
end(group?: string | number): number;
/**
* Returns the program size of this pattern.
*
* <p>
* Similar to the C++ implementation, the program size is a very approximate measure of a regexp's
* "cost". Larger numbers are more expensive than smaller numbers.
* </p>
*
* @returns {number} the program size of this pattern
*/
programSize(): number;
/**
* Returns the named group of the most recent match, or {@code null} if the group was not matched.
* @param {string|number} [group=0]
* @returns {string|null}
*/
group(group?: string | number): string | null;
/**
* Returns a dictionary map of all named capturing groups and their matched values.
* If a group was not matched, its value will be `null`.
* @returns {Record<string, string|null>}
*/
getNamedGroups(): Record<string, string | null>;
/**
* Returns the number of subgroups in this pattern.
*
* @returns {number} the number of subgroups; the overall match (group 0) does not count
*/
groupCount(): number;
/**
* Helper: finds subgroup information if needed for group.
* @param {number} group
* @private
*/
private loadGroup;
/**
* Matches the entire input against the pattern (anchored start and end). If there is a match,
* {@code matches} sets the match state to describe it.
*
* @returns {boolean} true if the entire input matches the pattern
*/
matches(): boolean;
/**
* Matches the beginning of input against the pattern (anchored start). If there is a match,
* {@code lookingAt} sets the match state to describe it.
*
* @returns {boolean} true if the beginning of the input matches the pattern
*/
lookingAt(): boolean;
/**
* Matches the input against the pattern (unanchored), starting at a specified position. If there
* is a match, {@code find} sets the match state to describe it.
*
* @param {number|null} [start=null] the input position where the search begins
* @returns {boolean} if it finds a match
* @throws IndexOutOfBoundsException if start is not a valid input position
*/
find(start?: number | null): boolean;
/**
* Helper: does match starting at start, with RE2 anchor flag.
* @param {number} startByte
* @param {number} anchor
* @returns {boolean}
* @private
*/
private genMatch;
/**
* Helper: return substring for [start, end).
* @param {number} start
* @param {number} end
* @returns {string}
*/
substring(start: number, end: number): string;
/**
* Helper for Pattern: return input length.
* @returns {number}
*/
inputLength(): number;
/**
* Appends to result two strings: the text from the append position up to the beginning of the
* most recent match, and then the replacement with submatch groups substituted for references of
* the form {@code $n}, where {@code n} is the group number in decimal. It advances the append
* position to where the most recent match ended.
*
* To embed a literal {@code $}, use \$ (actually {@code "\\$"} with string escapes). The escape
* is only necessary when {@code $} is followed by a digit, but it is always allowed. Only
* {@code $} and {@code \} need escaping, but any character can be escaped.
*
* The group number {@code n} in {@code $n} is always at least one digit and expands to use more
* digits as long as the resulting number is a valid group number for this pattern. To cut it off
* earlier, escape the first digit that should not be used.
*
* @param {string} replacement the replacement string
* @param {boolean} [javaMode=false] activate java mode (different behaviour for capture groups and special characters)
* @returns {string}
* @throws IllegalStateException if there was no most recent match
* @throws IndexOutOfBoundsException if replacement refers to an invalid group
* @private
*/
private appendReplacement;
/**
* @param {string} replacement - the replacement string
* @returns {string}
* @private
*/
private appendReplacementInternalJava;
/**
* @param {string} replacement - the replacement string
* @returns {string}
* @private
*/
private appendReplacementInternalJs;
/**
* Return the substring of the input from the append position to the end of the
* input.
* @returns {string}
*/
appendTail(): string;
/**
* Returns the input with all matches replaced by {@code replacement}, interpreted as for
* {@code appendReplacement}.
*
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string} the input string with the matches replaced
* @throws IndexOutOfBoundsException if replacement refers to an invalid group and javaMode is true
*/
replaceAll(replacement: string | ((...args: any[]) => string), javaMode?: boolean): string;
/**
* Returns the input with the first match replaced by {@code replacement}, interpreted as for
* {@code appendReplacement}.
*
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string} the input string with the first match replaced
* @throws IndexOutOfBoundsException if replacement refers to an invalid group and javaMode is true
*/
replaceFirst(replacement: string | ((...args: any[]) => string), javaMode?: boolean): string;
/**
* Helper: replaceAll/replaceFirst hybrid.
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [all=true] - replace all matches
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string}
* @private
*/
private replace;
/**
* Evaluates a replacer function for the current match and appends the result,
* along with any un-matched preceding text, advancing the append position.
* @param {Function} replacer - the replacer function
* @param {boolean} hasNamedGroups - cached flag if pattern has named groups
* @param {string|Uint8Array|number[]} originalInput - the cached original input reference
* @returns {string} the evaluated string to append
* @private
*/
private appendReplacementFunc;
/**
* Builds the argument array for the replacer function matching the standard
* JS String.prototype.replace(regex, replacer) signature.
* @param {number} matchStart - the start index of the match
* @param {boolean} hasNamedGroups - cached flag if pattern has named groups
* @param {string|Uint8Array|number[]} originalInput - the cached original input reference
* @returns {Array} array of arguments
* @private
*/
private buildReplacerArgs;
}
export class RE2JSException extends Error {
/** @param {string} message */
constructor(message: string);
}
/**
* An exception thrown by the parser if the pattern was invalid.
*/
export class RE2JSSyntaxException extends RE2JSException {
/**
* @param {string} error
* @param {string|null} [input=null]
*/
constructor(error: string, input?: string | null);
/** @type {string} */
error: string;
/** @type {string|null} */
input: string | null;
/**
* Retrieves the description of the error.
* @returns {string}
*/
getDescription(): string;
/**
* Retrieves the erroneous regular-expression pattern.
* @returns {string|null}
*/
getPattern(): string | null;
}
/**
* An exception thrown by the compiler
*/
export class RE2JSCompileException extends RE2JSException {
}
/**
* An exception thrown by using groups
*/
export class RE2JSGroupException extends RE2JSException {
}
/**
* An exception thrown by flags
*/
export class RE2JSFlagsException extends RE2JSException {
}
/**
* An exception thrown for internal engine errors, such as corrupted bytecodes.
*/
export class RE2JSInternalException extends RE2JSException {
}
declare class RE2 {
static initTest(expr: any): RE2;
/**
* Parses a regular expression and returns, if successful, an {@code RE2} instance that can be
* used to match against text.
*
* When matching against text, the regexp returns a match that begins as early as possible in the
* input (leftmost), and among those it chooses the one that a backtracking search would have
* found first. This so-called leftmost-first matching is the same semantics that Perl, Python,
* and other implementations use, although this package implements it without the expense of
* backtracking. For POSIX leftmost-longest matching, see {@link #compilePOSIX}.
*/
static compile(expr: any): RE2;
/**
* {@code compilePOSIX} is like {@link #compile} but restricts the regular expression to POSIX ERE
* (egrep) syntax and changes the match semantics to leftmost-longest.
*
* That is, when matching against text, the regexp returns a match that begins as early as
* possible in the input (leftmost), and among those it chooses a match that is as long as
* possible. This so-called leftmost-longest matching is the same semantics that early regular
* expression implementations used and that POSIX specifies.
*
* However, there can be multiple leftmost-longest matches, with different submatch choices, and
* here this package diverges from POSIX. Among the possible leftmost-longest matches, this
* package chooses the one that a backtracking search would have found first, while POSIX
* specifies that the match be chosen to maximize the length of the first subexpression, then the
* second, and so on from left to right. The POSIX rule is computationally prohibitive and not
* even well-defined. See http://swtch.com/~rsc/regexp/regexp2.html#posix
*/
static compilePOSIX(expr: any): RE2;
static compileImpl(expr: any, mode: any, longest: any): RE2;
/**
* Returns true iff textual regular expression {@code pattern} matches string {@code s}.
*
* More complicated queries need to use {@link #compile} and the full {@code RE2} interface.
*/
static match(pattern: any, s: any): boolean;
constructor(expr: any, prog: any, numSubexp?: number, longest?: number);
expr: any;
prog: any;
numSubexp: number;
longest: number;
cond: any;
prefix: any;
prefixUTF8: any;
prefixComplete: boolean;
prefixRune: number;
machinePool: any[];
dfa: DFA;
onepass: {
start: any;
numCap: any;
inst: any[];
};
prefilter: any;
matchPrefixComplete(input: any, pos: any, anchor: any, ncap: any): number[];
executeEngine(input: any, pos: any, anchor: any, ncap: any): any;
/**
* Returns the number of parenthesized subexpressions in this regular expression.
*/
numberOfCapturingGroups(): number;
/**
* Returns the number of instructions in this compiled regular expression program.
*/
numberOfInstructions(): any;
get(): any;
reset(): void;
put(m: any): void;
toString(): any;
doExecuteNFA(input: any, pos: any, anchor: any, ncap: any): any;
match(s: any): boolean;
/**
* Matches the regular expression against input starting at position start and ending at position
* end, with the given anchoring. Records the submatch boundaries in group, which is [start, end)
* pairs of byte offsets. The number of boundaries needed is inferred from the size of the group
* array. It is most efficient not to ask for submatch boundaries.
*
* @param input the input byte array
* @param start the beginning position in the input
* @param end the end position in the input
* @param anchor the anchoring flag (UNANCHORED, ANCHOR_START, ANCHOR_BOTH)
* @param group the array to fill with submatch positions
* @param ngroup the number of array pairs to fill in
* @returns true if a match was found
*/
matchWithGroup(input: any, start: any, end: any, anchor: any, ngroup: any): any[];
matchMachineInput(input: any, start: any, end: any, anchor: any, ngroup: any): any[];
/**
* Returns true iff this regexp matches the UTF-8 byte array {@code b}.
*/
matchUTF8(b: any): boolean;
/**
* Returns a copy of {@code src} in which all matches for this regexp have been replaced by
* {@code repl}. No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceAll(src: any, repl: any): string;
/**
* Returns a copy of {@code src} in which only the first match for this regexp has been replaced
* by {@code repl}. No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceFirst(src: any, repl: any): string;
/**
* Returns a copy of {@code src} in which at most {@code maxReplaces} matches for this regexp have
* been replaced by the return value of of function {@code repl} (whose first argument is the
* matched string). No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceAllFunc(src: any, replFunc: any, maxReplaces: any): string;
pad(a: any): any;
allMatches(input: any, n: any, deliverFun?: (v: any) => any): any[];
/**
* Returns an array holding the text of the leftmost match in {@code b} of this regular
* expression.
*
* A return value of null indicates no match.
*/
findUTF8(b: any): any;
/**
* Returns a two-element array of integers defining the location of the leftmost match in
* {@code b} of this regular expression. The match itself is at {@code b[loc[0]...loc[1]]}.
*
* A return value of null indicates no match.
*/
findUTF8Index(b: any): any;
/**
* Returns a string holding the text of the leftmost match in {@code s} of this regular
* expression.
*
* If there is no match, the return value is an empty string, but it will also be empty if the
* regular expression successfully matches an empty string. Use {@link #findIndex} or
* {@link #findSubmatch} if it is necessary to distinguish these cases.
*/
find(s: any): any;
/**
* Returns a two-element array of integers defining the location of the leftmost match in
* {@code s} of this regular expression. The match itself is at
* {@code s.substring(loc[0], loc[1])}.
*
* A return value of null indicates no match.
*/
findIndex(s: any): any;
/**
* Returns an array of arrays the text of the leftmost match of the regular expression in
* {@code b} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findUTF8Submatch(b: any): any[];
/**
* Returns an array holding the index pairs identifying the leftmost match of this regular
* expression in {@code b} and the matches, if any, of its subexpressions, as defined by the the
* <a href='#submatch'>Submatch</a> and <a href='#index'>Index</a> descriptions above.
*
* A return value of null indicates no match.
*/
findUTF8SubmatchIndex(b: any): any;
/**
* Returns an array of strings holding the text of the leftmost match of the regular expression in
* {@code s} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findSubmatch(s: any): any[];
/**
* Returns an array holding the index pairs identifying the leftmost match of this regular
* expression in {@code s} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findSubmatchIndex(s: any): any;
/**
* {@code findAllUTF8()} is the <a href='#all'>All</a> version of {@link #findUTF8}; it returns a
* list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*
* TODO(adonovan): think about defining a byte slice view class, like a read-only Go slice backed
* by |b|.
*/
findAllUTF8(b: any, n: any): any[];
/**
* {@code findAllUTF8Index} is the <a href='#all'>All</a> version of {@link #findUTF8Index}; it
* returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8Index(b: any, n: any): any[];
/**
* {@code findAll} is the <a href='#all'>All</a> version of {@link #find}; it returns a list of up
* to {@code n} successive matches of the expression, as defined by the <a href='#all'>All</a>
* description above.
*
* A return value of null indicates no match.
*/
findAll(s: any, n: any): any[];
/**
* {@code findAllIndex} is the <a href='#all'>All</a> version of {@link #findIndex}; it returns a
* list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllIndex(s: any, n: any): any[];
/**
* {@code findAllUTF8Submatch} is the <a href='#all'>All</a> version of {@link #findUTF8Submatch};
* it returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8Submatch(b: any, n: any): any[];
/**
* {@code findAllUTF8SubmatchIndex} is the <a href='#all'>All</a> version of
* {@link #findUTF8SubmatchIndex}; it returns a list of up to {@code n} successive matches of the
* expression, as defined by the <a href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8SubmatchIndex(b: any, n: any): any[];
/**
* {@code findAllSubmatch} is the <a href='#all'>All</a> version of {@link #findSubmatch}; it
* returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllSubmatch(s: any, n: any): any[];
/**
* {@code findAllSubmatchIndex} is the <a href='#all'>All</a> version of
* {@link #findSubmatchIndex}; it returns a list of up to {@code n} successive matches of the
* expression, as defined by the <a href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllSubmatchIndex(s: any, n: any): any[];
}
/**
* Creates an RE2JS regex directly from a template literal.
* @overload
* @param {TemplateStringsArray} stringsOrFlags - The raw string segments of the template literal.
* @param {...any} values - The interpolated values.
* @returns {RE2JS}
*/
export function re(stringsOrFlags: TemplateStringsArray, ...values: any[]): RE2JS;
/**
* Creates a template literal tag function with specific RE2JS flags.
* @overload
* @param {number} stringsOrFlags - The RE2JS flags to apply (e.g., RE2JS.CASE_INSENSITIVE).
* @returns {(strings: TemplateStringsArray, ...tagValues: any[]) => RE2JS}
*/
export function re(stringsOrFlags: number): (strings: TemplateStringsArray, ...tagValues: any[]) => RE2JS;
/**
* A compiled representation of an RE2 regular expression
*
* The matching functions take {@code String} arguments instead of the more general Java
* {@code CharSequence} since the latter doesn't provide UTF-16 decoding.
*
*
* @author rsc@google.com (Russ Cox)
* @class
*/
export class RE2JS {
/**
* Flag: case insensitive matching.
*/
static CASE_INSENSITIVE: number;
/**
* Flag: dot ({@code .}) matches all characters, including newline.
*/
static DOTALL: number;
/**
* Flag: multiline matching: {@code ^} and {@code $} match at beginning and end of line, not just
* beginning and end of input.
*/
static MULTILINE: number;
/**
* Flag: Unicode groups (e.g. {@code \p\ Greek\} ) will be syntax errors.
*/
static DISABLE_UNICODE_GROUPS: number;
/**
* Flag: matches longest possible string.
*/
static LONGEST_MATCH: number;
/**
* Flag: enable linear-time captureless lookbehinds.
*/
static LOOKBEHINDS: number;
/**
* Returns a literal pattern string for the specified string.
*
* This method produces a string that can be used to create a <code>RE2JS</code> that would
* match the string <code>s</code> as if it were a literal pattern.
*
* Metacharacters or escape sequences in the input sequence will be given no special meaning.
*
* @param {string} str The string to be literalized
* @returns {string} A literal string replacement
*/
static quote(str: string): string;
/**
* Quotes '\' and '$' in {@code str}, so that the returned string could be used in
* replacement methods as a literal replacement of {@code str}.
*
* This is a convenience delegation to {@link Matcher.quoteReplacement}.
*
* @param {string} str the string to be quoted
* @param {boolean} [javaMode=false] whether the replacement will be used in javaMode
* @returns {string} the quoted string
*/
static quoteReplacement(str: string, javaMode?: boolean): string;
/**
* Translates a given regular expression string to ensure compatibility with RE2JS.
*
* This function preprocesses the input regex string by applying necessary transformations,
* such as escaping special characters (e.g., `/`), converting named capture groups to
* RE2JS-compatible syntax, and handling Unicode sequences properly. It ensures that the
* resulting regex is safe and properly formatted before compilation.
*
* @param {string|RegExp} expr - The regular expression string to be translated.
* @returns {string} - The transformed regular expression string, ready for compilation.
*/
static translateRegExp(expr: string | RegExp): string;
/**
* Helper: create new RE2JS with given regex and flags. Flregex is the regex with flags applied.
* @param {string} regex
* @param {number} [flags=0]
* @returns {RE2JS}
*/
static compile(regex: string, flags?: number): RE2JS;
/**
* Matches a string against a regular expression.
*
* @param {string} regex the regular expression
* @param {string|number[]|Uint8Array} input the input
* @returns {boolean} true if the regular expression matches the entire input
* @throws RE2JSSyntaxException if the regular expression is malformed
*/
static matches(regex: string, input: string | number[] | Uint8Array): boolean;
/**
* This is visible for testing.
* @private
*/
private static initTest;
/**
*
* @param {string} pattern
* @param {number} flags
*/
constructor(pattern: string, flags: number);
patternInput: string;
flagsInput: number;
/** @type {import('./RE2.js').RE2} */
re2Input: RE2;
/**
* Releases memory used by internal caches associated with this pattern. Does not change the
* observable behaviour. Useful for tests that detect memory leaks via allocation tracking.
*/
reset(): void;
/**
* Returns the flags used in the constructor.
* @returns {number}
*/
flags(): number;
/**
* Returns the pattern used in the constructor.
* @returns {string}
*/
pattern(): string;
re2(): RE2;
/**
* Matches a string against a regular expression.
*
* @param {string|number[]|Uint8Array} input the input
* @returns {boolean} true if the regular expression matches the entire input
*/
matches(input: string | number[] | Uint8Array): boolean;
/**
* Creates a new {@code Matcher} matching the pattern against the input.
*
* @param {string|number[]|Uint8Array|MatcherInputBase} input the input string
* @returns {Matcher}
*/
matcher(input: string | number[] | Uint8Array | MatcherInputBase): Matcher;
/**
* Tests whether the regular expression matches any part of the input string.
* Performance Note: This method is highly optimized. Because it only returns
* a boolean and does not extract capture groups, it bypasses the `Matcher` overhead
* and guarantees execution on the high-speed DFA engine whenever possible.
*
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to test against.
* @returns {boolean} `true` if the pattern is found anywhere in the input, `false` otherwise.
*/
test(input: string | number[] | Uint8Array): boolean;
/**
* Tests whether the regular expression matches the ENTIRE input string.
* * **Performance Note:** This operates identically to `.matches()`, but is significantly
* faster because it does not request capture group data. By requesting 0 capture groups,
* it securely routes execution through the DFA fast-path.
*
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to test against.
* @returns {boolean} `true` if the exact input string fully matches the pattern, `false` otherwise.
*/
testExact(input: string | number[] | Uint8Array): boolean;
/**
* Executes a search for a match in a specified string.
* Returns a result array, or null if no match is found.
* The returned array perfectly mirrors standard JavaScript `RegExpExecArray`,
* including `.index`, `.input`, and `.groups` properties.
*
* @param {string|number[]|Uint8Array} input the input string or byte array
* @returns {Array|null} the match array with index, input, and groups properties, or null
*/
exec(input: string | number[] | Uint8Array): any[] | null;
/**
* Splits input around instances of the regular expression. It returns an array giving the strings
* that occur before, between, and after instances of the regular expression.
*
* If {@code limit <= 0}, there is no limit on the size of the returned array. If
* {@code limit == 0}, empty strings that would occur at the end of the array are omitted. If
* {@code limit > 0}, at most limit strings are returned. The final string contains the remainder
* of the input, possibly including additional matches of the pattern.
*
* @param {string} input the input string to be split
* @param {number} [limit=0] the limit
* @returns {string[]} the split strings
*/
split(input: string, limit?: number): string[];
/**
* Returns an iterator of all results matching a string against the regular expression,
* including capturing groups.
*
* @param {string|number[]|Uint8Array} input the input string or byte array
* @returns {IterableIterator<RegExpMatchArray>}
*/
matchAll(input: string | number[] | Uint8Array): IterableIterator<RegExpMatchArray>;
/**
*
* @returns {string}
*/
toString(): string;
/**
* Returns the program size of this pattern.
*
* <p>
* Similar to the C++ implementation, the program size is a very approximate measure of a regexp's
* "cost". Larger numbers are more expensive than smaller numbers.
* </p>
*
* @returns {number} the program size of this pattern
*/
programSize(): number;
/**
* Returns the number of capturing groups in this matcher's pattern. Group zero denotes the entire
* pattern and is excluded from this count.
*
* @returns {number} the number of capturing groups in this pattern
*/
groupCount(): number;
/**
* Return a map of the capturing groups in this matcher's pattern, where key is the name and value
* is the index of the group in the pattern.
* @returns {Record<string, number>}
*/
namedGroups(): Record<string, number>;
/**
*
* @param {*} other
* @returns {boolean}
*/
equals(other: any): boolean;
}
export {};

10
package-lock.json generated
View File

@@ -10,6 +10,7 @@
"license": "MIT",
"dependencies": {
"@anthropic-ai/claude-agent-sdk": "^0.2.84",
"re2js": "^2.8.6",
"ws": "^8.21.0"
},
"bin": {
@@ -4531,6 +4532,15 @@
"node": ">= 0.10"
}
},
"node_modules/re2js": {
"version": "2.8.6",
"resolved": "https://registry.npmjs.org/re2js/-/re2js-2.8.6.tgz",
"integrity": "sha512-xLgQil4kIUCrAzVk9fRSkxkFNwmygLFjVxXrLc65aE1F0+Zsb8rxumFBy4XKyvgMCTL6kilDq3EZ0piE2dP/Dg==",
"license": "MIT",
"engines": {
"node": ">=18.0.0"
}
},
"node_modules/require-directory": {
"version": "2.1.1",
"resolved": "https://registry.npmjs.org/require-directory/-/require-directory-2.1.1.tgz",

View File

@@ -74,6 +74,7 @@
"fast-check": "^4.8.0",
"globals": "^16.5.0",
"js-yaml": "^4.3.1",
"re2js": "^2.8.6",
"typescript": "^6.0.3",
"typescript-eslint": "^8.60.0"
},
@@ -116,7 +117,7 @@
"lint:table-schema-drift": "node scripts/lint-table-schema-drift.cjs",
"lint:frontmatter-scalar-broad-grep": "node scripts/lint-frontmatter-scalar-broad-grep.cjs",
"lint:removed-but-needed": "node scripts/lint-removed-but-needed.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs && node scripts/lint-vendored-deps.cjs",
"lint:allow-test-rule-refs": "node scripts/lint-allow-test-rule-refs.cjs",
"lint:regression-names": "node scripts/lint-regression-test-names.cjs",
"lint:descriptions": "node scripts/lint-descriptions.cjs",

View File

@@ -1,4 +1,12 @@
[
{
"path": "gsd-core/bin/lib/vendor/re2js.d.cts",
"reason": "Vendored third-party type declaration (#3477): a .d.cts carries no executable code, so no ESLint rule is meaningful; the vendored .cjs it describes is globally ignored as verbatim upstream output. Freshness is enforced by scripts/lint-vendored-deps.cjs, not by lint rules."
},
{
"path": "src/vendor/re2js.d.cts",
"reason": "Vendored third-party type declaration (#3477): a .d.cts carries no executable code, so no ESLint rule is meaningful; the vendored .cjs it describes is globally ignored as verbatim upstream output. Freshness is enforced by scripts/lint-vendored-deps.cjs, not by lint rules."
},
{
"path": "tests/fixtures/brand-typing/bad-calibrated-as-sample-basis.cts",
"reason": "Deliberate MUST-NOT-COMPILE type-error fixture (#3059): type-aware linting would fail by design; it exists to prove the compiler rejects it."

View File

@@ -0,0 +1,124 @@
#!/usr/bin/env node
'use strict';
/**
* lint-vendored-deps.cjs — freshness gate for gsd-core/bin/lib/vendor/.
*
* #3477 follow-up: gsd-core/bin/** is copied by the installer into trees
* that have NO node_modules, so it must carry zero external requires
* (local/no-external-require-in-bin, eslint-rules/no-external-require-in-bin.cjs).
* `re2js` (src/pattern.cts's RE2 engine) is vendored verbatim under
* gsd-core/bin/lib/vendor/ instead — see gsd-core/bin/lib/vendor/README.md.
*
* A vendored artifact that silently drifts from its upstream package is
* just as dangerous as never vendoring it in the first place (a stale
* copy ships a different engine than the one actually reviewed/audited).
* This guard fails CI when:
* 1. gsd-core/bin/lib/vendor/re2js.cjs no longer matches
* node_modules/re2js/build/index.cjs byte-for-byte.
* 2. gsd-core/bin/lib/vendor/re2js.d.cts no longer matches
* node_modules/re2js/build/index.d.cts byte-for-byte.
* 3. src/vendor/re2js.d.cts (the source-side twin tsc needs to resolve
* types for src/pattern.cts's relative './vendor/re2js.cjs' import —
* module resolution for a .cts source is relative to src/, not the
* output dir) no longer matches gsd-core/bin/lib/vendor/re2js.d.cts.
* 4. The `re2js` version pinned in package.json `devDependencies` no
* longer matches the version actually installed at
* node_modules/re2js/package.json (read there, per the dispatch
* brief, rather than duplicating a second pin).
*
* Usage: node scripts/lint-vendored-deps.cjs
* Exit 0 when every vendored copy is fresh; 1 otherwise.
*/
const fs = require('node:fs');
const path = require('node:path');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
const ROOT = path.join(__dirname, '..');
const REFRESH_COMMAND =
'cp node_modules/re2js/build/index.cjs gsd-core/bin/lib/vendor/re2js.cjs && '
+ 'cp node_modules/re2js/build/index.d.cts gsd-core/bin/lib/vendor/re2js.d.cts && '
+ 'cp node_modules/re2js/build/index.d.cts src/vendor/re2js.d.cts';
/**
* Compare two files byte-for-byte. Returns null when equal, or a short
* mismatch description (missing file / byte-length delta) otherwise.
* @param {string} relA
* @param {string} relB
* @returns {string | null}
*/
function compareFiles(relA, relB) {
const absA = path.join(ROOT, relA);
const absB = path.join(ROOT, relB);
if (!fs.existsSync(absA)) return `${relA} does not exist`;
if (!fs.existsSync(absB)) return `${relB} does not exist`;
const a = fs.readFileSync(absA);
const b = fs.readFileSync(absB);
if (a.equals(b)) return null;
return `${relA} (${a.length} bytes) != ${relB} (${b.length} bytes)`;
}
/**
* Strip a leading semver range operator (^, ~, >=, >, <=, <, =) from a
* package.json dependency spec, leaving a bare version.
* @param {string} spec
* @returns {string}
*/
function stripRangeOperator(spec) {
return String(spec || '').trim().replace(/^[\^~]|^>=|^<=|^>|^<|^=/, '').trim();
}
function main() {
const findings = [];
const cjsDrift = compareFiles('gsd-core/bin/lib/vendor/re2js.cjs', 'node_modules/re2js/build/index.cjs');
if (cjsDrift) findings.push(cjsDrift);
const dctsDrift = compareFiles('gsd-core/bin/lib/vendor/re2js.d.cts', 'node_modules/re2js/build/index.d.cts');
if (dctsDrift) findings.push(dctsDrift);
const srcTwinDrift = compareFiles('src/vendor/re2js.d.cts', 'gsd-core/bin/lib/vendor/re2js.d.cts');
if (srcTwinDrift) findings.push(srcTwinDrift);
const pkgPath = path.join(ROOT, 'package.json');
const pkg = JSON.parse(fs.readFileSync(pkgPath, 'utf8'));
const pinnedSpec = pkg.devDependencies && pkg.devDependencies.re2js;
if (!pinnedSpec) {
findings.push('package.json devDependencies.re2js is missing');
} else {
const installedPkgPath = path.join(ROOT, 'node_modules', 're2js', 'package.json');
if (!fs.existsSync(installedPkgPath)) {
findings.push('node_modules/re2js/package.json does not exist (run npm install)');
} else {
const installed = JSON.parse(fs.readFileSync(installedPkgPath, 'utf8'));
const pinned = stripRangeOperator(pinnedSpec);
if (pinned !== installed.version) {
findings.push(
`package.json devDependencies.re2js ("${pinnedSpec}" -> "${pinned}") != `
+ `node_modules/re2js/package.json version ("${installed.version}")`,
);
}
}
}
if (findings.length > 0) {
const detail = findings.map((f) => ` ${f}`).join('\n');
throw new ExitError(
1,
'lint-vendored-deps: gsd-core/bin/lib/vendor/re2js.* has drifted from its\n'
+ 'upstream package (or its version pin). Refresh with:\n'
+ ` ${REFRESH_COMMAND}\n`
+ 'Findings:\n'
+ detail,
);
}
process.stdout.write('ok lint-vendored-deps: gsd-core/bin/lib/vendor/re2js.* matches node_modules/re2js and its pinned version\n');
return 0;
}
if (require.main === module) runMain(main);
module.exports = { compareFiles, stripRangeOperator };

View File

@@ -41,6 +41,12 @@
* migration-equivalence property sweep in tests/pattern.test.cjs (rows 15-17).
*/
// re2js is vendored, not an npm dependency at runtime: gsd-core/bin/** is
// copied into installed trees that have no node_modules, so this module must
// carry zero external requires (eslint-rules/no-external-require-in-bin.cjs
// enforces it). See gsd-core/bin/lib/vendor/README.md.
import { RE2JS } from './vendor/re2js.cjs';
// #3498: RegExp.escape is ES2026 (first shipped in Node 24). The gsd-test
// matrix still runs a linux-node22 lane, and the build itself consumes this
// module (scripts/gen-loop-host-contract.cjs), so a hard dependency breaks
@@ -64,3 +70,66 @@ export function escapeRegex(value: string): string {
export function literalPattern(value: string, flags?: string): RegExp {
return new RegExp(escapeRegex(value), flags);
}
/** Max length for a user-supplied regex pattern before it is refused (ReDoS/compile-cost mitigation). */
export const MAX_USER_PATTERN_LEN = 512;
/** Reason an untrusted pattern was refused and never compiled, or `null` if it compiled. */
export type UserPatternNeutralization = 'empty' | 'too-long' | 'unsupported';
export interface UserPatternResult {
/** Linear-time match via RE2. Returns false for every neutralized pattern, by construction. */
test(input: string): boolean;
/** null when the pattern compiled; otherwise why it was refused. */
neutralized: UserPatternNeutralization | null;
}
/** Always-false matcher shared by every neutralization path — a refused pattern must never be able to report a match. */
const NEVER_MATCH: UserPatternResult['test'] = () => false;
/**
* Compile an UNTRUSTED, user-supplied pattern (e.g. plan frontmatter) via RE2
* (re2js), whose matching is linear-time in input length by construction —
* there is no backtracking engine here to exploit, so the vulnerability class
* (catastrophic/exponential backtracking) is closed by the engine rather than
* detected by a heuristic scan of the pattern text.
*
* Never throws. A pattern that is empty, too long, or not valid RE2 syntax
* (backreferences and look-around are unsupported by RE2 — those are exactly
* the constructs that require backtracking) is REFUSED: `test()` always
* returns `false`, and `neutralized` reports why so callers can surface the
* refusal (#3477 follow-up: a neutralized pattern must not look like a plain
* "not found"). Restores the guards lost with
* `sdk/src/query/validate.ts:regexForKeyLinkPattern` (#3477); this revision
* (post-#3477-follow-up) replaces the hand-rolled backtracking-shape scanner
* with RE2's linear-time guarantee — a refused pattern is never re-attempted
* as a literal-escaped match, since guessing at a pattern we could not
* compile is what produced the prior false-pass regression.
*
* `pattern` is `unknown`, not `string`, because callers pull this straight off
* parsed plan frontmatter (untrusted YAML) — a non-string value must reach the
* `'empty'`/never-match branch rather than being force-cast by the caller.
*/
export function compileUserPattern(pattern: unknown): UserPatternResult {
if (typeof pattern !== 'string' || pattern.length === 0) {
return { test: NEVER_MATCH, neutralized: 'empty' };
}
if (pattern.length > MAX_USER_PATTERN_LEN) {
// The cap now bounds compile cost/memory, not backtracking (RE2 has none) —
// an over-long pattern is refused outright rather than truncated-and-compiled.
return { test: NEVER_MATCH, neutralized: 'too-long' };
}
try {
// translateRegExp accepts JS-flavored syntax (named groups, `/`-escaping,
// etc.) that RE2's own grammar doesn't, reducing spurious refusals of
// otherwise-safe, JS-authored patterns before compiling under RE2's
// linear-time engine.
const compiled = RE2JS.compile(RE2JS.translateRegExp(pattern));
return { test: (input: string) => compiled.test(input), neutralized: null };
} catch {
// Backreferences, look-around, or any other RE2-unsupported/malformed
// syntax. No literal-escape fallback: a pattern we could not compile is
// never guessed at — guessing produced the #3477 false-pass regression.
return { test: NEVER_MATCH, neutralized: 'unsupported' };
}
}

938
src/vendor/re2js.d.cts vendored Normal file
View File

@@ -0,0 +1,938 @@
// Generated by dts-bundle-generator v9.5.1
declare class DFA {
static MAX_CACHE_CLEARS: number;
static STATE_MEMORY_ESTIMATE: number;
constructor(prog: any, maxMem?: number);
prog: any;
stateCache: Map<any, any>;
stateCount: number;
startState: any;
stateLimit: number;
cacheClears: number;
failed: boolean;
clock: number;
computeClosure(pcs: any): {
pcs: Int32Array<ArrayBuffer>;
isMatch: boolean;
matchIDs: any[];
};
getState(pcs: any): any;
evictCache(): void;
step(state: any, charCode: any, anchor: any): any;
match(input: any, pos: any, anchor: any): boolean;
matchSet(input: any, pos: any, anchor: any): any[];
}
declare class Prog {
inst: any[];
start: number;
numCap: number;
lbStarts: any[];
numLb: number;
getInst(pc: any): any;
numInst(): number;
addInst(op: any): void;
skipNop(pc: any): any;
prefix(): (string | boolean)[];
startCond(): number;
patch(l: any, val: any): void;
append(l1: any, l2: any): any;
/**
*
* @returns {string}
*/
toString(): string;
}
export class RE2Set {
/** @type {number} */
static UNANCHORED: number;
/** @type {number} */
static ANCHOR_START: number;
/** @type {number} */
static ANCHOR_BOTH: number;
/**
* Constructs a new RE2Set with the specified anchor mode and flags.
* @param {number} [anchor=RE2Set.UNANCHORED] - The anchoring mode (e.g., RE2Set.UNANCHORED).
* @param {number} [flags=0] - The public flags to apply to all patterns in the set.
* @param {number} [maxMem=8388608] - The maximum memory in bytes to use for the DFA (default 8MB).
*/
constructor(anchor?: number, flags?: number, maxMem?: number);
anchor: number;
jsFlags: number;
maxMem: number;
re2Flags: number;
regexps: any[];
prog: Prog;
dfa: DFA;
dummyRe2: {
prog: Prog;
cond: number;
prefix: string;
prefixRune: number;
longest: boolean;
};
/**
* Adds a new regular expression pattern to the set.
* Patterns cannot be added after the set has been compiled.
* @param {string} pattern - The regular expression pattern to add.
* @returns {number} The integer index assigned to the added pattern.
* @throws {RE2JSCompileException} If patterns are added after compilation.
*/
add(pattern: string): number;
/**
* Compiles the added patterns into a single state machine.
* This is automatically called on the first match if not called explicitly.
* @returns {void}
*/
compile(): void;
/**
* Matches the input against the compiled set of regular expressions.
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to match against.
* @returns {number[]} An array of indices representing the patterns that successfully matched the input.
*/
match(input: string | number[] | Uint8Array): number[];
}
export class MatcherInput {
/**
* Return the MatcherInput for UTF_16 encoding.
* @returns {Utf16MatcherInput}
*/
static utf16(charSequence: any): Utf16MatcherInput;
/**
* Return the MatcherInput for UTF_8 encoding.
* @returns {Utf8MatcherInput}
*/
static utf8(input: any): Utf8MatcherInput;
}
/**
* Abstract the representations of input text supplied to Matcher.
*/
export class MatcherInputBase {
static Encoding: any;
getEncoding(): void;
/** @returns {string} */
asCharSequence(): string;
/** @returns {Uint8Array|number[]} */
asBytes(): Uint8Array | number[];
/** @returns {number} */
length(): number;
/**
*
* @returns {boolean}
*/
isUTF8Encoding(): boolean;
/**
*
* @returns {boolean}
*/
isUTF16Encoding(): boolean;
}
declare class Utf16MatcherInput extends MatcherInputBase {
/** @param {string|null} charSequence */
constructor(charSequence?: string | null);
charSequence: string;
getEncoding(): any;
/**
*
* @returns {number[]}
*/
asBytes(): number[];
}
declare class Utf8MatcherInput extends MatcherInputBase {
/** @param {Uint8Array|number[]|null} bytes */
constructor(bytes?: Uint8Array | number[] | null);
bytes: number[] | Uint8Array<ArrayBufferLike>;
getEncoding(): any;
}
/**
* A stateful iterator that interprets a regex {@code RE2JS} on a specific input.
*
* Conceptually, a Matcher consists of four parts:
* <ol>
* <li>A compiled regular expression {@code RE2JS}, set at construction and fixed for the lifetime
* of the matcher.</li>
*
* <li>The remainder of the input string, set at construction or {@link #reset()} and advanced by
* each match operation such as {@link #find}, {@link #matches} or {@link #lookingAt}.</li>
*
* <li>The current match information, accessible via {@link #start}, {@link #end}, and
* {@link #group}, and updated by each match operation.</li>
*
* <li>The append position, used and advanced by {@link #appendReplacement} and {@link #appendTail}
* if performing a search and replace from the input to an external {@code StringBuffer}.
*
* </ol>
*
*
* @author rsc@google.com (Russ Cox)
*/
export class Matcher {
/**
* V8 and WebKit have historical hard limits on the number of arguments
* that can be passed to a function. We cap replacer arguments to prevent
* Call Stack Overflow (DoS) vulnerabilities on massive ASTs.
*/
static MAX_REPLACER_ARGS: number;
/**
* Quotes '\' and '$' in {@code s}, so that the returned string could be used in
* {@link #appendReplacement} as a literal replacement of {@code s}.
*
* @param {string} str the string to be quoted
* @param {boolean} [javaMode=false] whether the replacement will be used in javaMode
* @returns {string} the quoted string
*/
static quoteReplacement(str: string, javaMode?: boolean): string;
/**
*
* @param {import('./index.js').RE2JS} pattern
* @param {string|number[]|Uint8Array|MatcherInputBase} input
*/
constructor(pattern: RE2JS, input: string | number[] | Uint8Array | MatcherInputBase);
/**
* The pattern being matched.
* @type {import('./index.js').RE2JS}
*/
patternInput: RE2JS;
/** @type {number} */
patternGroupCount: number;
/** @type {number[]} */
groups: number[];
/** @type {Record<string, number>} */
namedGroups: Record<string, number>;
/** @type {number} */
numberOfInstructions: number;
/**
* Returns the {@code RE2JS} associated with this {@code Matcher}.
* @returns {import('./index.js').RE2JS}
*/
pattern(): RE2JS;
/**
* Resets the {@code Matcher}, rewinding input and discarding any match information.
*
* @returns {Matcher} the {@code Matcher} itself, for chained method calls
*/
reset(): Matcher;
/** @type {number} */
matcherInputLength: number;
/** @type {number} */
appendPos: number;
hasMatch: boolean;
hasGroups: boolean;
anchorFlag: number;
/**
* Resets the {@code Matcher} and changes the input.
* @param {string|number[]|Uint8Array|MatcherInputBase} input
* @returns {Matcher} the {@code Matcher} itself, for chained method calls
*/
resetMatcherInput(input: string | number[] | Uint8Array | MatcherInputBase): Matcher;
matcherInput: MatcherInputBase;
/**
* Returns the start of the named group of the most recent match, or -1 if the group was not
* matched.
* @param {string|number} [group=0]
* @returns {number}
*/
start(group?: string | number): number;
/**
* Returns the end of the named group of the most recent match, or -1 if the group was not
* matched.
* @param {string|number} [group=0]
* @returns {number}
*/
end(group?: string | number): number;
/**
* Returns the program size of this pattern.
*
* <p>
* Similar to the C++ implementation, the program size is a very approximate measure of a regexp's
* "cost". Larger numbers are more expensive than smaller numbers.
* </p>
*
* @returns {number} the program size of this pattern
*/
programSize(): number;
/**
* Returns the named group of the most recent match, or {@code null} if the group was not matched.
* @param {string|number} [group=0]
* @returns {string|null}
*/
group(group?: string | number): string | null;
/**
* Returns a dictionary map of all named capturing groups and their matched values.
* If a group was not matched, its value will be `null`.
* @returns {Record<string, string|null>}
*/
getNamedGroups(): Record<string, string | null>;
/**
* Returns the number of subgroups in this pattern.
*
* @returns {number} the number of subgroups; the overall match (group 0) does not count
*/
groupCount(): number;
/**
* Helper: finds subgroup information if needed for group.
* @param {number} group
* @private
*/
private loadGroup;
/**
* Matches the entire input against the pattern (anchored start and end). If there is a match,
* {@code matches} sets the match state to describe it.
*
* @returns {boolean} true if the entire input matches the pattern
*/
matches(): boolean;
/**
* Matches the beginning of input against the pattern (anchored start). If there is a match,
* {@code lookingAt} sets the match state to describe it.
*
* @returns {boolean} true if the beginning of the input matches the pattern
*/
lookingAt(): boolean;
/**
* Matches the input against the pattern (unanchored), starting at a specified position. If there
* is a match, {@code find} sets the match state to describe it.
*
* @param {number|null} [start=null] the input position where the search begins
* @returns {boolean} if it finds a match
* @throws IndexOutOfBoundsException if start is not a valid input position
*/
find(start?: number | null): boolean;
/**
* Helper: does match starting at start, with RE2 anchor flag.
* @param {number} startByte
* @param {number} anchor
* @returns {boolean}
* @private
*/
private genMatch;
/**
* Helper: return substring for [start, end).
* @param {number} start
* @param {number} end
* @returns {string}
*/
substring(start: number, end: number): string;
/**
* Helper for Pattern: return input length.
* @returns {number}
*/
inputLength(): number;
/**
* Appends to result two strings: the text from the append position up to the beginning of the
* most recent match, and then the replacement with submatch groups substituted for references of
* the form {@code $n}, where {@code n} is the group number in decimal. It advances the append
* position to where the most recent match ended.
*
* To embed a literal {@code $}, use \$ (actually {@code "\\$"} with string escapes). The escape
* is only necessary when {@code $} is followed by a digit, but it is always allowed. Only
* {@code $} and {@code \} need escaping, but any character can be escaped.
*
* The group number {@code n} in {@code $n} is always at least one digit and expands to use more
* digits as long as the resulting number is a valid group number for this pattern. To cut it off
* earlier, escape the first digit that should not be used.
*
* @param {string} replacement the replacement string
* @param {boolean} [javaMode=false] activate java mode (different behaviour for capture groups and special characters)
* @returns {string}
* @throws IllegalStateException if there was no most recent match
* @throws IndexOutOfBoundsException if replacement refers to an invalid group
* @private
*/
private appendReplacement;
/**
* @param {string} replacement - the replacement string
* @returns {string}
* @private
*/
private appendReplacementInternalJava;
/**
* @param {string} replacement - the replacement string
* @returns {string}
* @private
*/
private appendReplacementInternalJs;
/**
* Return the substring of the input from the append position to the end of the
* input.
* @returns {string}
*/
appendTail(): string;
/**
* Returns the input with all matches replaced by {@code replacement}, interpreted as for
* {@code appendReplacement}.
*
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string} the input string with the matches replaced
* @throws IndexOutOfBoundsException if replacement refers to an invalid group and javaMode is true
*/
replaceAll(replacement: string | ((...args: any[]) => string), javaMode?: boolean): string;
/**
* Returns the input with the first match replaced by {@code replacement}, interpreted as for
* {@code appendReplacement}.
*
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string} the input string with the first match replaced
* @throws IndexOutOfBoundsException if replacement refers to an invalid group and javaMode is true
*/
replaceFirst(replacement: string | ((...args: any[]) => string), javaMode?: boolean): string;
/**
* Helper: replaceAll/replaceFirst hybrid.
* @param {string|((...args: any[]) => string)} replacement - the replacement string or a replacer function
* @param {boolean} [all=true] - replace all matches
* @param {boolean} [javaMode=false] - activate java mode (different behaviour for capture groups and special characters)
* @returns {string}
* @private
*/
private replace;
/**
* Evaluates a replacer function for the current match and appends the result,
* along with any un-matched preceding text, advancing the append position.
* @param {Function} replacer - the replacer function
* @param {boolean} hasNamedGroups - cached flag if pattern has named groups
* @param {string|Uint8Array|number[]} originalInput - the cached original input reference
* @returns {string} the evaluated string to append
* @private
*/
private appendReplacementFunc;
/**
* Builds the argument array for the replacer function matching the standard
* JS String.prototype.replace(regex, replacer) signature.
* @param {number} matchStart - the start index of the match
* @param {boolean} hasNamedGroups - cached flag if pattern has named groups
* @param {string|Uint8Array|number[]} originalInput - the cached original input reference
* @returns {Array} array of arguments
* @private
*/
private buildReplacerArgs;
}
export class RE2JSException extends Error {
/** @param {string} message */
constructor(message: string);
}
/**
* An exception thrown by the parser if the pattern was invalid.
*/
export class RE2JSSyntaxException extends RE2JSException {
/**
* @param {string} error
* @param {string|null} [input=null]
*/
constructor(error: string, input?: string | null);
/** @type {string} */
error: string;
/** @type {string|null} */
input: string | null;
/**
* Retrieves the description of the error.
* @returns {string}
*/
getDescription(): string;
/**
* Retrieves the erroneous regular-expression pattern.
* @returns {string|null}
*/
getPattern(): string | null;
}
/**
* An exception thrown by the compiler
*/
export class RE2JSCompileException extends RE2JSException {
}
/**
* An exception thrown by using groups
*/
export class RE2JSGroupException extends RE2JSException {
}
/**
* An exception thrown by flags
*/
export class RE2JSFlagsException extends RE2JSException {
}
/**
* An exception thrown for internal engine errors, such as corrupted bytecodes.
*/
export class RE2JSInternalException extends RE2JSException {
}
declare class RE2 {
static initTest(expr: any): RE2;
/**
* Parses a regular expression and returns, if successful, an {@code RE2} instance that can be
* used to match against text.
*
* When matching against text, the regexp returns a match that begins as early as possible in the
* input (leftmost), and among those it chooses the one that a backtracking search would have
* found first. This so-called leftmost-first matching is the same semantics that Perl, Python,
* and other implementations use, although this package implements it without the expense of
* backtracking. For POSIX leftmost-longest matching, see {@link #compilePOSIX}.
*/
static compile(expr: any): RE2;
/**
* {@code compilePOSIX} is like {@link #compile} but restricts the regular expression to POSIX ERE
* (egrep) syntax and changes the match semantics to leftmost-longest.
*
* That is, when matching against text, the regexp returns a match that begins as early as
* possible in the input (leftmost), and among those it chooses a match that is as long as
* possible. This so-called leftmost-longest matching is the same semantics that early regular
* expression implementations used and that POSIX specifies.
*
* However, there can be multiple leftmost-longest matches, with different submatch choices, and
* here this package diverges from POSIX. Among the possible leftmost-longest matches, this
* package chooses the one that a backtracking search would have found first, while POSIX
* specifies that the match be chosen to maximize the length of the first subexpression, then the
* second, and so on from left to right. The POSIX rule is computationally prohibitive and not
* even well-defined. See http://swtch.com/~rsc/regexp/regexp2.html#posix
*/
static compilePOSIX(expr: any): RE2;
static compileImpl(expr: any, mode: any, longest: any): RE2;
/**
* Returns true iff textual regular expression {@code pattern} matches string {@code s}.
*
* More complicated queries need to use {@link #compile} and the full {@code RE2} interface.
*/
static match(pattern: any, s: any): boolean;
constructor(expr: any, prog: any, numSubexp?: number, longest?: number);
expr: any;
prog: any;
numSubexp: number;
longest: number;
cond: any;
prefix: any;
prefixUTF8: any;
prefixComplete: boolean;
prefixRune: number;
machinePool: any[];
dfa: DFA;
onepass: {
start: any;
numCap: any;
inst: any[];
};
prefilter: any;
matchPrefixComplete(input: any, pos: any, anchor: any, ncap: any): number[];
executeEngine(input: any, pos: any, anchor: any, ncap: any): any;
/**
* Returns the number of parenthesized subexpressions in this regular expression.
*/
numberOfCapturingGroups(): number;
/**
* Returns the number of instructions in this compiled regular expression program.
*/
numberOfInstructions(): any;
get(): any;
reset(): void;
put(m: any): void;
toString(): any;
doExecuteNFA(input: any, pos: any, anchor: any, ncap: any): any;
match(s: any): boolean;
/**
* Matches the regular expression against input starting at position start and ending at position
* end, with the given anchoring. Records the submatch boundaries in group, which is [start, end)
* pairs of byte offsets. The number of boundaries needed is inferred from the size of the group
* array. It is most efficient not to ask for submatch boundaries.
*
* @param input the input byte array
* @param start the beginning position in the input
* @param end the end position in the input
* @param anchor the anchoring flag (UNANCHORED, ANCHOR_START, ANCHOR_BOTH)
* @param group the array to fill with submatch positions
* @param ngroup the number of array pairs to fill in
* @returns true if a match was found
*/
matchWithGroup(input: any, start: any, end: any, anchor: any, ngroup: any): any[];
matchMachineInput(input: any, start: any, end: any, anchor: any, ngroup: any): any[];
/**
* Returns true iff this regexp matches the UTF-8 byte array {@code b}.
*/
matchUTF8(b: any): boolean;
/**
* Returns a copy of {@code src} in which all matches for this regexp have been replaced by
* {@code repl}. No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceAll(src: any, repl: any): string;
/**
* Returns a copy of {@code src} in which only the first match for this regexp has been replaced
* by {@code repl}. No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceFirst(src: any, repl: any): string;
/**
* Returns a copy of {@code src} in which at most {@code maxReplaces} matches for this regexp have
* been replaced by the return value of of function {@code repl} (whose first argument is the
* matched string). No support is provided for expressions (e.g. {@code \1} or {@code $1}) in the
* replacement string.
*/
replaceAllFunc(src: any, replFunc: any, maxReplaces: any): string;
pad(a: any): any;
allMatches(input: any, n: any, deliverFun?: (v: any) => any): any[];
/**
* Returns an array holding the text of the leftmost match in {@code b} of this regular
* expression.
*
* A return value of null indicates no match.
*/
findUTF8(b: any): any;
/**
* Returns a two-element array of integers defining the location of the leftmost match in
* {@code b} of this regular expression. The match itself is at {@code b[loc[0]...loc[1]]}.
*
* A return value of null indicates no match.
*/
findUTF8Index(b: any): any;
/**
* Returns a string holding the text of the leftmost match in {@code s} of this regular
* expression.
*
* If there is no match, the return value is an empty string, but it will also be empty if the
* regular expression successfully matches an empty string. Use {@link #findIndex} or
* {@link #findSubmatch} if it is necessary to distinguish these cases.
*/
find(s: any): any;
/**
* Returns a two-element array of integers defining the location of the leftmost match in
* {@code s} of this regular expression. The match itself is at
* {@code s.substring(loc[0], loc[1])}.
*
* A return value of null indicates no match.
*/
findIndex(s: any): any;
/**
* Returns an array of arrays the text of the leftmost match of the regular expression in
* {@code b} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findUTF8Submatch(b: any): any[];
/**
* Returns an array holding the index pairs identifying the leftmost match of this regular
* expression in {@code b} and the matches, if any, of its subexpressions, as defined by the the
* <a href='#submatch'>Submatch</a> and <a href='#index'>Index</a> descriptions above.
*
* A return value of null indicates no match.
*/
findUTF8SubmatchIndex(b: any): any;
/**
* Returns an array of strings holding the text of the leftmost match of the regular expression in
* {@code s} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findSubmatch(s: any): any[];
/**
* Returns an array holding the index pairs identifying the leftmost match of this regular
* expression in {@code s} and the matches, if any, of its subexpressions, as defined by the <a
* href='#submatch'>Submatch</a> description above.
*
* A return value of null indicates no match.
*/
findSubmatchIndex(s: any): any;
/**
* {@code findAllUTF8()} is the <a href='#all'>All</a> version of {@link #findUTF8}; it returns a
* list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*
* TODO(adonovan): think about defining a byte slice view class, like a read-only Go slice backed
* by |b|.
*/
findAllUTF8(b: any, n: any): any[];
/**
* {@code findAllUTF8Index} is the <a href='#all'>All</a> version of {@link #findUTF8Index}; it
* returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8Index(b: any, n: any): any[];
/**
* {@code findAll} is the <a href='#all'>All</a> version of {@link #find}; it returns a list of up
* to {@code n} successive matches of the expression, as defined by the <a href='#all'>All</a>
* description above.
*
* A return value of null indicates no match.
*/
findAll(s: any, n: any): any[];
/**
* {@code findAllIndex} is the <a href='#all'>All</a> version of {@link #findIndex}; it returns a
* list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllIndex(s: any, n: any): any[];
/**
* {@code findAllUTF8Submatch} is the <a href='#all'>All</a> version of {@link #findUTF8Submatch};
* it returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8Submatch(b: any, n: any): any[];
/**
* {@code findAllUTF8SubmatchIndex} is the <a href='#all'>All</a> version of
* {@link #findUTF8SubmatchIndex}; it returns a list of up to {@code n} successive matches of the
* expression, as defined by the <a href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllUTF8SubmatchIndex(b: any, n: any): any[];
/**
* {@code findAllSubmatch} is the <a href='#all'>All</a> version of {@link #findSubmatch}; it
* returns a list of up to {@code n} successive matches of the expression, as defined by the <a
* href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllSubmatch(s: any, n: any): any[];
/**
* {@code findAllSubmatchIndex} is the <a href='#all'>All</a> version of
* {@link #findSubmatchIndex}; it returns a list of up to {@code n} successive matches of the
* expression, as defined by the <a href='#all'>All</a> description above.
*
* A return value of null indicates no match.
*/
findAllSubmatchIndex(s: any, n: any): any[];
}
/**
* Creates an RE2JS regex directly from a template literal.
* @overload
* @param {TemplateStringsArray} stringsOrFlags - The raw string segments of the template literal.
* @param {...any} values - The interpolated values.
* @returns {RE2JS}
*/
export function re(stringsOrFlags: TemplateStringsArray, ...values: any[]): RE2JS;
/**
* Creates a template literal tag function with specific RE2JS flags.
* @overload
* @param {number} stringsOrFlags - The RE2JS flags to apply (e.g., RE2JS.CASE_INSENSITIVE).
* @returns {(strings: TemplateStringsArray, ...tagValues: any[]) => RE2JS}
*/
export function re(stringsOrFlags: number): (strings: TemplateStringsArray, ...tagValues: any[]) => RE2JS;
/**
* A compiled representation of an RE2 regular expression
*
* The matching functions take {@code String} arguments instead of the more general Java
* {@code CharSequence} since the latter doesn't provide UTF-16 decoding.
*
*
* @author rsc@google.com (Russ Cox)
* @class
*/
export class RE2JS {
/**
* Flag: case insensitive matching.
*/
static CASE_INSENSITIVE: number;
/**
* Flag: dot ({@code .}) matches all characters, including newline.
*/
static DOTALL: number;
/**
* Flag: multiline matching: {@code ^} and {@code $} match at beginning and end of line, not just
* beginning and end of input.
*/
static MULTILINE: number;
/**
* Flag: Unicode groups (e.g. {@code \p\ Greek\} ) will be syntax errors.
*/
static DISABLE_UNICODE_GROUPS: number;
/**
* Flag: matches longest possible string.
*/
static LONGEST_MATCH: number;
/**
* Flag: enable linear-time captureless lookbehinds.
*/
static LOOKBEHINDS: number;
/**
* Returns a literal pattern string for the specified string.
*
* This method produces a string that can be used to create a <code>RE2JS</code> that would
* match the string <code>s</code> as if it were a literal pattern.
*
* Metacharacters or escape sequences in the input sequence will be given no special meaning.
*
* @param {string} str The string to be literalized
* @returns {string} A literal string replacement
*/
static quote(str: string): string;
/**
* Quotes '\' and '$' in {@code str}, so that the returned string could be used in
* replacement methods as a literal replacement of {@code str}.
*
* This is a convenience delegation to {@link Matcher.quoteReplacement}.
*
* @param {string} str the string to be quoted
* @param {boolean} [javaMode=false] whether the replacement will be used in javaMode
* @returns {string} the quoted string
*/
static quoteReplacement(str: string, javaMode?: boolean): string;
/**
* Translates a given regular expression string to ensure compatibility with RE2JS.
*
* This function preprocesses the input regex string by applying necessary transformations,
* such as escaping special characters (e.g., `/`), converting named capture groups to
* RE2JS-compatible syntax, and handling Unicode sequences properly. It ensures that the
* resulting regex is safe and properly formatted before compilation.
*
* @param {string|RegExp} expr - The regular expression string to be translated.
* @returns {string} - The transformed regular expression string, ready for compilation.
*/
static translateRegExp(expr: string | RegExp): string;
/**
* Helper: create new RE2JS with given regex and flags. Flregex is the regex with flags applied.
* @param {string} regex
* @param {number} [flags=0]
* @returns {RE2JS}
*/
static compile(regex: string, flags?: number): RE2JS;
/**
* Matches a string against a regular expression.
*
* @param {string} regex the regular expression
* @param {string|number[]|Uint8Array} input the input
* @returns {boolean} true if the regular expression matches the entire input
* @throws RE2JSSyntaxException if the regular expression is malformed
*/
static matches(regex: string, input: string | number[] | Uint8Array): boolean;
/**
* This is visible for testing.
* @private
*/
private static initTest;
/**
*
* @param {string} pattern
* @param {number} flags
*/
constructor(pattern: string, flags: number);
patternInput: string;
flagsInput: number;
/** @type {import('./RE2.js').RE2} */
re2Input: RE2;
/**
* Releases memory used by internal caches associated with this pattern. Does not change the
* observable behaviour. Useful for tests that detect memory leaks via allocation tracking.
*/
reset(): void;
/**
* Returns the flags used in the constructor.
* @returns {number}
*/
flags(): number;
/**
* Returns the pattern used in the constructor.
* @returns {string}
*/
pattern(): string;
re2(): RE2;
/**
* Matches a string against a regular expression.
*
* @param {string|number[]|Uint8Array} input the input
* @returns {boolean} true if the regular expression matches the entire input
*/
matches(input: string | number[] | Uint8Array): boolean;
/**
* Creates a new {@code Matcher} matching the pattern against the input.
*
* @param {string|number[]|Uint8Array|MatcherInputBase} input the input string
* @returns {Matcher}
*/
matcher(input: string | number[] | Uint8Array | MatcherInputBase): Matcher;
/**
* Tests whether the regular expression matches any part of the input string.
* Performance Note: This method is highly optimized. Because it only returns
* a boolean and does not extract capture groups, it bypasses the `Matcher` overhead
* and guarantees execution on the high-speed DFA engine whenever possible.
*
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to test against.
* @returns {boolean} `true` if the pattern is found anywhere in the input, `false` otherwise.
*/
test(input: string | number[] | Uint8Array): boolean;
/**
* Tests whether the regular expression matches the ENTIRE input string.
* * **Performance Note:** This operates identically to `.matches()`, but is significantly
* faster because it does not request capture group data. By requesting 0 capture groups,
* it securely routes execution through the DFA fast-path.
*
* @param {string|number[]|Uint8Array} input - The input string or UTF-8 byte array to test against.
* @returns {boolean} `true` if the exact input string fully matches the pattern, `false` otherwise.
*/
testExact(input: string | number[] | Uint8Array): boolean;
/**
* Executes a search for a match in a specified string.
* Returns a result array, or null if no match is found.
* The returned array perfectly mirrors standard JavaScript `RegExpExecArray`,
* including `.index`, `.input`, and `.groups` properties.
*
* @param {string|number[]|Uint8Array} input the input string or byte array
* @returns {Array|null} the match array with index, input, and groups properties, or null
*/
exec(input: string | number[] | Uint8Array): any[] | null;
/**
* Splits input around instances of the regular expression. It returns an array giving the strings
* that occur before, between, and after instances of the regular expression.
*
* If {@code limit <= 0}, there is no limit on the size of the returned array. If
* {@code limit == 0}, empty strings that would occur at the end of the array are omitted. If
* {@code limit > 0}, at most limit strings are returned. The final string contains the remainder
* of the input, possibly including additional matches of the pattern.
*
* @param {string} input the input string to be split
* @param {number} [limit=0] the limit
* @returns {string[]} the split strings
*/
split(input: string, limit?: number): string[];
/**
* Returns an iterator of all results matching a string against the regular expression,
* including capturing groups.
*
* @param {string|number[]|Uint8Array} input the input string or byte array
* @returns {IterableIterator<RegExpMatchArray>}
*/
matchAll(input: string | number[] | Uint8Array): IterableIterator<RegExpMatchArray>;
/**
*
* @returns {string}
*/
toString(): string;
/**
* Returns the program size of this pattern.
*
* <p>
* Similar to the C++ implementation, the program size is a very approximate measure of a regexp's
* "cost". Larger numbers are more expensive than smaller numbers.
* </p>
*
* @returns {number} the program size of this pattern
*/
programSize(): number;
/**
* Returns the number of capturing groups in this matcher's pattern. Group zero denotes the entire
* pattern and is excluded from this count.
*
* @returns {number} the number of capturing groups in this pattern
*/
groupCount(): number;
/**
* Return a map of the capturing groups in this matcher's pattern, where key is the name and value
* is the index of the group in the pattern.
* @returns {Record<string, number>}
*/
namedGroups(): Record<string, number>;
/**
*
* @param {*} other
* @returns {boolean}
*/
equals(other: any): boolean;
}
export {};

View File

@@ -30,6 +30,7 @@ import { execGit, platformReadSync as safeReadFile } from './shell-command-proje
import { formatGsdSlash, resolveRuntime } from './runtime-slash.cjs';
import { detectSchemaFiles, checkSchemaDrift } from './schema-detect.cjs';
import { extractTaggedBlocks } from './markdown-sectionizer.cjs';
import { compileUserPattern, MAX_USER_PATTERN_LEN } from './pattern.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports -- agent-install-check.cjs is an export= CommonJS module
import agentInstallCheck = require('./agent-install-check.cjs');
const { checkAgentsInstalled, checkCodexModelPosture } = agentInstallCheck;
@@ -1235,22 +1236,54 @@ function cmdVerifyKeyLinks(cwd: string, planFilePath: string, raw: boolean): voi
check['detail'] = 'Source file not found (from: must be a relative file path; describe components/endpoints in via:)';
}
} else if (link['pattern']) {
try {
const regex = new RegExp(link['pattern'] as string);
if (regex.test(sourceContent)) {
check['verified'] = true;
check['detail'] = 'Pattern found in source';
} else {
const targetContent = safeReadFile(path.join(cwd, (link['to'] as string) || ''));
if (targetContent && regex.test(targetContent)) {
check['verified'] = true;
check['detail'] = 'Pattern found in target';
} else {
check['detail'] = `Pattern "${link['pattern'] as string}" not found in source or target`;
}
const pat = compileUserPattern(link['pattern']);
if (pat.neutralized !== null) {
// A neutralized pattern was refused and never compiled — that is NOT
// the check the plan author wrote, so it must never report verified
// regardless of what an unrelated fallback might otherwise have
// matched (#3477 regression: pattern "(" previously neutralized to a
// literal-escaped match that matched nearly any source file,
// producing a false verified: true / all_verified: true). The engine
// itself now guarantees `test()` returns false for a refused
// pattern; this explicit branch exists to produce the good message.
// The match-and-report path below is skipped entirely rather than
// run and then overwritten, which used to leave a misleading
// "Pattern found in source" detail alongside the neutralization note.
check['pattern_neutralized'] = pat.neutralized;
let reason: string;
switch (pat.neutralized) {
case 'empty':
reason = 'pattern is not a usable string — no match attempted';
break;
case 'too-long':
reason = `pattern exceeded ${MAX_USER_PATTERN_LEN} chars — not evaluated`;
break;
case 'unsupported':
reason = 'pattern is not valid RE2 syntax — backreferences and look-around are not supported';
break;
}
check['detail'] = `Pattern not verified (${reason})`;
} else {
try {
if (pat.test(sourceContent)) {
check['verified'] = true;
check['detail'] = 'Pattern found in source';
} else {
const targetContent = safeReadFile(path.join(cwd, (link['to'] as string) || ''));
if (targetContent && pat.test(targetContent)) {
check['verified'] = true;
check['detail'] = 'Pattern found in target';
} else {
check['detail'] = `Pattern "${link['pattern'] as string}" not found in source or target`;
}
}
} catch (err) {
// Report the errno only — never the full error/message, which for a
// re-thrown non-ENOENT platformReadSync failure (e.g. EISDIR from an
// untrusted `to:` like "../..") embeds an absolute filesystem path.
const code = (err as NodeJS.ErrnoException)?.code ?? 'unknown';
check['detail'] = `Pattern check failed: ${code}`;
}
} catch {
check['detail'] = `Invalid regex pattern: ${link['pattern'] as string}`;
}
} else {
if (sourceContent.includes((link['to'] as string) || '')) {

View File

@@ -24,7 +24,7 @@ const assert = require('node:assert/strict');
// other *.property.test.cjs file. seed: 42, overridable via GSD_FC_SEED.
const fc = require('./helpers/fast-check-setup.cjs');
const { escapeRegex, literalPattern } = require('../gsd-core/bin/lib/pattern.cjs');
const { escapeRegex, literalPattern, compileUserPattern, MAX_USER_PATTERN_LEN } = require('../gsd-core/bin/lib/pattern.cjs');
// ─── Section 1: escapeRegex — rows 1-10 ───────────────────────────────────
@@ -249,3 +249,164 @@ describe('escapeRegex without RegExp.escape (#3498 Node-22 fallback)', () => {
assert.strictEqual(r.exitCode, 0, `post-load capture failed: ${r.stdout}\n${r.stderr}`);
});
});
// ─── Section 5: compileUserPattern — RE2 linear-time engine (#3477) ────────
//
// re2js guarantees match time linear in input length — there is no
// backtracking engine to exploit, so the vulnerability class is closed by
// the engine rather than detected by a heuristic scan of the pattern text
// (the hand-rolled `hasRiskyBacktrackingGroup` scanner this seam used to run
// is deleted). Assertions here are BEHAVIORAL ONLY — never against
// `RegExp.source` or any RegExp-shaped property, because the result is no
// longer a RegExp at all.
describe('compileUserPattern', () => {
test('a plain safe pattern compiles and reports neutralized: null', () => {
const { test: matches, neutralized } = compileUserPattern('fetch.*api/feed');
assert.strictEqual(matches('do a fetch of the api/feed endpoint'), true);
assert.strictEqual(matches('completely unrelated text'), false);
assert.strictEqual(neutralized, null);
});
test('boundary: a pattern at MAX_USER_PATTERN_LEN - 1 compiles', () => {
const p = 'a'.repeat(MAX_USER_PATTERN_LEN - 1);
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(matches(p), true);
assert.strictEqual(matches('b'.repeat(MAX_USER_PATTERN_LEN - 1)), false);
assert.strictEqual(neutralized, null);
});
test('boundary: a pattern at exactly MAX_USER_PATTERN_LEN compiles', () => {
const p = 'a'.repeat(MAX_USER_PATTERN_LEN);
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(matches(p), true);
assert.strictEqual(matches('b'.repeat(MAX_USER_PATTERN_LEN)), false);
assert.strictEqual(neutralized, null);
});
test('boundary: a pattern at MAX_USER_PATTERN_LEN + 1 (513 chars) is refused, never matches', () => {
const p = 'a'.repeat(MAX_USER_PATTERN_LEN + 1);
assert.strictEqual(p.length, 513);
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(matches(p), false);
assert.strictEqual(matches(p.slice(0, MAX_USER_PATTERN_LEN)), false);
assert.strictEqual(neutralized, 'too-long');
});
test('empty string input is refused, never matches', () => {
const { test: matches, neutralized } = compileUserPattern('');
assert.strictEqual(matches(''), false);
assert.strictEqual(matches('anything'), false);
assert.strictEqual(neutralized, 'empty');
});
test('non-string input is refused, never matches', () => {
for (const bad of [null, undefined, 42, {}, []]) {
const { test: matches, neutralized } = compileUserPattern(bad);
assert.strictEqual(matches(''), false);
assert.strictEqual(matches('anything'), false);
assert.strictEqual(neutralized, 'empty');
}
});
test('property: for any string input, compileUserPattern returns a well-shaped result and never throws', () => {
// maxLength: 600 straddles MAX_USER_PATTERN_LEN (512) so fast-check actually
// exercises the >512 'too-long' branch. A neutralized result must never be
// able to report a match against arbitrary input — the security invariant.
const validNeutralizations = new Set(['empty', 'too-long', 'unsupported', null]);
fc.assert(
fc.property(fc.string({ maxLength: 600 }), fc.string({ maxLength: 50 }), (s, probe) => {
let result;
assert.doesNotThrow(() => {
result = compileUserPattern(s);
});
assert.strictEqual(typeof result.test, 'function');
assert.ok(validNeutralizations.has(result.neutralized));
if (result.neutralized !== null) {
let matched;
assert.doesNotThrow(() => {
matched = result.test(probe);
});
assert.strictEqual(matched, false, `neutralized (${result.neutralized}) result must never match: s=${JSON.stringify(s)} probe=${JSON.stringify(probe)}`);
}
})
);
});
test('a known-valid regex round-trips through the >512-length property test unmangled', () => {
// Companion assertion for the property test above: a legitimate long-ish
// regex must still compile (neutralized: null), not just "didn't throw" —
// pins that the length straddle doesn't accidentally neutralize everything
// under 512 chars too.
const p = 'fetch\\(.*\\)\\.then\\(' + 'x'.repeat(400) + '\\)';
assert.ok(p.length < MAX_USER_PATTERN_LEN, 'sanity: fixture must stay under the length threshold');
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(neutralized, null);
assert.strictEqual(matches('fetch(url).then(' + 'x'.repeat(400) + ')'), true);
assert.strictEqual(matches('unrelated text'), false);
});
});
// ─── Section 6: RE2 engine acceptance table (#3477 follow-up) ─────────────
//
// The prior hand-rolled screen either hung (patterns it missed, run live
// through the JS backtracking engine) or wrongly refused (patterns it
// flagged as risky-shaped that are actually linear) on the rows below. RE2
// closes both failure modes: every MUST_COMPILE row below is evaluated for
// real, in linear time, with a correct match verdict — no heuristic,
// no false refusal.
describe('compileUserPattern — RE2 engine acceptance table', () => {
// [pattern, subjectA, subjectB, expectedA, expectedB]. Every one of these
// patterns is plain JS-valid ERE syntax (no backreferences/look-around), so
// the expected/expectedB booleans are exactly what `new RegExp(p).test(...)`
// would report — computed offline rather than hand-assumed, since a
// `*`/`?`-quantified pattern with nothing required outside the optional
// part (e.g. `(a|a)*$`, `(abc)?`, `(abc)*`) trivially matches an unanchored
// `.test()` against ANY subject (zero-width match), which is correct
// JS-regex semantics, not a defect. The expectations are hardcoded rather
// than oracle-derived because running these patterns through the JS
// backtracking engine (`new RegExp(p).test(...)`) is the exact
// catastrophic-backtracking vulnerability this issue is about — the suite
// must never execute that engine against them, even with short subjects.
const MUST_COMPILE = [
['(a+)+$', 'aaaa', 'bbbb', true, false],
['(a|a)*$', 'aaaa', 'aaab', true, true],
['((a+))+$', 'aaaa', 'aaab', true, false],
['(a+){2,}$', 'aaaa', 'a', true, false],
['(a{1,3})+$', 'aaaaaa', 'bbbbbb', true, false],
['^(\\s*\\w+)+$', 'foo bar baz', 'foo, bar', true, false],
['fetch.*api/feed', 'do a fetch of the api/feed endpoint', 'completely unrelated text', true, false],
['prisma\\.message\\.(find|create)', 'call prisma.message.find(x)', 'call prisma.other.find(x)', true, false],
['(abc)?', 'abc', '', true, true],
['(abc)*', 'abcabc', 'xyz', true, true],
['^\\s*export\\s+function\\s+\\w+', ' export function foo() {}', 'const foo = 1', true, false],
];
const MUST_REFUSE = ['(\\w+)\\1', '(?!x)a'];
for (const [p, subjectA, subjectB, expectedA, expectedB] of MUST_COMPILE) {
test(`compiles and evaluates correctly: ${JSON.stringify(p)}`, () => {
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(neutralized, null, `expected null for ${JSON.stringify(p)}, got ${JSON.stringify(neutralized)}`);
assert.strictEqual(
matches(subjectA),
expectedA,
`expected ${JSON.stringify(p)} against ${JSON.stringify(subjectA)} to match real regex semantics`
);
assert.strictEqual(
matches(subjectB),
expectedB,
`expected ${JSON.stringify(p)} against ${JSON.stringify(subjectB)} to match real regex semantics`
);
});
}
for (const p of MUST_REFUSE) {
test(`is refused (unsupported RE2 syntax): ${JSON.stringify(p)}`, () => {
const { test: matches, neutralized } = compileUserPattern(p);
assert.strictEqual(neutralized, 'unsupported', `expected 'unsupported' for ${JSON.stringify(p)}, got ${JSON.stringify(neutralized)}`);
assert.strictEqual(matches('anything'), false);
});
}
});

View File

@@ -1760,6 +1760,60 @@ describe('verify key-links command', () => {
);
});
test('a formerly-ReDoS-shaped pattern (nested quantifiers) is evaluated normally via RE2 (#3477)', () => {
// Pre-RE2 this pattern was neutralized (hand-rolled screen) to avoid
// catastrophic backtracking in the JS regex engine. RE2 (re2js) matches
// in linear time by construction, so this is no longer a neutralization
// case at all — the pattern is compiled and evaluated for real.
writePlanWithKeyLinks(tmpDir, [
'- from: "src/a.js"',
' to: "src/b.js"',
' pattern: "(a+)+$"',
]);
fs.writeFileSync(path.join(tmpDir, 'src', 'a.js'), 'a'.repeat(25) + 'b\n');
fs.writeFileSync(path.join(tmpDir, 'src', 'b.js'), 'module.exports = {};\n');
const result = runGsdTools('verify key-links .planning/phases/01-test/01-01-PLAN.md', tmpDir);
assert.ok(result.success, `Command failed: ${result.error}`);
const output = JSON.parse(result.output);
assert.strictEqual(output.links[0].verified, false, 'link should not be verified — subject does not end in "a"');
assert.strictEqual(output.all_verified, false, `Expected all_verified false: ${JSON.stringify(output)}`);
assert.strictEqual(
output.links[0].pattern_neutralized,
undefined,
`RE2 evaluates this pattern normally — pattern_neutralized must be absent: ${JSON.stringify(output.links[0])}`
);
});
test('a refused pattern (unsupported RE2 syntax) never reports verified: true (#3477 regression)', () => {
// pattern: "(?!x)a" is a negative lookahead — RE2 has no backtracking
// engine and does not support look-around, so this is refused outright
// (neutralized: 'unsupported') rather than guessed at via a literal
// fallback. Pre-#3477-fix, a similarly unparseable pattern neutralized to
// a literal-escaped match that happened to match nearly any source file,
// producing a false verified: true / all_verified: true.
writePlanWithKeyLinks(tmpDir, [
'- from: "src/a.js"',
' to: "src/b.js"',
' pattern: "(?!x)a"',
]);
fs.writeFileSync(path.join(tmpDir, 'src', 'a.js'), 'function f(x) { return x; }\n');
fs.writeFileSync(path.join(tmpDir, 'src', 'b.js'), 'module.exports = {};\n');
const result = runGsdTools('verify key-links .planning/phases/01-test/01-01-PLAN.md', tmpDir);
assert.ok(result.success, `Command failed: ${result.error}`);
const output = JSON.parse(result.output);
assert.strictEqual(output.links[0].verified, false, 'a neutralized pattern must never report verified: true');
assert.strictEqual(output.all_verified, false, `Expected all_verified false: ${JSON.stringify(output)}`);
assert.strictEqual(
output.links[0].pattern_neutralized,
'unsupported',
`Expected pattern_neutralized: 'unsupported': ${JSON.stringify(output.links[0])}`
);
});
test('returns error when no key_links in frontmatter', () => {
const content = [
'---',