* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags an unbounded */+/{n,} quantifier over a broad character class ([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the exact #2128-fixed shape) applied to a regex whose match target is data-flow-traced to readFileSync content. eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer shared with no-crlf-fragile-split (Phase 2) rather than a second copy — no-crlf-fragile-split refactored onto it with zero behavior change, parity-tested. Real triage, not 798 mechanical edits: the ADR's census (2026-08-08) screened every unbounded quantifier in the tree unscoped. Correctly scoped to readFileSync-derived content (matching Phase 2's own G2/G3 scoping), the rule found 162 real hits across two detection waves — the second wave (93) surfaced only after a genuine off-by-one bug in this rule's own first draft was caught while writing its RuleTester tests and fixed (the bug silently missed every directly-quantified [\s\S]* with no gap before the quantifier — exactly the class this rule exists to catch). 3 hits landed in production src/ (commands.cts, milestone.cts, roadmap.cts) and were each empirically timed against adversarial input (matching #2128's own measured-not-assumed precedent) — all confirmed linear-time/benign, left unbounded with a measured-evidence comment rather than mechanically bounded. The remaining 159 are test-file fixture parsing (test-author-controlled, fixed-size content, not adversarial input) — each suppressed with a specific, non-generic reason. Zero functional behavior changed anywhere in this diff. tests/no-pending-3212-markers.test.cjs locks the epic's own closing invariant (ADR §7: "assert zero pending #3212 markers remain") — ground truth confirmed trivially true today (no phase left any such marker behind), now regression-locked going forward. Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): correct rule category mislabel, add CI test-scope entry An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs mistakenly carried meta.docs.category: 'Portability', copied from a sibling rule without realizing what that implied: docs/contributing/cross-platform- portability-rules.md governs an ADR-1703 rule family under a hard "zero escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic — and its eslint-disable-next-line suppressions (159 of them, added earlier this same phase after empirical benign-verification) are an intentional, correct design, not a bypass. Corrected to category: 'Best Practices', matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the same epic, which is also correctly outside PROTECTED_RULES), and the rule's own docstring now states this explicitly so a future reader doesn't have to re-derive it. Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their own test suites under targeted CI selection — was previously unregistered and invisible to that fast-path (this PR's own gsd-test checkpoint runs the full suite regardless, so this only affects future narrowly-scoped PRs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic) Security review found the rule meant to catch algorithmic-complexity bugs had one of its own: hasUnboundedBroadQuantifier's negated-class inner scan walked from each `[^` occurrence to the next `]` (or EOF) with no bound, while the outer loop only ever advanced by one character — O(n²) total work on a pattern with many unclosed `[^` runs. Runs unconditionally inside checkPattern on any `new RegExp('literal string')` argument in any linted file, before the (cheap) readFileSync data-flow gate — so a single crafted string literal, no valid regex syntax required, could make `npm run lint` / CI hang. Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/ 16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n, quadratic); extrapolated, the 300000-char repro from the finding would run ~165s. Post-fix (bail the inner scan once units exceeds the rule's own 1-2-unit scope, rather than continuing to hunt for a closing `]`), the same 300000-char input runs in 8.7ms via the real rule module, independently reconfirmed at 18ms via a fresh Linter.verify() call. New regression row in tests/no-unbounded-quantifier.rule.test.cjs asserts the RuleTester run on a 50000-char adversarial pattern completes and returns a defined result — no wall-clock assertion (CLAUDE.md Clock Seams / local/no-elapsed-assertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch next merged 12 more PRs during this PR's review. Two consequences: - tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own workflow .md content — the same Class A pattern as the ~159 sites already triaged elsewhere in this PR. Suppressed with the same established reason. - lint-allow-test-rule-refs' ratchet ceiling needed re-raising again (301 -> 303) for the same reason as the two prior bumps: organic growth from unrelated, already-reviewed PRs landing concurrently, not a defect in this branch's own diff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
280 lines
10 KiB
JavaScript
280 lines
10 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* no-crlf-fragile-split
|
|
*
|
|
* Flag CRLF-fragile file-content splitting and regex patterns in test files.
|
|
* Windows git-autocrlf causes readFileSync to return \r\n line endings; code
|
|
* that splits on bare `\n` or uses regexes with bare `\n` will silently
|
|
* mismatch on Windows.
|
|
*
|
|
* ## What this enforces
|
|
*
|
|
* G1 — a `.split('\n')` / `.split("\n")` CallExpression whose receiver is
|
|
* (transitively) a `readFileSync`/`fs.readFileSync` result — directly,
|
|
* via a chain, or via an Identifier that scope-resolves to a variable
|
|
* initialized from readFileSync.
|
|
* Message: use `splitLines()` from `src/text-lines.cts`.
|
|
*
|
|
* G2/G3 — a RegExpLiteral whose pattern contains a bare `\n` (a `\n` not
|
|
* part of `\r?\n` / `\r\n` / `[\r\n]` etc.) used as the pattern of a
|
|
* `.match`/`.test`/`.exec`/`.replace`/`.replaceAll`/`.split`/`.matchAll`
|
|
* call on a readFileSync-derived receiver. ALSO flags a RegExpLiteral
|
|
* with a bare `\n` whose source contains a markdown fence (```) or a
|
|
* frontmatter anchor (`^---`), since those shapes target file content.
|
|
* Message: use `splitLines()` from `src/text-lines.cts` (or the raw
|
|
* `\r?\n` regex, for a file that cannot import the compiled seam).
|
|
*
|
|
* ## Known boundaries
|
|
*
|
|
* The data-flow is scope-based: a readFileSync result is tracked via the
|
|
* immediate call-chain or a single variable binding initialized from
|
|
* readFileSync in the same file scope. A regex stored far from its use, or
|
|
* content obtained via a non-readFileSync read (e.g. fs.readFile callback,
|
|
* streams), may not be caught. G2/G3 additionally fires on fence/frontmatter
|
|
* regex shapes even when data-flow is indirect, to catch the most common
|
|
* markdown parsing patterns.
|
|
*
|
|
* DEFECT category: DEFECT.WINDOWS-CRLF-TEST-PORTABILITY
|
|
*/
|
|
|
|
const { isWindowsExcludedNode } = require('./lib/platform-guard.cjs');
|
|
const { isReadFileSyncDerived, isPatternUsedOnFileContent } = require('./lib/readfilesync-trace.cjs');
|
|
|
|
/** @type {import('eslint').Rule.RuleModule} */
|
|
const rule = {
|
|
meta: {
|
|
type: 'problem',
|
|
docs: {
|
|
description:
|
|
'Disallow CRLF-fragile file-content split and regex patterns in tests (fails on Windows with git-autocrlf)',
|
|
category: 'Portability',
|
|
},
|
|
schema: [],
|
|
messages: {
|
|
crlfFragileSplit:
|
|
'Splitting on literal "\\n" on readFileSync content is CRLF-fragile ' +
|
|
'(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY): Windows git-autocrlf yields "\\r\\n" ' +
|
|
'line endings. Use splitLines() from src/text-lines.cts instead.',
|
|
crlfFragileRegex:
|
|
'RegExp with a bare "\\n" on readFileSync content is CRLF-fragile ' +
|
|
'(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY): Windows git-autocrlf yields "\\r\\n" ' +
|
|
'line endings. Use splitLines() from src/text-lines.cts instead.',
|
|
},
|
|
},
|
|
|
|
create(context) {
|
|
const sourceCode = context.sourceCode ?? context.getSourceCode();
|
|
|
|
// ── Helpers ─────────────────────────────────────────────────────────────
|
|
|
|
/**
|
|
* Returns the string value of a Literal node, or null.
|
|
* @param {import('eslint').Rule.Node} node
|
|
* @returns {string|null}
|
|
*/
|
|
function stringValue(node) {
|
|
if (node && node.type === 'Literal' && typeof node.value === 'string') {
|
|
return node.value;
|
|
}
|
|
return null;
|
|
}
|
|
|
|
/**
|
|
* Returns true if a RegExpLiteral has at least one FRAGILE bare \n — a \n
|
|
* that is not adequately protected against CRLF.
|
|
*
|
|
* Per-occurrence classification: every \n in the pattern is inspected
|
|
* individually. A \n is SAFE when ANY of these hold:
|
|
* 1. Immediately preceded by \r? (part of \r?\n)
|
|
* 2. Immediately preceded by \r (part of \r\n)
|
|
* 3. Inside a character class [...] that also contains \r
|
|
* (e.g. [\r\n], [^\r\n], [\n\r])
|
|
*
|
|
* Everything else is FRAGILE: [^\n], [\n], or a bare \n in the main pattern.
|
|
*
|
|
* @param {import('eslint').Rule.Node} node — Literal with regex
|
|
* @returns {boolean}
|
|
*/
|
|
function hasBareLiteralNewline(node) {
|
|
if (!node || node.type !== 'Literal' || !node.regex) return false;
|
|
const pattern = node.regex.pattern;
|
|
if (!pattern.includes('\\n')) return false;
|
|
|
|
// Walk the pattern, find every \n occurrence and classify it.
|
|
let i = 0;
|
|
// Track whether we are inside a [...] character class and whether
|
|
// the current class contains \r.
|
|
let inClass = false;
|
|
let classHasCarriageReturn = false;
|
|
let foundFragile = false;
|
|
|
|
while (i < pattern.length) {
|
|
// Entering a character class
|
|
if (pattern[i] === '[' && !inClass) {
|
|
inClass = true;
|
|
classHasCarriageReturn = false;
|
|
i++;
|
|
// Skip optional ^ negation
|
|
if (i < pattern.length && pattern[i] === '^') i++;
|
|
// Skip ] if it appears immediately after [ or [^, where it is literal
|
|
if (i < pattern.length && pattern[i] === ']') i++;
|
|
continue;
|
|
}
|
|
|
|
// Exiting a character class
|
|
if (pattern[i] === ']' && inClass) {
|
|
inClass = false;
|
|
i++;
|
|
continue;
|
|
}
|
|
|
|
// Escape sequences inside the pattern
|
|
if (pattern[i] === '\\' && i + 1 < pattern.length) {
|
|
const next = pattern[i + 1];
|
|
if (next === 'r') {
|
|
// \r — if inside a class, note it contains \r
|
|
if (inClass) classHasCarriageReturn = true;
|
|
i += 2;
|
|
continue;
|
|
}
|
|
if (next === 'n') {
|
|
// \n found — classify it
|
|
// Check if preceded by \r? or \r (look back in the raw pattern string)
|
|
// "preceded by" means the two chars before the current \\ are \r or \r?
|
|
const before2 = pattern.slice(Math.max(0, i - 2), i); // up to 2 chars before \\
|
|
|
|
// Re-check with a wider window for \r?\n (pattern chars: \r?\n = 5 chars)
|
|
const before3 = pattern.slice(Math.max(0, i - 3), i);
|
|
const safeByPrefixFull =
|
|
before3 === '\\r?' || // \r?\n
|
|
before2 === '\\r'; // \r\n
|
|
|
|
if (inClass) {
|
|
// Inside a class: safe only if the class itself contains \r
|
|
if (!classHasCarriageReturn) {
|
|
foundFragile = true;
|
|
}
|
|
} else if (!safeByPrefixFull) {
|
|
foundFragile = true;
|
|
}
|
|
i += 2;
|
|
continue;
|
|
}
|
|
// Any other escape: skip both chars
|
|
i += 2;
|
|
continue;
|
|
}
|
|
|
|
i++;
|
|
}
|
|
|
|
return foundFragile;
|
|
}
|
|
|
|
/**
|
|
* Returns true if a RegExpLiteral with a bare \n is used on a readFileSync-
|
|
* derived receiver via .match/.test/.exec/.replace/.replaceAll/.split/.matchAll.
|
|
*
|
|
* Two AST shapes:
|
|
* Shape A: str.match(/regex/) — regex is an ARG to the call.
|
|
* regex.parent = CallExpression (arg), callee.object = str
|
|
* Shape B: /regex/.test(str) — regex is the callee object.
|
|
* regex.parent = MemberExpression (the .test callee)
|
|
* regex.parent.parent = CallExpression, first arg = str
|
|
*
|
|
* @param {import('eslint').Rule.Node} regexNode — the RegExpLiteral
|
|
* @returns {boolean}
|
|
*/
|
|
function isRegexUsedOnFileContent(regexNode) {
|
|
return isPatternUsedOnFileContent(regexNode, sourceCode);
|
|
}
|
|
|
|
/**
|
|
* Returns true if a RegExpLiteral pattern:
|
|
* - has a bare \n, AND
|
|
* - contains a markdown fence (```) or frontmatter anchor (^---)
|
|
*
|
|
* These shapes target file content by convention even without direct
|
|
* data-flow tracking.
|
|
*
|
|
* @param {import('eslint').Rule.Node} node — Literal with regex
|
|
* @returns {boolean}
|
|
*/
|
|
function isMarkdownOrFrontmatterRegex(node) {
|
|
if (!node || node.type !== 'Literal' || !node.regex) return false;
|
|
if (!hasBareLiteralNewline(node)) return false;
|
|
const pattern = node.regex.pattern;
|
|
// Markdown fence: ```
|
|
if (pattern.includes('```')) return true;
|
|
// Frontmatter anchor: ^---
|
|
if (/\^---/.test(pattern)) return true;
|
|
return false;
|
|
}
|
|
|
|
// ── Per-file state ──────────────────────────────────────────────────────
|
|
|
|
/** Collected G1 violations: {node} */
|
|
const g1Violations = [];
|
|
/** Collected G2/G3 violations: {node} */
|
|
const g2g3Violations = [];
|
|
|
|
return {
|
|
// G1: .split('\n') on readFileSync-derived content
|
|
CallExpression(node) {
|
|
const callee = node.callee;
|
|
if (
|
|
callee &&
|
|
callee.type === 'MemberExpression' &&
|
|
!callee.computed &&
|
|
callee.property.type === 'Identifier' &&
|
|
callee.property.name === 'split'
|
|
) {
|
|
const args = node.arguments;
|
|
if (args && args.length >= 1) {
|
|
const argVal = stringValue(args[0]);
|
|
if (argVal === '\n') {
|
|
// Is the receiver derived from readFileSync?
|
|
if (isReadFileSyncDerived(callee.object, sourceCode)) {
|
|
g1Violations.push(node);
|
|
}
|
|
}
|
|
}
|
|
}
|
|
},
|
|
|
|
// G2/G3: RegExpLiteral with bare \n
|
|
Literal(node) {
|
|
if (!node.regex) return;
|
|
if (!hasBareLiteralNewline(node)) return;
|
|
|
|
// Check G2/G3 via data-flow (receiver is readFileSync-derived)
|
|
if (isRegexUsedOnFileContent(node)) {
|
|
g2g3Violations.push(node);
|
|
return;
|
|
}
|
|
|
|
// Also check G2/G3 via content shape (markdown fence or frontmatter)
|
|
if (isMarkdownOrFrontmatterRegex(node)) {
|
|
g2g3Violations.push(node);
|
|
}
|
|
},
|
|
|
|
'Program:exit'() {
|
|
for (const node of g1Violations) {
|
|
if (!isWindowsExcludedNode(node, sourceCode)) {
|
|
context.report({ node, messageId: 'crlfFragileSplit' });
|
|
}
|
|
}
|
|
for (const node of g2g3Violations) {
|
|
if (!isWindowsExcludedNode(node, sourceCode)) {
|
|
context.report({ node, messageId: 'crlfFragileRegex' });
|
|
}
|
|
}
|
|
},
|
|
};
|
|
},
|
|
};
|
|
|
|
module.exports = rule;
|