Files
msd-core/eslint-rules/no-crlf-fragile-split.cjs
Tom Boucher 69e7afd0c7 chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class

Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags
an unbounded */+/{n,} quantifier over a broad character class
([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the
exact #2128-fixed shape) applied to a regex whose match target is
data-flow-traced to readFileSync content.

eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer
shared with no-crlf-fragile-split (Phase 2) rather than a second copy
— no-crlf-fragile-split refactored onto it with zero behavior change,
parity-tested.

Real triage, not 798 mechanical edits: the ADR's census (2026-08-08)
screened every unbounded quantifier in the tree unscoped. Correctly
scoped to readFileSync-derived content (matching Phase 2's own G2/G3
scoping), the rule found 162 real hits across two detection waves — the
second wave (93) surfaced only after a genuine off-by-one bug in this
rule's own first draft was caught while writing its RuleTester tests
and fixed (the bug silently missed every directly-quantified [\s\S]*
with no gap before the quantifier — exactly the class this rule exists
to catch). 3 hits landed in production src/ (commands.cts, milestone.cts,
roadmap.cts) and were each empirically timed against adversarial input
(matching #2128's own measured-not-assumed precedent) — all confirmed
linear-time/benign, left unbounded with a measured-evidence comment
rather than mechanically bounded. The remaining 159 are test-file
fixture parsing (test-author-controlled, fixed-size content, not
adversarial input) — each suppressed with a specific, non-generic
reason. Zero functional behavior changed anywhere in this diff.

tests/no-pending-3212-markers.test.cjs locks the epic's own closing
invariant (ADR §7: "assert zero pending #3212 markers remain") — ground
truth confirmed trivially true today (no phase left any such marker
behind), now regression-locked going forward.

Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md
Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): correct rule category mislabel, add CI test-scope entry

An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs
mistakenly carried meta.docs.category: 'Portability', copied from a sibling
rule without realizing what that implied: docs/contributing/cross-platform-
portability-rules.md governs an ADR-1703 rule family under a hard "zero
escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's
PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is
not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic —
and its eslint-disable-next-line suppressions (159 of them, added earlier
this same phase after empirical benign-verification) are an intentional,
correct design, not a bypass. Corrected to category: 'Best Practices',
matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the
same epic, which is also correctly outside PROTECTED_RULES), and the rule's
own docstring now states this explicitly so a future reader doesn't have to
re-derive it.

Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule
or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their
own test suites under targeted CI selection — was previously unregistered
and invisible to that fast-path (this PR's own gsd-test checkpoint runs the
full suite regardless, so this only affects future narrowly-scoped PRs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic)

Security review found the rule meant to catch algorithmic-complexity bugs
had one of its own: hasUnboundedBroadQuantifier's negated-class inner
scan walked from each `[^` occurrence to the next `]` (or EOF) with no
bound, while the outer loop only ever advanced by one character — O(n²)
total work on a pattern with many unclosed `[^` runs. Runs unconditionally
inside checkPattern on any `new RegExp('literal string')` argument in any
linted file, before the (cheap) readFileSync data-flow gate — so a single
crafted string literal, no valid regex syntax required, could make
`npm run lint` / CI hang.

Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/
16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n,
quadratic); extrapolated, the 300000-char repro from the finding would
run ~165s. Post-fix (bail the inner scan once units exceeds the rule's
own 1-2-unit scope, rather than continuing to hunt for a closing `]`),
the same 300000-char input runs in 8.7ms via the real rule module,
independently reconfirmed at 18ms via a fresh Linter.verify() call.

New regression row in tests/no-unbounded-quantifier.rule.test.cjs
asserts the RuleTester run on a 50000-char adversarial pattern
completes and returns a defined result — no wall-clock assertion
(CLAUDE.md Clock Seams / local/no-elapsed-assertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch

next merged 12 more PRs during this PR's review. Two consequences:

- tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new
  content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own
  workflow .md content — the same Class A pattern as the ~159 sites
  already triaged elsewhere in this PR. Suppressed with the same
  established reason.
- lint-allow-test-rule-refs' ratchet ceiling needed re-raising again
  (301 -> 303) for the same reason as the two prior bumps: organic
  growth from unrelated, already-reviewed PRs landing concurrently,
  not a defect in this branch's own diff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 10:02:28 -04:00

280 lines
10 KiB
JavaScript

'use strict';
/**
* no-crlf-fragile-split
*
* Flag CRLF-fragile file-content splitting and regex patterns in test files.
* Windows git-autocrlf causes readFileSync to return \r\n line endings; code
* that splits on bare `\n` or uses regexes with bare `\n` will silently
* mismatch on Windows.
*
* ## What this enforces
*
* G1 — a `.split('\n')` / `.split("\n")` CallExpression whose receiver is
* (transitively) a `readFileSync`/`fs.readFileSync` result — directly,
* via a chain, or via an Identifier that scope-resolves to a variable
* initialized from readFileSync.
* Message: use `splitLines()` from `src/text-lines.cts`.
*
* G2/G3 — a RegExpLiteral whose pattern contains a bare `\n` (a `\n` not
* part of `\r?\n` / `\r\n` / `[\r\n]` etc.) used as the pattern of a
* `.match`/`.test`/`.exec`/`.replace`/`.replaceAll`/`.split`/`.matchAll`
* call on a readFileSync-derived receiver. ALSO flags a RegExpLiteral
* with a bare `\n` whose source contains a markdown fence (```) or a
* frontmatter anchor (`^---`), since those shapes target file content.
* Message: use `splitLines()` from `src/text-lines.cts` (or the raw
* `\r?\n` regex, for a file that cannot import the compiled seam).
*
* ## Known boundaries
*
* The data-flow is scope-based: a readFileSync result is tracked via the
* immediate call-chain or a single variable binding initialized from
* readFileSync in the same file scope. A regex stored far from its use, or
* content obtained via a non-readFileSync read (e.g. fs.readFile callback,
* streams), may not be caught. G2/G3 additionally fires on fence/frontmatter
* regex shapes even when data-flow is indirect, to catch the most common
* markdown parsing patterns.
*
* DEFECT category: DEFECT.WINDOWS-CRLF-TEST-PORTABILITY
*/
const { isWindowsExcludedNode } = require('./lib/platform-guard.cjs');
const { isReadFileSyncDerived, isPatternUsedOnFileContent } = require('./lib/readfilesync-trace.cjs');
/** @type {import('eslint').Rule.RuleModule} */
const rule = {
meta: {
type: 'problem',
docs: {
description:
'Disallow CRLF-fragile file-content split and regex patterns in tests (fails on Windows with git-autocrlf)',
category: 'Portability',
},
schema: [],
messages: {
crlfFragileSplit:
'Splitting on literal "\\n" on readFileSync content is CRLF-fragile ' +
'(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY): Windows git-autocrlf yields "\\r\\n" ' +
'line endings. Use splitLines() from src/text-lines.cts instead.',
crlfFragileRegex:
'RegExp with a bare "\\n" on readFileSync content is CRLF-fragile ' +
'(DEFECT.WINDOWS-CRLF-TEST-PORTABILITY): Windows git-autocrlf yields "\\r\\n" ' +
'line endings. Use splitLines() from src/text-lines.cts instead.',
},
},
create(context) {
const sourceCode = context.sourceCode ?? context.getSourceCode();
// ── Helpers ─────────────────────────────────────────────────────────────
/**
* Returns the string value of a Literal node, or null.
* @param {import('eslint').Rule.Node} node
* @returns {string|null}
*/
function stringValue(node) {
if (node && node.type === 'Literal' && typeof node.value === 'string') {
return node.value;
}
return null;
}
/**
* Returns true if a RegExpLiteral has at least one FRAGILE bare \n — a \n
* that is not adequately protected against CRLF.
*
* Per-occurrence classification: every \n in the pattern is inspected
* individually. A \n is SAFE when ANY of these hold:
* 1. Immediately preceded by \r? (part of \r?\n)
* 2. Immediately preceded by \r (part of \r\n)
* 3. Inside a character class [...] that also contains \r
* (e.g. [\r\n], [^\r\n], [\n\r])
*
* Everything else is FRAGILE: [^\n], [\n], or a bare \n in the main pattern.
*
* @param {import('eslint').Rule.Node} node — Literal with regex
* @returns {boolean}
*/
function hasBareLiteralNewline(node) {
if (!node || node.type !== 'Literal' || !node.regex) return false;
const pattern = node.regex.pattern;
if (!pattern.includes('\\n')) return false;
// Walk the pattern, find every \n occurrence and classify it.
let i = 0;
// Track whether we are inside a [...] character class and whether
// the current class contains \r.
let inClass = false;
let classHasCarriageReturn = false;
let foundFragile = false;
while (i < pattern.length) {
// Entering a character class
if (pattern[i] === '[' && !inClass) {
inClass = true;
classHasCarriageReturn = false;
i++;
// Skip optional ^ negation
if (i < pattern.length && pattern[i] === '^') i++;
// Skip ] if it appears immediately after [ or [^, where it is literal
if (i < pattern.length && pattern[i] === ']') i++;
continue;
}
// Exiting a character class
if (pattern[i] === ']' && inClass) {
inClass = false;
i++;
continue;
}
// Escape sequences inside the pattern
if (pattern[i] === '\\' && i + 1 < pattern.length) {
const next = pattern[i + 1];
if (next === 'r') {
// \r — if inside a class, note it contains \r
if (inClass) classHasCarriageReturn = true;
i += 2;
continue;
}
if (next === 'n') {
// \n found — classify it
// Check if preceded by \r? or \r (look back in the raw pattern string)
// "preceded by" means the two chars before the current \\ are \r or \r?
const before2 = pattern.slice(Math.max(0, i - 2), i); // up to 2 chars before \\
// Re-check with a wider window for \r?\n (pattern chars: \r?\n = 5 chars)
const before3 = pattern.slice(Math.max(0, i - 3), i);
const safeByPrefixFull =
before3 === '\\r?' || // \r?\n
before2 === '\\r'; // \r\n
if (inClass) {
// Inside a class: safe only if the class itself contains \r
if (!classHasCarriageReturn) {
foundFragile = true;
}
} else if (!safeByPrefixFull) {
foundFragile = true;
}
i += 2;
continue;
}
// Any other escape: skip both chars
i += 2;
continue;
}
i++;
}
return foundFragile;
}
/**
* Returns true if a RegExpLiteral with a bare \n is used on a readFileSync-
* derived receiver via .match/.test/.exec/.replace/.replaceAll/.split/.matchAll.
*
* Two AST shapes:
* Shape A: str.match(/regex/) — regex is an ARG to the call.
* regex.parent = CallExpression (arg), callee.object = str
* Shape B: /regex/.test(str) — regex is the callee object.
* regex.parent = MemberExpression (the .test callee)
* regex.parent.parent = CallExpression, first arg = str
*
* @param {import('eslint').Rule.Node} regexNode — the RegExpLiteral
* @returns {boolean}
*/
function isRegexUsedOnFileContent(regexNode) {
return isPatternUsedOnFileContent(regexNode, sourceCode);
}
/**
* Returns true if a RegExpLiteral pattern:
* - has a bare \n, AND
* - contains a markdown fence (```) or frontmatter anchor (^---)
*
* These shapes target file content by convention even without direct
* data-flow tracking.
*
* @param {import('eslint').Rule.Node} node — Literal with regex
* @returns {boolean}
*/
function isMarkdownOrFrontmatterRegex(node) {
if (!node || node.type !== 'Literal' || !node.regex) return false;
if (!hasBareLiteralNewline(node)) return false;
const pattern = node.regex.pattern;
// Markdown fence: ```
if (pattern.includes('```')) return true;
// Frontmatter anchor: ^---
if (/\^---/.test(pattern)) return true;
return false;
}
// ── Per-file state ──────────────────────────────────────────────────────
/** Collected G1 violations: {node} */
const g1Violations = [];
/** Collected G2/G3 violations: {node} */
const g2g3Violations = [];
return {
// G1: .split('\n') on readFileSync-derived content
CallExpression(node) {
const callee = node.callee;
if (
callee &&
callee.type === 'MemberExpression' &&
!callee.computed &&
callee.property.type === 'Identifier' &&
callee.property.name === 'split'
) {
const args = node.arguments;
if (args && args.length >= 1) {
const argVal = stringValue(args[0]);
if (argVal === '\n') {
// Is the receiver derived from readFileSync?
if (isReadFileSyncDerived(callee.object, sourceCode)) {
g1Violations.push(node);
}
}
}
}
},
// G2/G3: RegExpLiteral with bare \n
Literal(node) {
if (!node.regex) return;
if (!hasBareLiteralNewline(node)) return;
// Check G2/G3 via data-flow (receiver is readFileSync-derived)
if (isRegexUsedOnFileContent(node)) {
g2g3Violations.push(node);
return;
}
// Also check G2/G3 via content shape (markdown fence or frontmatter)
if (isMarkdownOrFrontmatterRegex(node)) {
g2g3Violations.push(node);
}
},
'Program:exit'() {
for (const node of g1Violations) {
if (!isWindowsExcludedNode(node, sourceCode)) {
context.report({ node, messageId: 'crlfFragileSplit' });
}
}
for (const node of g2g3Violations) {
if (!isWindowsExcludedNode(node, sourceCode)) {
context.report({ node, messageId: 'crlfFragileRegex' });
}
}
},
};
},
};
module.exports = rule;