Files
msd-core/tests/verifier-behavior-unverified.test.cjs
Tom Boucher 69e7afd0c7 chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class

Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags
an unbounded */+/{n,} quantifier over a broad character class
([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the
exact #2128-fixed shape) applied to a regex whose match target is
data-flow-traced to readFileSync content.

eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer
shared with no-crlf-fragile-split (Phase 2) rather than a second copy
— no-crlf-fragile-split refactored onto it with zero behavior change,
parity-tested.

Real triage, not 798 mechanical edits: the ADR's census (2026-08-08)
screened every unbounded quantifier in the tree unscoped. Correctly
scoped to readFileSync-derived content (matching Phase 2's own G2/G3
scoping), the rule found 162 real hits across two detection waves — the
second wave (93) surfaced only after a genuine off-by-one bug in this
rule's own first draft was caught while writing its RuleTester tests
and fixed (the bug silently missed every directly-quantified [\s\S]*
with no gap before the quantifier — exactly the class this rule exists
to catch). 3 hits landed in production src/ (commands.cts, milestone.cts,
roadmap.cts) and were each empirically timed against adversarial input
(matching #2128's own measured-not-assumed precedent) — all confirmed
linear-time/benign, left unbounded with a measured-evidence comment
rather than mechanically bounded. The remaining 159 are test-file
fixture parsing (test-author-controlled, fixed-size content, not
adversarial input) — each suppressed with a specific, non-generic
reason. Zero functional behavior changed anywhere in this diff.

tests/no-pending-3212-markers.test.cjs locks the epic's own closing
invariant (ADR §7: "assert zero pending #3212 markers remain") — ground
truth confirmed trivially true today (no phase left any such marker
behind), now regression-locked going forward.

Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md
Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): correct rule category mislabel, add CI test-scope entry

An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs
mistakenly carried meta.docs.category: 'Portability', copied from a sibling
rule without realizing what that implied: docs/contributing/cross-platform-
portability-rules.md governs an ADR-1703 rule family under a hard "zero
escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's
PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is
not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic —
and its eslint-disable-next-line suppressions (159 of them, added earlier
this same phase after empirical benign-verification) are an intentional,
correct design, not a bypass. Corrected to category: 'Best Practices',
matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the
same epic, which is also correctly outside PROTECTED_RULES), and the rule's
own docstring now states this explicitly so a future reader doesn't have to
re-derive it.

Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule
or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their
own test suites under targeted CI selection — was previously unregistered
and invisible to that fast-path (this PR's own gsd-test checkpoint runs the
full suite regardless, so this only affects future narrowly-scoped PRs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic)

Security review found the rule meant to catch algorithmic-complexity bugs
had one of its own: hasUnboundedBroadQuantifier's negated-class inner
scan walked from each `[^` occurrence to the next `]` (or EOF) with no
bound, while the outer loop only ever advanced by one character — O(n²)
total work on a pattern with many unclosed `[^` runs. Runs unconditionally
inside checkPattern on any `new RegExp('literal string')` argument in any
linted file, before the (cheap) readFileSync data-flow gate — so a single
crafted string literal, no valid regex syntax required, could make
`npm run lint` / CI hang.

Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/
16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n,
quadratic); extrapolated, the 300000-char repro from the finding would
run ~165s. Post-fix (bail the inner scan once units exceeds the rule's
own 1-2-unit scope, rather than continuing to hunt for a closing `]`),
the same 300000-char input runs in 8.7ms via the real rule module,
independently reconfirmed at 18ms via a fresh Linter.verify() call.

New regression row in tests/no-unbounded-quantifier.rule.test.cjs
asserts the RuleTester run on a 50000-char adversarial pattern
completes and returns a defined result — no wall-clock assertion
(CLAUDE.md Clock Seams / local/no-elapsed-assertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch

next merged 12 more PRs during this PR's review. Two consequences:

- tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new
  content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own
  workflow .md content — the same Class A pattern as the ~159 sites
  already triaged elsewhere in this PR. Suppressed with the same
  established reason.
- lint-allow-test-rule-refs' ratchet ceiling needed re-raising again
  (301 -> 303) for the same reason as the two prior bumps: organic
  growth from unrelated, already-reviewed PRs landing concurrently,
  not a defect in this branch's own diff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 10:02:28 -04:00

199 lines
9.9 KiB
JavaScript

'use strict';
// Issue #966 — behavior-dependent must-haves must not pass on symbol presence.
// Content-assertion contract for the gsd-verifier agent: the
// PRESENT_BEHAVIOR_UNVERIFIED per-truth state, its routing to human_needed,
// the behavior-verified score split, and the parity invariant that the new
// per-truth state never leaks into the overall-status vocabulary.
const test = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const ROOT = path.join(__dirname, '..');
const verifierPath = path.join(ROOT, 'agents', 'gsd-verifier.md');
const verifier = fs.readFileSync(verifierPath, 'utf-8');
const standaloneTemplatePath = path.join(ROOT, 'gsd-core', 'templates', 'verification-report.md');
const standalone = fs.readFileSync(standaloneTemplatePath, 'utf-8');
test('Step 3 defines the PRESENT_BEHAVIOR_UNVERIFIED per-truth state', () => {
assert.match(verifier, /PRESENT_BEHAVIOR_UNVERIFIED/);
assert.match(verifier, /present[^\n]*wired|wired[^\n]*present/i);
});
test('behavior-dependent trigger names transition + cancellation/cleanup/ordering invariants', () => {
assert.match(verifier, /state transition/i);
assert.match(verifier, /cancellation|cleanup|ordering/i);
assert.match(verifier, /invariant/i);
});
test('PRESENT_BEHAVIOR_UNVERIFIED routes to human verification', () => {
assert.match(verifier, /PRESENT_BEHAVIOR_UNVERIFIED[\s\S]{0,400}?human/i);
});
test('PRESENT_BEHAVIOR_UNVERIFIED is the only truth excluded from the verified score', () => {
assert.match(
verifier,
/(do not count it toward the verified score)|(only[^\n]*excluded from `verified_truths`)|(only truths excluded)/i,
);
});
test('score still credits PASSED (override) truths (override contract preserved)', () => {
// The Step 9 score definition must count override-passed truths in verified_truths.
assert.match(
verifier,
/verified_truths[\s\S]{0,200}?PASSED \(override\)/,
'Step 9 score must count PASSED (override) truths in verified_truths',
);
});
test('behavior-unverified truths get a structured frontmatter list that survives gaps_found', () => {
assert.match(verifier, /behavior_unverified_items/);
// and it must NOT be gated only to human_needed (must mention it is emitted regardless of status / when count > 0)
assert.match(
verifier,
/behavior_unverified_items[\s\S]{0,160}?(regardless of (overall )?status|count > 0)/i,
);
});
test('Step 9 / template carry the behavior_unverified score-split field', () => {
assert.match(verifier, /behavior_unverified/);
});
test('critical_rules calibrates "presence is not behavior" without dropping the speed guard', () => {
assert.match(verifier, /presence is not behavior/i);
assert.match(verifier, /Keep verification fast/);
});
test('PARITY: per-truth state never leaks into the overall-status vocabulary', () => {
assert.doesNotMatch(verifier, /→ \*\*status:\s*present_behavior_unverified\*\*/i);
const unionLines = verifier.match(/^status:\s+[a-z_]+(?:\s*\|\s*[a-z_]+)+\s*$/gm) || [];
assert.ok(unionLines.length > 0, 'expected at least one status union line');
for (const line of unionLines) {
assert.doesNotMatch(line, /present_behavior_unverified/i);
// The real invariant is that the per-truth state is NOT in the union (above).
// Membership (order-independent) avoids brittleness on a future legitimate reorder.
for (const s of ['passed', 'gaps_found', 'human_needed']) {
assert.ok(line.includes(s), `status union must still contain ${s}: ${line}`);
}
}
});
test('overall-status enum in verification.cts is unchanged (no per-truth leak)', () => {
const cts = fs.readFileSync(path.join(ROOT, 'src', 'verification.cts'), 'utf-8');
// eslint-disable-next-line local/no-unbounded-quantifier -- parses this repo's own bounded src/verification.cts source, not adversarial input
const m = cts.match(/VERIFIER_STATUSES[^=]*=\s*\[([^\]]*)\]/);
assert.ok(m, 'VERIFIER_STATUSES array must be present');
assert.doesNotMatch(m[1], /present_behavior_unverified/i);
for (const s of ['passed', 'gaps_found', 'human_needed']) {
assert.match(m[1], new RegExp(`'${s}'`));
}
});
test('VERIFICATION.md templates carry behavior_unverified + the new truth-state', () => {
assert.match(verifier, /behavior_unverified/);
assert.match(standalone, /PRESENT_BEHAVIOR_UNVERIFIED/);
assert.match(standalone, /behavior_unverified/);
assert.match(verifier, /behavior_unverified_items/);
assert.match(standalone, /behavior_unverified_items/);
});
const verifyPhase = fs.readFileSync(path.join(ROOT, 'gsd-core', 'references', 'verifier-phase-gates.md'), 'utf-8');
const planningArtifacts = fs.readFileSync(path.join(ROOT, 'docs', 'reference', 'planning-artifacts.md'), 'utf-8');
test('shipped verifier-phase-gates reference mirrors the behavior-unverified calibration', () => {
assert.match(verifyPhase, /PRESENT_BEHAVIOR_UNVERIFIED/);
assert.match(verifyPhase, /behavior_unverified/);
assert.match(verifyPhase, /state transition/i);
});
test('planning-artifacts reference documents the behavior-unverified calibration', () => {
assert.match(planningArtifacts, /PRESENT_BEHAVIOR_UNVERIFIED/);
assert.match(planningArtifacts, /behavior_unverified/);
});
test('Step 9 keeps gaps_found precedence and preserves behavior-unverified items', () => {
assert.match(verifier, /gaps_found's precedence|gaps_found[\s\S]{0,160}?precedence/i);
assert.match(verifier, /behavior_unverified_items[\s\S]{0,120}?(never lost|survive|regardless)/i);
});
test('shipped workflow flags behavior-unverified truths even on infrastructure phases', () => {
assert.match(
verifyPhase,
/PRESENT_BEHAVIOR_UNVERIFIED[\s\S]{0,400}?infrastructure|infrastructure[\s\S]{0,400}?PRESENT_BEHAVIOR_UNVERIFIED/i,
);
});
test('standalone template per-truth guideline respects gaps_found precedence', () => {
assert.match(standalone, /becomes `human_needed`[\s\S]{0,80}?gaps_found/i);
});
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-3321-verifier-runs-probes.test.cjs — consolidation epic #1969 (B7 #1976)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-3321-verifier-runs-probes (consolidation epic #1969 B7 #1976)", () => {
'use strict';
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const REPO_ROOT = path.join(__dirname, '..');
const VERIFIER_AGENT = path.join(REPO_ROOT, 'agents', 'gsd-verifier.md');
function verifierProbeContract(content) {
const sectionStart = content.indexOf('## Step 7c: Probe Execution');
const sectionEnd = content.indexOf('## Step 8:', sectionStart);
assert.notEqual(sectionStart, -1, 'verifier must define Step 7c');
assert.notEqual(sectionEnd, -1, 'verifier must close Step 7c before Step 8');
const section = content.slice(sectionStart, sectionEnd);
const codeBlocks = [...section.matchAll(/```bash\r?\n([\s\S]*?)\r?\n```/g)].map((match) => match[1].split(/\r?\n/).join('\n'));
const executionSteps = [...section.matchAll(/^\d+\.\s+(.+)$/gm)].map((match) => match[1]);
return {
title: 'Step 7c: Probe Execution',
conventionalDiscoveryCommand: codeBlocks[0]?.split('\n').find((line) => line.startsWith('find scripts')) || null,
declaredDiscoveryCommand: codeBlocks[0]?.split('\n').find((line) => line.startsWith('grep -R')) || null,
executionCommand: codeBlocks[1] || '',
executionSteps,
statusRows: [...section.matchAll(/^\|\s*`([^`]+)`\s*\|\s*`([^`]+)`\s*\|[^|]+\|\s*([^|]+)\|$/gm)]
.map((match) => ({ probe: match[1], command: match[2], statuses: match[3].trim() })),
summaryClaimsRejected: section.includes('SUMMARY.md probe pass claims are not evidence'),
};
}
describe('bug #3321: gsd-verifier runs probes instead of trusting SUMMARY claims', () => {
test('verifier prompt requires direct probe discovery and execution', () => {
const content = fs.readFileSync(VERIFIER_AGENT, 'utf8');
const contract = verifierProbeContract(content);
assert.equal(contract.title, 'Step 7c: Probe Execution');
assert.equal(contract.conventionalDiscoveryCommand, "find scripts -path '*/tests/probe-*.sh' -type f 2>/dev/null | sort");
assert.equal(
contract.declaredDiscoveryCommand,
"grep -R -n -E 'probe-[^[:space:]]+\\.sh|scripts/.*/tests/probe-.*\\.sh' \"$PHASE_DIR\"/*-PLAN.md \"$PHASE_DIR\"/*-SUMMARY.md 2>/dev/null",
);
assert.deepEqual(contract.executionSteps, [
'Build the `PROBES` list from explicit PLAN declarations first; include conventional `scripts/*/tests/probe-*.sh` when the phase is a migration/tooling phase or the success criteria mention probes.',
'For every documented probe path, if the file is missing or unreadable, mark `MISSING_PROBE` and set `status: gaps_found`. Do not require the executable bit because probes run through `bash "$probe"`.',
'Run each probe from the built `PROBES` list from the repository root:',
'Exit code 0 is PASS. Any non-zero exit is FAILED and must include stdout/stderr evidence in VERIFICATION.md.',
'Do not substitute executor narration, SUMMARY.md PASS-marker counts, or a different dry-run driver command for the probe result.',
]);
assert.equal(contract.executionCommand, 'for probe in "${PROBES[@]}"; do\n gsd_run run-with-timeout 30 -- bash "$probe"\ndone');
assert.deepEqual(contract.statusRows, [{
probe: 'scripts/.../probe-name.sh',
command: 'bash "$probe"',
statuses: 'PASS / FAILED / MISSING_PROBE',
}]);
assert.equal(contract.summaryClaimsRejected, true);
});
});
});
}