From 31b822b3122020d7e68aa2ada70cb5bde3b8640d Mon Sep 17 00:00:00 2001 From: Dave Date: Mon, 15 Jun 2026 13:38:15 -0400 Subject: [PATCH] =?UTF-8?q?fix(1259-01):=20close=20adversarial=20review=20?= =?UTF-8?q?findings=20=E2=80=94=20genuine=20enforcement,=20honest=20fail-f?= =?UTF-8?q?irst=20scope?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adversarial pre-submission review found the injected-runCheck tests masked a non-functional real runner. Fixes: - BL-01 (false green on vacuous test): the node-test runner now parses the TAP summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a reported test named distinctly from the file — node --test counts an empty file as one passing test, so counts alone could not catch it. - SF-01 (lint anchor never greened): the lint-rule runner now runs the project eslint as --format json and filters by ruleId, so plugin rules (local/*) load via the flat config — bare --rule cannot load a plugin. local/no-source-grep now genuinely greens (covered by a real, non-injected test). - BL-02 (tautological fail-first): the runner no longer echoes the caller's failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the producer requires attestation + a genuine non-vacuous pass. Machine-proven fail-first (needs a violation fixture) is flagged as a tracked follow-up in ADR-550, the changeset, FEATURES, the reference doc, and verify-phase. - SF-02: added real-runner end-to-end tests (no injected runCheck) + pure, exported parse/filter helpers (parseNodeTestSummary, tapTestNames, eslintJsonHasRule, eslintFileResultCount) so the shipping branches are mutation-pinned. - NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds. - Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an ambient test-runner context cannot corrupt a verify-time result. - Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim). --- .changeset/1259-test-tier-enforcement.md | 6 +- docs/FEATURES.md | 2 +- docs/adr/550-spec-phase-probe-contract.md | 5 +- gsd-core/references/prohibition-probe.md | 13 +- gsd-core/workflows/verify-phase.md | 6 +- src/prohibition-enforcement.cts | 265 ++++++++++++++----- tests/prohibition-enforcement.test.cjs | 154 +++++++++-- tests/prohibition-probe.verify-tier.test.cjs | 10 +- 8 files changed, 345 insertions(+), 116 deletions(-) diff --git a/.changeset/1259-test-tier-enforcement.md b/.changeset/1259-test-tier-enforcement.md index a0b02d905..e1568dd27 100644 --- a/.changeset/1259-test-tier-enforcement.md +++ b/.changeset/1259-test-tier-enforcement.md @@ -3,6 +3,8 @@ type: Changed pr: 1259 --- -**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, confirms it is fail-first (`regression-must-fail-first`), runs it, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a PASSING wired check disposes `green` and can reach `passed`, while a missing, failing, or non-fail-first check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test AND a lint/AST rule, anchored on the in-tree `no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259) +**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, runs it for a genuine **non-vacuous** pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a genuinely-passing wired check disposes `green` and can reach `passed`, while a missing, non-attested, or non-passing check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test (requiring a real reported test — an empty file, which `node --test` counts as one passing "test", does NOT green) AND a lint/AST rule run as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot), anchored on the in-tree `local/no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259) -**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. The definitive count is **26 invalid cases / 23 error-expectation objects** across all rule-test files; the `no-source-grep` anchor itself has exactly **2 invalid cases** (`tests/eslint-rules.test.cjs`, the `.includes()` and `.match()` invalid blocks). Those 2 cases ARE genuine `regression-must-fail-first` proofs, so the anchor argument is unaffected — but the count is "the 2 `no-source-grep` invalid cases," not 96. +**Honest scope — `failFirst` is caller-attested, not yet machine-proven.** This lands the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and the check to genuinely run and pass. It does NOT yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property) — cheap proof of that at verify time needs running the check against a known violation fixture, which is a **tracked follow-up**. The red-first property currently rests on caller attestation, surfaced transparently in the evidence record. + +**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. For the `no-source-grep` anchor specifically, the genuine `regression-must-fail-first` proofs are its **two `invalid` cases** (the `.includes()` and `.match()` blocks) in `tests/eslint-rules.test.cjs` — not 96. The anchor argument is unaffected (those two cases ARE real fail-first proofs); only the count was off. diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 396079842..293ec3343 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3180,6 +3180,6 @@ The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, s - REQ-PROHIB-04: `--auto` MUST never auto-dismiss. - REQ-PROHIB-05: `plan-phase` MUST lift resolved prohibitions into `must_haves.prohibitions` (never `truths`). - REQ-PROHIB-06: A well-formed but unwired `test`-tier prohibition MUST fail closed at verify time — never a silent pass. -- REQ-PROHIB-07: A `test`-tier prohibition with a PASSING wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing or failing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half). +- REQ-PROHIB-07: A `test`-tier prohibition with a caller-attested, genuinely-passing (non-vacuous) wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, non-attested, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half; machine-proven fail-first is a tracked follow-up). **Reference:** [Prohibition Probe](../gsd-core/references/prohibition-probe.md) diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md index 015b5ef0d..8dba554b0 100644 --- a/docs/adr/550-spec-phase-probe-contract.md +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -90,8 +90,9 @@ The capability system (ADR-857) classifies **predicate-generation as core verifi Decision 4 describes the `test`-tier as a "**Hard gate in both interactive and autonomous modes.**" The #644 implementation revises that to a **fail-closed-now / deferred-enforcement** resolution (the "B-with-guard" maintainer decision of 2026-06-12), so the architecture-of-record matches the shipped code: - A well-formed but **unwired** `test`-tier prohibition resolves via `dispositionForProhibition()` to `{ status: 'unverified', flagged: true }` — **provably never green** without explicit evidence (REQ-PROHIB-06). This is the load-bearing safety half and it holds today. -- The **heavy negative-test enforcement mechanism** — locating the wired mechanical check, confirming it is fail-first, running it, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test OR a lint/AST rule — and is anchored on the in-tree `no-source-grep` AST rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, failing, or non-fail-first check hard-gates (flagged, non-green) in both interactive and autonomous modes. +- The **negative-test enforcement mechanism** — locating the wired mechanical check, running it for a genuine **non-vacuous** pass, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test (requiring a real reported test, not the empty file `node --test` would count as one passing "test") OR a lint/AST rule run through the project flat config as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot) — and is anchored on the in-tree `local/no-source-grep` rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, non-attested, or genuinely-non-passing check hard-gates (flagged, non-green) in both interactive and autonomous modes. +- **Honest scope — `failFirst` is caller-ATTESTED, not machine-proven (tracked follow-up).** What #1259 lands is the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and requires the check to genuinely run and pass. It does **not** yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property), because cheap proof of that at verify time needs running the check against a known **violation fixture** — deferred as a follow-up. Until then the red-first property rests on caller attestation, surfaced transparently in the evidence record. This closes the permanent-`gaps_found` dead-end with a genuinely-executed gate without overclaiming machine-proven fail-first. -Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the mechanical-pass half — a test-tier prohibition with a passing wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The unqualified "hard gate" wording of Decision 4 is now fully realized for `test`-tier items: the previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets. +Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the genuine-execution half — a test-tier prohibition with a passing, non-vacuous wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The one remaining gap to D4's literal intent — *machine-proven* fail-first — is documented above as a tracked follow-up. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets. This enforcement seam is the concrete instance of **ADR-857 open-question §147** — the deferred "deterministic CI conformance test for the verifier↔predicate contract." Per D6 it lands on the **core verify rail** (non-toggleable substrate), never in `capabilities/`: the verifier consuming a contract-shaped, deterministic predicate is core, not an opt-in capability. diff --git a/gsd-core/references/prohibition-probe.md b/gsd-core/references/prohibition-probe.md index ccb34de67..2cebfa403 100644 --- a/gsd-core/references/prohibition-probe.md +++ b/gsd-core/references/prohibition-probe.md @@ -113,11 +113,14 @@ lifecycle is identical to the edge-probe, the verification tiers differ): At verify time these tiers are routed differently (ADR-550 D4): - A **test**-tier prohibition is enforced + hard-gates via the deterministic `check prohibition-enforcement` sub-command (#1259, ADR-550 D5d): it locates the wired - mechanical check (a `node --test` negative test OR a lint/AST rule), confirms it is - fail-first, runs it, and emits the `dispositionForProhibition()` verdict. A passing wired - check disposes **green** (satisfiable → can reach `passed`); a missing, failing, or - non-fail-first check **hard-gates** (flagged, never green → `gaps_found`) in BOTH - interactive and autonomous modes — never a silent pass. + mechanical check (a `node --test` negative test OR a lint/AST rule run as + `eslint --format json` and filtered by `ruleId`), requires the caller-attested `failFirst` + marker, runs it for a genuine **non-vacuous** pass, and emits the + `dispositionForProhibition()` verdict. A passing wired check disposes **green** (satisfiable + → can reach `passed`); a missing, non-attested, or genuinely-non-passing check **hard-gates** + (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes — never a + silent pass. (`failFirst` is caller-attested, not yet machine-proven against a violation + fixture — a tracked follow-up; see ADR-550 D5d.) - A **judgment**-tier prohibition routes to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag). diff --git a/gsd-core/workflows/verify-phase.md b/gsd-core/workflows/verify-phase.md index 589bd4770..993b411ba 100644 --- a/gsd-core/workflows/verify-phase.md +++ b/gsd-core/workflows/verify-phase.md @@ -76,9 +76,9 @@ Aggregate all must_haves across plans for phase-level verification. gsd_run check prohibition-enforcement ``` - where `` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, CONFIRMS it is fail-first (`regression-must-fail-first`), RUNS it, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). Route the result by its typed fields: - - **`status: 'green'`, `flagged: false`** (a passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**. - - **missing, failing, or non-fail-first check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). + where `` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, requires the caller-attested `failFirst` marker, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). (`failFirst` is caller-attested, not yet machine-proven against a violation fixture — a tracked follow-up.) Route the result by its typed fields: + - **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**. + - **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). **Option B: Use Success Criteria from ROADMAP.md** diff --git a/src/prohibition-enforcement.cts b/src/prohibition-enforcement.cts index 76833d816..d23b4298c 100644 --- a/src/prohibition-enforcement.cts +++ b/src/prohibition-enforcement.cts @@ -8,14 +8,22 @@ * branch at probe-core 420-427); with empty evidence it fails closed to flagged-unverified. But * NOTHING in the live pipeline ever produced `enforcementEvidence`, so the green branch was * unreachable. This module is the missing producer: it LOCATES the wired mechanical check from a - * check descriptor, CONFIRMS it is fail-first (regression-must-fail-first), RUNS it, builds a typed + * check descriptor, RUNS it and requires a genuine NON-VACUOUS pass, builds a typed * `enforcementEvidence` array on PASS, and emits the `dispositionForProhibition` verdict as JSON. * The green/fail-closed policy itself is untouched (no src/probe-core.cts edit). * - * Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR an existing - * lint/AST rule (e.g. the in-tree `no-source-grep` rule — the D4 dogfood anchor). A missing OR - * failing OR non-fail-first check hard-gates (flagged, non-green) in BOTH interactive and - * autonomous modes (ADR-550 D4 / D3) — never a silent green. + * Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR a lint/AST rule + * (e.g. the in-tree `local/no-source-grep` rule — the D4 dogfood anchor, run via the project flat + * config so the plugin loads). A missing, non-attested, or genuinely-non-passing check hard-gates + * (flagged, non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3) — never a silent + * green. + * + * FAIL-FIRST IS CALLER-ATTESTED (honest scope, #1259): the producer requires the caller to ATTEST + * `failFirst: true` AND requires the runner to observe a real non-vacuous pass — but it does NOT + * independently prove the check fails-on-violation. Genuine fail-first confirmation needs running + * the check against a known violation fixture and is a tracked follow-up (recorded in ADR-550's D5d + * note). What ships here closes the permanent-`gaps_found` dead-end with a genuinely-executed gate; + * it does not yet replace caller attestation with machine proof of the red-first property. * * Authored as strict TypeScript (`src/prohibition-enforcement.cts`) and compiled by * `tsc -p tsconfig.build.json` (`npm run build:lib`) to the gitignored runtime artifact @@ -43,7 +51,8 @@ export type CheckKind = 'node-test' | 'lint-rule'; * runner family; `target` is the negative-test file path (node-test) or the PATH to lint * (lint-rule); `rule` is the eslint rule id (lint-rule only — e.g. `local/no-source-grep`) and is * REQUIRED for the lint-rule kind (a lint-rule descriptor without it is not a valid wired check); - * `failFirst` records whether the check is a genuine `regression-must-fail-first` proof. + * `failFirst` is the caller's ATTESTATION that the check is a genuine `regression-must-fail-first` + * proof (caller-declared — see the module docstring; not independently confirmed at verify time). */ export interface CheckDescriptor { kind: CheckKind; @@ -52,9 +61,14 @@ export interface CheckDescriptor { failFirst?: boolean; } -/** The result a check-runner returns: whether the check is fail-first and whether it passed. */ +/** + * The result a check-runner returns: whether the check genuinely, non-vacuously PASSED. The runner + * reports only what it can OBSERVE (a real pass) — it does NOT determine fail-first. Whether the + * check is a `regression-must-fail-first` proof is a CALLER-ATTESTED property of the descriptor + * (`CheckDescriptor.failFirst`); the producer cannot independently confirm it at verify time without + * a violation fixture (a tracked follow-up — see the module docstring). + */ export interface CheckRunResult { - failFirst: boolean; passed: boolean; } @@ -85,60 +99,171 @@ export interface EnforcementResult extends ProhibitionDisposition { mode?: string; } -/** - * The default REAL check runner (used when no `runCheck` is injected). Deterministic per - * environment and guarded so a missing tool yields a non-passing result, NEVER an uncaught throw - * (the no-throw contract). A real run is fail-first by construction here — the descriptor's - * `failFirst` marker is the authoritative regression-must-fail-first signal the producer confirms. - * - node-test: runs `node --test `; exit 0 = passed. - * - lint-rule: runs `eslint --rule ': error' `, exit 0 = passed. - */ +/** node --test argv. Forces the TAP reporter so the summary counts are parseable + version-stable. */ +export function buildNodeTestArgs(check: CheckDescriptor): string[] { + return ['--test', '--test-reporter=tap', check.target]; +} + +/** eslint argv (the args AFTER `npx`). Runs the project flat config so plugin rules (e.g. `local/*`) + * load — `--rule` CANNOT load a plugin, so we lint the TARGET path as JSON and filter by rule id. */ +export function buildLintArgs(check: CheckDescriptor): string[] { + return ['eslint', '--format', 'json', check.target]; +} /** - * Pure mapper from a lint-rule descriptor to the eslint argv (the args AFTER `npx`). The rule id - * (`check.rule`, forced to `error`) and the lint TARGET path (`check.target`) are DISTINCT tokens — - * reusing `target` as both (the #1259 pre-fix bug) makes eslint try to lint a file named after the - * rule, which can never pass. Exported so the mapping is unit-testable without spawning eslint. + * Pure parser for the `node --test` TAP summary. A genuine pass is NON-VACUOUS: exit 0 is NOT enough + * (an empty / all-skipped / deleted-negative-test file exits 0 with `# tests 0`). Mutation-pinned by + * unit tests so a threshold flip is caught. */ -export function buildLintArgs(check: CheckDescriptor): string[] { - return ['eslint', '--rule', `${check.rule}: error`, check.target]; +export function parseNodeTestSummary(out: string): { tests: number; pass: number; fail: number } { + const num = (re: RegExp): number => { + const m = typeof out === 'string' ? out.match(re) : null; + return m ? Number(m[1]) : 0; + }; + return { + tests: num(/^# tests (\d+)/m), + pass: num(/^# pass (\d+)/m), + fail: num(/^# fail (\d+)/m), + }; +} + +/** The names from TAP `ok N - ` / `not ok N - ` lines (directives like `# SKIP` stripped). */ +export function tapTestNames(out: string): string[] { + if (typeof out !== 'string') return []; + const names: string[] = []; + const re = /^(?:not )?ok \d+ - (.+)$/gm; + let m: RegExpExecArray | null; + while ((m = re.exec(out)) !== null) { + names.push(m[1].replace(/\s+#\s.*$/, '').trim()); + } + return names; +} + +/** + * A non-vacuous node-test pass: at least one test, at least one pass, zero failures — AND at least + * one reported test whose name is NOT merely the target file. `node --test ` counts a file + * with ZERO `test()` calls as one passing "test" named after the file, so the counts alone cannot + * tell an empty/deleted negative test from a real one (the #1259 BL-01 false-green). Requiring a + * named test distinct from the file closes that hole. + */ +export function isNonVacuousNodeTestPass(out: string, target: string): boolean { + const s = parseNodeTestSummary(out); + if (!(s.tests >= 1 && s.pass >= 1 && s.fail === 0)) return false; + const base = typeof target === 'string' ? (target.split(/[\\/]/).pop() ?? target) : ''; + return tapTestNames(out).some((n) => n !== base && n !== target); +} + +/** Number of file results in an eslint `--format json` report (0 if unparseable / not an array). */ +export function eslintFileResultCount(jsonText: string): number { + try { + const parsed: unknown = JSON.parse(jsonText); + return Array.isArray(parsed) ? parsed.length : 0; + } catch { + return 0; + } +} + +/** True if the eslint `--format json` report has ANY message for `rule`. Unparseable -> true + * (fail-closed: treat an unreadable report as a violation rather than a silent pass). */ +export function eslintJsonHasRule(jsonText: string, rule: string): boolean { + let parsed: unknown; + try { + parsed = JSON.parse(jsonText); + } catch { + return true; + } + if (!Array.isArray(parsed)) return true; + for (const file of parsed) { + const messages = file && typeof file === 'object' && Array.isArray((file as { messages?: unknown }).messages) + ? (file as { messages: Array<{ ruleId?: unknown }> }).messages + : []; + for (const msg of messages) { + if (msg && typeof msg === 'object' && msg.ruleId === rule) return true; + } + } + return false; +} + +/** + * The default REAL check runner (used when no `runCheck` is injected). Reports only an OBSERVED, + * genuinely-non-vacuous pass; guarded so a missing tool / non-zero exit yields a non-passing result, + * NEVER an uncaught throw (the no-throw contract). It does NOT determine fail-first (caller-attested). + * - node-test: runs `node --test` (TAP) and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail + * AND a reported test named distinctly from the file). A bare exit 0 for an empty/zero-test file + * — which `node --test` counts as one passing "test" named after the file — is NOT a pass (the + * #1259 BL-01 false-green fix). + * - lint-rule: runs the project `eslint --format json ` (flat config loads `local/*` + * plugins) and requires the target to actually lint (>=1 file result) AND ZERO messages for the + * specific rule id. `--rule` cannot load a plugin rule, so we filter the structured report by + * `ruleId` instead (the #1259 SF-01 fix). + */ +/** + * Env for spawned checks: strip `NODE_TEST_CONTEXT` and `NODE_OPTIONS` so an AMBIENT test-runner + * context (e.g. running verify under `node --test`, which sets `NODE_TEST_CONTEXT=child-v8`) cannot + * turn the child `node --test` into a silent v8-reporter worker that emits no parseable TAP — which + * would otherwise corrupt the verdict. Deterministic, environment-independent execution. + */ +function childEnv(): NodeJS.ProcessEnv { + const env = { ...process.env }; + delete env.NODE_TEST_CONTEXT; + delete env.NODE_OPTIONS; + return env; } function defaultRunCheck(check: CheckDescriptor, cwd: string): CheckRunResult { - const failFirst = check.failFirst === true; try { if (check.kind === 'node-test') { - execFileSync('node', ['--test', check.target], { - cwd, - encoding: 'utf-8', - stdio: 'ignore', - windowsHide: true, - }); - return { failFirst, passed: true }; + let out = ''; + try { + out = execFileSync('node', buildNodeTestArgs(check), { + cwd, + encoding: 'utf-8', + stdio: ['ignore', 'pipe', 'pipe'], + windowsHide: true, + env: childEnv(), + }); + } catch (e) { + // A failing test run exits non-zero (TAP still on stdout). Parse it: a real failure has + // `# fail >= 1` -> non-vacuous check returns false. Missing `node` -> no stdout -> false. + const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : ''; + out = typeof stdout === 'string' ? stdout : ''; + } + return { passed: isNonVacuousNodeTestPass(out, check.target) }; } - // lint-rule: force the rule to error and lint the TARGET path (distinct from the rule id). A - // clean exit (0) means no violation surfaced -> the must-NOT holds across the target. - execFileSync('npx', buildLintArgs(check), { - cwd, - encoding: 'utf-8', - stdio: 'ignore', - windowsHide: true, - }); - return { failFirst, passed: true }; + if (check.kind === 'lint-rule') { + let json = ''; + try { + json = execFileSync('npx', buildLintArgs(check), { + cwd, + encoding: 'utf-8', + stdio: ['ignore', 'pipe', 'pipe'], + windowsHide: true, + env: childEnv(), + }); + } catch (e) { + // eslint exits non-zero when ANY error is present; the JSON report is still on stdout. + const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : ''; + json = typeof stdout === 'string' ? stdout : ''; + } + const lintedSomething = eslintFileResultCount(json) >= 1; + return { passed: lintedSomething && !eslintJsonHasRule(json, check.rule as string) }; + } + // Unknown kind — defensive; the LOCATE guard already rejects it. + return { passed: false }; } catch { - // Non-zero exit (violation surfaced) OR missing tool -> not passing. Never throw. - return { failFirst, passed: false }; + return { passed: false }; } } /** * LOCATE -> CONFIRM fail-first -> RUN -> build enforcementEvidence -> dispositionForProhibition. * - * (1) LOCATE: if no check descriptor is locatable -> fail-closed (`dispositionForProhibition` with - * empty evidence) plus `{ located: false, kind: null, evidence: [] }`. - * (2) CONFIRM + RUN: confirm the descriptor is fail-first and run it via `runCheck`. A check that - * is not fail-first, that the runner reports not fail-first, or that FAILS -> fail-closed - * disposition with `located: true` (a real located miss, non-green, flagged) in BOTH modes. + * (1) LOCATE: if no well-formed check descriptor is locatable -> fail-closed + * (`dispositionForProhibition` with empty evidence) plus `{ located: false, kind: null, evidence: [] }`. + * (2) ATTEST + RUN: require the caller to attest `failFirst: true` and run it via `runCheck`. A check + * the caller does not attest as fail-first, or that does not genuinely (non-vacuously) PASS -> + * fail-closed disposition with `located: true` (a real located miss, non-green, flagged) in BOTH + * modes. Fail-first is caller-attested, not independently proven (see module docstring). * (3) PASS: build a typed `enforcementEvidence` array and call `dispositionForProhibition` — the * non-empty array flips a test-tier item to green (the previously-unreachable branch). * @@ -151,46 +276,48 @@ export function runProhibitionEnforcement( ): EnforcementResult { const mode = options.mode; - // (1) LOCATE — no locatable wired check -> fail-closed, located: false. A lint-rule descriptor - // MUST also carry a string `rule` id (its target is the lint PATH, not the rule) — an - // under-specified lint-rule is not a valid wired check, so it is not locatable. - if ( - !check || - typeof check !== 'object' || - typeof check.kind !== 'string' || - typeof check.target !== 'string' || - (check.kind === 'lint-rule' && typeof check.rule !== 'string') - ) { + // (1) LOCATE — no locatable, well-formed wired check -> fail-closed, located: false. The kind must + // be one of the two known kinds; the target must be a non-empty string; a lint-rule descriptor MUST + // also carry a non-empty `rule` id (its target is the lint PATH, not the rule). An under-specified + // descriptor is not a valid wired check, so it is not locatable (does not rely on the runner failing). + const c = check && typeof check === 'object' ? check : null; + const validKind = !!c && (c.kind === 'node-test' || c.kind === 'lint-rule'); + const validTarget = !!c && typeof c.target === 'string' && c.target.trim().length > 0; + const validRule = !!c && (c.kind !== 'lint-rule' || (typeof c.rule === 'string' && c.rule.trim().length > 0)); + if (!c || !validKind || !validTarget || !validRule) { const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] }); return { ...disposition, located: false, kind: null, evidence: [], ...(mode ? { mode } : {}) }; } - const runCheck = options.runCheck ?? ((c: CheckDescriptor) => defaultRunCheck(c, options.cwd ?? process.cwd())); + const runCheck = options.runCheck ?? ((toRun: CheckDescriptor) => defaultRunCheck(toRun, options.cwd ?? process.cwd())); - // (2) CONFIRM fail-first + RUN. Descriptor must declare fail-first AND the runner must agree. - const descriptorFailFirst = check.failFirst === true; - const run = runCheck(check); - const failFirstConfirmed = descriptorFailFirst && run.failFirst === true; - const passed = failFirstConfirmed && run.passed === true; + // (2) ATTEST fail-first (CALLER-DECLARED) + RUN. The caller must attest `failFirst: true` AND the + // runner must observe a genuine NON-VACUOUS pass. The producer does NOT independently prove + // fail-first (that needs a violation fixture — tracked follow-up, ADR-550 D5d). A non-attested or + // non-passing check hard-gates (never green) in BOTH modes. + const attestedFailFirst = c.failFirst === true; + const run = runCheck(c); + const passed = attestedFailFirst && run.passed === true; if (!passed) { - // FAIL / not-fail-first -> fail-closed, located: true (an actual located miss/fail). Hard-gate - // applies in BOTH modes; the disposition stays non-green / flagged. + // NOT attested fail-first OR did not genuinely pass -> fail-closed, located: true (an actual + // located miss/fail). Hard-gate applies in BOTH modes; the disposition stays non-green / flagged. const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] }); return { ...disposition, located: true, - kind: check.kind, + kind: c.kind, evidence: [], ...(mode ? { mode } : {}), }; } // (3) PASS -> build typed enforcementEvidence and let the policy flip a test-tier item green. + // `failFirst` here is the caller's attestation (recorded for provenance), not a machine proof. const evidence: EnforcementEvidence[] = [{ - kind: check.kind, - target: check.target, - ...(typeof check.rule === 'string' ? { rule: check.rule } : {}), + kind: c.kind, + target: c.target, + ...(typeof c.rule === 'string' ? { rule: c.rule } : {}), failFirst: true, passed: true, }]; @@ -198,7 +325,7 @@ export function runProhibitionEnforcement( return { ...disposition, located: true, - kind: check.kind, + kind: c.kind, evidence, ...(mode ? { mode } : {}), }; diff --git a/tests/prohibition-enforcement.test.cjs b/tests/prohibition-enforcement.test.cjs index 2c2ba10e0..142a1e539 100644 --- a/tests/prohibition-enforcement.test.cjs +++ b/tests/prohibition-enforcement.test.cjs @@ -39,7 +39,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR test('locate-miss (no check descriptor) -> fail-closed, located:false, no evidence', () => { const enforce = require(ENFORCEMENT_LIB); const result = enforce.runProhibitionEnforcement(TEST_TIER, null, { - runCheck: () => ({ failFirst: true, passed: true }), + runCheck: () => ({ passed: true }), }); assert.equal(result.located, false, 'no locatable check'); assert.notEqual(result.status, 'green', 'locate-miss must never be green'); @@ -51,7 +51,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR test('malformed check descriptor (missing target) -> treated as locate-miss', () => { const enforce = require(ENFORCEMENT_LIB); const result = enforce.runProhibitionEnforcement(TEST_TIER, { kind: 'node-test' }, { - runCheck: () => ({ failFirst: true, passed: true }), + runCheck: () => ({ passed: true }), }); assert.equal(result.located, false, 'a descriptor without a target is not locatable'); assert.notEqual(result.status, 'green'); @@ -63,7 +63,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.equal(result.status, 'green'); assert.equal(result.flagged, false); @@ -83,7 +83,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.equal(result.status, 'green'); assert.equal(result.flagged, false); @@ -93,18 +93,19 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR assert.equal(result.evidence[0].target, 'tests/', 'evidence records the linted target path, not the rule id'); }); - test('buildLintArgs maps the rule id and lint target to DISTINCT eslint args (#1259 runner fix)', () => { + test('buildLintArgs runs the project eslint as JSON over the target (plugins load via flat config; #1259 SF-01)', () => { const enforce = require(ENFORCEMENT_LIB); assert.equal(typeof enforce.buildLintArgs, 'function', - 'must export buildLintArgs — the pure argv mapper for the lint-rule real runner'); + 'must export buildLintArgs — the eslint argv builder for the lint-rule real runner'); const argv = enforce.buildLintArgs({ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/' }); assert.ok(Array.isArray(argv), 'argv is an array'); - const ruleIdx = argv.indexOf('--rule'); - assert.ok(ruleIdx !== -1, 'forces a specific rule via --rule'); - assert.equal(argv[ruleIdx + 1], 'local/no-source-grep: error', 'the rule id is forced to error'); + assert.equal(argv[0], 'eslint'); + const fmtIdx = argv.indexOf('--format'); + assert.ok(fmtIdx !== -1 && argv[fmtIdx + 1] === 'json', + 'emits --format json so the report can be filtered by ruleId'); + assert.ok(!argv.includes('--rule'), + 'must NOT use --rule — it cannot load a plugin rule like local/no-source-grep (the SF-01 bug)'); assert.equal(argv[argv.length - 1], 'tests/', 'the LAST arg is the lint target path'); - assert.notEqual(argv[ruleIdx + 1], argv[argv.length - 1], - 'the rule id must NOT be reused as the lint target (the bug this guards)'); }); test('lint-rule descriptor missing its rule id -> locate-miss, never green', () => { @@ -112,7 +113,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'lint-rule', target: 'tests/', failFirst: true }, // no `rule` - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.notEqual(result.status, 'green', 'a lint-rule with no rule id is not a valid wired check'); assert.equal(result.flagged, true); @@ -124,7 +125,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: false }) }, + { runCheck: () => ({ passed: false }) }, ); assert.notEqual(result.status, 'green'); assert.equal(result.flagged, true); @@ -132,24 +133,17 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR assert.equal(result.evidence.length, 0, 'a failing check builds no evidence'); }); - test('fail-first NOT satisfied (descriptor or runner) -> hard-gate, never green', () => { + test('caller does NOT attest fail-first (descriptor failFirst:false) -> hard-gate, never green', () => { const enforce = require(ENFORCEMENT_LIB); - // descriptor declares failFirst:false - const a = enforce.runProhibitionEnforcement( + // fail-first is caller-attested (#1259 BL-02): a check the caller does not attest as fail-first + // is not a valid regression proof and must never green, even if the run passes. + const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: false }, - { runCheck: () => ({ failFirst: false, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); - assert.notEqual(a.status, 'green', 'not-fail-first is not a valid regression proof'); - assert.equal(a.flagged, true); - // descriptor says failFirst:true but runner reports failFirst:false -> still hard-gate - const b = enforce.runProhibitionEnforcement( - TEST_TIER, - { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: false, passed: true }) }, - ); - assert.notEqual(b.status, 'green', 'runner-reported not-fail-first must also hard-gate'); - assert.equal(b.flagged, true); + assert.notEqual(result.status, 'green', 'a non-attested check is not a valid regression proof'); + assert.equal(result.flagged, true); }); test('hard-gates in BOTH modes on a failing check (ADR-550 D4)', () => { @@ -158,7 +152,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: false }), mode }, + { runCheck: () => ({ passed: false }), mode }, ); assert.notEqual(result.status, 'green', `non-green in ${mode}`); assert.equal(result.flagged, true, `flagged in ${mode}`); @@ -171,7 +165,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR const result = enforce.runProhibitionEnforcement( TEST_TIER, { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: true }), mode: 'autonomous' }, + { runCheck: () => ({ passed: true }), mode: 'autonomous' }, ); assert.equal(result.status, 'green', 'a passing wired check is green in autonomous mode too'); assert.equal(result.mode, 'autonomous'); @@ -208,3 +202,105 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR assert.ok(Array.isArray(parsed.evidence), 'evidence is an array'); }); }); + +// ─── Real-runner helpers (mutation-pinned; #1259 BL-01 / SF-01) ───────────────── +// These pin the deterministic parsing/threshold logic of the REAL runner so a Stryker mutant that +// weakens "non-vacuous pass" or the ruleId filter is caught — the contract the injected-runner tests +// above deliberately bypass. +describe('prohibition-enforcement real-runner helpers (#1259)', () => { + test('parseNodeTestSummary extracts the TAP tests/pass/fail counts', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.deepEqual(enforce.parseNodeTestSummary('# tests 3\n# pass 2\n# fail 1\n'), { tests: 3, pass: 2, fail: 1 }); + assert.deepEqual(enforce.parseNodeTestSummary('no summary here'), { tests: 0, pass: 0, fail: 0 }); + }); + + test('tapTestNames extracts ok/not-ok names, stripping directives', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.deepEqual(enforce.tapTestNames('ok 1 - guards the must-NOT\nnot ok 2 - other # SKIP\n'), + ['guards the must-NOT', 'other']); + }); + + test('isNonVacuousNodeTestPass: an empty file (node names the test after the file) is NOT a pass (BL-01)', () => { + const enforce = require(ENFORCEMENT_LIB); + // node --test of a zero-test file: `ok 1 - empty.test.cjs`, `# tests 1 # pass 1` — counts alone + // cannot distinguish it from a real test, so the file-named result must NOT count as a pass. + const empty = 'ok 1 - empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(empty, 'empty.test.cjs'), false, + 'a file-named-only result is vacuous — the BL-01 false-green guard'); + const real = 'ok 1 - guards the must-NOT\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(real, 'neg.test.cjs'), true, + 'a real named test distinct from the file is a genuine pass'); + const failing = 'not ok 1 - guards\n# tests 1\n# pass 0\n# fail 1\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false, + 'any failure means not a pass'); + }); + + test('eslintJsonHasRule detects a ruleId; unparseable report -> true (fail-closed)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'local/no-source-grep' }] }]), 'local/no-source-grep'), true); + assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'other' }] }]), 'local/no-source-grep'), false); + assert.equal(enforce.eslintJsonHasRule('not json', 'local/no-source-grep'), true, + 'an unreadable report must be treated as a violation, never a silent pass'); + }); + + test('eslintFileResultCount: 0 when nothing linted (vacuity guard)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(enforce.eslintFileResultCount(JSON.stringify([{}, {}])), 2); + assert.equal(enforce.eslintFileResultCount('[]'), 0); + assert.equal(enforce.eslintFileResultCount('garbage'), 0); + }); +}); + +// ─── Real runner end-to-end (NO injected runCheck; #1259 SF-02 / BL-01 / SF-01) ── +// Spawns real subprocesses so the SHIPPING default runner is exercised — the gap that let BL-01 and +// SF-01 slip past the injected-double tests. Typed-field assertions only. +describe('prohibition-enforcement REAL runner end-to-end (#1259)', () => { + const fs = require('node:fs'); + + test('a genuine non-vacuous passing node-test greens via the real runner', (t) => { + const enforce = require(ENFORCEMENT_LIB); + const dir = createTempDir('prohib-real-pass-'); + t.after(() => cleanup(dir)); + const tf = path.join(dir, 'neg.test.cjs'); + fs.writeFileSync(tf, + "const { test } = require('node:test');\nconst assert = require('node:assert');\ntest('guards the must-NOT', () => { assert.ok(true); });\n"); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: tf, failFirst: true }, + { cwd: dir }, + ); + assert.equal(result.status, 'green', 'a real, passing, non-vacuous negative test must green'); + assert.equal(result.located, true); + assert.equal(result.evidence.length, 1); + }); + + test('an EMPTY node-test file (exit 0, zero tests) does NOT green via the real runner (BL-01)', (t) => { + const enforce = require(ENFORCEMENT_LIB); + const dir = createTempDir('prohib-real-empty-'); + t.after(() => cleanup(dir)); + const tf = path.join(dir, 'empty.test.cjs'); + fs.writeFileSync(tf, '// intentionally empty — no test cases\n'); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: tf, failFirst: true }, + { cwd: dir }, + ); + assert.notEqual(result.status, 'green', 'an empty (zero-test) file must NEVER green — fail-closed'); + assert.equal(result.located, true, 'the check was located; it just did not genuinely pass'); + assert.equal(result.evidence.length, 0); + }); + + test('a clean in-tree target greens the lint-rule kind via the real eslint runner (SF-01: plugin loads)', () => { + const enforce = require(ENFORCEMENT_LIB); + // Runs real `npx eslint --format json src/clock.cts` under the project flat config (so the + // `local` plugin loads). src/clock.cts is a clean source with no no-source-grep violation. + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'src/clock.cts', failFirst: true }, + { cwd: process.cwd() }, + ); + assert.equal(result.status, 'green', 'a clean target with no no-source-grep violation must green via real eslint'); + assert.equal(result.kind, 'lint-rule'); + assert.equal(result.evidence[0].rule, 'local/no-source-grep'); + }); +}); diff --git a/tests/prohibition-probe.verify-tier.test.cjs b/tests/prohibition-probe.verify-tier.test.cjs index 2ca7ace8e..e07f8116d 100644 --- a/tests/prohibition-probe.verify-tier.test.cjs +++ b/tests/prohibition-probe.verify-tier.test.cjs @@ -88,7 +88,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / const result = enforce.runProhibitionEnforcement( testTierProhibition, { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.ok(result && typeof result === 'object', 'result must be a structured object'); assert.equal(result.status, 'green', 'a passing wired node-test check must dispose green'); @@ -105,7 +105,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / const result = enforce.runProhibitionEnforcement( testTierProhibition, { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.equal(result.status, 'green', 'a passing wired lint-rule check must dispose green'); assert.equal(result.flagged, false, 'a green disposition must not be flagged'); @@ -120,7 +120,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / const result = enforce.runProhibitionEnforcement( testTierProhibition, null, - { runCheck: () => ({ failFirst: true, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.notEqual(result.status, 'green', 'a missing wired check must NEVER be green (fail-closed)'); assert.equal(result.flagged, true, 'a missing wired check must be flagged unverified'); @@ -136,7 +136,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / const result = enforce.runProhibitionEnforcement( testTierProhibition, { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true }, - { runCheck: () => ({ failFirst: true, passed: false }), mode }, + { runCheck: () => ({ passed: false }), mode }, ); assert.notEqual(result.status, 'green', `a failing wired check must NEVER be green (mode=${mode})`); @@ -152,7 +152,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / const result = enforce.runProhibitionEnforcement( testTierProhibition, { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: false }, - { runCheck: () => ({ failFirst: false, passed: true }) }, + { runCheck: () => ({ passed: true }) }, ); assert.notEqual(result.status, 'green', 'a check that is not fail-first is not a valid regression-must-fail-first proof — never green');