fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a non-functional real runner. Fixes: - BL-01 (false green on vacuous test): the node-test runner now parses the TAP summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a reported test named distinctly from the file — node --test counts an empty file as one passing test, so counts alone could not catch it. - SF-01 (lint anchor never greened): the lint-rule runner now runs the project eslint as --format json and filters by ruleId, so plugin rules (local/*) load via the flat config — bare --rule cannot load a plugin. local/no-source-grep now genuinely greens (covered by a real, non-injected test). - BL-02 (tautological fail-first): the runner no longer echoes the caller's failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the producer requires attestation + a genuine non-vacuous pass. Machine-proven fail-first (needs a violation fixture) is flagged as a tracked follow-up in ADR-550, the changeset, FEATURES, the reference doc, and verify-phase. - SF-02: added real-runner end-to-end tests (no injected runCheck) + pure, exported parse/filter helpers (parseNodeTestSummary, tapTestNames, eslintJsonHasRule, eslintFileResultCount) so the shipping branches are mutation-pinned. - NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds. - Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an ambient test-runner context cannot corrupt a verify-time result. - Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
This commit is contained in:
@@ -3,6 +3,8 @@ type: Changed
|
||||
pr: 1259
|
||||
---
|
||||
|
||||
**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, confirms it is fail-first (`regression-must-fail-first`), runs it, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a PASSING wired check disposes `green` and can reach `passed`, while a missing, failing, or non-fail-first check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test AND a lint/AST rule, anchored on the in-tree `no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259)
|
||||
**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, runs it for a genuine **non-vacuous** pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a genuinely-passing wired check disposes `green` and can reach `passed`, while a missing, non-attested, or non-passing check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test (requiring a real reported test — an empty file, which `node --test` counts as one passing "test", does NOT green) AND a lint/AST rule run as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot), anchored on the in-tree `local/no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259)
|
||||
|
||||
**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. The definitive count is **26 invalid cases / 23 error-expectation objects** across all rule-test files; the `no-source-grep` anchor itself has exactly **2 invalid cases** (`tests/eslint-rules.test.cjs`, the `.includes()` and `.match()` invalid blocks). Those 2 cases ARE genuine `regression-must-fail-first` proofs, so the anchor argument is unaffected — but the count is "the 2 `no-source-grep` invalid cases," not 96.
|
||||
**Honest scope — `failFirst` is caller-attested, not yet machine-proven.** This lands the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and the check to genuinely run and pass. It does NOT yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property) — cheap proof of that at verify time needs running the check against a known violation fixture, which is a **tracked follow-up**. The red-first property currently rests on caller attestation, surfaced transparently in the evidence record.
|
||||
|
||||
**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. For the `no-source-grep` anchor specifically, the genuine `regression-must-fail-first` proofs are its **two `invalid` cases** (the `.includes()` and `.match()` blocks) in `tests/eslint-rules.test.cjs` — not 96. The anchor argument is unaffected (those two cases ARE real fail-first proofs); only the count was off.
|
||||
|
||||
@@ -3180,6 +3180,6 @@ The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, s
|
||||
- REQ-PROHIB-04: `--auto` MUST never auto-dismiss.
|
||||
- REQ-PROHIB-05: `plan-phase` MUST lift resolved prohibitions into `must_haves.prohibitions` (never `truths`).
|
||||
- REQ-PROHIB-06: A well-formed but unwired `test`-tier prohibition MUST fail closed at verify time — never a silent pass.
|
||||
- REQ-PROHIB-07: A `test`-tier prohibition with a PASSING wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing or failing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half).
|
||||
- REQ-PROHIB-07: A `test`-tier prohibition with a caller-attested, genuinely-passing (non-vacuous) wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, non-attested, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half; machine-proven fail-first is a tracked follow-up).
|
||||
|
||||
**Reference:** [Prohibition Probe](../gsd-core/references/prohibition-probe.md)
|
||||
|
||||
@@ -90,8 +90,9 @@ The capability system (ADR-857) classifies **predicate-generation as core verifi
|
||||
Decision 4 describes the `test`-tier as a "**Hard gate in both interactive and autonomous modes.**" The #644 implementation revises that to a **fail-closed-now / deferred-enforcement** resolution (the "B-with-guard" maintainer decision of 2026-06-12), so the architecture-of-record matches the shipped code:
|
||||
|
||||
- A well-formed but **unwired** `test`-tier prohibition resolves via `dispositionForProhibition()` to `{ status: 'unverified', flagged: true }` — **provably never green** without explicit evidence (REQ-PROHIB-06). This is the load-bearing safety half and it holds today.
|
||||
- The **heavy negative-test enforcement mechanism** — locating the wired mechanical check, confirming it is fail-first, running it, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test OR a lint/AST rule — and is anchored on the in-tree `no-source-grep` AST rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, failing, or non-fail-first check hard-gates (flagged, non-green) in both interactive and autonomous modes.
|
||||
- The **negative-test enforcement mechanism** — locating the wired mechanical check, running it for a genuine **non-vacuous** pass, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test (requiring a real reported test, not the empty file `node --test` would count as one passing "test") OR a lint/AST rule run through the project flat config as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot) — and is anchored on the in-tree `local/no-source-grep` rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, non-attested, or genuinely-non-passing check hard-gates (flagged, non-green) in both interactive and autonomous modes.
|
||||
- **Honest scope — `failFirst` is caller-ATTESTED, not machine-proven (tracked follow-up).** What #1259 lands is the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and requires the check to genuinely run and pass. It does **not** yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property), because cheap proof of that at verify time needs running the check against a known **violation fixture** — deferred as a follow-up. Until then the red-first property rests on caller attestation, surfaced transparently in the evidence record. This closes the permanent-`gaps_found` dead-end with a genuinely-executed gate without overclaiming machine-proven fail-first.
|
||||
|
||||
Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the mechanical-pass half — a test-tier prohibition with a passing wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The unqualified "hard gate" wording of Decision 4 is now fully realized for `test`-tier items: the previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets.
|
||||
Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the genuine-execution half — a test-tier prohibition with a passing, non-vacuous wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The one remaining gap to D4's literal intent — *machine-proven* fail-first — is documented above as a tracked follow-up. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets.
|
||||
|
||||
This enforcement seam is the concrete instance of **ADR-857 open-question §147** — the deferred "deterministic CI conformance test for the verifier↔predicate contract." Per D6 it lands on the **core verify rail** (non-toggleable substrate), never in `capabilities/`: the verifier consuming a contract-shaped, deterministic predicate is core, not an opt-in capability.
|
||||
|
||||
@@ -113,11 +113,14 @@ lifecycle is identical to the edge-probe, the verification tiers differ):
|
||||
At verify time these tiers are routed differently (ADR-550 D4):
|
||||
- A **test**-tier prohibition is enforced + hard-gates via the deterministic
|
||||
`check prohibition-enforcement` sub-command (#1259, ADR-550 D5d): it locates the wired
|
||||
mechanical check (a `node --test` negative test OR a lint/AST rule), confirms it is
|
||||
fail-first, runs it, and emits the `dispositionForProhibition()` verdict. A passing wired
|
||||
check disposes **green** (satisfiable → can reach `passed`); a missing, failing, or
|
||||
non-fail-first check **hard-gates** (flagged, never green → `gaps_found`) in BOTH
|
||||
interactive and autonomous modes — never a silent pass.
|
||||
mechanical check (a `node --test` negative test OR a lint/AST rule run as
|
||||
`eslint --format json` and filtered by `ruleId`), requires the caller-attested `failFirst`
|
||||
marker, runs it for a genuine **non-vacuous** pass, and emits the
|
||||
`dispositionForProhibition()` verdict. A passing wired check disposes **green** (satisfiable
|
||||
→ can reach `passed`); a missing, non-attested, or genuinely-non-passing check **hard-gates**
|
||||
(flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes — never a
|
||||
silent pass. (`failFirst` is caller-attested, not yet machine-proven against a violation
|
||||
fixture — a tracked follow-up; see ADR-550 D5d.)
|
||||
- A **judgment**-tier prohibition routes to a never-silent / never-hard-halt soft gate
|
||||
(autonomous emits an `unverified-prohibition — human review recommended` flag).
|
||||
|
||||
|
||||
@@ -76,9 +76,9 @@ Aggregate all must_haves across plans for phase-level verification.
|
||||
gsd_run check prohibition-enforcement <request.json>
|
||||
```
|
||||
|
||||
where `<request.json>` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, CONFIRMS it is fail-first (`regression-must-fail-first`), RUNS it, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). Route the result by its typed fields:
|
||||
- **`status: 'green'`, `flagged: false`** (a passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**.
|
||||
- **missing, failing, or non-fail-first check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`).
|
||||
where `<request.json>` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, requires the caller-attested `failFirst` marker, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). (`failFirst` is caller-attested, not yet machine-proven against a violation fixture — a tracked follow-up.) Route the result by its typed fields:
|
||||
- **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**.
|
||||
- **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`).
|
||||
|
||||
**Option B: Use Success Criteria from ROADMAP.md**
|
||||
|
||||
|
||||
@@ -8,14 +8,22 @@
|
||||
* branch at probe-core 420-427); with empty evidence it fails closed to flagged-unverified. But
|
||||
* NOTHING in the live pipeline ever produced `enforcementEvidence`, so the green branch was
|
||||
* unreachable. This module is the missing producer: it LOCATES the wired mechanical check from a
|
||||
* check descriptor, CONFIRMS it is fail-first (regression-must-fail-first), RUNS it, builds a typed
|
||||
* check descriptor, RUNS it and requires a genuine NON-VACUOUS pass, builds a typed
|
||||
* `enforcementEvidence` array on PASS, and emits the `dispositionForProhibition` verdict as JSON.
|
||||
* The green/fail-closed policy itself is untouched (no src/probe-core.cts edit).
|
||||
*
|
||||
* Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR an existing
|
||||
* lint/AST rule (e.g. the in-tree `no-source-grep` rule — the D4 dogfood anchor). A missing OR
|
||||
* failing OR non-fail-first check hard-gates (flagged, non-green) in BOTH interactive and
|
||||
* autonomous modes (ADR-550 D4 / D3) — never a silent green.
|
||||
* Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR a lint/AST rule
|
||||
* (e.g. the in-tree `local/no-source-grep` rule — the D4 dogfood anchor, run via the project flat
|
||||
* config so the plugin loads). A missing, non-attested, or genuinely-non-passing check hard-gates
|
||||
* (flagged, non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3) — never a silent
|
||||
* green.
|
||||
*
|
||||
* FAIL-FIRST IS CALLER-ATTESTED (honest scope, #1259): the producer requires the caller to ATTEST
|
||||
* `failFirst: true` AND requires the runner to observe a real non-vacuous pass — but it does NOT
|
||||
* independently prove the check fails-on-violation. Genuine fail-first confirmation needs running
|
||||
* the check against a known violation fixture and is a tracked follow-up (recorded in ADR-550's D5d
|
||||
* note). What ships here closes the permanent-`gaps_found` dead-end with a genuinely-executed gate;
|
||||
* it does not yet replace caller attestation with machine proof of the red-first property.
|
||||
*
|
||||
* Authored as strict TypeScript (`src/prohibition-enforcement.cts`) and compiled by
|
||||
* `tsc -p tsconfig.build.json` (`npm run build:lib`) to the gitignored runtime artifact
|
||||
@@ -43,7 +51,8 @@ export type CheckKind = 'node-test' | 'lint-rule';
|
||||
* runner family; `target` is the negative-test file path (node-test) or the PATH to lint
|
||||
* (lint-rule); `rule` is the eslint rule id (lint-rule only — e.g. `local/no-source-grep`) and is
|
||||
* REQUIRED for the lint-rule kind (a lint-rule descriptor without it is not a valid wired check);
|
||||
* `failFirst` records whether the check is a genuine `regression-must-fail-first` proof.
|
||||
* `failFirst` is the caller's ATTESTATION that the check is a genuine `regression-must-fail-first`
|
||||
* proof (caller-declared — see the module docstring; not independently confirmed at verify time).
|
||||
*/
|
||||
export interface CheckDescriptor {
|
||||
kind: CheckKind;
|
||||
@@ -52,9 +61,14 @@ export interface CheckDescriptor {
|
||||
failFirst?: boolean;
|
||||
}
|
||||
|
||||
/** The result a check-runner returns: whether the check is fail-first and whether it passed. */
|
||||
/**
|
||||
* The result a check-runner returns: whether the check genuinely, non-vacuously PASSED. The runner
|
||||
* reports only what it can OBSERVE (a real pass) — it does NOT determine fail-first. Whether the
|
||||
* check is a `regression-must-fail-first` proof is a CALLER-ATTESTED property of the descriptor
|
||||
* (`CheckDescriptor.failFirst`); the producer cannot independently confirm it at verify time without
|
||||
* a violation fixture (a tracked follow-up — see the module docstring).
|
||||
*/
|
||||
export interface CheckRunResult {
|
||||
failFirst: boolean;
|
||||
passed: boolean;
|
||||
}
|
||||
|
||||
@@ -85,60 +99,171 @@ export interface EnforcementResult extends ProhibitionDisposition {
|
||||
mode?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* The default REAL check runner (used when no `runCheck` is injected). Deterministic per
|
||||
* environment and guarded so a missing tool yields a non-passing result, NEVER an uncaught throw
|
||||
* (the no-throw contract). A real run is fail-first by construction here — the descriptor's
|
||||
* `failFirst` marker is the authoritative regression-must-fail-first signal the producer confirms.
|
||||
* - node-test: runs `node --test <target>`; exit 0 = passed.
|
||||
* - lint-rule: runs `eslint --rule '<rule>: error' <target>`, exit 0 = passed.
|
||||
*/
|
||||
/** node --test argv. Forces the TAP reporter so the summary counts are parseable + version-stable. */
|
||||
export function buildNodeTestArgs(check: CheckDescriptor): string[] {
|
||||
return ['--test', '--test-reporter=tap', check.target];
|
||||
}
|
||||
|
||||
/** eslint argv (the args AFTER `npx`). Runs the project flat config so plugin rules (e.g. `local/*`)
|
||||
* load — `--rule` CANNOT load a plugin, so we lint the TARGET path as JSON and filter by rule id. */
|
||||
export function buildLintArgs(check: CheckDescriptor): string[] {
|
||||
return ['eslint', '--format', 'json', check.target];
|
||||
}
|
||||
|
||||
/**
|
||||
* Pure mapper from a lint-rule descriptor to the eslint argv (the args AFTER `npx`). The rule id
|
||||
* (`check.rule`, forced to `error`) and the lint TARGET path (`check.target`) are DISTINCT tokens —
|
||||
* reusing `target` as both (the #1259 pre-fix bug) makes eslint try to lint a file named after the
|
||||
* rule, which can never pass. Exported so the mapping is unit-testable without spawning eslint.
|
||||
* Pure parser for the `node --test` TAP summary. A genuine pass is NON-VACUOUS: exit 0 is NOT enough
|
||||
* (an empty / all-skipped / deleted-negative-test file exits 0 with `# tests 0`). Mutation-pinned by
|
||||
* unit tests so a threshold flip is caught.
|
||||
*/
|
||||
export function buildLintArgs(check: CheckDescriptor): string[] {
|
||||
return ['eslint', '--rule', `${check.rule}: error`, check.target];
|
||||
export function parseNodeTestSummary(out: string): { tests: number; pass: number; fail: number } {
|
||||
const num = (re: RegExp): number => {
|
||||
const m = typeof out === 'string' ? out.match(re) : null;
|
||||
return m ? Number(m[1]) : 0;
|
||||
};
|
||||
return {
|
||||
tests: num(/^# tests (\d+)/m),
|
||||
pass: num(/^# pass (\d+)/m),
|
||||
fail: num(/^# fail (\d+)/m),
|
||||
};
|
||||
}
|
||||
|
||||
/** The names from TAP `ok N - <name>` / `not ok N - <name>` lines (directives like `# SKIP` stripped). */
|
||||
export function tapTestNames(out: string): string[] {
|
||||
if (typeof out !== 'string') return [];
|
||||
const names: string[] = [];
|
||||
const re = /^(?:not )?ok \d+ - (.+)$/gm;
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(out)) !== null) {
|
||||
names.push(m[1].replace(/\s+#\s.*$/, '').trim());
|
||||
}
|
||||
return names;
|
||||
}
|
||||
|
||||
/**
|
||||
* A non-vacuous node-test pass: at least one test, at least one pass, zero failures — AND at least
|
||||
* one reported test whose name is NOT merely the target file. `node --test <file>` counts a file
|
||||
* with ZERO `test()` calls as one passing "test" named after the file, so the counts alone cannot
|
||||
* tell an empty/deleted negative test from a real one (the #1259 BL-01 false-green). Requiring a
|
||||
* named test distinct from the file closes that hole.
|
||||
*/
|
||||
export function isNonVacuousNodeTestPass(out: string, target: string): boolean {
|
||||
const s = parseNodeTestSummary(out);
|
||||
if (!(s.tests >= 1 && s.pass >= 1 && s.fail === 0)) return false;
|
||||
const base = typeof target === 'string' ? (target.split(/[\\/]/).pop() ?? target) : '';
|
||||
return tapTestNames(out).some((n) => n !== base && n !== target);
|
||||
}
|
||||
|
||||
/** Number of file results in an eslint `--format json` report (0 if unparseable / not an array). */
|
||||
export function eslintFileResultCount(jsonText: string): number {
|
||||
try {
|
||||
const parsed: unknown = JSON.parse(jsonText);
|
||||
return Array.isArray(parsed) ? parsed.length : 0;
|
||||
} catch {
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
|
||||
/** True if the eslint `--format json` report has ANY message for `rule`. Unparseable -> true
|
||||
* (fail-closed: treat an unreadable report as a violation rather than a silent pass). */
|
||||
export function eslintJsonHasRule(jsonText: string, rule: string): boolean {
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = JSON.parse(jsonText);
|
||||
} catch {
|
||||
return true;
|
||||
}
|
||||
if (!Array.isArray(parsed)) return true;
|
||||
for (const file of parsed) {
|
||||
const messages = file && typeof file === 'object' && Array.isArray((file as { messages?: unknown }).messages)
|
||||
? (file as { messages: Array<{ ruleId?: unknown }> }).messages
|
||||
: [];
|
||||
for (const msg of messages) {
|
||||
if (msg && typeof msg === 'object' && msg.ruleId === rule) return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* The default REAL check runner (used when no `runCheck` is injected). Reports only an OBSERVED,
|
||||
* genuinely-non-vacuous pass; guarded so a missing tool / non-zero exit yields a non-passing result,
|
||||
* NEVER an uncaught throw (the no-throw contract). It does NOT determine fail-first (caller-attested).
|
||||
* - node-test: runs `node --test` (TAP) and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail
|
||||
* AND a reported test named distinctly from the file). A bare exit 0 for an empty/zero-test file
|
||||
* — which `node --test` counts as one passing "test" named after the file — is NOT a pass (the
|
||||
* #1259 BL-01 false-green fix).
|
||||
* - lint-rule: runs the project `eslint --format json <target>` (flat config loads `local/*`
|
||||
* plugins) and requires the target to actually lint (>=1 file result) AND ZERO messages for the
|
||||
* specific rule id. `--rule` cannot load a plugin rule, so we filter the structured report by
|
||||
* `ruleId` instead (the #1259 SF-01 fix).
|
||||
*/
|
||||
/**
|
||||
* Env for spawned checks: strip `NODE_TEST_CONTEXT` and `NODE_OPTIONS` so an AMBIENT test-runner
|
||||
* context (e.g. running verify under `node --test`, which sets `NODE_TEST_CONTEXT=child-v8`) cannot
|
||||
* turn the child `node --test` into a silent v8-reporter worker that emits no parseable TAP — which
|
||||
* would otherwise corrupt the verdict. Deterministic, environment-independent execution.
|
||||
*/
|
||||
function childEnv(): NodeJS.ProcessEnv {
|
||||
const env = { ...process.env };
|
||||
delete env.NODE_TEST_CONTEXT;
|
||||
delete env.NODE_OPTIONS;
|
||||
return env;
|
||||
}
|
||||
|
||||
function defaultRunCheck(check: CheckDescriptor, cwd: string): CheckRunResult {
|
||||
const failFirst = check.failFirst === true;
|
||||
try {
|
||||
if (check.kind === 'node-test') {
|
||||
execFileSync('node', ['--test', check.target], {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: 'ignore',
|
||||
windowsHide: true,
|
||||
});
|
||||
return { failFirst, passed: true };
|
||||
let out = '';
|
||||
try {
|
||||
out = execFileSync('node', buildNodeTestArgs(check), {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
windowsHide: true,
|
||||
env: childEnv(),
|
||||
});
|
||||
} catch (e) {
|
||||
// A failing test run exits non-zero (TAP still on stdout). Parse it: a real failure has
|
||||
// `# fail >= 1` -> non-vacuous check returns false. Missing `node` -> no stdout -> false.
|
||||
const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : '';
|
||||
out = typeof stdout === 'string' ? stdout : '';
|
||||
}
|
||||
return { passed: isNonVacuousNodeTestPass(out, check.target) };
|
||||
}
|
||||
// lint-rule: force the rule to error and lint the TARGET path (distinct from the rule id). A
|
||||
// clean exit (0) means no violation surfaced -> the must-NOT holds across the target.
|
||||
execFileSync('npx', buildLintArgs(check), {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: 'ignore',
|
||||
windowsHide: true,
|
||||
});
|
||||
return { failFirst, passed: true };
|
||||
if (check.kind === 'lint-rule') {
|
||||
let json = '';
|
||||
try {
|
||||
json = execFileSync('npx', buildLintArgs(check), {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
windowsHide: true,
|
||||
env: childEnv(),
|
||||
});
|
||||
} catch (e) {
|
||||
// eslint exits non-zero when ANY error is present; the JSON report is still on stdout.
|
||||
const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : '';
|
||||
json = typeof stdout === 'string' ? stdout : '';
|
||||
}
|
||||
const lintedSomething = eslintFileResultCount(json) >= 1;
|
||||
return { passed: lintedSomething && !eslintJsonHasRule(json, check.rule as string) };
|
||||
}
|
||||
// Unknown kind — defensive; the LOCATE guard already rejects it.
|
||||
return { passed: false };
|
||||
} catch {
|
||||
// Non-zero exit (violation surfaced) OR missing tool -> not passing. Never throw.
|
||||
return { failFirst, passed: false };
|
||||
return { passed: false };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* LOCATE -> CONFIRM fail-first -> RUN -> build enforcementEvidence -> dispositionForProhibition.
|
||||
*
|
||||
* (1) LOCATE: if no check descriptor is locatable -> fail-closed (`dispositionForProhibition` with
|
||||
* empty evidence) plus `{ located: false, kind: null, evidence: [] }`.
|
||||
* (2) CONFIRM + RUN: confirm the descriptor is fail-first and run it via `runCheck`. A check that
|
||||
* is not fail-first, that the runner reports not fail-first, or that FAILS -> fail-closed
|
||||
* disposition with `located: true` (a real located miss, non-green, flagged) in BOTH modes.
|
||||
* (1) LOCATE: if no well-formed check descriptor is locatable -> fail-closed
|
||||
* (`dispositionForProhibition` with empty evidence) plus `{ located: false, kind: null, evidence: [] }`.
|
||||
* (2) ATTEST + RUN: require the caller to attest `failFirst: true` and run it via `runCheck`. A check
|
||||
* the caller does not attest as fail-first, or that does not genuinely (non-vacuously) PASS ->
|
||||
* fail-closed disposition with `located: true` (a real located miss, non-green, flagged) in BOTH
|
||||
* modes. Fail-first is caller-attested, not independently proven (see module docstring).
|
||||
* (3) PASS: build a typed `enforcementEvidence` array and call `dispositionForProhibition` — the
|
||||
* non-empty array flips a test-tier item to green (the previously-unreachable branch).
|
||||
*
|
||||
@@ -151,46 +276,48 @@ export function runProhibitionEnforcement(
|
||||
): EnforcementResult {
|
||||
const mode = options.mode;
|
||||
|
||||
// (1) LOCATE — no locatable wired check -> fail-closed, located: false. A lint-rule descriptor
|
||||
// MUST also carry a string `rule` id (its target is the lint PATH, not the rule) — an
|
||||
// under-specified lint-rule is not a valid wired check, so it is not locatable.
|
||||
if (
|
||||
!check ||
|
||||
typeof check !== 'object' ||
|
||||
typeof check.kind !== 'string' ||
|
||||
typeof check.target !== 'string' ||
|
||||
(check.kind === 'lint-rule' && typeof check.rule !== 'string')
|
||||
) {
|
||||
// (1) LOCATE — no locatable, well-formed wired check -> fail-closed, located: false. The kind must
|
||||
// be one of the two known kinds; the target must be a non-empty string; a lint-rule descriptor MUST
|
||||
// also carry a non-empty `rule` id (its target is the lint PATH, not the rule). An under-specified
|
||||
// descriptor is not a valid wired check, so it is not locatable (does not rely on the runner failing).
|
||||
const c = check && typeof check === 'object' ? check : null;
|
||||
const validKind = !!c && (c.kind === 'node-test' || c.kind === 'lint-rule');
|
||||
const validTarget = !!c && typeof c.target === 'string' && c.target.trim().length > 0;
|
||||
const validRule = !!c && (c.kind !== 'lint-rule' || (typeof c.rule === 'string' && c.rule.trim().length > 0));
|
||||
if (!c || !validKind || !validTarget || !validRule) {
|
||||
const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] });
|
||||
return { ...disposition, located: false, kind: null, evidence: [], ...(mode ? { mode } : {}) };
|
||||
}
|
||||
|
||||
const runCheck = options.runCheck ?? ((c: CheckDescriptor) => defaultRunCheck(c, options.cwd ?? process.cwd()));
|
||||
const runCheck = options.runCheck ?? ((toRun: CheckDescriptor) => defaultRunCheck(toRun, options.cwd ?? process.cwd()));
|
||||
|
||||
// (2) CONFIRM fail-first + RUN. Descriptor must declare fail-first AND the runner must agree.
|
||||
const descriptorFailFirst = check.failFirst === true;
|
||||
const run = runCheck(check);
|
||||
const failFirstConfirmed = descriptorFailFirst && run.failFirst === true;
|
||||
const passed = failFirstConfirmed && run.passed === true;
|
||||
// (2) ATTEST fail-first (CALLER-DECLARED) + RUN. The caller must attest `failFirst: true` AND the
|
||||
// runner must observe a genuine NON-VACUOUS pass. The producer does NOT independently prove
|
||||
// fail-first (that needs a violation fixture — tracked follow-up, ADR-550 D5d). A non-attested or
|
||||
// non-passing check hard-gates (never green) in BOTH modes.
|
||||
const attestedFailFirst = c.failFirst === true;
|
||||
const run = runCheck(c);
|
||||
const passed = attestedFailFirst && run.passed === true;
|
||||
|
||||
if (!passed) {
|
||||
// FAIL / not-fail-first -> fail-closed, located: true (an actual located miss/fail). Hard-gate
|
||||
// applies in BOTH modes; the disposition stays non-green / flagged.
|
||||
// NOT attested fail-first OR did not genuinely pass -> fail-closed, located: true (an actual
|
||||
// located miss/fail). Hard-gate applies in BOTH modes; the disposition stays non-green / flagged.
|
||||
const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] });
|
||||
return {
|
||||
...disposition,
|
||||
located: true,
|
||||
kind: check.kind,
|
||||
kind: c.kind,
|
||||
evidence: [],
|
||||
...(mode ? { mode } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
// (3) PASS -> build typed enforcementEvidence and let the policy flip a test-tier item green.
|
||||
// `failFirst` here is the caller's attestation (recorded for provenance), not a machine proof.
|
||||
const evidence: EnforcementEvidence[] = [{
|
||||
kind: check.kind,
|
||||
target: check.target,
|
||||
...(typeof check.rule === 'string' ? { rule: check.rule } : {}),
|
||||
kind: c.kind,
|
||||
target: c.target,
|
||||
...(typeof c.rule === 'string' ? { rule: c.rule } : {}),
|
||||
failFirst: true,
|
||||
passed: true,
|
||||
}];
|
||||
@@ -198,7 +325,7 @@ export function runProhibitionEnforcement(
|
||||
return {
|
||||
...disposition,
|
||||
located: true,
|
||||
kind: check.kind,
|
||||
kind: c.kind,
|
||||
evidence,
|
||||
...(mode ? { mode } : {}),
|
||||
};
|
||||
|
||||
@@ -39,7 +39,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
test('locate-miss (no check descriptor) -> fail-closed, located:false, no evidence', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(TEST_TIER, null, {
|
||||
runCheck: () => ({ failFirst: true, passed: true }),
|
||||
runCheck: () => ({ passed: true }),
|
||||
});
|
||||
assert.equal(result.located, false, 'no locatable check');
|
||||
assert.notEqual(result.status, 'green', 'locate-miss must never be green');
|
||||
@@ -51,7 +51,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
test('malformed check descriptor (missing target) -> treated as locate-miss', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(TEST_TIER, { kind: 'node-test' }, {
|
||||
runCheck: () => ({ failFirst: true, passed: true }),
|
||||
runCheck: () => ({ passed: true }),
|
||||
});
|
||||
assert.equal(result.located, false, 'a descriptor without a target is not locatable');
|
||||
assert.notEqual(result.status, 'green');
|
||||
@@ -63,7 +63,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green');
|
||||
assert.equal(result.flagged, false);
|
||||
@@ -83,7 +83,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green');
|
||||
assert.equal(result.flagged, false);
|
||||
@@ -93,18 +93,19 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
assert.equal(result.evidence[0].target, 'tests/', 'evidence records the linted target path, not the rule id');
|
||||
});
|
||||
|
||||
test('buildLintArgs maps the rule id and lint target to DISTINCT eslint args (#1259 runner fix)', () => {
|
||||
test('buildLintArgs runs the project eslint as JSON over the target (plugins load via flat config; #1259 SF-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(typeof enforce.buildLintArgs, 'function',
|
||||
'must export buildLintArgs — the pure argv mapper for the lint-rule real runner');
|
||||
'must export buildLintArgs — the eslint argv builder for the lint-rule real runner');
|
||||
const argv = enforce.buildLintArgs({ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/' });
|
||||
assert.ok(Array.isArray(argv), 'argv is an array');
|
||||
const ruleIdx = argv.indexOf('--rule');
|
||||
assert.ok(ruleIdx !== -1, 'forces a specific rule via --rule');
|
||||
assert.equal(argv[ruleIdx + 1], 'local/no-source-grep: error', 'the rule id is forced to error');
|
||||
assert.equal(argv[0], 'eslint');
|
||||
const fmtIdx = argv.indexOf('--format');
|
||||
assert.ok(fmtIdx !== -1 && argv[fmtIdx + 1] === 'json',
|
||||
'emits --format json so the report can be filtered by ruleId');
|
||||
assert.ok(!argv.includes('--rule'),
|
||||
'must NOT use --rule — it cannot load a plugin rule like local/no-source-grep (the SF-01 bug)');
|
||||
assert.equal(argv[argv.length - 1], 'tests/', 'the LAST arg is the lint target path');
|
||||
assert.notEqual(argv[ruleIdx + 1], argv[argv.length - 1],
|
||||
'the rule id must NOT be reused as the lint target (the bug this guards)');
|
||||
});
|
||||
|
||||
test('lint-rule descriptor missing its rule id -> locate-miss, never green', () => {
|
||||
@@ -112,7 +113,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', target: 'tests/', failFirst: true }, // no `rule`
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a lint-rule with no rule id is not a valid wired check');
|
||||
assert.equal(result.flagged, true);
|
||||
@@ -124,7 +125,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: false }) },
|
||||
{ runCheck: () => ({ passed: false }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green');
|
||||
assert.equal(result.flagged, true);
|
||||
@@ -132,24 +133,17 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
assert.equal(result.evidence.length, 0, 'a failing check builds no evidence');
|
||||
});
|
||||
|
||||
test('fail-first NOT satisfied (descriptor or runner) -> hard-gate, never green', () => {
|
||||
test('caller does NOT attest fail-first (descriptor failFirst:false) -> hard-gate, never green', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// descriptor declares failFirst:false
|
||||
const a = enforce.runProhibitionEnforcement(
|
||||
// fail-first is caller-attested (#1259 BL-02): a check the caller does not attest as fail-first
|
||||
// is not a valid regression proof and must never green, even if the run passes.
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: false },
|
||||
{ runCheck: () => ({ failFirst: false, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(a.status, 'green', 'not-fail-first is not a valid regression proof');
|
||||
assert.equal(a.flagged, true);
|
||||
// descriptor says failFirst:true but runner reports failFirst:false -> still hard-gate
|
||||
const b = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: false, passed: true }) },
|
||||
);
|
||||
assert.notEqual(b.status, 'green', 'runner-reported not-fail-first must also hard-gate');
|
||||
assert.equal(b.flagged, true);
|
||||
assert.notEqual(result.status, 'green', 'a non-attested check is not a valid regression proof');
|
||||
assert.equal(result.flagged, true);
|
||||
});
|
||||
|
||||
test('hard-gates in BOTH modes on a failing check (ADR-550 D4)', () => {
|
||||
@@ -158,7 +152,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: false }), mode },
|
||||
{ runCheck: () => ({ passed: false }), mode },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', `non-green in ${mode}`);
|
||||
assert.equal(result.flagged, true, `flagged in ${mode}`);
|
||||
@@ -171,7 +165,7 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }), mode: 'autonomous' },
|
||||
{ runCheck: () => ({ passed: true }), mode: 'autonomous' },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a passing wired check is green in autonomous mode too');
|
||||
assert.equal(result.mode, 'autonomous');
|
||||
@@ -208,3 +202,105 @@ describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR
|
||||
assert.ok(Array.isArray(parsed.evidence), 'evidence is an array');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Real-runner helpers (mutation-pinned; #1259 BL-01 / SF-01) ─────────────────
|
||||
// These pin the deterministic parsing/threshold logic of the REAL runner so a Stryker mutant that
|
||||
// weakens "non-vacuous pass" or the ruleId filter is caught — the contract the injected-runner tests
|
||||
// above deliberately bypass.
|
||||
describe('prohibition-enforcement real-runner helpers (#1259)', () => {
|
||||
test('parseNodeTestSummary extracts the TAP tests/pass/fail counts', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.deepEqual(enforce.parseNodeTestSummary('# tests 3\n# pass 2\n# fail 1\n'), { tests: 3, pass: 2, fail: 1 });
|
||||
assert.deepEqual(enforce.parseNodeTestSummary('no summary here'), { tests: 0, pass: 0, fail: 0 });
|
||||
});
|
||||
|
||||
test('tapTestNames extracts ok/not-ok names, stripping directives', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.deepEqual(enforce.tapTestNames('ok 1 - guards the must-NOT\nnot ok 2 - other # SKIP\n'),
|
||||
['guards the must-NOT', 'other']);
|
||||
});
|
||||
|
||||
test('isNonVacuousNodeTestPass: an empty file (node names the test after the file) is NOT a pass (BL-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// node --test of a zero-test file: `ok 1 - empty.test.cjs`, `# tests 1 # pass 1` — counts alone
|
||||
// cannot distinguish it from a real test, so the file-named result must NOT count as a pass.
|
||||
const empty = 'ok 1 - empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(empty, 'empty.test.cjs'), false,
|
||||
'a file-named-only result is vacuous — the BL-01 false-green guard');
|
||||
const real = 'ok 1 - guards the must-NOT\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(real, 'neg.test.cjs'), true,
|
||||
'a real named test distinct from the file is a genuine pass');
|
||||
const failing = 'not ok 1 - guards\n# tests 1\n# pass 0\n# fail 1\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false,
|
||||
'any failure means not a pass');
|
||||
});
|
||||
|
||||
test('eslintJsonHasRule detects a ruleId; unparseable report -> true (fail-closed)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'local/no-source-grep' }] }]), 'local/no-source-grep'), true);
|
||||
assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'other' }] }]), 'local/no-source-grep'), false);
|
||||
assert.equal(enforce.eslintJsonHasRule('not json', 'local/no-source-grep'), true,
|
||||
'an unreadable report must be treated as a violation, never a silent pass');
|
||||
});
|
||||
|
||||
test('eslintFileResultCount: 0 when nothing linted (vacuity guard)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(enforce.eslintFileResultCount(JSON.stringify([{}, {}])), 2);
|
||||
assert.equal(enforce.eslintFileResultCount('[]'), 0);
|
||||
assert.equal(enforce.eslintFileResultCount('garbage'), 0);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Real runner end-to-end (NO injected runCheck; #1259 SF-02 / BL-01 / SF-01) ──
|
||||
// Spawns real subprocesses so the SHIPPING default runner is exercised — the gap that let BL-01 and
|
||||
// SF-01 slip past the injected-double tests. Typed-field assertions only.
|
||||
describe('prohibition-enforcement REAL runner end-to-end (#1259)', () => {
|
||||
const fs = require('node:fs');
|
||||
|
||||
test('a genuine non-vacuous passing node-test greens via the real runner', (t) => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const dir = createTempDir('prohib-real-pass-');
|
||||
t.after(() => cleanup(dir));
|
||||
const tf = path.join(dir, 'neg.test.cjs');
|
||||
fs.writeFileSync(tf,
|
||||
"const { test } = require('node:test');\nconst assert = require('node:assert');\ntest('guards the must-NOT', () => { assert.ok(true); });\n");
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: tf, failFirst: true },
|
||||
{ cwd: dir },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a real, passing, non-vacuous negative test must green');
|
||||
assert.equal(result.located, true);
|
||||
assert.equal(result.evidence.length, 1);
|
||||
});
|
||||
|
||||
test('an EMPTY node-test file (exit 0, zero tests) does NOT green via the real runner (BL-01)', (t) => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const dir = createTempDir('prohib-real-empty-');
|
||||
t.after(() => cleanup(dir));
|
||||
const tf = path.join(dir, 'empty.test.cjs');
|
||||
fs.writeFileSync(tf, '// intentionally empty — no test cases\n');
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: tf, failFirst: true },
|
||||
{ cwd: dir },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'an empty (zero-test) file must NEVER green — fail-closed');
|
||||
assert.equal(result.located, true, 'the check was located; it just did not genuinely pass');
|
||||
assert.equal(result.evidence.length, 0);
|
||||
});
|
||||
|
||||
test('a clean in-tree target greens the lint-rule kind via the real eslint runner (SF-01: plugin loads)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// Runs real `npx eslint --format json src/clock.cts` under the project flat config (so the
|
||||
// `local` plugin loads). src/clock.cts is a clean source with no no-source-grep violation.
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'src/clock.cts', failFirst: true },
|
||||
{ cwd: process.cwd() },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a clean target with no no-source-grep violation must green via real eslint');
|
||||
assert.equal(result.kind, 'lint-rule');
|
||||
assert.equal(result.evidence[0].rule, 'local/no-source-grep');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -88,7 +88,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 /
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.ok(result && typeof result === 'object', 'result must be a structured object');
|
||||
assert.equal(result.status, 'green', 'a passing wired node-test check must dispose green');
|
||||
@@ -105,7 +105,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 /
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a passing wired lint-rule check must dispose green');
|
||||
assert.equal(result.flagged, false, 'a green disposition must not be flagged');
|
||||
@@ -120,7 +120,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 /
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
null,
|
||||
{ runCheck: () => ({ failFirst: true, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a missing wired check must NEVER be green (fail-closed)');
|
||||
assert.equal(result.flagged, true, 'a missing wired check must be flagged unverified');
|
||||
@@ -136,7 +136,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 /
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ failFirst: true, passed: false }), mode },
|
||||
{ runCheck: () => ({ passed: false }), mode },
|
||||
);
|
||||
assert.notEqual(result.status, 'green',
|
||||
`a failing wired check must NEVER be green (mode=${mode})`);
|
||||
@@ -152,7 +152,7 @@ describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 /
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: false },
|
||||
{ runCheck: () => ({ failFirst: false, passed: true }) },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green',
|
||||
'a check that is not fail-first is not a valid regression-must-fail-first proof — never green');
|
||||
|
||||
Reference in New Issue
Block a user