diff --git a/.changeset/1259-test-tier-enforcement.md b/.changeset/1259-test-tier-enforcement.md new file mode 100644 index 000000000..b9ef51535 --- /dev/null +++ b/.changeset/1259-test-tier-enforcement.md @@ -0,0 +1,10 @@ +--- +type: Changed +pr: 1273 +--- + +**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, runs it for a genuine **non-vacuous** pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a genuinely-passing wired check disposes `green` and can reach `passed`, while a missing, non-attested, or non-passing check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test (requiring a real reported test — an empty file, which `node --test` counts as one passing "test", does NOT green) AND a lint/AST rule run as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot), anchored on the in-tree `local/no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259) + +**Honest scope — `failFirst` is caller-attested, not yet machine-proven.** This lands the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and the check to genuinely run and pass. It does NOT yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property) — cheap proof of that at verify time needs running the check against a known violation fixture, which is a **tracked follow-up (#1279)**. The red-first property currently rests on caller attestation, surfaced transparently in the evidence record. + +**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. For the `no-source-grep` anchor specifically, the genuine `regression-must-fail-first` proofs are its **two `invalid` cases** (the `.includes()` and `.match()` blocks) in `tests/eslint-rules.test.cjs` — not 96. The anchor argument is unaffected (those two cases ARE real fail-first proofs); only the count was off. diff --git a/.gitignore b/.gitignore index 38fc86117..c10433094 100644 --- a/.gitignore +++ b/.gitignore @@ -74,6 +74,7 @@ build/ /gsd-core/bin/lib/plan-drift-guard.cjs /gsd-core/bin/lib/edge-probe.cjs /gsd-core/bin/lib/probe-core.cjs +/gsd-core/bin/lib/prohibition-enforcement.cjs /gsd-core/bin/lib/config-types.cjs /gsd-core/bin/lib/cli-exit.cjs /gsd-core/bin/lib/code-review-flags.cjs diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 8f6e038c0..3d2ad7a5a 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3169,7 +3169,7 @@ Each surfaced prohibition is resolved to exactly one of three states: | `dismissed` | Not a genuine prohibition (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected | | `unresolved` | Deferred | Soft-gates the spec; surfaced as a planner assumption | -Each resolved prohibition carries a `verification` tier — `test` (a negative test can enforce it) or `judgment` (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag); test-tier prohibitions fail closed when unwired (never silently green). Under `--auto`, the probe **never auto-dismisses**. Canon-bound concerns (OWASP / GDPR / fairness) are referred to `/gsd:secure-phase` rather than minting SPEC prohibitions (ADR-550 D6). +Each resolved prohibition carries a `verification` tier — `test` (a negative test can enforce it) or `judgment` (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag); test-tier prohibitions are enforced via the deterministic `check prohibition-enforcement` gate — green when the wired negative test / lint rule passes, hard-gate (flagged, non-green) when missing or failing, in both interactive and autonomous modes (#1259, ADR-550 D5d). Under `--auto`, the probe **never auto-dismisses**. Canon-bound concerns (OWASP / GDPR / fairness) are referred to `/gsd:secure-phase` rather than minting SPEC prohibitions (ADR-550 D6). The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, so the section is not merely documentation. @@ -3180,5 +3180,6 @@ The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, s - REQ-PROHIB-04: `--auto` MUST never auto-dismiss. - REQ-PROHIB-05: `plan-phase` MUST lift resolved prohibitions into `must_haves.prohibitions` (never `truths`). - REQ-PROHIB-06: A well-formed but unwired `test`-tier prohibition MUST fail closed at verify time — never a silent pass. +- REQ-PROHIB-07: A `test`-tier prohibition with a caller-attested, genuinely-passing (non-vacuous) wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, non-attested, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half; machine-proven fail-first is tracked in #1279, deterministic descriptor auto-locate in #1278). **Reference:** [Prohibition Probe](../gsd-core/references/prohibition-probe.md) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 7510ac2d3..d045279e7 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -345,6 +345,7 @@ "profile-output.cjs", "profile-pipeline-command-router.cjs", "profile-pipeline.cjs", + "prohibition-enforcement.cjs", "project-root.cjs", "prompt-budget.cjs", "research-provider.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index eb7dfba87..e60245bf4 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -459,6 +459,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `research-provider.cjs` | Research provider waterfall, confidence tiers, and planResearch (cache-hits + fetch plan) | | `research-store.cjs` | Content-addressed research cache: sha256 keys, per-source TTL staleness, two-tier (user ~/.gsd / project .planning) store | | `probe-core.cjs` | Generic spec-phase probe resolution model (compiled from `src/probe-core.cts`, gitignored; ADR-550 Decision 7) — the status×verification re-cut (`status: resolved/dismissed/unresolved` × per-probe `verification`), `validateResolution`/`validateRequirement`, `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject, the `byVerification` rollup, and the `runProbeCli` I/O scaffold; the shared seam consumed by `edge-probe` (and the prohibition probe #644); exports `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` (#550) | +| `prohibition-enforcement.cjs` | Deterministic test-tier prohibition PRODUCER/gate (compiled from `src/prohibition-enforcement.cts`, gitignored; #1259, ADR-550 D5d "heavy half") — locates the wired mechanical check (`node-test` or `lint-rule`), confirms it is fail-first, runs it via an injectable runner, builds typed `enforcementEvidence`, and emits the `dispositionForProhibition` verdict; a passing wired check disposes green, a missing/failing/non-fail-first check hard-gates (flagged, non-green) in both interactive and autonomous modes; exports `runProhibitionEnforcement`, `routeProhibitionEnforcement`; CLI surface `gsd_run check prohibition-enforcement ` | | `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence | | `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` | | `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) | diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md index 0567658ef..b76ebf637 100644 --- a/docs/adr/550-spec-phase-probe-contract.md +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -85,11 +85,14 @@ The capability system (ADR-857) classifies **predicate-generation as core verifi - **The probe *adapters* are the core-default *generator*** — `edge-probe`'s `classifyShape`/`proposeEdges` and the prohibition probe's adversarial LLM-propose, the surfaces that *propose* predicates — default-on and non-removable, but **independently versionable**. The classifier's measured recall gap (the prose→shape under-fire on terse prose) is the reason they stay their own modules under this ADR rather than being folded into the slow core rail. **`probe-core` is not the generator:** per Decision 7b it ingests already-proposed items, and its deterministic validators are the **contract**'s CI-testable surface (Decision 5) — so `probe-core` sits on the contract side, the adapters on the generator side. ADR-857 phase 6 wires these modules onto the core predicate rail; it does **not** relabel them as a `capabilities/edge-probe/` plug-in. - Decision 5's rule — *the CI-testable surface is the contract, not the classifier* — extends to ADR-857's core rail: the deterministic conformance test is the contract shape the verifier consumes, never the LLM's judgment. -## Addendum (2026-06-12): test-tier disposition — fail-closed now, heavy enforcement deferred +## Addendum (2026-06-12; updated 2026-06-15): test-tier disposition — fail-closed safety half (#644) + enforcement half LANDED (#1259) Decision 4 describes the `test`-tier as a "**Hard gate in both interactive and autonomous modes.**" The #644 implementation revises that to a **fail-closed-now / deferred-enforcement** resolution (the "B-with-guard" maintainer decision of 2026-06-12), so the architecture-of-record matches the shipped code: - A well-formed but **unwired** `test`-tier prohibition resolves via `dispositionForProhibition()` to `{ status: 'unverified', flagged: true }` — **provably never green** without explicit evidence (REQ-PROHIB-06). This is the load-bearing safety half and it holds today. -- The **heavy negative-test enforcement mechanism** — a contrived `test`-tier consumer that mechanically runs the negative test — is **deferred to a follow-up**, because the #644 corpus is entirely `judgment`-tier and wiring a synthetic test-tier consumer now would be gold-plating. The fail-closed disposition guards the gap in the meantime: the gate cannot silently pass, it can only report `unverified`/flagged until enforcement lands. +- The **negative-test enforcement mechanism** — locating the wired mechanical check, running it for a genuine **non-vacuous** pass, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test (requiring a real reported test, not the empty file `node --test` would count as one passing "test") OR a lint/AST rule run through the project flat config as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot) — and is anchored on the in-tree `local/no-source-grep` rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, non-attested, or genuinely-non-passing check hard-gates (flagged, non-green) in both interactive and autonomous modes. +- **Honest scope — `failFirst` is caller-ATTESTED, not machine-proven (tracked follow-up).** What #1259 lands is the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and requires the check to genuinely run and pass. It does **not** yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property), because cheap proof of that at verify time needs running the check against a known **violation fixture** — deferred as a follow-up (#1279; the descriptor auto-locate half is #1278). Until then the red-first property rests on caller attestation, surfaced transparently in the evidence record. This closes the permanent-`gaps_found` dead-end with a genuinely-executed gate without overclaiming machine-proven fail-first. -Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") is preserved exactly; what is deferred is only the mechanical-pass half of the gate, not the never-green safety property. This addendum supersedes the unqualified "hard gate" wording of Decision 4 for `test`-tier items until the enforcement follow-up lands. The decision also lives in `src/probe-core.cts` comments, `verify-phase.md`, and the #644 changeset. +Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the genuine-execution half — a test-tier prohibition with a passing, non-vacuous wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The one remaining gap to D4's literal intent — *machine-proven* fail-first — is documented above as a tracked follow-up. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets. + +This enforcement seam is the concrete instance of **ADR-857 open-question §147** — the deferred "deterministic CI conformance test for the verifier↔predicate contract." Per D6 it lands on the **core verify rail** (non-toggleable substrate), never in `capabilities/`: the verifier consuming a contract-shaped, deterministic predicate is core, not an opt-in capability. diff --git a/eslint.config.mjs b/eslint.config.mjs index 16f729a43..1266924a7 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -41,6 +41,7 @@ export default tseslint.config( 'gsd-core/bin/lib/cli-exit.cjs', 'gsd-core/bin/lib/edge-probe.cjs', 'gsd-core/bin/lib/probe-core.cjs', + 'gsd-core/bin/lib/prohibition-enforcement.cjs', 'gsd-core/bin/lib/code-review-flags.cjs', 'gsd-core/bin/lib/context-utilization.cjs', 'gsd-core/bin/lib/artifacts.cjs', diff --git a/gsd-core/references/prohibition-probe.md b/gsd-core/references/prohibition-probe.md index 1915d2d23..2cebfa403 100644 --- a/gsd-core/references/prohibition-probe.md +++ b/gsd-core/references/prohibition-probe.md @@ -110,6 +110,20 @@ lifecycle is identical to the edge-probe, the verification tiers differ): human/LLM judgment that the framing is not manipulative). It records intent and routes to a judgment-based review rather than a green/red test. + At verify time these tiers are routed differently (ADR-550 D4): + - A **test**-tier prohibition is enforced + hard-gates via the deterministic + `check prohibition-enforcement` sub-command (#1259, ADR-550 D5d): it locates the wired + mechanical check (a `node --test` negative test OR a lint/AST rule run as + `eslint --format json` and filtered by `ruleId`), requires the caller-attested `failFirst` + marker, runs it for a genuine **non-vacuous** pass, and emits the + `dispositionForProhibition()` verdict. A passing wired check disposes **green** (satisfiable + → can reach `passed`); a missing, non-attested, or genuinely-non-passing check **hard-gates** + (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes — never a + silent pass. (`failFirst` is caller-attested, not yet machine-proven against a violation + fixture — a tracked follow-up; see ADR-550 D5d.) + - A **judgment**-tier prohibition routes to a never-silent / never-hard-halt soft gate + (autonomous emits an `unverified-prohibition — human review recommended` flag). + Splitting these axes keeps the lifecycle enum free of a verification fact and lets the prohibition adapter declare `test | judgment` without forking the shared lifecycle enum that the edge-probe's `explicit | backstop` also uses. diff --git a/gsd-core/workflows/verify-phase.md b/gsd-core/workflows/verify-phase.md index 5f6b37629..fa556c2a8 100644 --- a/gsd-core/workflows/verify-phase.md +++ b/gsd-core/workflows/verify-phase.md @@ -70,7 +70,17 @@ Aggregate all must_haves across plans for phase-level verification. **Prohibitions (`must_haves.prohibitions`, ADR-550 D3 — the must-NOT sibling block):** When a plan carries `must_haves.prohibitions`, extract each `{ statement, status, verification }` item and route it by `verification` tier in verdict assembly (ADR-550 D4, "B-with-guard", 2026-06-12 maintainer decision). These are NEGATIVE checks (the must-NOT must NOT have happened), distinct from positive `truths`: - **judgment-tier → mode-dependent soft-gate.** Interactive verify defers each item to the end-of-phase human checkpoint (`human_verify_mode: end-of-phase`). Autonomous verify records a NON-AUTHORITATIVE LLM-judge verdict + a prominent `unverified-prohibition — human review recommended` flag (autonomous completion reads "complete with N flagged prohibitions"). NEVER a silent pass; NEVER a hard halt of an AFK run. -- **test-tier → FAIL CLOSED (accept-and-flag).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later), but a well-formed test-tier item reaching verify with NO wired enforcement disposes as UNVERIFIED, flagged like an unresolved judgment item, NEVER green. The deterministic fail-closed default is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). The real fail-first negative-test enforcement MECHANISM defers to a follow-up PR (#644's corpus is entirely judgment-tier; a contrived test-tier fixture here would be the gold-plating failure mode). +- **test-tier → ENFORCED via `check prohibition-enforcement` (green on pass, hard-gate on miss/fail).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later). For each test-tier item, the verifier invokes the deterministic producer: + + ```bash + gsd_run check prohibition-enforcement + ``` + + where `` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, requires the caller-attested `failFirst` marker, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). (`failFirst` is caller-attested, not yet machine-proven against a violation fixture — a tracked follow-up.) Route the result by its typed fields: + - **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**. + - **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). + + > **Descriptor authoring — current scope (#1259).** The `check` descriptor is **supplied by the phase author / verifier**; there is no projection field yet that deterministically derives `{ kind, target, rule, failFirst }` from a prohibition in `must_haves.prohibitions` (which carries only `{ statement, status, verification }`). So #1259 lands the **deterministic run+verdict half** (locate→run→evidence→disposition, all CI-testable) while the **locate→descriptor half is author-provided** for now. Deterministic auto-locate — so a wired passing test closes the gap with zero manual descriptor authoring — is a **tracked follow-up: #1278**. Until then, the green path requires the author to wire the descriptor explicitly. **Option B: Use Success Criteria from ROADMAP.md** @@ -474,7 +484,7 @@ Classify status using this decision tree IN ORDER (most restrictive first): → **gaps_found** 2. IF any `must_haves.prohibitions` item disposes as flagged-unverified (ADR-550 D4): - - **test-tier, fail-closed** (no wired enforcement — `dispositionForProhibition()` returns `status: 'unverified'`, `flagged: true`): → **gaps_found** (never green; the unwired test-tier item is an unverified gap). + - **test-tier, fail-closed when the wired check is MISSING OR FAILS** (now run via `check prohibition-enforcement` — `located: false`, or `dispositionForProhibition()` returns `status: 'unverified'`, `flagged: true`): → **gaps_found** in both interactive and autonomous modes (never green; a missing/failing mechanical check is an unverified gap). A test-tier item whose wired check PASSES disposes `status: 'green'`, `flagged: false` and is NOT a gap — it can reach **passed**. - **judgment-tier, autonomous run** (non-authoritative LLM-judge verdict): emit the `unverified-prohibition — human review recommended` flag and classify → **human_needed** (autonomous completion reads "complete with N flagged prohibitions"; never a silent pass, never a hard halt). - **judgment-tier, interactive run**: route to the end-of-phase human checkpoint → **human_needed**. diff --git a/src/check-command-router.cts b/src/check-command-router.cts index fbcc07806..b93c59b8a 100644 --- a/src/check-command-router.cts +++ b/src/check-command-router.cts @@ -24,6 +24,7 @@ const { getRoadmapPhaseWithFallback } = roadmapModule; // eslint-disable-next-line @typescript-eslint/no-require-imports import gapCheckerModule = require('./gap-checker.cjs'); const { runGapAnalysis } = gapCheckerModule; +import { routeProhibitionEnforcement } from './prohibition-enforcement.cjs'; // ─── Helpers ────────────────────────────────────────────────────────────────── @@ -887,7 +888,15 @@ function routeCheckCommand({ args, cwd, raw }: RouteCheckCommandOptions): void { cmdVerifyCodebaseDrift(cwd, raw); return; } - error('Unknown check subcommand. Available: auto-mode, decision-coverage-plan, decision-coverage-verify, gap-analysis-plan-post, tdd-review-checkpoint, ui-plan-gate, ui-safety-gate, verify-schema-drift, verify-codebase-drift', ERROR_REASON.SDK_UNKNOWN_COMMAND); + if (subcommand === 'prohibition-enforcement') { + // The deterministic test-tier prohibition PRODUCER/gate (#1259, ADR-550 D5d). Locates the + // wired mechanical check (node-test or lint-rule), confirms fail-first, runs it, builds + // enforcementEvidence, and emits the dispositionForProhibition verdict. Invocable as + // `gsd_run check prohibition-enforcement `. + routeProhibitionEnforcement(args, raw); + return; + } + error('Unknown check subcommand. Available: auto-mode, decision-coverage-plan, decision-coverage-verify, gap-analysis-plan-post, prohibition-enforcement, tdd-review-checkpoint, ui-plan-gate, ui-safety-gate, verify-schema-drift, verify-codebase-drift', ERROR_REASON.SDK_UNKNOWN_COMMAND); } export = { diff --git a/src/probe-core.cts b/src/probe-core.cts index 6e89f98df..85a39f395 100644 --- a/src/probe-core.cts +++ b/src/probe-core.cts @@ -376,11 +376,10 @@ export interface ProhibitionDispositionContext { * This is the cheap safety guarantee: a well-formed prohibition that reaches verify-phase with NO * wired enforcement evidence can NEVER be a silent pass. It is `{ status: 'unverified', flagged: * true }` — never `green` — exactly like an unresolved judgment item. The HEAVY half (a real - * fail-first negative-test enforcement mechanism that, given evidence, would flip a test-tier item - * to green) is OUT of #644 scope and defers to a follow-up PR: #644's corpus is entirely - * judgment-tier, so wiring a contrived test-tier consumer here would be the delete-bad-tests / - * gold-plating failure mode. Until that follow-up lands, ANY prohibition without enforcement - * evidence — test- or judgment-tier — disposes as flagged-unverified. + * negative-test enforcement mechanism that, given evidence, flips a test-tier item to green) was OUT + * of #644 scope and LANDED in #1259 as the `prohibition-enforcement` producer (it builds the + * `enforcementEvidence` this helper reads). This helper's policy is unchanged: ANY prohibition + * without enforcement evidence — test- or judgment-tier — disposes as flagged-unverified. * * The function is pure: same input always yields the same disposition (no LLM judgment, ADR-550 * D5). The LLM-judge soft-gate for judgment-tier items is a verify-phase PROSE concern (the @@ -398,9 +397,9 @@ export function dispositionForProhibition( const hasEnforcement = evidence.length > 0; // FAIL CLOSED: no wired enforcement evidence -> flagged unverified, never green. This holds for - // every tier today (the real enforcement mechanism that could flip a test-tier item to green is - // deferred to a follow-up PR). The guard the safety assertion proves: an unwired item can never - // be silently skipped. + // every tier (the producer that builds enforcement evidence for a test-tier item — the + // `prohibition-enforcement` module — landed in #1259). The guard the safety assertion proves: an + // unwired item can never be silently skipped. if (!hasEnforcement) { return { status: 'unverified', @@ -408,15 +407,15 @@ export function dispositionForProhibition( tier, reason: tier === 'test' - ? 'test-tier prohibition has no wired enforcement evidence — flagged unverified (fail-closed; real negative-test enforcement deferred to a follow-up PR, ADR-550 D5d)' + ? 'test-tier prohibition has no passing wired enforcement check — flagged unverified (fail-closed; never a silent pass, ADR-550 D5d)' : 'prohibition has no enforcement evidence — flagged unverified (fail-closed; never a silent pass, ADR-550 D5d)', }; } // D4 GUARD: a judgment-tier (or unknown-tier) prohibition is NEVER a silent green from this // deterministic helper — it always routes to human/LLM judgment review (ADR-550 D4; verify-phase.md). - // Only a test-tier item with wired enforcement evidence may go green, and even that is the deferred - // heavy half until the real negative-test enforcement mechanism lands (no #644 caller passes evidence). + // Only a test-tier item with wired enforcement evidence may go green; the producer that supplies + // that evidence (`prohibition-enforcement`, #1259) runs the wired check and requires a genuine pass. if (tier === 'test') { return { status: 'green', diff --git a/src/prohibition-enforcement.cts b/src/prohibition-enforcement.cts new file mode 100644 index 000000000..d3284462d --- /dev/null +++ b/src/prohibition-enforcement.cts @@ -0,0 +1,496 @@ +/** + * prohibition-enforcement — the deterministic PRODUCER for test-tier prohibition verification + * (#1259, ADR-550 Decision 5d "heavy half"; the D1 seam — a NEW deterministic gsd-tools + * sub-command, NOT free-form workflow prose). + * + * Today `dispositionForProhibition()` (src/probe-core.cts) already carries the POLICY seam: with + * non-empty `enforcementEvidence` AND `tier === 'test'` it returns `{ status: 'green' }` (the + * branch at probe-core 420-427); with empty evidence it fails closed to flagged-unverified. But + * NOTHING in the live pipeline ever produced `enforcementEvidence`, so the green branch was + * unreachable. This module is the missing producer: it LOCATES the wired mechanical check from a + * check descriptor, RUNS it and requires a genuine NON-VACUOUS pass, builds a typed + * `enforcementEvidence` array on PASS, and emits the `dispositionForProhibition` verdict as JSON. + * The green/fail-closed policy itself is untouched (no src/probe-core.cts edit). + * + * Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR a lint/AST rule + * (e.g. the in-tree `local/no-source-grep` rule — the D4 dogfood anchor, run via the project flat + * config so the plugin loads). A missing, non-attested, or genuinely-non-passing check hard-gates + * (flagged, non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3) — never a silent + * green. + * + * FAIL-FIRST IS CALLER-ATTESTED (honest scope, #1259): the producer requires the caller to ATTEST + * `failFirst: true` AND requires the runner to observe a real non-vacuous pass — but it does NOT + * independently prove the check fails-on-violation. Genuine fail-first confirmation needs running + * the check against a known violation fixture and is a tracked follow-up (recorded in ADR-550's D5d + * note). What ships here closes the permanent-`gaps_found` dead-end with a genuinely-executed gate; + * it does not yet replace caller attestation with machine proof of the red-first property. + * + * Authored as strict TypeScript (`src/prohibition-enforcement.cts`) and compiled by + * `tsc -p tsconfig.build.json` (`npm run build:lib`) to the gitignored runtime artifact + * `gsd-core/bin/lib/prohibition-enforcement.cjs`. Do NOT hand-write the `.cjs`; it is emitted. + * + * DETERMINISM SCOPE: the DECISION layer is pure/deterministic and no-LLM — given a `runCheck` result + * the disposition is same-input-same-output and mutation-survivable, and the parse/filter helpers + * (`parseNodeTestSummary`, `tapTestNames`, `eslintJsonHasRule`, `eslintHasFatalError`, …) are pure. + * The DEFAULT REAL runner is NOT pure — it spawns `node --test` / eslint, so its result depends on the + * environment (eslint version + flat config, node version, the target file). That is why the runner is + * an injectable seam (`runCheck`): the contract is unit-tested against injected results, mirroring the + * injectable I/O pattern in `runProbeCli` / `ProbeCliOptions`. + */ + +import fs from 'node:fs'; +import path from 'node:path'; +import { execFileSync } from 'node:child_process'; +// Import the leaf I/O module directly, not the core.cjs re-export spine (being retired, #1268). +// eslint-disable-next-line @typescript-eslint/no-require-imports +import io = require('./io.cjs'); +const { output, error, ERROR_REASON } = io; +import { dispositionForProhibition } from './probe-core.cjs'; +import type { ProhibitionDisposition } from './probe-core.cjs'; + +/** The two accepted wired-check kinds (ADR-550 D2). */ +export type CheckKind = 'node-test' | 'lint-rule'; + +/** + * A descriptor of the wired mechanical check that asserts the must-NOT. `kind` selects the + * runner family; `target` is the negative-test file path (node-test) or the PATH to lint + * (lint-rule); `rule` is the eslint rule id (lint-rule only — e.g. `local/no-source-grep`) and is + * REQUIRED for the lint-rule kind (a lint-rule descriptor without it is not a valid wired check); + * `failFirst` is the caller's ATTESTATION that the check is a genuine `regression-must-fail-first` + * proof (caller-declared — see the module docstring; not independently confirmed at verify time). + */ +export interface CheckDescriptor { + kind: CheckKind; + target: string; + rule?: string; + failFirst?: boolean; +} + +/** + * The result a check-runner returns: whether the check genuinely, non-vacuously PASSED. The runner + * reports only what it can OBSERVE (a real pass) — it does NOT determine fail-first. Whether the + * check is a `regression-must-fail-first` proof is a CALLER-ATTESTED property of the descriptor + * (`CheckDescriptor.failFirst`); the producer cannot independently confirm it at verify time without + * a violation fixture (a tracked follow-up — see the module docstring). + */ +export interface CheckRunResult { + passed: boolean; +} + +/** A single typed enforcement-evidence record (the array `dispositionForProhibition` reads). */ +export interface EnforcementEvidence { + kind: CheckKind; + target: string; + rule?: string; + failFirst: boolean; + passed: boolean; +} + +/** Injectable options for `runProhibitionEnforcement` (defaults wire to the real runner). */ +export interface EnforcementOptions { + /** Runs the located check; injected in tests so no real subprocess is spawned. */ + runCheck?: (check: CheckDescriptor) => CheckRunResult; + /** Verify mode — recorded for transparency; the hard-gate applies in BOTH modes (ADR-550 D4). */ + mode?: string; + /** Project root for the default real runner (defaults to process.cwd()). */ + cwd?: string; + /** Override the per-kind subprocess timeout (ms); defaults to 30s (node-test) / 60s (eslint). + * Injected in tests to prove the fail-closed-on-timeout bound without a 30s wait. */ + timeoutMs?: number; +} + +/** The producer's verdict: the disposition PLUS the located/kind/evidence provenance. */ +export interface EnforcementResult extends ProhibitionDisposition { + located: boolean; + kind: CheckKind | null; + evidence: EnforcementEvidence[]; + mode?: string; +} + +/** node --test argv. Forces the TAP reporter so the summary counts are parseable + version-stable; + * `--` before the target so a target starting with `-` is not parsed as a flag (option-injection). */ +export function buildNodeTestArgs(check: CheckDescriptor): string[] { + return ['--test', '--test-reporter=tap', '--', check.target]; +} + +/** eslint argv (the args AFTER the eslint CLI path). Runs the project flat config so plugin rules + * (e.g. `local/*`) load — `--rule` CANNOT load a plugin, so we lint the TARGET path as JSON and + * filter by rule id. `--no-warn-ignored` makes an eslint-IGNORED target return `[]` (not a length-1 + * "File ignored" result) so an ignored path fails closed via the vacuity guard, not a false green. + * `--` before the target so a target starting with `-` is not parsed as a flag (option-injection). */ +export function buildLintArgs(check: CheckDescriptor): string[] { + return ['--no-warn-ignored', '--format', 'json', '--', check.target]; +} + +/** + * Resolve the project's eslint CLI entry portably (no `npx` — not spawnable via `execFileSync` on + * Windows). Resolves eslint's package.json from the target project's `node_modules` and derives + * `bin/eslint.js`, so it is run as `node ` (portable). Returns null if eslint is not installed + * (→ the lint-rule check fails closed, never throws). + */ +function resolveEslintCli(cwd: string): string | null { + try { + const pkg = require.resolve('eslint/package.json', { paths: [cwd] }); + const cli = path.join(path.dirname(pkg), 'bin', 'eslint.js'); + return fs.existsSync(cli) ? cli : null; + } catch { + return null; + } +} + +/** Basename of a path, separator-agnostic (handles `\` and `/` so node-test names compare stably + * across OSes / node versions that report the file-test by differing path forms). */ +function baseOf(p: string): string { + return typeof p === 'string' ? (p.split(/[\\/]/).pop() ?? p) : ''; +} + +/** + * Pure parser for the `node --test` TAP summary. A genuine pass is NON-VACUOUS: exit 0 is NOT enough + * (an empty / all-skipped / deleted-negative-test file exits 0 with `# tests 0`). Mutation-pinned by + * unit tests so a threshold flip is caught. + */ +export function parseNodeTestSummary(out: string): { tests: number; pass: number; fail: number; cancelled: number } { + const num = (re: RegExp): number => { + const m = typeof out === 'string' ? out.match(re) : null; + return m ? Number(m[1]) : 0; + }; + return { + tests: num(/^# tests (\d+)/m), + pass: num(/^# pass (\d+)/m), + fail: num(/^# fail (\d+)/m), + cancelled: num(/^# cancelled (\d+)/m), + }; +} + +/** The names of REAL (run) tests from TAP `ok N - ` / `not ok N - ` lines. A line with a + * `# SKIP` / `# TODO` directive is EXCLUDED — a skipped/todo negative test never executed, so it must + * not count toward non-vacuity (#1259 m1). */ +export function tapTestNames(out: string): string[] { + if (typeof out !== 'string') return []; + const names: string[] = []; + const re = /^(?:not )?ok \d+ - (.+)$/gm; + let m: RegExpExecArray | null; + while ((m = re.exec(out)) !== null) { + const rest = m[1]; + if (/\s#\s*(?:SKIP|TODO)\b/i.test(rest)) continue; // skipped/todo did not run + names.push(rest.replace(/\s+#\s.*$/, '').trim()); + } + return names; +} + +/** + * A non-vacuous node-test pass: at least one test, at least one pass, zero failures — AND at least + * one reported test whose name is NOT merely the target file. `node --test ` counts a file + * with ZERO `test()` calls as one passing "test" named after the file, so the counts alone cannot + * tell an empty/deleted negative test from a real one (the #1259 BL-01 false-green). Requiring a + * named test distinct from the file closes that hole. + * + * KNOWN CONSTRAINT (fail-closed, not a hole): a real test whose `test('...')` name is EXACTLY the + * target file's basename emits TAP indistinguishable from an empty file and is conservatively + * rejected (non-green). A wired negative test must carry a descriptive name, not be named after its + * own file — a benign authoring constraint, and the safe direction if violated. + */ +export function isNonVacuousNodeTestPass(out: string, target: string): boolean { + const s = parseNodeTestSummary(out); + // >=1 test, >=1 pass, ZERO failures AND ZERO cancelled (a cancelled run is not a clean pass, m1). + if (!(s.tests >= 1 && s.pass >= 1 && s.fail === 0 && s.cancelled === 0)) return false; + // Compare BASENAMES: node reports the file-test by varying path forms across OS / node version + // (absolute, relative, normalized), so an exact-string compare misfires. A real test name (e.g. + // "guards the must-NOT") has no separators, so its basename never equals the target file's. + const tgtBase = baseOf(target); + return tapTestNames(out).some((n) => baseOf(n) !== tgtBase); +} + +/** Number of file results in an eslint `--format json` report (0 if unparseable / not an array). */ +export function eslintFileResultCount(jsonText: string): number { + try { + const parsed: unknown = JSON.parse(jsonText); + return Array.isArray(parsed) ? parsed.length : 0; + } catch { + return 0; + } +} + +/** Messages array of a single eslint file-result (empty if absent / wrong shape). */ +function eslintMessages(file: unknown, key: 'messages' | 'suppressedMessages'): Array<{ ruleId?: unknown; fatal?: unknown }> { + return file && typeof file === 'object' && Array.isArray((file as Record)[key]) + ? ((file as Record)[key] as Array<{ ruleId?: unknown; fatal?: unknown }>) + : []; +} + +/** + * True if the eslint `--format json` report has a FATAL / parse error — meaning the rule never got + * to run on the target. A prohibition gate must fail closed on "the rule didn't execute" (#1259 B1), + * NOT treat a length-1 fatal result as "clean". Unparseable report -> true (fail-closed). + */ +export function eslintHasFatalError(jsonText: string): boolean { + let parsed: unknown; + try { + parsed = JSON.parse(jsonText); + } catch { + return true; + } + if (!Array.isArray(parsed)) return true; + for (const file of parsed) { + const fec = file && typeof file === 'object' ? (file as { fatalErrorCount?: unknown }).fatalErrorCount : undefined; + if (typeof fec === 'number' && fec > 0) return true; + if (eslintMessages(file, 'messages').some((m) => m && m.fatal === true)) return true; + } + return false; +} + +/** True if the eslint `--format json` report has ANY message for `rule` — in EITHER `messages` or + * `suppressedMessages` (an inline `// eslint-disable` of the rule is still a violation, #1259 B1). + * Unparseable -> true (fail-closed: treat an unreadable report as a violation, not a silent pass). */ +export function eslintJsonHasRule(jsonText: string, rule: string): boolean { + let parsed: unknown; + try { + parsed = JSON.parse(jsonText); + } catch { + return true; + } + if (!Array.isArray(parsed)) return true; + for (const file of parsed) { + for (const key of ['messages', 'suppressedMessages'] as const) { + for (const msg of eslintMessages(file, key)) { + if (msg && typeof msg === 'object' && msg.ruleId === rule) return true; + } + } + } + return false; +} + +/** + * The default REAL check runner (used when no `runCheck` is injected). Reports only an OBSERVED, + * genuinely-non-vacuous pass; guarded so a missing tool / non-zero exit yields a non-passing result, + * NEVER an uncaught throw (the no-throw contract). It does NOT determine fail-first (caller-attested). + * - node-test: runs `node --test` (TAP) and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail + * AND a reported test named distinctly from the file). A bare exit 0 for an empty/zero-test file + * — which `node --test` counts as one passing "test" named after the file — is NOT a pass (the + * #1259 BL-01 false-green fix). + * - lint-rule: runs the project eslint as `node --format json ` (flat config + * loads `local/*` plugins) and requires the target to actually lint (>=1 file result) AND ZERO + * messages for the specific rule id. `--rule` cannot load a plugin rule, so we filter the + * structured report by `ruleId` instead (the #1259 SF-01 fix). + * + * Both kinds spawn via `process.execPath` (never bare `node`/`npx` — not portably spawnable via + * `execFileSync` on Windows) with arg arrays (no shell → no injection from a caller-supplied target). + */ +/** + * Env for spawned checks: strip `NODE_TEST_CONTEXT` and `NODE_OPTIONS` so an AMBIENT test-runner + * context (e.g. running verify under `node --test`, which sets `NODE_TEST_CONTEXT=child-v8`) cannot + * turn the child `node --test` into a silent v8-reporter worker that emits no parseable TAP — which + * would otherwise corrupt the verdict. Deterministic, environment-independent execution. + */ +function childEnv(): NodeJS.ProcessEnv { + const env = { ...process.env }; + delete env.NODE_TEST_CONTEXT; + delete env.NODE_OPTIONS; + return env; +} + +// Bounded subprocess limits (DEFECT.UNBOUNDED-SUBPROCESS): a stuck wired test / eslint must not hang +// verify forever. On timeout `execFileSync` throws -> caught -> fail-closed (degraded, non-passing). +// `maxBuffer` caps output so a runaway producer throws (safe direction) rather than OOMs the verifier. +const NODE_TEST_TIMEOUT_MS = 30_000; +const ESLINT_TIMEOUT_MS = 60_000; +const CHECK_MAX_BUFFER = 16 * 1024 * 1024; + +/** Resolve the effective timeout: only a POSITIVE override is honored — `0` (which Node treats as + * "no timeout") or a negative value falls back to the bounded default, so the subprocess is ALWAYS + * bounded (a `timeoutMs: 0` injection can never disable the bound). */ +function posTimeout(timeoutMs: number | undefined, def: number): number { + return typeof timeoutMs === 'number' && timeoutMs > 0 ? timeoutMs : def; +} + +function defaultRunCheck(check: CheckDescriptor, cwd: string, timeoutMs?: number): CheckRunResult { + try { + if (check.kind === 'node-test') { + let out = ''; + try { + out = execFileSync(process.execPath, buildNodeTestArgs(check), { + cwd, + encoding: 'utf-8', + stdio: ['ignore', 'pipe', 'pipe'], + windowsHide: true, + env: childEnv(), + timeout: posTimeout(timeoutMs, NODE_TEST_TIMEOUT_MS), + maxBuffer: CHECK_MAX_BUFFER, + }); + } catch (e) { + // A failing/timed-out run exits non-zero or is killed (partial TAP on stdout, no `# pass` + // summary). Parse what we have: a real failure or timeout -> not a non-vacuous pass -> false. + const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : ''; + out = typeof stdout === 'string' ? stdout : ''; + } + return { passed: isNonVacuousNodeTestPass(out, check.target) }; + } + if (check.kind === 'lint-rule') { + const eslintCli = resolveEslintCli(cwd); + if (!eslintCli) return { passed: false }; // eslint not installed -> fail closed, never throw + let json = ''; + try { + json = execFileSync(process.execPath, [eslintCli, ...buildLintArgs(check)], { + cwd, + encoding: 'utf-8', + stdio: ['ignore', 'pipe', 'pipe'], + windowsHide: true, + env: childEnv(), + timeout: posTimeout(timeoutMs, ESLINT_TIMEOUT_MS), + maxBuffer: CHECK_MAX_BUFFER, + }); + } catch (e) { + // eslint exits non-zero when ANY error is present; the JSON report is still on stdout. + // A timeout/kill leaves no parseable JSON -> eslintHasFatalError(unparseable) -> fail-closed. + const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : ''; + json = typeof stdout === 'string' ? stdout : ''; + } + // PASS requires: the target actually linted (>=1 file result), NO fatal/parse error (the rule + // must have RUN — #1259 B1), and ZERO messages for the rule (in messages OR suppressedMessages). + const lintedSomething = eslintFileResultCount(json) >= 1; + return { + passed: lintedSomething && !eslintHasFatalError(json) && !eslintJsonHasRule(json, check.rule as string), + }; + } + // Unknown kind — defensive; the LOCATE guard already rejects it. + return { passed: false }; + } catch { + return { passed: false }; + } +} + +/** + * LOCATE -> CONFIRM fail-first -> RUN -> build enforcementEvidence -> dispositionForProhibition. + * + * (1) LOCATE: if no well-formed check descriptor is locatable -> fail-closed + * (`dispositionForProhibition` with empty evidence) plus `{ located: false, kind: null, evidence: [] }`. + * (2) ATTEST + RUN: require the caller to attest `failFirst: true` and run it via `runCheck`. A check + * the caller does not attest as fail-first, or that does not genuinely (non-vacuously) PASS -> + * fail-closed disposition with `located: true` (a real located miss, non-green, flagged) in BOTH + * modes. Fail-first is caller-attested, not independently proven (see module docstring). + * (3) PASS: build a typed `enforcementEvidence` array and call `dispositionForProhibition` — the + * non-empty array flips a test-tier item to green (the previously-unreachable branch). + * + * Pure/deterministic: same (prohibition, check, runCheck) -> same result. + */ +export function runProhibitionEnforcement( + prohibition: unknown, + check: CheckDescriptor | null | undefined, + options: EnforcementOptions = {}, +): EnforcementResult { + const mode = options.mode; + + // (1) LOCATE — no locatable, well-formed wired check -> fail-closed, located: false. The kind must + // be one of the two known kinds; the target must be a non-empty string; a lint-rule descriptor MUST + // also carry a non-empty `rule` id (its target is the lint PATH, not the rule). An under-specified + // descriptor is not a valid wired check, so it is not locatable (does not rely on the runner failing). + const c = check && typeof check === 'object' ? check : null; + const validKind = !!c && (c.kind === 'node-test' || c.kind === 'lint-rule'); + const validTarget = !!c && typeof c.target === 'string' && c.target.trim().length > 0; + const validRule = !!c && (c.kind !== 'lint-rule' || (typeof c.rule === 'string' && c.rule.trim().length > 0)); + if (!c || !validKind || !validTarget || !validRule) { + const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] }); + return { ...disposition, located: false, kind: null, evidence: [], ...(mode ? { mode } : {}) }; + } + + const runCheck = options.runCheck ?? ((toRun: CheckDescriptor) => defaultRunCheck(toRun, options.cwd ?? process.cwd(), options.timeoutMs)); + + // (2) ATTEST fail-first (CALLER-DECLARED) + RUN. The caller must attest `failFirst: true` AND the + // runner must observe a genuine NON-VACUOUS pass. The producer does NOT independently prove + // fail-first (that needs a violation fixture — tracked follow-up, ADR-550 D5d). A non-attested or + // non-passing check hard-gates (never green) in BOTH modes. + const attestedFailFirst = c.failFirst === true; + // No-throw contract end-to-end: even a (test-injected) runCheck that throws must fail closed, + // never propagate. The default real runner already never throws. + let run: CheckRunResult; + try { + run = runCheck(c); + } catch { + run = { passed: false }; + } + const passed = attestedFailFirst && run.passed === true; + + if (!passed) { + // NOT attested fail-first OR did not genuinely pass -> fail-closed, located: true (an actual + // located miss/fail). Hard-gate applies in BOTH modes; the disposition stays non-green / flagged. + const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] }); + return { + ...disposition, + located: true, + kind: c.kind, + evidence: [], + ...(mode ? { mode } : {}), + }; + } + + // (3) PASS -> build typed enforcementEvidence and let the policy flip a test-tier item green. + // `failFirst` here is the caller's attestation (recorded for provenance), not a machine proof. + const evidence: EnforcementEvidence[] = [{ + kind: c.kind, + target: c.target, + ...(typeof c.rule === 'string' ? { rule: c.rule } : {}), + failFirst: true, + passed: true, + }]; + const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: evidence }); + return { + ...disposition, + located: true, + kind: c.kind, + evidence, + ...(mode ? { mode } : {}), + }; +} + +/** + * Parse a `{ prohibition, check, mode }` request from a JSON file path or inline `--json` string. + * Returns null on any parse failure (the caller surfaces a structured error, never a throw). + */ +function parseRequest(args: string[]): { prohibition: unknown; check: CheckDescriptor | null; mode?: string } | null { + // args[0] = 'check', args[1] = 'prohibition-enforcement', args[2] = + const jsonFlagIdx = args.indexOf('--json'); + let payload = ''; + if (jsonFlagIdx !== -1 && typeof args[jsonFlagIdx + 1] === 'string') { + payload = args[jsonFlagIdx + 1]; + } else if (typeof args[2] === 'string' && args[2]) { + try { + payload = fs.readFileSync(args[2], 'utf-8'); + } catch { + return null; + } + } else { + return null; + } + try { + const parsed = JSON.parse(payload) as Record; + const checkRaw = parsed['check']; + const check: CheckDescriptor | null = (checkRaw && typeof checkRaw === 'object') + ? (checkRaw as CheckDescriptor) + : null; + const modeRaw = parsed['mode']; + const mode = typeof modeRaw === 'string' ? modeRaw : undefined; + return { prohibition: parsed['prohibition'] ?? null, check, ...(mode ? { mode } : {}) }; + } catch { + return null; + } +} + +/** + * CLI surface: `gsd_run check prohibition-enforcement ` (or `--json ''`). + * Parses the request, runs the producer, and emits the result as JSON. Honors the no-throw + * contract: malformed input -> structured `error(...)`, never an uncaught throw. + */ +export function routeProhibitionEnforcement(args: string[], raw: boolean): void { + const req = parseRequest(args); + if (!req) { + error( + 'prohibition-enforcement requires a JSON request: check prohibition-enforcement | --json \'{"prohibition":{...},"check":{...}}\'', + ERROR_REASON.SDK_MISSING_ARG, + ); + return; + } + const result = runProhibitionEnforcement(req.prohibition, req.check, req.mode ? { mode: req.mode } : {}); + output(result, raw, undefined); +} + +export {}; diff --git a/tests/prohibition-enforcement.property.test.cjs b/tests/prohibition-enforcement.property.test.cjs new file mode 100644 index 000000000..6d07e5ade --- /dev/null +++ b/tests/prohibition-enforcement.property.test.cjs @@ -0,0 +1,60 @@ +// Property-based tests for the prohibition-enforcement parsing/transform helpers (#1259, ADR-550 D5d). +// RULESET.TESTS.property-based-testing: the producer is a parsing module (parseNodeTestSummary, +// tapTestNames, eslintJsonHasRule, eslintHasFatalError, eslintFileResultCount), so it carries +// fast-check invariants — especially the fail-closed safety invariants of the verify-time gate. +'use strict'; +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const path = require('node:path'); +const fc = require('./helpers/fast-check-setup.cjs'); + +const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs'); + +describe('prohibition-enforcement properties (#1259)', () => { + test('parseNodeTestSummary never throws and returns non-negative integer counts', () => { + const enforce = require(ENFORCEMENT_LIB); + fc.assert(fc.property(fc.string(), (s) => { + const r = enforce.parseNodeTestSummary(s); + for (const k of ['tests', 'pass', 'fail', 'cancelled']) { + assert.ok(Number.isInteger(r[k]) && r[k] >= 0, `${k} is a non-negative integer`); + } + })); + }); + + test('isNonVacuousNodeTestPass FAIL-CLOSED: a run with any failure or cancellation never greens', () => { + const enforce = require(ENFORCEMENT_LIB); + fc.assert(fc.property( + fc.nat({ max: 50 }), fc.integer({ min: 1, max: 50 }), fc.nat({ max: 50 }), fc.string(), + (tests, failOrCancel, pass, name) => { + // A summary with fail>=1 (and, separately, cancelled>=1) must NEVER be a non-vacuous pass, + // regardless of the reported test name. + const failing = `ok 1 - ${name}\n# tests ${tests + 1}\n# pass ${pass}\n# fail ${failOrCancel}\n# cancelled 0\n`; + assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false); + const cancelled = `ok 1 - ${name}\n# tests ${tests + 1}\n# pass ${pass}\n# fail 0\n# cancelled ${failOrCancel}\n`; + assert.equal(enforce.isNonVacuousNodeTestPass(cancelled, 'neg.test.cjs'), false); + }, + )); + }); + + test('eslintJsonHasRule / eslintHasFatalError FAIL-CLOSED on any non-JSON / non-array input', () => { + const enforce = require(ENFORCEMENT_LIB); + fc.assert(fc.property(fc.string(), (s) => { + // Only exercise strings that are NOT a valid JSON array (the unreadable-report branch). + let isArray = false; + try { isArray = Array.isArray(JSON.parse(s)); } catch { isArray = false; } + fc.pre(!isArray); + assert.equal(enforce.eslintJsonHasRule(s, 'local/no-source-grep'), true, 'unreadable report -> violation (fail-closed)'); + assert.equal(enforce.eslintHasFatalError(s), true, 'unreadable report -> fatal (fail-closed)'); + })); + }); + + test('eslintFileResultCount never throws and is non-negative', () => { + const enforce = require(ENFORCEMENT_LIB); + fc.assert(fc.property(fc.string(), (s) => { + const n = enforce.eslintFileResultCount(s); + assert.ok(Number.isInteger(n) && n >= 0); + })); + }); +}); diff --git a/tests/prohibition-enforcement.test.cjs b/tests/prohibition-enforcement.test.cjs new file mode 100644 index 000000000..014133523 --- /dev/null +++ b/tests/prohibition-enforcement.test.cjs @@ -0,0 +1,386 @@ +// Behavioral tests for the deterministic prohibition-enforcement producer (#1259, ADR-550 D5d +// "heavy half"). Requires the BUILT gsd-core/bin/lib/prohibition-enforcement.cjs — authored as +// src/prohibition-enforcement.cts and compiled by `npm run build:lib` (mirrors how the verify-tier +// suite requires the built probe-core.cjs). Typed-field assertions only; the check-runner is +// injected so no real subprocess is spawned. No source-grep. +'use strict'; +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const path = require('node:path'); +const { createTempDir, cleanup } = require('./helpers.cjs'); + +const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs'); + +const TEST_TIER = Object.freeze({ + requirement_id: 'R1', + category: 'safety', + status: 'resolved', + verification: 'test', + resolution: null, + reason: null, + statement: 'MUST NOT read source files and text-search them in tests', +}); + +describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR-550 D5d)', () => { + test('exports the producer + route functions', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(typeof enforce.runProhibitionEnforcement, 'function', + 'must export runProhibitionEnforcement (the deterministic producer)'); + assert.equal(typeof enforce.routeProhibitionEnforcement, 'function', + 'must export routeProhibitionEnforcement (the CLI surface)'); + }); + + test('locate-miss (no check descriptor) -> fail-closed, located:false, no evidence', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement(TEST_TIER, null, { + runCheck: () => ({ passed: true }), + }); + assert.equal(result.located, false, 'no locatable check'); + assert.notEqual(result.status, 'green', 'locate-miss must never be green'); + assert.equal(result.flagged, true, 'locate-miss must be flagged'); + assert.equal(result.kind, null, 'no kind when nothing located'); + assert.ok(Array.isArray(result.evidence) && result.evidence.length === 0, 'no evidence on locate-miss'); + }); + + test('malformed check descriptor (missing target) -> treated as locate-miss', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement(TEST_TIER, { kind: 'node-test' }, { + runCheck: () => ({ passed: true }), + }); + assert.equal(result.located, false, 'a descriptor without a target is not locatable'); + assert.notEqual(result.status, 'green'); + assert.equal(result.flagged, true); + }); + + test('node-test check that passes -> green + non-empty typed evidence', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: true }) }, + ); + assert.equal(result.status, 'green'); + assert.equal(result.flagged, false); + assert.equal(result.tier, 'test'); + assert.equal(result.located, true); + assert.equal(result.kind, 'node-test'); + assert.equal(result.evidence.length, 1, 'one evidence record built'); + const ev = result.evidence[0]; + assert.equal(ev.kind, 'node-test'); + assert.equal(ev.target, 'tests/neg.test.cjs'); + assert.equal(ev.failFirst, true); + assert.equal(ev.passed, true); + }); + + test('lint-rule (no-source-grep) check that passes -> green, evidence carries rule id', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true }, + { runCheck: () => ({ passed: true }) }, + ); + assert.equal(result.status, 'green'); + assert.equal(result.flagged, false); + assert.equal(result.kind, 'lint-rule'); + assert.equal(result.evidence[0].kind, 'lint-rule'); + assert.equal(result.evidence[0].rule, 'local/no-source-grep', 'evidence records which rule asserted the must-NOT'); + assert.equal(result.evidence[0].target, 'tests/', 'evidence records the linted target path, not the rule id'); + }); + + test('buildLintArgs runs the project eslint as JSON over the target (plugins load via flat config; #1259 SF-01)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(typeof enforce.buildLintArgs, 'function', + 'must export buildLintArgs — the eslint argv builder for the lint-rule real runner'); + const argv = enforce.buildLintArgs({ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/' }); + assert.ok(Array.isArray(argv), 'argv is an array'); + const fmtIdx = argv.indexOf('--format'); + assert.ok(fmtIdx !== -1 && argv[fmtIdx + 1] === 'json', + 'emits --format json so the report can be filtered by ruleId'); + assert.ok(argv.includes('--no-warn-ignored'), + 'must pass --no-warn-ignored so an eslint-ignored target returns [] (fails closed), not a length-1 warning result'); + assert.ok(!argv.includes('--rule'), + 'must NOT use --rule — it cannot load a plugin rule like local/no-source-grep (the SF-01 bug)'); + assert.equal(argv[argv.length - 1], 'tests/', 'the LAST arg is the lint target path'); + }); + + test('lint-rule descriptor missing its rule id -> locate-miss, never green', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'lint-rule', target: 'tests/', failFirst: true }, // no `rule` + { runCheck: () => ({ passed: true }) }, + ); + assert.notEqual(result.status, 'green', 'a lint-rule with no rule id is not a valid wired check'); + assert.equal(result.flagged, true); + assert.equal(result.located, false, 'an under-specified lint-rule descriptor is not locatable'); + }); + + test('check that FAILS -> hard-gate (non-green, flagged), located:true, no evidence', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: false }) }, + ); + assert.notEqual(result.status, 'green'); + assert.equal(result.flagged, true); + assert.equal(result.located, true, 'the check was located even though it failed'); + assert.equal(result.evidence.length, 0, 'a failing check builds no evidence'); + }); + + test('caller does NOT attest fail-first (descriptor failFirst:false) -> hard-gate, never green', () => { + const enforce = require(ENFORCEMENT_LIB); + // fail-first is caller-attested (#1259 BL-02): a check the caller does not attest as fail-first + // is not a valid regression proof and must never green, even if the run passes. + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: false }, + { runCheck: () => ({ passed: true }) }, + ); + assert.notEqual(result.status, 'green', 'a non-attested check is not a valid regression proof'); + assert.equal(result.flagged, true); + }); + + test('a runCheck that THROWS fails closed, never propagates (no-throw contract, NEW-WR-01)', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + { runCheck: () => { throw new Error('runner blew up'); } }, + ); + assert.notEqual(result.status, 'green', 'a throwing runner must never green'); + assert.equal(result.flagged, true); + assert.equal(result.located, true); + }); + + test('hard-gates in BOTH modes on a failing check (ADR-550 D4)', () => { + const enforce = require(ENFORCEMENT_LIB); + for (const mode of ['interactive', 'autonomous']) { + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: false }), mode }, + ); + assert.notEqual(result.status, 'green', `non-green in ${mode}`); + assert.equal(result.flagged, true, `flagged in ${mode}`); + assert.equal(result.mode, mode, 'mode echoed for transparency'); + } + }); + + test('passing run echoes the requested mode without changing the green verdict', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: true }), mode: 'autonomous' }, + ); + assert.equal(result.status, 'green', 'a passing wired check is green in autonomous mode too'); + assert.equal(result.mode, 'autonomous'); + }); + + test('routeProhibitionEnforcement parses a JSON request file and emits a structured result', (t) => { + const fs = require('node:fs'); + const { execFileSync } = require('node:child_process'); + // Write a request file; the route reads it and runs the node-test descriptor's default runner + // (its target does not exist, so it fail-closes deterministically — we assert the JSON SHAPE, + // not a green verdict). We invoke the built CLI surface in a child process so output() + // (writeAllSync to fd 1) is captured on stdout — no source-grep (we parse our own emitted JSON). + const dir = createTempDir('prohib-enf-'); + const reqPath = path.join(dir, 'req.json'); + const runnerPath = path.join(dir, 'runner.cjs'); + fs.writeFileSync(reqPath, JSON.stringify({ + prohibition: TEST_TIER, + check: { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true }, + mode: 'autonomous', + })); + // A tiny runner that requires the BUILT module and invokes the route — output() writes to fd 1. + fs.writeFileSync(runnerPath, + "require(" + JSON.stringify(ENFORCEMENT_LIB) + ")" + + ".routeProhibitionEnforcement(['check','prohibition-enforcement'," + JSON.stringify(reqPath) + "], false);\n"); + t.after(() => cleanup(dir)); + + const captured = execFileSync('node', [runnerPath], { encoding: 'utf-8' }); + const parsed = JSON.parse(captured); + assert.equal(typeof parsed, 'object', 'route emits a JSON object'); + assert.equal(parsed.tier, 'test', 'tier is preserved through the CLI surface'); + assert.equal(parsed.located, true, 'the check descriptor was located'); + assert.equal(parsed.mode, 'autonomous', 'mode flows through the CLI surface'); + assert.equal(typeof parsed.flagged, 'boolean', 'flagged is a typed boolean'); + assert.ok(Array.isArray(parsed.evidence), 'evidence is an array'); + }); +}); + +// ─── Real-runner helpers (mutation-pinned; #1259 BL-01 / SF-01) ───────────────── +// These pin the deterministic parsing/threshold logic of the REAL runner so a Stryker mutant that +// weakens "non-vacuous pass" or the ruleId filter is caught — the contract the injected-runner tests +// above deliberately bypass. +describe('prohibition-enforcement real-runner helpers (#1259)', () => { + test('parseNodeTestSummary extracts the TAP tests/pass/fail/cancelled counts', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.deepEqual(enforce.parseNodeTestSummary('# tests 3\n# pass 2\n# fail 1\n# cancelled 1\n'), + { tests: 3, pass: 2, fail: 1, cancelled: 1 }); + assert.deepEqual(enforce.parseNodeTestSummary('no summary here'), { tests: 0, pass: 0, fail: 0, cancelled: 0 }); + }); + + test('tapTestNames EXCLUDES skipped/todo tests (they never ran, m1)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.deepEqual(enforce.tapTestNames('ok 1 - guards the must-NOT\nok 2 - other # SKIP\nok 3 - later # TODO\n'), + ['guards the must-NOT'], 'a # SKIP / # TODO test is not a real run and must not count'); + }); + + test('isNonVacuousNodeTestPass: a SKIPPED negative test (file wrapper passes) is NOT a pass (m1)', () => { + const enforce = require(ENFORCEMENT_LIB); + // file wrapper + a skipped negative test: pass>=1 but the only named test is skipped -> vacuous. + const skipped = 'ok 1 - empty.test.cjs\nok 2 - the negative test # SKIP\n# tests 2\n# pass 2\n# fail 0\n# cancelled 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(skipped, 'empty.test.cjs'), false, + 'a skipped negative test never executed -> must not green'); + }); + + test('isNonVacuousNodeTestPass: a CANCELLED run is not a pass (m1)', () => { + const enforce = require(ENFORCEMENT_LIB); + const cancelled = 'ok 1 - guards\n# tests 1\n# pass 1\n# fail 0\n# cancelled 1\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(cancelled, 'neg.test.cjs'), false, + 'a cancelled run is not a clean pass'); + }); + + test('isNonVacuousNodeTestPass: an empty file (node names the test after the file) is NOT a pass (BL-01)', () => { + const enforce = require(ENFORCEMENT_LIB); + // node --test of a zero-test file: `ok 1 - empty.test.cjs`, `# tests 1 # pass 1` — counts alone + // cannot distinguish it from a real test, so the file-named result must NOT count as a pass. + const empty = 'ok 1 - empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(empty, 'empty.test.cjs'), false, + 'a file-named-only result is vacuous — the BL-01 false-green guard'); + // BASENAME-NORMALIZED: node may report the file-test by an ABSOLUTE/normalized path while the + // descriptor target is relative (cross-OS / node-version). The basenames must still match → vacuous. + const emptyAbs = 'ok 1 - /tmp/x/empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(emptyAbs, 'empty.test.cjs'), false, + 'an absolute-path file-test name must still be recognized as vacuous (basename compare, WR-02)'); + // Mirror case (pins the TARGET-side basename): relative TAP name vs ABSOLUTE descriptor target. + const emptyRelName = 'ok 1 - neg.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(emptyRelName, '/abs/path/neg.test.cjs'), false, + 'a relative file-test name vs an absolute target must still be vacuous — both sides basename-normalized (WR-R4-01)'); + const real = 'ok 1 - guards the must-NOT\n1..1\n# tests 1\n# pass 1\n# fail 0\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(real, '/abs/path/neg.test.cjs'), true, + 'a real named test distinct from the file is a genuine pass (even vs an absolute target)'); + const failing = 'not ok 1 - guards\n# tests 1\n# pass 0\n# fail 1\n'; + assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false, + 'any failure means not a pass'); + }); + + test('eslintJsonHasRule detects a ruleId; unparseable report -> true (fail-closed)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'local/no-source-grep' }] }]), 'local/no-source-grep'), true); + assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'other' }] }]), 'local/no-source-grep'), false); + assert.equal(enforce.eslintJsonHasRule('not json', 'local/no-source-grep'), true, + 'an unreadable report must be treated as a violation, never a silent pass'); + }); + + test('eslintFileResultCount: 0 when nothing linted (vacuity guard)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(enforce.eslintFileResultCount(JSON.stringify([{}, {}])), 2); + assert.equal(enforce.eslintFileResultCount('[]'), 0); + assert.equal(enforce.eslintFileResultCount('garbage'), 0); + }); + + test('eslintHasFatalError: a parse/fatal error must fail closed (B1)', () => { + const enforce = require(ENFORCEMENT_LIB); + const fatal = JSON.stringify([{ messages: [{ ruleId: null, fatal: true, severity: 2, message: 'Parsing error' }], fatalErrorCount: 1 }]); + assert.equal(enforce.eslintHasFatalError(fatal), true, 'a fatal/parse error means the rule never ran -> fail closed'); + const clean = JSON.stringify([{ messages: [], fatalErrorCount: 0 }]); + assert.equal(enforce.eslintHasFatalError(clean), false, 'a clean lint has no fatal error'); + assert.equal(enforce.eslintHasFatalError('not json'), true, 'an unreadable report is treated as fatal (fail closed)'); + }); + + test('eslintJsonHasRule also reads suppressedMessages — an inline-disabled violation still counts (B1)', () => { + const enforce = require(ENFORCEMENT_LIB); + const suppressed = JSON.stringify([{ messages: [], suppressedMessages: [{ ruleId: 'local/no-source-grep' }] }]); + assert.equal(enforce.eslintJsonHasRule(suppressed, 'local/no-source-grep'), true, + 'a violation suppressed via // eslint-disable must NOT be treated as clean'); + }); +}); + +// ─── Real runner end-to-end (NO injected runCheck; #1259 SF-02 / BL-01 / SF-01) ── +// Spawns real subprocesses so the SHIPPING default runner is exercised — the gap that let BL-01 and +// SF-01 slip past the injected-double tests. Typed-field assertions only. +describe('prohibition-enforcement REAL runner end-to-end (#1259)', () => { + const fs = require('node:fs'); + + test('a genuine non-vacuous passing node-test greens via the real runner', (t) => { + const enforce = require(ENFORCEMENT_LIB); + const dir = createTempDir('prohib-real-pass-'); + t.after(() => cleanup(dir)); + const tf = path.join(dir, 'neg.test.cjs'); + fs.writeFileSync(tf, + "const { test } = require('node:test');\nconst assert = require('node:assert');\ntest('guards the must-NOT', () => { assert.ok(true); });\n"); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: tf, failFirst: true }, + { cwd: dir }, + ); + assert.equal(result.status, 'green', 'a real, passing, non-vacuous negative test must green'); + assert.equal(result.located, true); + assert.equal(result.evidence.length, 1); + }); + + test('a HANGING node-test fails closed via the bounded timeout (B2: no unbounded subprocess)', (t) => { + const enforce = require(ENFORCEMENT_LIB); + const dir = createTempDir('prohib-hang-'); + t.after(() => cleanup(dir)); + const tf = path.join(dir, 'hang.test.cjs'); + // A test that never returns; the bounded timeout must kill it and dispose non-green. + fs.writeFileSync(tf, + "const { test } = require('node:test');\ntest('hangs forever', () => { while (true) {} });\n"); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: tf, failFirst: true }, + { cwd: dir, timeoutMs: 1500 }, + ); + assert.notEqual(result.status, 'green', 'a hung check must be killed and fail closed — never hang verify or green'); + assert.equal(result.located, true); + }); + + test('an EMPTY node-test file (exit 0, zero tests) does NOT green via the real runner (BL-01)', (t) => { + const enforce = require(ENFORCEMENT_LIB); + const dir = createTempDir('prohib-real-empty-'); + t.after(() => cleanup(dir)); + const tf = path.join(dir, 'empty.test.cjs'); + fs.writeFileSync(tf, '// intentionally empty — no test cases\n'); + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'node-test', target: tf, failFirst: true }, + { cwd: dir }, + ); + assert.notEqual(result.status, 'green', 'an empty (zero-test) file must NEVER green — fail-closed'); + assert.equal(result.located, true, 'the check was located; it just did not genuinely pass'); + assert.equal(result.evidence.length, 0); + }); + + test('a clean in-tree target greens the lint-rule kind via the real eslint runner (SF-01: plugin loads)', () => { + const enforce = require(ENFORCEMENT_LIB); + // Runs real `npx eslint --format json src/clock.cts` under the project flat config (so the + // `local` plugin loads). src/clock.cts is a clean source with no no-source-grep violation. + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'src/clock.cts', failFirst: true }, + { cwd: process.cwd() }, + ); + assert.equal(result.status, 'green', 'a clean target with no no-source-grep violation must green via real eslint'); + assert.equal(result.kind, 'lint-rule'); + assert.equal(result.evidence[0].rule, 'local/no-source-grep'); + }); + + test('an eslint-IGNORED target does NOT green the lint-rule kind (vacuous-green guard, NEW-BL-01)', () => { + const enforce = require(ENFORCEMENT_LIB); + // The generated bin/lib artifact is eslint-ignored. Without --no-warn-ignored, eslint returns a + // length-1 "File ignored" result that would falsely pass the vacuity guard. It must fail closed. + const result = enforce.runProhibitionEnforcement( + TEST_TIER, + { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'gsd-core/bin/lib/prohibition-enforcement.cjs', failFirst: true }, + { cwd: process.cwd() }, + ); + assert.notEqual(result.status, 'green', 'an ignored path lints nothing — must NEVER green'); + assert.equal(result.located, true, 'the descriptor was well-formed; it just did not genuinely pass'); + }); +}); diff --git a/tests/prohibition-probe.verify-tier.test.cjs b/tests/prohibition-probe.verify-tier.test.cjs index 664a2db80..e07f8116d 100644 --- a/tests/prohibition-probe.verify-tier.test.cjs +++ b/tests/prohibition-probe.verify-tier.test.cjs @@ -21,6 +21,10 @@ const assert = require('node:assert/strict'); const path = require('node:path'); const PROBE_CORE_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs'); +// The ENFORCEMENT-half producer (#1259, ADR-550 D5d heavy half). Authored as +// src/prohibition-enforcement.cts and compiled by `npm run build:lib` to this gitignored +// artifact — mirroring how PROBE_CORE_LIB requires the BUILT probe-core.cjs above. +const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs'); describe('prohibition-probe verify-tier: test-tier fail-closed safety (PROB-12 / ADR-550 D5d)', () => { test('probe-core exports a deterministic prohibition-disposition helper', () => { @@ -55,3 +59,103 @@ describe('prohibition-probe verify-tier: test-tier fail-closed safety (PROB-12 / 'an unwired test-tier item must be flagged unverified — it can never be silently skipped'); }); }); + +// ─── ENFORCEMENT HALF (#1259, ADR-550 D5d heavy half) ────────────────────────── +// +// RED-first: these assertions require the BUILT gsd-core/bin/lib/prohibition-enforcement.cjs, +// which does not exist until Task 2 authors src/prohibition-enforcement.cts and runs build:lib. +// They prove the previously-unreachable green branch in dispositionForProhibition (probe-core +// 420-427) becomes reachable from a real PRODUCER: a passing wired test-tier check builds +// non-empty enforcementEvidence -> green; a missing or failing check hard-gates (flagged, +// non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). +// +// Typed-field assertions ONLY (status / flagged / tier / evidence / located / kind / mode-block). +// The check-runner is injected via options.runCheck so no real subprocess is spawned. +describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / ADR-550 D5d)', () => { + const testTierProhibition = { + requirement_id: 'R1', + category: 'safety', + status: 'resolved', + verification: 'test', + resolution: null, + reason: null, + statement: 'MUST NOT read source files and text-search them in tests', + }; + + // Test A — node --test negative test that PASSES -> green + non-empty evidence. + test('A: wired node-test check that passes disposes green with non-empty enforcement evidence', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + testTierProhibition, + { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: true }) }, + ); + assert.ok(result && typeof result === 'object', 'result must be a structured object'); + assert.equal(result.status, 'green', 'a passing wired node-test check must dispose green'); + assert.equal(result.flagged, false, 'a green test-tier disposition must not be flagged'); + assert.equal(result.tier, 'test', 'tier must be preserved as test'); + assert.equal(result.located, true, 'the wired check was locatable'); + assert.ok(Array.isArray(result.evidence) && result.evidence.length >= 1, + 'a passing check must build non-empty enforcementEvidence (the array dispositionForProhibition reads)'); + }); + + // Test B — lint/AST-rule (no-source-grep) check that PASSES -> green (D4 dogfood anchor). + test('B: wired lint-rule (no-source-grep) check that passes disposes green', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + testTierProhibition, + { kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true }, + { runCheck: () => ({ passed: true }) }, + ); + assert.equal(result.status, 'green', 'a passing wired lint-rule check must dispose green'); + assert.equal(result.flagged, false, 'a green disposition must not be flagged'); + assert.equal(result.kind, 'lint-rule', 'the located check kind must be the lint-rule kind'); + assert.ok(Array.isArray(result.evidence) && result.evidence.length >= 1, + 'a passing lint-rule check must build non-empty enforcementEvidence'); + }); + + // Test C — MISSING check (no locatable wired check) -> hard-gate (non-green, flagged). + test('C: a test-tier prohibition with NO locatable wired check hard-gates (non-green, flagged)', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + testTierProhibition, + null, + { runCheck: () => ({ passed: true }) }, + ); + assert.notEqual(result.status, 'green', 'a missing wired check must NEVER be green (fail-closed)'); + assert.equal(result.flagged, true, 'a missing wired check must be flagged unverified'); + assert.equal(result.located, false, 'no check was locatable'); + assert.ok(Array.isArray(result.evidence) && result.evidence.length === 0, + 'a missing check builds no enforcement evidence'); + }); + + // Test D — FAILING check -> hard-gate in BOTH modes (interactive + autonomous). + test('D: a wired check that FAILS hard-gates (non-green, flagged) in both modes', () => { + const enforce = require(ENFORCEMENT_LIB); + for (const mode of ['interactive', 'autonomous']) { + const result = enforce.runProhibitionEnforcement( + testTierProhibition, + { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true }, + { runCheck: () => ({ passed: false }), mode }, + ); + assert.notEqual(result.status, 'green', + `a failing wired check must NEVER be green (mode=${mode})`); + assert.equal(result.flagged, true, + `a failing wired check must be flagged in both modes (mode=${mode})`); + assert.equal(result.located, true, 'the check was located even though it failed'); + } + }); + + // D (fail-first not satisfied) — a check that is NOT fail-first is not a valid regression proof. + test('D2: a wired check that is not fail-first hard-gates (non-green, flagged)', () => { + const enforce = require(ENFORCEMENT_LIB); + const result = enforce.runProhibitionEnforcement( + testTierProhibition, + { kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: false }, + { runCheck: () => ({ passed: true }) }, + ); + assert.notEqual(result.status, 'green', + 'a check that is not fail-first is not a valid regression-must-fail-first proof — never green'); + assert.equal(result.flagged, true, 'a non-fail-first check must be flagged'); + }); +}); diff --git a/tests/workflow-size-baseline.json b/tests/workflow-size-baseline.json index 9a6c19ee4..47d96e563 100644 --- a/tests/workflow-size-baseline.json +++ b/tests/workflow-size-baseline.json @@ -85,6 +85,6 @@ "undo.md": 10431, "update.md": 21053, "validate-phase.md": 10745, - "verify-phase.md": 33161, + "verify-phase.md": 35362, "verify-work.md": 31157 }