Merge pull request #1273 from davesienkowski/feat/1259-test-tier-enforcement
enhance(verify-phase): enforce test-tier prohibitions — wire the deferred negative-test gate (#1259, ADR-550 D5d)
This commit is contained in:
10
.changeset/1259-test-tier-enforcement.md
Normal file
10
.changeset/1259-test-tier-enforcement.md
Normal file
@@ -0,0 +1,10 @@
|
||||
---
|
||||
type: Changed
|
||||
pr: 1273
|
||||
---
|
||||
|
||||
**Test-tier prohibitions are now a real, provable gate instead of a permanent, unsatisfiable `gaps_found`** — the deferred ENFORCEMENT half of ADR-550 Decision 5d (the "heavy half" that #644 / PR #1149 deferred) has landed. A new deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`) is the missing PRODUCER: it locates the wired mechanical check, runs it for a genuine **non-vacuous** pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict. The previously-unreachable green branch in `dispositionForProhibition()` is now reachable from the live pipeline — a test-tier prohibition with a genuinely-passing wired check disposes `green` and can reach `passed`, while a missing, non-attested, or non-passing check hard-gates (flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes (ADR-550 D4 / D3). `verify-phase.md` wires the consumer; the green/fail-closed policy in `src/probe-core.cts` is untouched. Both wired-check kinds are accepted (ADR-550 D2): a `node --test` negative test (requiring a real reported test — an empty file, which `node --test` counts as one passing "test", does NOT green) AND a lint/AST rule run as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot), anchored on the in-tree `local/no-source-grep` rule (dogfooding, ADR-550 D4). This enforcement seam is the concrete instance of ADR-857 open-question §147 and lands on the core verify rail, never in `capabilities/` (D6). (#1259)
|
||||
|
||||
**Honest scope — `failFirst` is caller-attested, not yet machine-proven.** This lands the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and the check to genuinely run and pass. It does NOT yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property) — cheap proof of that at verify time needs running the check against a known violation fixture, which is a **tracked follow-up (#1279)**. The red-first property currently rests on caller attestation, surfaced transparently in the evidence record.
|
||||
|
||||
**Correction to the issue body (#1259):** the issue's "96 invalid/error negative-proof cases" figure is wrong. For the `no-source-grep` anchor specifically, the genuine `regression-must-fail-first` proofs are its **two `invalid` cases** (the `.includes()` and `.match()` blocks) in `tests/eslint-rules.test.cjs` — not 96. The anchor argument is unaffected (those two cases ARE real fail-first proofs); only the count was off.
|
||||
1
.gitignore
vendored
1
.gitignore
vendored
@@ -74,6 +74,7 @@ build/
|
||||
/gsd-core/bin/lib/plan-drift-guard.cjs
|
||||
/gsd-core/bin/lib/edge-probe.cjs
|
||||
/gsd-core/bin/lib/probe-core.cjs
|
||||
/gsd-core/bin/lib/prohibition-enforcement.cjs
|
||||
/gsd-core/bin/lib/config-types.cjs
|
||||
/gsd-core/bin/lib/cli-exit.cjs
|
||||
/gsd-core/bin/lib/code-review-flags.cjs
|
||||
|
||||
@@ -3169,7 +3169,7 @@ Each surfaced prohibition is resolved to exactly one of three states:
|
||||
| `dismissed` | Not a genuine prohibition (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
|
||||
| `unresolved` | Deferred | Soft-gates the spec; surfaced as a planner assumption |
|
||||
|
||||
Each resolved prohibition carries a `verification` tier — `test` (a negative test can enforce it) or `judgment` (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag); test-tier prohibitions fail closed when unwired (never silently green). Under `--auto`, the probe **never auto-dismisses**. Canon-bound concerns (OWASP / GDPR / fairness) are referred to `/gsd:secure-phase` rather than minting SPEC prohibitions (ADR-550 D6).
|
||||
Each resolved prohibition carries a `verification` tier — `test` (a negative test can enforce it) or `judgment` (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an `unverified-prohibition — human review recommended` flag); test-tier prohibitions are enforced via the deterministic `check prohibition-enforcement` gate — green when the wired negative test / lint rule passes, hard-gate (flagged, non-green) when missing or failing, in both interactive and autonomous modes (#1259, ADR-550 D5d). Under `--auto`, the probe **never auto-dismisses**. Canon-bound concerns (OWASP / GDPR / fairness) are referred to `/gsd:secure-phase` rather than minting SPEC prohibitions (ADR-550 D6).
|
||||
|
||||
The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, so the section is not merely documentation.
|
||||
|
||||
@@ -3180,5 +3180,6 @@ The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, s
|
||||
- REQ-PROHIB-04: `--auto` MUST never auto-dismiss.
|
||||
- REQ-PROHIB-05: `plan-phase` MUST lift resolved prohibitions into `must_haves.prohibitions` (never `truths`).
|
||||
- REQ-PROHIB-06: A well-formed but unwired `test`-tier prohibition MUST fail closed at verify time — never a silent pass.
|
||||
- REQ-PROHIB-07: A `test`-tier prohibition with a caller-attested, genuinely-passing (non-vacuous) wired mechanical check (a `node --test` negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, non-attested, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes (#1259, ADR-550 D5d — the enforcement half; machine-proven fail-first is tracked in #1279, deterministic descriptor auto-locate in #1278).
|
||||
|
||||
**Reference:** [Prohibition Probe](../gsd-core/references/prohibition-probe.md)
|
||||
|
||||
@@ -345,6 +345,7 @@
|
||||
"profile-output.cjs",
|
||||
"profile-pipeline-command-router.cjs",
|
||||
"profile-pipeline.cjs",
|
||||
"prohibition-enforcement.cjs",
|
||||
"project-root.cjs",
|
||||
"prompt-budget.cjs",
|
||||
"research-provider.cjs",
|
||||
|
||||
@@ -459,6 +459,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
| `research-provider.cjs` | Research provider waterfall, confidence tiers, and planResearch (cache-hits + fetch plan) |
|
||||
| `research-store.cjs` | Content-addressed research cache: sha256 keys, per-source TTL staleness, two-tier (user ~/.gsd / project .planning) store |
|
||||
| `probe-core.cjs` | Generic spec-phase probe resolution model (compiled from `src/probe-core.cts`, gitignored; ADR-550 Decision 7) — the status×verification re-cut (`status: resolved/dismissed/unresolved` × per-probe `verification`), `validateResolution`/`validateRequirement`, `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject, the `byVerification` rollup, and the `runProbeCli` I/O scaffold; the shared seam consumed by `edge-probe` (and the prohibition probe #644); exports `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` (#550) |
|
||||
| `prohibition-enforcement.cjs` | Deterministic test-tier prohibition PRODUCER/gate (compiled from `src/prohibition-enforcement.cts`, gitignored; #1259, ADR-550 D5d "heavy half") — locates the wired mechanical check (`node-test` or `lint-rule`), confirms it is fail-first, runs it via an injectable runner, builds typed `enforcementEvidence`, and emits the `dispositionForProhibition` verdict; a passing wired check disposes green, a missing/failing/non-fail-first check hard-gates (flagged, non-green) in both interactive and autonomous modes; exports `runProhibitionEnforcement`, `routeProhibitionEnforcement`; CLI surface `gsd_run check prohibition-enforcement <request.json>` |
|
||||
| `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence |
|
||||
| `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` |
|
||||
| `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) |
|
||||
|
||||
@@ -85,11 +85,14 @@ The capability system (ADR-857) classifies **predicate-generation as core verifi
|
||||
- **The probe *adapters* are the core-default *generator*** — `edge-probe`'s `classifyShape`/`proposeEdges` and the prohibition probe's adversarial LLM-propose, the surfaces that *propose* predicates — default-on and non-removable, but **independently versionable**. The classifier's measured recall gap (the prose→shape under-fire on terse prose) is the reason they stay their own modules under this ADR rather than being folded into the slow core rail. **`probe-core` is not the generator:** per Decision 7b it ingests already-proposed items, and its deterministic validators are the **contract**'s CI-testable surface (Decision 5) — so `probe-core` sits on the contract side, the adapters on the generator side. ADR-857 phase 6 wires these modules onto the core predicate rail; it does **not** relabel them as a `capabilities/edge-probe/` plug-in.
|
||||
- Decision 5's rule — *the CI-testable surface is the contract, not the classifier* — extends to ADR-857's core rail: the deterministic conformance test is the contract shape the verifier consumes, never the LLM's judgment.
|
||||
|
||||
## Addendum (2026-06-12): test-tier disposition — fail-closed now, heavy enforcement deferred
|
||||
## Addendum (2026-06-12; updated 2026-06-15): test-tier disposition — fail-closed safety half (#644) + enforcement half LANDED (#1259)
|
||||
|
||||
Decision 4 describes the `test`-tier as a "**Hard gate in both interactive and autonomous modes.**" The #644 implementation revises that to a **fail-closed-now / deferred-enforcement** resolution (the "B-with-guard" maintainer decision of 2026-06-12), so the architecture-of-record matches the shipped code:
|
||||
|
||||
- A well-formed but **unwired** `test`-tier prohibition resolves via `dispositionForProhibition()` to `{ status: 'unverified', flagged: true }` — **provably never green** without explicit evidence (REQ-PROHIB-06). This is the load-bearing safety half and it holds today.
|
||||
- The **heavy negative-test enforcement mechanism** — a contrived `test`-tier consumer that mechanically runs the negative test — is **deferred to a follow-up**, because the #644 corpus is entirely `judgment`-tier and wiring a synthetic test-tier consumer now would be gold-plating. The fail-closed disposition guards the gap in the meantime: the gate cannot silently pass, it can only report `unverified`/flagged until enforcement lands.
|
||||
- The **negative-test enforcement mechanism** — locating the wired mechanical check, running it for a genuine **non-vacuous** pass, and building the `enforcementEvidence` that flips a passing test-tier item green — **landed in #1259** as the deterministic `check prohibition-enforcement` sub-command (authored as `src/prohibition-enforcement.cts`, compiled by `build:lib` to the gitignored `gsd-core/bin/lib/prohibition-enforcement.cjs`). It accepts BOTH wired-check kinds — a `node --test` negative test (requiring a real reported test, not the empty file `node --test` would count as one passing "test") OR a lint/AST rule run through the project flat config as `eslint --format json` filtered by `ruleId` (so plugin rules like `local/*` load — bare `--rule` cannot) — and is anchored on the in-tree `local/no-source-grep` rule (dogfooding the existing must-NOT proof, ADR-550 D4; the #644 corpus had zero authored test-tier prohibitions, so no contrived consumer was minted). A passing wired check disposes green; a missing, non-attested, or genuinely-non-passing check hard-gates (flagged, non-green) in both interactive and autonomous modes.
|
||||
- **Honest scope — `failFirst` is caller-ATTESTED, not machine-proven (tracked follow-up).** What #1259 lands is the *execution + non-vacuous-pass* half: the producer requires the caller to attest `failFirst: true` and requires the check to genuinely run and pass. It does **not** yet independently prove the check *fails-on-violation* (the literal `regression-must-fail-first` property), because cheap proof of that at verify time needs running the check against a known **violation fixture** — deferred as a follow-up (#1279; the descriptor auto-locate half is #1278). Until then the red-first property rests on caller attestation, surfaced transparently in the evidence record. This closes the permanent-`gaps_found` dead-end with a genuinely-executed gate without overclaiming machine-proven fail-first.
|
||||
|
||||
Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") is preserved exactly; what is deferred is only the mechanical-pass half of the gate, not the never-green safety property. This addendum supersedes the unqualified "hard gate" wording of Decision 4 for `test`-tier items until the enforcement follow-up lands. The decision also lives in `src/probe-core.cts` comments, `verify-phase.md`, and the #644 changeset.
|
||||
Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved through the fail-closed-now half and is now joined by the genuine-execution half — a test-tier prohibition with a passing, non-vacuous wired check can reach `green`/`passed`, and a missing/failing one hard-gates. The previously-unreachable green branch in `dispositionForProhibition()` is reachable from the live pipeline, and the fail-closed default backs every miss/fail. The one remaining gap to D4's literal intent — *machine-proven* fail-first — is documented above as a tracked follow-up. The decision also lives in `src/probe-core.cts` comments, `src/prohibition-enforcement.cts`, `verify-phase.md`, and the #644 / #1259 changesets.
|
||||
|
||||
This enforcement seam is the concrete instance of **ADR-857 open-question §147** — the deferred "deterministic CI conformance test for the verifier↔predicate contract." Per D6 it lands on the **core verify rail** (non-toggleable substrate), never in `capabilities/`: the verifier consuming a contract-shaped, deterministic predicate is core, not an opt-in capability.
|
||||
|
||||
@@ -41,6 +41,7 @@ export default tseslint.config(
|
||||
'gsd-core/bin/lib/cli-exit.cjs',
|
||||
'gsd-core/bin/lib/edge-probe.cjs',
|
||||
'gsd-core/bin/lib/probe-core.cjs',
|
||||
'gsd-core/bin/lib/prohibition-enforcement.cjs',
|
||||
'gsd-core/bin/lib/code-review-flags.cjs',
|
||||
'gsd-core/bin/lib/context-utilization.cjs',
|
||||
'gsd-core/bin/lib/artifacts.cjs',
|
||||
|
||||
@@ -110,6 +110,20 @@ lifecycle is identical to the edge-probe, the verification tiers differ):
|
||||
human/LLM judgment that the framing is not manipulative). It records intent and routes
|
||||
to a judgment-based review rather than a green/red test.
|
||||
|
||||
At verify time these tiers are routed differently (ADR-550 D4):
|
||||
- A **test**-tier prohibition is enforced + hard-gates via the deterministic
|
||||
`check prohibition-enforcement` sub-command (#1259, ADR-550 D5d): it locates the wired
|
||||
mechanical check (a `node --test` negative test OR a lint/AST rule run as
|
||||
`eslint --format json` and filtered by `ruleId`), requires the caller-attested `failFirst`
|
||||
marker, runs it for a genuine **non-vacuous** pass, and emits the
|
||||
`dispositionForProhibition()` verdict. A passing wired check disposes **green** (satisfiable
|
||||
→ can reach `passed`); a missing, non-attested, or genuinely-non-passing check **hard-gates**
|
||||
(flagged, never green → `gaps_found`) in BOTH interactive and autonomous modes — never a
|
||||
silent pass. (`failFirst` is caller-attested, not yet machine-proven against a violation
|
||||
fixture — a tracked follow-up; see ADR-550 D5d.)
|
||||
- A **judgment**-tier prohibition routes to a never-silent / never-hard-halt soft gate
|
||||
(autonomous emits an `unverified-prohibition — human review recommended` flag).
|
||||
|
||||
Splitting these axes keeps the lifecycle enum free of a verification fact and lets the
|
||||
prohibition adapter declare `test | judgment` without forking the shared lifecycle enum that
|
||||
the edge-probe's `explicit | backstop` also uses.
|
||||
|
||||
@@ -70,7 +70,17 @@ Aggregate all must_haves across plans for phase-level verification.
|
||||
**Prohibitions (`must_haves.prohibitions`, ADR-550 D3 — the must-NOT sibling block):** When a plan carries `must_haves.prohibitions`, extract each `{ statement, status, verification }` item and route it by `verification` tier in verdict assembly (ADR-550 D4, "B-with-guard", 2026-06-12 maintainer decision). These are NEGATIVE checks (the must-NOT must NOT have happened), distinct from positive `truths`:
|
||||
|
||||
- **judgment-tier → mode-dependent soft-gate.** Interactive verify defers each item to the end-of-phase human checkpoint (`human_verify_mode: end-of-phase`). Autonomous verify records a NON-AUTHORITATIVE LLM-judge verdict + a prominent `unverified-prohibition — human review recommended` flag (autonomous completion reads "complete with N flagged prohibitions"). NEVER a silent pass; NEVER a hard halt of an AFK run.
|
||||
- **test-tier → FAIL CLOSED (accept-and-flag).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later), but a well-formed test-tier item reaching verify with NO wired enforcement disposes as UNVERIFIED, flagged like an unresolved judgment item, NEVER green. The deterministic fail-closed default is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). The real fail-first negative-test enforcement MECHANISM defers to a follow-up PR (#644's corpus is entirely judgment-tier; a contrived test-tier fixture here would be the gold-plating failure mode).
|
||||
- **test-tier → ENFORCED via `check prohibition-enforcement` (green on pass, hard-gate on miss/fail).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later). For each test-tier item, the verifier invokes the deterministic producer:
|
||||
|
||||
```bash
|
||||
gsd_run check prohibition-enforcement <request.json>
|
||||
```
|
||||
|
||||
where `<request.json>` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, failFirst: true }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). The producer LOCATES the wired check, requires the caller-attested `failFirst` marker, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259, ADR-550 D5d). (`failFirst` is caller-attested, not yet machine-proven against a violation fixture — a tracked follow-up.) Route the result by its typed fields:
|
||||
- **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**.
|
||||
- **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`).
|
||||
|
||||
> **Descriptor authoring — current scope (#1259).** The `check` descriptor is **supplied by the phase author / verifier**; there is no projection field yet that deterministically derives `{ kind, target, rule, failFirst }` from a prohibition in `must_haves.prohibitions` (which carries only `{ statement, status, verification }`). So #1259 lands the **deterministic run+verdict half** (locate→run→evidence→disposition, all CI-testable) while the **locate→descriptor half is author-provided** for now. Deterministic auto-locate — so a wired passing test closes the gap with zero manual descriptor authoring — is a **tracked follow-up: #1278**. Until then, the green path requires the author to wire the descriptor explicitly.
|
||||
|
||||
**Option B: Use Success Criteria from ROADMAP.md**
|
||||
|
||||
@@ -474,7 +484,7 @@ Classify status using this decision tree IN ORDER (most restrictive first):
|
||||
→ **gaps_found**
|
||||
|
||||
2. IF any `must_haves.prohibitions` item disposes as flagged-unverified (ADR-550 D4):
|
||||
- **test-tier, fail-closed** (no wired enforcement — `dispositionForProhibition()` returns `status: 'unverified'`, `flagged: true`): → **gaps_found** (never green; the unwired test-tier item is an unverified gap).
|
||||
- **test-tier, fail-closed when the wired check is MISSING OR FAILS** (now run via `check prohibition-enforcement` — `located: false`, or `dispositionForProhibition()` returns `status: 'unverified'`, `flagged: true`): → **gaps_found** in both interactive and autonomous modes (never green; a missing/failing mechanical check is an unverified gap). A test-tier item whose wired check PASSES disposes `status: 'green'`, `flagged: false` and is NOT a gap — it can reach **passed**.
|
||||
- **judgment-tier, autonomous run** (non-authoritative LLM-judge verdict): emit the `unverified-prohibition — human review recommended` flag and classify → **human_needed** (autonomous completion reads "complete with N flagged prohibitions"; never a silent pass, never a hard halt).
|
||||
- **judgment-tier, interactive run**: route to the end-of-phase human checkpoint → **human_needed**.
|
||||
|
||||
|
||||
@@ -24,6 +24,7 @@ const { getRoadmapPhaseWithFallback } = roadmapModule;
|
||||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||||
import gapCheckerModule = require('./gap-checker.cjs');
|
||||
const { runGapAnalysis } = gapCheckerModule;
|
||||
import { routeProhibitionEnforcement } from './prohibition-enforcement.cjs';
|
||||
|
||||
// ─── Helpers ──────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -887,7 +888,15 @@ function routeCheckCommand({ args, cwd, raw }: RouteCheckCommandOptions): void {
|
||||
cmdVerifyCodebaseDrift(cwd, raw);
|
||||
return;
|
||||
}
|
||||
error('Unknown check subcommand. Available: auto-mode, decision-coverage-plan, decision-coverage-verify, gap-analysis-plan-post, tdd-review-checkpoint, ui-plan-gate, ui-safety-gate, verify-schema-drift, verify-codebase-drift', ERROR_REASON.SDK_UNKNOWN_COMMAND);
|
||||
if (subcommand === 'prohibition-enforcement') {
|
||||
// The deterministic test-tier prohibition PRODUCER/gate (#1259, ADR-550 D5d). Locates the
|
||||
// wired mechanical check (node-test or lint-rule), confirms fail-first, runs it, builds
|
||||
// enforcementEvidence, and emits the dispositionForProhibition verdict. Invocable as
|
||||
// `gsd_run check prohibition-enforcement <request.json>`.
|
||||
routeProhibitionEnforcement(args, raw);
|
||||
return;
|
||||
}
|
||||
error('Unknown check subcommand. Available: auto-mode, decision-coverage-plan, decision-coverage-verify, gap-analysis-plan-post, prohibition-enforcement, tdd-review-checkpoint, ui-plan-gate, ui-safety-gate, verify-schema-drift, verify-codebase-drift', ERROR_REASON.SDK_UNKNOWN_COMMAND);
|
||||
}
|
||||
|
||||
export = {
|
||||
|
||||
@@ -376,11 +376,10 @@ export interface ProhibitionDispositionContext {
|
||||
* This is the cheap safety guarantee: a well-formed prohibition that reaches verify-phase with NO
|
||||
* wired enforcement evidence can NEVER be a silent pass. It is `{ status: 'unverified', flagged:
|
||||
* true }` — never `green` — exactly like an unresolved judgment item. The HEAVY half (a real
|
||||
* fail-first negative-test enforcement mechanism that, given evidence, would flip a test-tier item
|
||||
* to green) is OUT of #644 scope and defers to a follow-up PR: #644's corpus is entirely
|
||||
* judgment-tier, so wiring a contrived test-tier consumer here would be the delete-bad-tests /
|
||||
* gold-plating failure mode. Until that follow-up lands, ANY prohibition without enforcement
|
||||
* evidence — test- or judgment-tier — disposes as flagged-unverified.
|
||||
* negative-test enforcement mechanism that, given evidence, flips a test-tier item to green) was OUT
|
||||
* of #644 scope and LANDED in #1259 as the `prohibition-enforcement` producer (it builds the
|
||||
* `enforcementEvidence` this helper reads). This helper's policy is unchanged: ANY prohibition
|
||||
* without enforcement evidence — test- or judgment-tier — disposes as flagged-unverified.
|
||||
*
|
||||
* The function is pure: same input always yields the same disposition (no LLM judgment, ADR-550
|
||||
* D5). The LLM-judge soft-gate for judgment-tier items is a verify-phase PROSE concern (the
|
||||
@@ -398,9 +397,9 @@ export function dispositionForProhibition(
|
||||
const hasEnforcement = evidence.length > 0;
|
||||
|
||||
// FAIL CLOSED: no wired enforcement evidence -> flagged unverified, never green. This holds for
|
||||
// every tier today (the real enforcement mechanism that could flip a test-tier item to green is
|
||||
// deferred to a follow-up PR). The guard the safety assertion proves: an unwired item can never
|
||||
// be silently skipped.
|
||||
// every tier (the producer that builds enforcement evidence for a test-tier item — the
|
||||
// `prohibition-enforcement` module — landed in #1259). The guard the safety assertion proves: an
|
||||
// unwired item can never be silently skipped.
|
||||
if (!hasEnforcement) {
|
||||
return {
|
||||
status: 'unverified',
|
||||
@@ -408,15 +407,15 @@ export function dispositionForProhibition(
|
||||
tier,
|
||||
reason:
|
||||
tier === 'test'
|
||||
? 'test-tier prohibition has no wired enforcement evidence — flagged unverified (fail-closed; real negative-test enforcement deferred to a follow-up PR, ADR-550 D5d)'
|
||||
? 'test-tier prohibition has no passing wired enforcement check — flagged unverified (fail-closed; never a silent pass, ADR-550 D5d)'
|
||||
: 'prohibition has no enforcement evidence — flagged unverified (fail-closed; never a silent pass, ADR-550 D5d)',
|
||||
};
|
||||
}
|
||||
|
||||
// D4 GUARD: a judgment-tier (or unknown-tier) prohibition is NEVER a silent green from this
|
||||
// deterministic helper — it always routes to human/LLM judgment review (ADR-550 D4; verify-phase.md).
|
||||
// Only a test-tier item with wired enforcement evidence may go green, and even that is the deferred
|
||||
// heavy half until the real negative-test enforcement mechanism lands (no #644 caller passes evidence).
|
||||
// Only a test-tier item with wired enforcement evidence may go green; the producer that supplies
|
||||
// that evidence (`prohibition-enforcement`, #1259) runs the wired check and requires a genuine pass.
|
||||
if (tier === 'test') {
|
||||
return {
|
||||
status: 'green',
|
||||
|
||||
496
src/prohibition-enforcement.cts
Normal file
496
src/prohibition-enforcement.cts
Normal file
@@ -0,0 +1,496 @@
|
||||
/**
|
||||
* prohibition-enforcement — the deterministic PRODUCER for test-tier prohibition verification
|
||||
* (#1259, ADR-550 Decision 5d "heavy half"; the D1 seam — a NEW deterministic gsd-tools
|
||||
* sub-command, NOT free-form workflow prose).
|
||||
*
|
||||
* Today `dispositionForProhibition()` (src/probe-core.cts) already carries the POLICY seam: with
|
||||
* non-empty `enforcementEvidence` AND `tier === 'test'` it returns `{ status: 'green' }` (the
|
||||
* branch at probe-core 420-427); with empty evidence it fails closed to flagged-unverified. But
|
||||
* NOTHING in the live pipeline ever produced `enforcementEvidence`, so the green branch was
|
||||
* unreachable. This module is the missing producer: it LOCATES the wired mechanical check from a
|
||||
* check descriptor, RUNS it and requires a genuine NON-VACUOUS pass, builds a typed
|
||||
* `enforcementEvidence` array on PASS, and emits the `dispositionForProhibition` verdict as JSON.
|
||||
* The green/fail-closed policy itself is untouched (no src/probe-core.cts edit).
|
||||
*
|
||||
* Accepts BOTH wired-check kinds (ADR-550 D2): a `node --test` negative test OR a lint/AST rule
|
||||
* (e.g. the in-tree `local/no-source-grep` rule — the D4 dogfood anchor, run via the project flat
|
||||
* config so the plugin loads). A missing, non-attested, or genuinely-non-passing check hard-gates
|
||||
* (flagged, non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3) — never a silent
|
||||
* green.
|
||||
*
|
||||
* FAIL-FIRST IS CALLER-ATTESTED (honest scope, #1259): the producer requires the caller to ATTEST
|
||||
* `failFirst: true` AND requires the runner to observe a real non-vacuous pass — but it does NOT
|
||||
* independently prove the check fails-on-violation. Genuine fail-first confirmation needs running
|
||||
* the check against a known violation fixture and is a tracked follow-up (recorded in ADR-550's D5d
|
||||
* note). What ships here closes the permanent-`gaps_found` dead-end with a genuinely-executed gate;
|
||||
* it does not yet replace caller attestation with machine proof of the red-first property.
|
||||
*
|
||||
* Authored as strict TypeScript (`src/prohibition-enforcement.cts`) and compiled by
|
||||
* `tsc -p tsconfig.build.json` (`npm run build:lib`) to the gitignored runtime artifact
|
||||
* `gsd-core/bin/lib/prohibition-enforcement.cjs`. Do NOT hand-write the `.cjs`; it is emitted.
|
||||
*
|
||||
* DETERMINISM SCOPE: the DECISION layer is pure/deterministic and no-LLM — given a `runCheck` result
|
||||
* the disposition is same-input-same-output and mutation-survivable, and the parse/filter helpers
|
||||
* (`parseNodeTestSummary`, `tapTestNames`, `eslintJsonHasRule`, `eslintHasFatalError`, …) are pure.
|
||||
* The DEFAULT REAL runner is NOT pure — it spawns `node --test` / eslint, so its result depends on the
|
||||
* environment (eslint version + flat config, node version, the target file). That is why the runner is
|
||||
* an injectable seam (`runCheck`): the contract is unit-tested against injected results, mirroring the
|
||||
* injectable I/O pattern in `runProbeCli` / `ProbeCliOptions`.
|
||||
*/
|
||||
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import { execFileSync } from 'node:child_process';
|
||||
// Import the leaf I/O module directly, not the core.cjs re-export spine (being retired, #1268).
|
||||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||||
import io = require('./io.cjs');
|
||||
const { output, error, ERROR_REASON } = io;
|
||||
import { dispositionForProhibition } from './probe-core.cjs';
|
||||
import type { ProhibitionDisposition } from './probe-core.cjs';
|
||||
|
||||
/** The two accepted wired-check kinds (ADR-550 D2). */
|
||||
export type CheckKind = 'node-test' | 'lint-rule';
|
||||
|
||||
/**
|
||||
* A descriptor of the wired mechanical check that asserts the must-NOT. `kind` selects the
|
||||
* runner family; `target` is the negative-test file path (node-test) or the PATH to lint
|
||||
* (lint-rule); `rule` is the eslint rule id (lint-rule only — e.g. `local/no-source-grep`) and is
|
||||
* REQUIRED for the lint-rule kind (a lint-rule descriptor without it is not a valid wired check);
|
||||
* `failFirst` is the caller's ATTESTATION that the check is a genuine `regression-must-fail-first`
|
||||
* proof (caller-declared — see the module docstring; not independently confirmed at verify time).
|
||||
*/
|
||||
export interface CheckDescriptor {
|
||||
kind: CheckKind;
|
||||
target: string;
|
||||
rule?: string;
|
||||
failFirst?: boolean;
|
||||
}
|
||||
|
||||
/**
|
||||
* The result a check-runner returns: whether the check genuinely, non-vacuously PASSED. The runner
|
||||
* reports only what it can OBSERVE (a real pass) — it does NOT determine fail-first. Whether the
|
||||
* check is a `regression-must-fail-first` proof is a CALLER-ATTESTED property of the descriptor
|
||||
* (`CheckDescriptor.failFirst`); the producer cannot independently confirm it at verify time without
|
||||
* a violation fixture (a tracked follow-up — see the module docstring).
|
||||
*/
|
||||
export interface CheckRunResult {
|
||||
passed: boolean;
|
||||
}
|
||||
|
||||
/** A single typed enforcement-evidence record (the array `dispositionForProhibition` reads). */
|
||||
export interface EnforcementEvidence {
|
||||
kind: CheckKind;
|
||||
target: string;
|
||||
rule?: string;
|
||||
failFirst: boolean;
|
||||
passed: boolean;
|
||||
}
|
||||
|
||||
/** Injectable options for `runProhibitionEnforcement` (defaults wire to the real runner). */
|
||||
export interface EnforcementOptions {
|
||||
/** Runs the located check; injected in tests so no real subprocess is spawned. */
|
||||
runCheck?: (check: CheckDescriptor) => CheckRunResult;
|
||||
/** Verify mode — recorded for transparency; the hard-gate applies in BOTH modes (ADR-550 D4). */
|
||||
mode?: string;
|
||||
/** Project root for the default real runner (defaults to process.cwd()). */
|
||||
cwd?: string;
|
||||
/** Override the per-kind subprocess timeout (ms); defaults to 30s (node-test) / 60s (eslint).
|
||||
* Injected in tests to prove the fail-closed-on-timeout bound without a 30s wait. */
|
||||
timeoutMs?: number;
|
||||
}
|
||||
|
||||
/** The producer's verdict: the disposition PLUS the located/kind/evidence provenance. */
|
||||
export interface EnforcementResult extends ProhibitionDisposition {
|
||||
located: boolean;
|
||||
kind: CheckKind | null;
|
||||
evidence: EnforcementEvidence[];
|
||||
mode?: string;
|
||||
}
|
||||
|
||||
/** node --test argv. Forces the TAP reporter so the summary counts are parseable + version-stable;
|
||||
* `--` before the target so a target starting with `-` is not parsed as a flag (option-injection). */
|
||||
export function buildNodeTestArgs(check: CheckDescriptor): string[] {
|
||||
return ['--test', '--test-reporter=tap', '--', check.target];
|
||||
}
|
||||
|
||||
/** eslint argv (the args AFTER the eslint CLI path). Runs the project flat config so plugin rules
|
||||
* (e.g. `local/*`) load — `--rule` CANNOT load a plugin, so we lint the TARGET path as JSON and
|
||||
* filter by rule id. `--no-warn-ignored` makes an eslint-IGNORED target return `[]` (not a length-1
|
||||
* "File ignored" result) so an ignored path fails closed via the vacuity guard, not a false green.
|
||||
* `--` before the target so a target starting with `-` is not parsed as a flag (option-injection). */
|
||||
export function buildLintArgs(check: CheckDescriptor): string[] {
|
||||
return ['--no-warn-ignored', '--format', 'json', '--', check.target];
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the project's eslint CLI entry portably (no `npx` — not spawnable via `execFileSync` on
|
||||
* Windows). Resolves eslint's package.json from the target project's `node_modules` and derives
|
||||
* `bin/eslint.js`, so it is run as `node <cli>` (portable). Returns null if eslint is not installed
|
||||
* (→ the lint-rule check fails closed, never throws).
|
||||
*/
|
||||
function resolveEslintCli(cwd: string): string | null {
|
||||
try {
|
||||
const pkg = require.resolve('eslint/package.json', { paths: [cwd] });
|
||||
const cli = path.join(path.dirname(pkg), 'bin', 'eslint.js');
|
||||
return fs.existsSync(cli) ? cli : null;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/** Basename of a path, separator-agnostic (handles `\` and `/` so node-test names compare stably
|
||||
* across OSes / node versions that report the file-test by differing path forms). */
|
||||
function baseOf(p: string): string {
|
||||
return typeof p === 'string' ? (p.split(/[\\/]/).pop() ?? p) : '';
|
||||
}
|
||||
|
||||
/**
|
||||
* Pure parser for the `node --test` TAP summary. A genuine pass is NON-VACUOUS: exit 0 is NOT enough
|
||||
* (an empty / all-skipped / deleted-negative-test file exits 0 with `# tests 0`). Mutation-pinned by
|
||||
* unit tests so a threshold flip is caught.
|
||||
*/
|
||||
export function parseNodeTestSummary(out: string): { tests: number; pass: number; fail: number; cancelled: number } {
|
||||
const num = (re: RegExp): number => {
|
||||
const m = typeof out === 'string' ? out.match(re) : null;
|
||||
return m ? Number(m[1]) : 0;
|
||||
};
|
||||
return {
|
||||
tests: num(/^# tests (\d+)/m),
|
||||
pass: num(/^# pass (\d+)/m),
|
||||
fail: num(/^# fail (\d+)/m),
|
||||
cancelled: num(/^# cancelled (\d+)/m),
|
||||
};
|
||||
}
|
||||
|
||||
/** The names of REAL (run) tests from TAP `ok N - <name>` / `not ok N - <name>` lines. A line with a
|
||||
* `# SKIP` / `# TODO` directive is EXCLUDED — a skipped/todo negative test never executed, so it must
|
||||
* not count toward non-vacuity (#1259 m1). */
|
||||
export function tapTestNames(out: string): string[] {
|
||||
if (typeof out !== 'string') return [];
|
||||
const names: string[] = [];
|
||||
const re = /^(?:not )?ok \d+ - (.+)$/gm;
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(out)) !== null) {
|
||||
const rest = m[1];
|
||||
if (/\s#\s*(?:SKIP|TODO)\b/i.test(rest)) continue; // skipped/todo did not run
|
||||
names.push(rest.replace(/\s+#\s.*$/, '').trim());
|
||||
}
|
||||
return names;
|
||||
}
|
||||
|
||||
/**
|
||||
* A non-vacuous node-test pass: at least one test, at least one pass, zero failures — AND at least
|
||||
* one reported test whose name is NOT merely the target file. `node --test <file>` counts a file
|
||||
* with ZERO `test()` calls as one passing "test" named after the file, so the counts alone cannot
|
||||
* tell an empty/deleted negative test from a real one (the #1259 BL-01 false-green). Requiring a
|
||||
* named test distinct from the file closes that hole.
|
||||
*
|
||||
* KNOWN CONSTRAINT (fail-closed, not a hole): a real test whose `test('...')` name is EXACTLY the
|
||||
* target file's basename emits TAP indistinguishable from an empty file and is conservatively
|
||||
* rejected (non-green). A wired negative test must carry a descriptive name, not be named after its
|
||||
* own file — a benign authoring constraint, and the safe direction if violated.
|
||||
*/
|
||||
export function isNonVacuousNodeTestPass(out: string, target: string): boolean {
|
||||
const s = parseNodeTestSummary(out);
|
||||
// >=1 test, >=1 pass, ZERO failures AND ZERO cancelled (a cancelled run is not a clean pass, m1).
|
||||
if (!(s.tests >= 1 && s.pass >= 1 && s.fail === 0 && s.cancelled === 0)) return false;
|
||||
// Compare BASENAMES: node reports the file-test by varying path forms across OS / node version
|
||||
// (absolute, relative, normalized), so an exact-string compare misfires. A real test name (e.g.
|
||||
// "guards the must-NOT") has no separators, so its basename never equals the target file's.
|
||||
const tgtBase = baseOf(target);
|
||||
return tapTestNames(out).some((n) => baseOf(n) !== tgtBase);
|
||||
}
|
||||
|
||||
/** Number of file results in an eslint `--format json` report (0 if unparseable / not an array). */
|
||||
export function eslintFileResultCount(jsonText: string): number {
|
||||
try {
|
||||
const parsed: unknown = JSON.parse(jsonText);
|
||||
return Array.isArray(parsed) ? parsed.length : 0;
|
||||
} catch {
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
|
||||
/** Messages array of a single eslint file-result (empty if absent / wrong shape). */
|
||||
function eslintMessages(file: unknown, key: 'messages' | 'suppressedMessages'): Array<{ ruleId?: unknown; fatal?: unknown }> {
|
||||
return file && typeof file === 'object' && Array.isArray((file as Record<string, unknown>)[key])
|
||||
? ((file as Record<string, unknown>)[key] as Array<{ ruleId?: unknown; fatal?: unknown }>)
|
||||
: [];
|
||||
}
|
||||
|
||||
/**
|
||||
* True if the eslint `--format json` report has a FATAL / parse error — meaning the rule never got
|
||||
* to run on the target. A prohibition gate must fail closed on "the rule didn't execute" (#1259 B1),
|
||||
* NOT treat a length-1 fatal result as "clean". Unparseable report -> true (fail-closed).
|
||||
*/
|
||||
export function eslintHasFatalError(jsonText: string): boolean {
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = JSON.parse(jsonText);
|
||||
} catch {
|
||||
return true;
|
||||
}
|
||||
if (!Array.isArray(parsed)) return true;
|
||||
for (const file of parsed) {
|
||||
const fec = file && typeof file === 'object' ? (file as { fatalErrorCount?: unknown }).fatalErrorCount : undefined;
|
||||
if (typeof fec === 'number' && fec > 0) return true;
|
||||
if (eslintMessages(file, 'messages').some((m) => m && m.fatal === true)) return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** True if the eslint `--format json` report has ANY message for `rule` — in EITHER `messages` or
|
||||
* `suppressedMessages` (an inline `// eslint-disable` of the rule is still a violation, #1259 B1).
|
||||
* Unparseable -> true (fail-closed: treat an unreadable report as a violation, not a silent pass). */
|
||||
export function eslintJsonHasRule(jsonText: string, rule: string): boolean {
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = JSON.parse(jsonText);
|
||||
} catch {
|
||||
return true;
|
||||
}
|
||||
if (!Array.isArray(parsed)) return true;
|
||||
for (const file of parsed) {
|
||||
for (const key of ['messages', 'suppressedMessages'] as const) {
|
||||
for (const msg of eslintMessages(file, key)) {
|
||||
if (msg && typeof msg === 'object' && msg.ruleId === rule) return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* The default REAL check runner (used when no `runCheck` is injected). Reports only an OBSERVED,
|
||||
* genuinely-non-vacuous pass; guarded so a missing tool / non-zero exit yields a non-passing result,
|
||||
* NEVER an uncaught throw (the no-throw contract). It does NOT determine fail-first (caller-attested).
|
||||
* - node-test: runs `node --test` (TAP) and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail
|
||||
* AND a reported test named distinctly from the file). A bare exit 0 for an empty/zero-test file
|
||||
* — which `node --test` counts as one passing "test" named after the file — is NOT a pass (the
|
||||
* #1259 BL-01 false-green fix).
|
||||
* - lint-rule: runs the project eslint as `node <eslint-cli> --format json <target>` (flat config
|
||||
* loads `local/*` plugins) and requires the target to actually lint (>=1 file result) AND ZERO
|
||||
* messages for the specific rule id. `--rule` cannot load a plugin rule, so we filter the
|
||||
* structured report by `ruleId` instead (the #1259 SF-01 fix).
|
||||
*
|
||||
* Both kinds spawn via `process.execPath` (never bare `node`/`npx` — not portably spawnable via
|
||||
* `execFileSync` on Windows) with arg arrays (no shell → no injection from a caller-supplied target).
|
||||
*/
|
||||
/**
|
||||
* Env for spawned checks: strip `NODE_TEST_CONTEXT` and `NODE_OPTIONS` so an AMBIENT test-runner
|
||||
* context (e.g. running verify under `node --test`, which sets `NODE_TEST_CONTEXT=child-v8`) cannot
|
||||
* turn the child `node --test` into a silent v8-reporter worker that emits no parseable TAP — which
|
||||
* would otherwise corrupt the verdict. Deterministic, environment-independent execution.
|
||||
*/
|
||||
function childEnv(): NodeJS.ProcessEnv {
|
||||
const env = { ...process.env };
|
||||
delete env.NODE_TEST_CONTEXT;
|
||||
delete env.NODE_OPTIONS;
|
||||
return env;
|
||||
}
|
||||
|
||||
// Bounded subprocess limits (DEFECT.UNBOUNDED-SUBPROCESS): a stuck wired test / eslint must not hang
|
||||
// verify forever. On timeout `execFileSync` throws -> caught -> fail-closed (degraded, non-passing).
|
||||
// `maxBuffer` caps output so a runaway producer throws (safe direction) rather than OOMs the verifier.
|
||||
const NODE_TEST_TIMEOUT_MS = 30_000;
|
||||
const ESLINT_TIMEOUT_MS = 60_000;
|
||||
const CHECK_MAX_BUFFER = 16 * 1024 * 1024;
|
||||
|
||||
/** Resolve the effective timeout: only a POSITIVE override is honored — `0` (which Node treats as
|
||||
* "no timeout") or a negative value falls back to the bounded default, so the subprocess is ALWAYS
|
||||
* bounded (a `timeoutMs: 0` injection can never disable the bound). */
|
||||
function posTimeout(timeoutMs: number | undefined, def: number): number {
|
||||
return typeof timeoutMs === 'number' && timeoutMs > 0 ? timeoutMs : def;
|
||||
}
|
||||
|
||||
function defaultRunCheck(check: CheckDescriptor, cwd: string, timeoutMs?: number): CheckRunResult {
|
||||
try {
|
||||
if (check.kind === 'node-test') {
|
||||
let out = '';
|
||||
try {
|
||||
out = execFileSync(process.execPath, buildNodeTestArgs(check), {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
windowsHide: true,
|
||||
env: childEnv(),
|
||||
timeout: posTimeout(timeoutMs, NODE_TEST_TIMEOUT_MS),
|
||||
maxBuffer: CHECK_MAX_BUFFER,
|
||||
});
|
||||
} catch (e) {
|
||||
// A failing/timed-out run exits non-zero or is killed (partial TAP on stdout, no `# pass`
|
||||
// summary). Parse what we have: a real failure or timeout -> not a non-vacuous pass -> false.
|
||||
const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : '';
|
||||
out = typeof stdout === 'string' ? stdout : '';
|
||||
}
|
||||
return { passed: isNonVacuousNodeTestPass(out, check.target) };
|
||||
}
|
||||
if (check.kind === 'lint-rule') {
|
||||
const eslintCli = resolveEslintCli(cwd);
|
||||
if (!eslintCli) return { passed: false }; // eslint not installed -> fail closed, never throw
|
||||
let json = '';
|
||||
try {
|
||||
json = execFileSync(process.execPath, [eslintCli, ...buildLintArgs(check)], {
|
||||
cwd,
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
windowsHide: true,
|
||||
env: childEnv(),
|
||||
timeout: posTimeout(timeoutMs, ESLINT_TIMEOUT_MS),
|
||||
maxBuffer: CHECK_MAX_BUFFER,
|
||||
});
|
||||
} catch (e) {
|
||||
// eslint exits non-zero when ANY error is present; the JSON report is still on stdout.
|
||||
// A timeout/kill leaves no parseable JSON -> eslintHasFatalError(unparseable) -> fail-closed.
|
||||
const stdout = e && typeof e === 'object' && 'stdout' in e ? (e as { stdout?: unknown }).stdout : '';
|
||||
json = typeof stdout === 'string' ? stdout : '';
|
||||
}
|
||||
// PASS requires: the target actually linted (>=1 file result), NO fatal/parse error (the rule
|
||||
// must have RUN — #1259 B1), and ZERO messages for the rule (in messages OR suppressedMessages).
|
||||
const lintedSomething = eslintFileResultCount(json) >= 1;
|
||||
return {
|
||||
passed: lintedSomething && !eslintHasFatalError(json) && !eslintJsonHasRule(json, check.rule as string),
|
||||
};
|
||||
}
|
||||
// Unknown kind — defensive; the LOCATE guard already rejects it.
|
||||
return { passed: false };
|
||||
} catch {
|
||||
return { passed: false };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* LOCATE -> CONFIRM fail-first -> RUN -> build enforcementEvidence -> dispositionForProhibition.
|
||||
*
|
||||
* (1) LOCATE: if no well-formed check descriptor is locatable -> fail-closed
|
||||
* (`dispositionForProhibition` with empty evidence) plus `{ located: false, kind: null, evidence: [] }`.
|
||||
* (2) ATTEST + RUN: require the caller to attest `failFirst: true` and run it via `runCheck`. A check
|
||||
* the caller does not attest as fail-first, or that does not genuinely (non-vacuously) PASS ->
|
||||
* fail-closed disposition with `located: true` (a real located miss, non-green, flagged) in BOTH
|
||||
* modes. Fail-first is caller-attested, not independently proven (see module docstring).
|
||||
* (3) PASS: build a typed `enforcementEvidence` array and call `dispositionForProhibition` — the
|
||||
* non-empty array flips a test-tier item to green (the previously-unreachable branch).
|
||||
*
|
||||
* Pure/deterministic: same (prohibition, check, runCheck) -> same result.
|
||||
*/
|
||||
export function runProhibitionEnforcement(
|
||||
prohibition: unknown,
|
||||
check: CheckDescriptor | null | undefined,
|
||||
options: EnforcementOptions = {},
|
||||
): EnforcementResult {
|
||||
const mode = options.mode;
|
||||
|
||||
// (1) LOCATE — no locatable, well-formed wired check -> fail-closed, located: false. The kind must
|
||||
// be one of the two known kinds; the target must be a non-empty string; a lint-rule descriptor MUST
|
||||
// also carry a non-empty `rule` id (its target is the lint PATH, not the rule). An under-specified
|
||||
// descriptor is not a valid wired check, so it is not locatable (does not rely on the runner failing).
|
||||
const c = check && typeof check === 'object' ? check : null;
|
||||
const validKind = !!c && (c.kind === 'node-test' || c.kind === 'lint-rule');
|
||||
const validTarget = !!c && typeof c.target === 'string' && c.target.trim().length > 0;
|
||||
const validRule = !!c && (c.kind !== 'lint-rule' || (typeof c.rule === 'string' && c.rule.trim().length > 0));
|
||||
if (!c || !validKind || !validTarget || !validRule) {
|
||||
const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] });
|
||||
return { ...disposition, located: false, kind: null, evidence: [], ...(mode ? { mode } : {}) };
|
||||
}
|
||||
|
||||
const runCheck = options.runCheck ?? ((toRun: CheckDescriptor) => defaultRunCheck(toRun, options.cwd ?? process.cwd(), options.timeoutMs));
|
||||
|
||||
// (2) ATTEST fail-first (CALLER-DECLARED) + RUN. The caller must attest `failFirst: true` AND the
|
||||
// runner must observe a genuine NON-VACUOUS pass. The producer does NOT independently prove
|
||||
// fail-first (that needs a violation fixture — tracked follow-up, ADR-550 D5d). A non-attested or
|
||||
// non-passing check hard-gates (never green) in BOTH modes.
|
||||
const attestedFailFirst = c.failFirst === true;
|
||||
// No-throw contract end-to-end: even a (test-injected) runCheck that throws must fail closed,
|
||||
// never propagate. The default real runner already never throws.
|
||||
let run: CheckRunResult;
|
||||
try {
|
||||
run = runCheck(c);
|
||||
} catch {
|
||||
run = { passed: false };
|
||||
}
|
||||
const passed = attestedFailFirst && run.passed === true;
|
||||
|
||||
if (!passed) {
|
||||
// NOT attested fail-first OR did not genuinely pass -> fail-closed, located: true (an actual
|
||||
// located miss/fail). Hard-gate applies in BOTH modes; the disposition stays non-green / flagged.
|
||||
const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: [] });
|
||||
return {
|
||||
...disposition,
|
||||
located: true,
|
||||
kind: c.kind,
|
||||
evidence: [],
|
||||
...(mode ? { mode } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
// (3) PASS -> build typed enforcementEvidence and let the policy flip a test-tier item green.
|
||||
// `failFirst` here is the caller's attestation (recorded for provenance), not a machine proof.
|
||||
const evidence: EnforcementEvidence[] = [{
|
||||
kind: c.kind,
|
||||
target: c.target,
|
||||
...(typeof c.rule === 'string' ? { rule: c.rule } : {}),
|
||||
failFirst: true,
|
||||
passed: true,
|
||||
}];
|
||||
const disposition = dispositionForProhibition(prohibition, { enforcementEvidence: evidence });
|
||||
return {
|
||||
...disposition,
|
||||
located: true,
|
||||
kind: c.kind,
|
||||
evidence,
|
||||
...(mode ? { mode } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a `{ prohibition, check, mode }` request from a JSON file path or inline `--json` string.
|
||||
* Returns null on any parse failure (the caller surfaces a structured error, never a throw).
|
||||
*/
|
||||
function parseRequest(args: string[]): { prohibition: unknown; check: CheckDescriptor | null; mode?: string } | null {
|
||||
// args[0] = 'check', args[1] = 'prohibition-enforcement', args[2] = <json-file-path | --json>
|
||||
const jsonFlagIdx = args.indexOf('--json');
|
||||
let payload = '';
|
||||
if (jsonFlagIdx !== -1 && typeof args[jsonFlagIdx + 1] === 'string') {
|
||||
payload = args[jsonFlagIdx + 1];
|
||||
} else if (typeof args[2] === 'string' && args[2]) {
|
||||
try {
|
||||
payload = fs.readFileSync(args[2], 'utf-8');
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
} else {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
const parsed = JSON.parse(payload) as Record<string, unknown>;
|
||||
const checkRaw = parsed['check'];
|
||||
const check: CheckDescriptor | null = (checkRaw && typeof checkRaw === 'object')
|
||||
? (checkRaw as CheckDescriptor)
|
||||
: null;
|
||||
const modeRaw = parsed['mode'];
|
||||
const mode = typeof modeRaw === 'string' ? modeRaw : undefined;
|
||||
return { prohibition: parsed['prohibition'] ?? null, check, ...(mode ? { mode } : {}) };
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* CLI surface: `gsd_run check prohibition-enforcement <request.json>` (or `--json '<inline>'`).
|
||||
* Parses the request, runs the producer, and emits the result as JSON. Honors the no-throw
|
||||
* contract: malformed input -> structured `error(...)`, never an uncaught throw.
|
||||
*/
|
||||
export function routeProhibitionEnforcement(args: string[], raw: boolean): void {
|
||||
const req = parseRequest(args);
|
||||
if (!req) {
|
||||
error(
|
||||
'prohibition-enforcement requires a JSON request: check prohibition-enforcement <request.json> | --json \'{"prohibition":{...},"check":{...}}\'',
|
||||
ERROR_REASON.SDK_MISSING_ARG,
|
||||
);
|
||||
return;
|
||||
}
|
||||
const result = runProhibitionEnforcement(req.prohibition, req.check, req.mode ? { mode: req.mode } : {});
|
||||
output(result, raw, undefined);
|
||||
}
|
||||
|
||||
export {};
|
||||
60
tests/prohibition-enforcement.property.test.cjs
Normal file
60
tests/prohibition-enforcement.property.test.cjs
Normal file
@@ -0,0 +1,60 @@
|
||||
// Property-based tests for the prohibition-enforcement parsing/transform helpers (#1259, ADR-550 D5d).
|
||||
// RULESET.TESTS.property-based-testing: the producer is a parsing module (parseNodeTestSummary,
|
||||
// tapTestNames, eslintJsonHasRule, eslintHasFatalError, eslintFileResultCount), so it carries
|
||||
// fast-check invariants — especially the fail-closed safety invariants of the verify-time gate.
|
||||
'use strict';
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const path = require('node:path');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
|
||||
const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs');
|
||||
|
||||
describe('prohibition-enforcement properties (#1259)', () => {
|
||||
test('parseNodeTestSummary never throws and returns non-negative integer counts', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
fc.assert(fc.property(fc.string(), (s) => {
|
||||
const r = enforce.parseNodeTestSummary(s);
|
||||
for (const k of ['tests', 'pass', 'fail', 'cancelled']) {
|
||||
assert.ok(Number.isInteger(r[k]) && r[k] >= 0, `${k} is a non-negative integer`);
|
||||
}
|
||||
}));
|
||||
});
|
||||
|
||||
test('isNonVacuousNodeTestPass FAIL-CLOSED: a run with any failure or cancellation never greens', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
fc.assert(fc.property(
|
||||
fc.nat({ max: 50 }), fc.integer({ min: 1, max: 50 }), fc.nat({ max: 50 }), fc.string(),
|
||||
(tests, failOrCancel, pass, name) => {
|
||||
// A summary with fail>=1 (and, separately, cancelled>=1) must NEVER be a non-vacuous pass,
|
||||
// regardless of the reported test name.
|
||||
const failing = `ok 1 - ${name}\n# tests ${tests + 1}\n# pass ${pass}\n# fail ${failOrCancel}\n# cancelled 0\n`;
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false);
|
||||
const cancelled = `ok 1 - ${name}\n# tests ${tests + 1}\n# pass ${pass}\n# fail 0\n# cancelled ${failOrCancel}\n`;
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(cancelled, 'neg.test.cjs'), false);
|
||||
},
|
||||
));
|
||||
});
|
||||
|
||||
test('eslintJsonHasRule / eslintHasFatalError FAIL-CLOSED on any non-JSON / non-array input', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
fc.assert(fc.property(fc.string(), (s) => {
|
||||
// Only exercise strings that are NOT a valid JSON array (the unreadable-report branch).
|
||||
let isArray = false;
|
||||
try { isArray = Array.isArray(JSON.parse(s)); } catch { isArray = false; }
|
||||
fc.pre(!isArray);
|
||||
assert.equal(enforce.eslintJsonHasRule(s, 'local/no-source-grep'), true, 'unreadable report -> violation (fail-closed)');
|
||||
assert.equal(enforce.eslintHasFatalError(s), true, 'unreadable report -> fatal (fail-closed)');
|
||||
}));
|
||||
});
|
||||
|
||||
test('eslintFileResultCount never throws and is non-negative', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
fc.assert(fc.property(fc.string(), (s) => {
|
||||
const n = enforce.eslintFileResultCount(s);
|
||||
assert.ok(Number.isInteger(n) && n >= 0);
|
||||
}));
|
||||
});
|
||||
});
|
||||
386
tests/prohibition-enforcement.test.cjs
Normal file
386
tests/prohibition-enforcement.test.cjs
Normal file
@@ -0,0 +1,386 @@
|
||||
// Behavioral tests for the deterministic prohibition-enforcement producer (#1259, ADR-550 D5d
|
||||
// "heavy half"). Requires the BUILT gsd-core/bin/lib/prohibition-enforcement.cjs — authored as
|
||||
// src/prohibition-enforcement.cts and compiled by `npm run build:lib` (mirrors how the verify-tier
|
||||
// suite requires the built probe-core.cjs). Typed-field assertions only; the check-runner is
|
||||
// injected so no real subprocess is spawned. No source-grep.
|
||||
'use strict';
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const path = require('node:path');
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
|
||||
const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs');
|
||||
|
||||
const TEST_TIER = Object.freeze({
|
||||
requirement_id: 'R1',
|
||||
category: 'safety',
|
||||
status: 'resolved',
|
||||
verification: 'test',
|
||||
resolution: null,
|
||||
reason: null,
|
||||
statement: 'MUST NOT read source files and text-search them in tests',
|
||||
});
|
||||
|
||||
describe('prohibition-enforcement: deterministic test-tier producer (#1259 / ADR-550 D5d)', () => {
|
||||
test('exports the producer + route functions', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(typeof enforce.runProhibitionEnforcement, 'function',
|
||||
'must export runProhibitionEnforcement (the deterministic producer)');
|
||||
assert.equal(typeof enforce.routeProhibitionEnforcement, 'function',
|
||||
'must export routeProhibitionEnforcement (the CLI surface)');
|
||||
});
|
||||
|
||||
test('locate-miss (no check descriptor) -> fail-closed, located:false, no evidence', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(TEST_TIER, null, {
|
||||
runCheck: () => ({ passed: true }),
|
||||
});
|
||||
assert.equal(result.located, false, 'no locatable check');
|
||||
assert.notEqual(result.status, 'green', 'locate-miss must never be green');
|
||||
assert.equal(result.flagged, true, 'locate-miss must be flagged');
|
||||
assert.equal(result.kind, null, 'no kind when nothing located');
|
||||
assert.ok(Array.isArray(result.evidence) && result.evidence.length === 0, 'no evidence on locate-miss');
|
||||
});
|
||||
|
||||
test('malformed check descriptor (missing target) -> treated as locate-miss', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(TEST_TIER, { kind: 'node-test' }, {
|
||||
runCheck: () => ({ passed: true }),
|
||||
});
|
||||
assert.equal(result.located, false, 'a descriptor without a target is not locatable');
|
||||
assert.notEqual(result.status, 'green');
|
||||
assert.equal(result.flagged, true);
|
||||
});
|
||||
|
||||
test('node-test check that passes -> green + non-empty typed evidence', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green');
|
||||
assert.equal(result.flagged, false);
|
||||
assert.equal(result.tier, 'test');
|
||||
assert.equal(result.located, true);
|
||||
assert.equal(result.kind, 'node-test');
|
||||
assert.equal(result.evidence.length, 1, 'one evidence record built');
|
||||
const ev = result.evidence[0];
|
||||
assert.equal(ev.kind, 'node-test');
|
||||
assert.equal(ev.target, 'tests/neg.test.cjs');
|
||||
assert.equal(ev.failFirst, true);
|
||||
assert.equal(ev.passed, true);
|
||||
});
|
||||
|
||||
test('lint-rule (no-source-grep) check that passes -> green, evidence carries rule id', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green');
|
||||
assert.equal(result.flagged, false);
|
||||
assert.equal(result.kind, 'lint-rule');
|
||||
assert.equal(result.evidence[0].kind, 'lint-rule');
|
||||
assert.equal(result.evidence[0].rule, 'local/no-source-grep', 'evidence records which rule asserted the must-NOT');
|
||||
assert.equal(result.evidence[0].target, 'tests/', 'evidence records the linted target path, not the rule id');
|
||||
});
|
||||
|
||||
test('buildLintArgs runs the project eslint as JSON over the target (plugins load via flat config; #1259 SF-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(typeof enforce.buildLintArgs, 'function',
|
||||
'must export buildLintArgs — the eslint argv builder for the lint-rule real runner');
|
||||
const argv = enforce.buildLintArgs({ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/' });
|
||||
assert.ok(Array.isArray(argv), 'argv is an array');
|
||||
const fmtIdx = argv.indexOf('--format');
|
||||
assert.ok(fmtIdx !== -1 && argv[fmtIdx + 1] === 'json',
|
||||
'emits --format json so the report can be filtered by ruleId');
|
||||
assert.ok(argv.includes('--no-warn-ignored'),
|
||||
'must pass --no-warn-ignored so an eslint-ignored target returns [] (fails closed), not a length-1 warning result');
|
||||
assert.ok(!argv.includes('--rule'),
|
||||
'must NOT use --rule — it cannot load a plugin rule like local/no-source-grep (the SF-01 bug)');
|
||||
assert.equal(argv[argv.length - 1], 'tests/', 'the LAST arg is the lint target path');
|
||||
});
|
||||
|
||||
test('lint-rule descriptor missing its rule id -> locate-miss, never green', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', target: 'tests/', failFirst: true }, // no `rule`
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a lint-rule with no rule id is not a valid wired check');
|
||||
assert.equal(result.flagged, true);
|
||||
assert.equal(result.located, false, 'an under-specified lint-rule descriptor is not locatable');
|
||||
});
|
||||
|
||||
test('check that FAILS -> hard-gate (non-green, flagged), located:true, no evidence', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: false }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green');
|
||||
assert.equal(result.flagged, true);
|
||||
assert.equal(result.located, true, 'the check was located even though it failed');
|
||||
assert.equal(result.evidence.length, 0, 'a failing check builds no evidence');
|
||||
});
|
||||
|
||||
test('caller does NOT attest fail-first (descriptor failFirst:false) -> hard-gate, never green', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// fail-first is caller-attested (#1259 BL-02): a check the caller does not attest as fail-first
|
||||
// is not a valid regression proof and must never green, even if the run passes.
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: false },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a non-attested check is not a valid regression proof');
|
||||
assert.equal(result.flagged, true);
|
||||
});
|
||||
|
||||
test('a runCheck that THROWS fails closed, never propagates (no-throw contract, NEW-WR-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => { throw new Error('runner blew up'); } },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a throwing runner must never green');
|
||||
assert.equal(result.flagged, true);
|
||||
assert.equal(result.located, true);
|
||||
});
|
||||
|
||||
test('hard-gates in BOTH modes on a failing check (ADR-550 D4)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
for (const mode of ['interactive', 'autonomous']) {
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: false }), mode },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', `non-green in ${mode}`);
|
||||
assert.equal(result.flagged, true, `flagged in ${mode}`);
|
||||
assert.equal(result.mode, mode, 'mode echoed for transparency');
|
||||
}
|
||||
});
|
||||
|
||||
test('passing run echoes the requested mode without changing the green verdict', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: true }), mode: 'autonomous' },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a passing wired check is green in autonomous mode too');
|
||||
assert.equal(result.mode, 'autonomous');
|
||||
});
|
||||
|
||||
test('routeProhibitionEnforcement parses a JSON request file and emits a structured result', (t) => {
|
||||
const fs = require('node:fs');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
// Write a request file; the route reads it and runs the node-test descriptor's default runner
|
||||
// (its target does not exist, so it fail-closes deterministically — we assert the JSON SHAPE,
|
||||
// not a green verdict). We invoke the built CLI surface in a child process so output()
|
||||
// (writeAllSync to fd 1) is captured on stdout — no source-grep (we parse our own emitted JSON).
|
||||
const dir = createTempDir('prohib-enf-');
|
||||
const reqPath = path.join(dir, 'req.json');
|
||||
const runnerPath = path.join(dir, 'runner.cjs');
|
||||
fs.writeFileSync(reqPath, JSON.stringify({
|
||||
prohibition: TEST_TIER,
|
||||
check: { kind: 'node-test', target: 'tests/neg.test.cjs', failFirst: true },
|
||||
mode: 'autonomous',
|
||||
}));
|
||||
// A tiny runner that requires the BUILT module and invokes the route — output() writes to fd 1.
|
||||
fs.writeFileSync(runnerPath,
|
||||
"require(" + JSON.stringify(ENFORCEMENT_LIB) + ")" +
|
||||
".routeProhibitionEnforcement(['check','prohibition-enforcement'," + JSON.stringify(reqPath) + "], false);\n");
|
||||
t.after(() => cleanup(dir));
|
||||
|
||||
const captured = execFileSync('node', [runnerPath], { encoding: 'utf-8' });
|
||||
const parsed = JSON.parse(captured);
|
||||
assert.equal(typeof parsed, 'object', 'route emits a JSON object');
|
||||
assert.equal(parsed.tier, 'test', 'tier is preserved through the CLI surface');
|
||||
assert.equal(parsed.located, true, 'the check descriptor was located');
|
||||
assert.equal(parsed.mode, 'autonomous', 'mode flows through the CLI surface');
|
||||
assert.equal(typeof parsed.flagged, 'boolean', 'flagged is a typed boolean');
|
||||
assert.ok(Array.isArray(parsed.evidence), 'evidence is an array');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Real-runner helpers (mutation-pinned; #1259 BL-01 / SF-01) ─────────────────
|
||||
// These pin the deterministic parsing/threshold logic of the REAL runner so a Stryker mutant that
|
||||
// weakens "non-vacuous pass" or the ruleId filter is caught — the contract the injected-runner tests
|
||||
// above deliberately bypass.
|
||||
describe('prohibition-enforcement real-runner helpers (#1259)', () => {
|
||||
test('parseNodeTestSummary extracts the TAP tests/pass/fail/cancelled counts', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.deepEqual(enforce.parseNodeTestSummary('# tests 3\n# pass 2\n# fail 1\n# cancelled 1\n'),
|
||||
{ tests: 3, pass: 2, fail: 1, cancelled: 1 });
|
||||
assert.deepEqual(enforce.parseNodeTestSummary('no summary here'), { tests: 0, pass: 0, fail: 0, cancelled: 0 });
|
||||
});
|
||||
|
||||
test('tapTestNames EXCLUDES skipped/todo tests (they never ran, m1)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.deepEqual(enforce.tapTestNames('ok 1 - guards the must-NOT\nok 2 - other # SKIP\nok 3 - later # TODO\n'),
|
||||
['guards the must-NOT'], 'a # SKIP / # TODO test is not a real run and must not count');
|
||||
});
|
||||
|
||||
test('isNonVacuousNodeTestPass: a SKIPPED negative test (file wrapper passes) is NOT a pass (m1)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// file wrapper + a skipped negative test: pass>=1 but the only named test is skipped -> vacuous.
|
||||
const skipped = 'ok 1 - empty.test.cjs\nok 2 - the negative test # SKIP\n# tests 2\n# pass 2\n# fail 0\n# cancelled 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(skipped, 'empty.test.cjs'), false,
|
||||
'a skipped negative test never executed -> must not green');
|
||||
});
|
||||
|
||||
test('isNonVacuousNodeTestPass: a CANCELLED run is not a pass (m1)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const cancelled = 'ok 1 - guards\n# tests 1\n# pass 1\n# fail 0\n# cancelled 1\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(cancelled, 'neg.test.cjs'), false,
|
||||
'a cancelled run is not a clean pass');
|
||||
});
|
||||
|
||||
test('isNonVacuousNodeTestPass: an empty file (node names the test after the file) is NOT a pass (BL-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// node --test of a zero-test file: `ok 1 - empty.test.cjs`, `# tests 1 # pass 1` — counts alone
|
||||
// cannot distinguish it from a real test, so the file-named result must NOT count as a pass.
|
||||
const empty = 'ok 1 - empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(empty, 'empty.test.cjs'), false,
|
||||
'a file-named-only result is vacuous — the BL-01 false-green guard');
|
||||
// BASENAME-NORMALIZED: node may report the file-test by an ABSOLUTE/normalized path while the
|
||||
// descriptor target is relative (cross-OS / node-version). The basenames must still match → vacuous.
|
||||
const emptyAbs = 'ok 1 - /tmp/x/empty.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(emptyAbs, 'empty.test.cjs'), false,
|
||||
'an absolute-path file-test name must still be recognized as vacuous (basename compare, WR-02)');
|
||||
// Mirror case (pins the TARGET-side basename): relative TAP name vs ABSOLUTE descriptor target.
|
||||
const emptyRelName = 'ok 1 - neg.test.cjs\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(emptyRelName, '/abs/path/neg.test.cjs'), false,
|
||||
'a relative file-test name vs an absolute target must still be vacuous — both sides basename-normalized (WR-R4-01)');
|
||||
const real = 'ok 1 - guards the must-NOT\n1..1\n# tests 1\n# pass 1\n# fail 0\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(real, '/abs/path/neg.test.cjs'), true,
|
||||
'a real named test distinct from the file is a genuine pass (even vs an absolute target)');
|
||||
const failing = 'not ok 1 - guards\n# tests 1\n# pass 0\n# fail 1\n';
|
||||
assert.equal(enforce.isNonVacuousNodeTestPass(failing, 'neg.test.cjs'), false,
|
||||
'any failure means not a pass');
|
||||
});
|
||||
|
||||
test('eslintJsonHasRule detects a ruleId; unparseable report -> true (fail-closed)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'local/no-source-grep' }] }]), 'local/no-source-grep'), true);
|
||||
assert.equal(enforce.eslintJsonHasRule(JSON.stringify([{ messages: [{ ruleId: 'other' }] }]), 'local/no-source-grep'), false);
|
||||
assert.equal(enforce.eslintJsonHasRule('not json', 'local/no-source-grep'), true,
|
||||
'an unreadable report must be treated as a violation, never a silent pass');
|
||||
});
|
||||
|
||||
test('eslintFileResultCount: 0 when nothing linted (vacuity guard)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
assert.equal(enforce.eslintFileResultCount(JSON.stringify([{}, {}])), 2);
|
||||
assert.equal(enforce.eslintFileResultCount('[]'), 0);
|
||||
assert.equal(enforce.eslintFileResultCount('garbage'), 0);
|
||||
});
|
||||
|
||||
test('eslintHasFatalError: a parse/fatal error must fail closed (B1)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const fatal = JSON.stringify([{ messages: [{ ruleId: null, fatal: true, severity: 2, message: 'Parsing error' }], fatalErrorCount: 1 }]);
|
||||
assert.equal(enforce.eslintHasFatalError(fatal), true, 'a fatal/parse error means the rule never ran -> fail closed');
|
||||
const clean = JSON.stringify([{ messages: [], fatalErrorCount: 0 }]);
|
||||
assert.equal(enforce.eslintHasFatalError(clean), false, 'a clean lint has no fatal error');
|
||||
assert.equal(enforce.eslintHasFatalError('not json'), true, 'an unreadable report is treated as fatal (fail closed)');
|
||||
});
|
||||
|
||||
test('eslintJsonHasRule also reads suppressedMessages — an inline-disabled violation still counts (B1)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const suppressed = JSON.stringify([{ messages: [], suppressedMessages: [{ ruleId: 'local/no-source-grep' }] }]);
|
||||
assert.equal(enforce.eslintJsonHasRule(suppressed, 'local/no-source-grep'), true,
|
||||
'a violation suppressed via // eslint-disable must NOT be treated as clean');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Real runner end-to-end (NO injected runCheck; #1259 SF-02 / BL-01 / SF-01) ──
|
||||
// Spawns real subprocesses so the SHIPPING default runner is exercised — the gap that let BL-01 and
|
||||
// SF-01 slip past the injected-double tests. Typed-field assertions only.
|
||||
describe('prohibition-enforcement REAL runner end-to-end (#1259)', () => {
|
||||
const fs = require('node:fs');
|
||||
|
||||
test('a genuine non-vacuous passing node-test greens via the real runner', (t) => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const dir = createTempDir('prohib-real-pass-');
|
||||
t.after(() => cleanup(dir));
|
||||
const tf = path.join(dir, 'neg.test.cjs');
|
||||
fs.writeFileSync(tf,
|
||||
"const { test } = require('node:test');\nconst assert = require('node:assert');\ntest('guards the must-NOT', () => { assert.ok(true); });\n");
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: tf, failFirst: true },
|
||||
{ cwd: dir },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a real, passing, non-vacuous negative test must green');
|
||||
assert.equal(result.located, true);
|
||||
assert.equal(result.evidence.length, 1);
|
||||
});
|
||||
|
||||
test('a HANGING node-test fails closed via the bounded timeout (B2: no unbounded subprocess)', (t) => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const dir = createTempDir('prohib-hang-');
|
||||
t.after(() => cleanup(dir));
|
||||
const tf = path.join(dir, 'hang.test.cjs');
|
||||
// A test that never returns; the bounded timeout must kill it and dispose non-green.
|
||||
fs.writeFileSync(tf,
|
||||
"const { test } = require('node:test');\ntest('hangs forever', () => { while (true) {} });\n");
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: tf, failFirst: true },
|
||||
{ cwd: dir, timeoutMs: 1500 },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a hung check must be killed and fail closed — never hang verify or green');
|
||||
assert.equal(result.located, true);
|
||||
});
|
||||
|
||||
test('an EMPTY node-test file (exit 0, zero tests) does NOT green via the real runner (BL-01)', (t) => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const dir = createTempDir('prohib-real-empty-');
|
||||
t.after(() => cleanup(dir));
|
||||
const tf = path.join(dir, 'empty.test.cjs');
|
||||
fs.writeFileSync(tf, '// intentionally empty — no test cases\n');
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'node-test', target: tf, failFirst: true },
|
||||
{ cwd: dir },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'an empty (zero-test) file must NEVER green — fail-closed');
|
||||
assert.equal(result.located, true, 'the check was located; it just did not genuinely pass');
|
||||
assert.equal(result.evidence.length, 0);
|
||||
});
|
||||
|
||||
test('a clean in-tree target greens the lint-rule kind via the real eslint runner (SF-01: plugin loads)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// Runs real `npx eslint --format json src/clock.cts` under the project flat config (so the
|
||||
// `local` plugin loads). src/clock.cts is a clean source with no no-source-grep violation.
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'src/clock.cts', failFirst: true },
|
||||
{ cwd: process.cwd() },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a clean target with no no-source-grep violation must green via real eslint');
|
||||
assert.equal(result.kind, 'lint-rule');
|
||||
assert.equal(result.evidence[0].rule, 'local/no-source-grep');
|
||||
});
|
||||
|
||||
test('an eslint-IGNORED target does NOT green the lint-rule kind (vacuous-green guard, NEW-BL-01)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
// The generated bin/lib artifact is eslint-ignored. Without --no-warn-ignored, eslint returns a
|
||||
// length-1 "File ignored" result that would falsely pass the vacuity guard. It must fail closed.
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
TEST_TIER,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'gsd-core/bin/lib/prohibition-enforcement.cjs', failFirst: true },
|
||||
{ cwd: process.cwd() },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'an ignored path lints nothing — must NEVER green');
|
||||
assert.equal(result.located, true, 'the descriptor was well-formed; it just did not genuinely pass');
|
||||
});
|
||||
});
|
||||
@@ -21,6 +21,10 @@ const assert = require('node:assert/strict');
|
||||
const path = require('node:path');
|
||||
|
||||
const PROBE_CORE_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
|
||||
// The ENFORCEMENT-half producer (#1259, ADR-550 D5d heavy half). Authored as
|
||||
// src/prohibition-enforcement.cts and compiled by `npm run build:lib` to this gitignored
|
||||
// artifact — mirroring how PROBE_CORE_LIB requires the BUILT probe-core.cjs above.
|
||||
const ENFORCEMENT_LIB = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs');
|
||||
|
||||
describe('prohibition-probe verify-tier: test-tier fail-closed safety (PROB-12 / ADR-550 D5d)', () => {
|
||||
test('probe-core exports a deterministic prohibition-disposition helper', () => {
|
||||
@@ -55,3 +59,103 @@ describe('prohibition-probe verify-tier: test-tier fail-closed safety (PROB-12 /
|
||||
'an unwired test-tier item must be flagged unverified — it can never be silently skipped');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── ENFORCEMENT HALF (#1259, ADR-550 D5d heavy half) ──────────────────────────
|
||||
//
|
||||
// RED-first: these assertions require the BUILT gsd-core/bin/lib/prohibition-enforcement.cjs,
|
||||
// which does not exist until Task 2 authors src/prohibition-enforcement.cts and runs build:lib.
|
||||
// They prove the previously-unreachable green branch in dispositionForProhibition (probe-core
|
||||
// 420-427) becomes reachable from a real PRODUCER: a passing wired test-tier check builds
|
||||
// non-empty enforcementEvidence -> green; a missing or failing check hard-gates (flagged,
|
||||
// non-green) in BOTH interactive and autonomous modes (ADR-550 D4 / D3).
|
||||
//
|
||||
// Typed-field assertions ONLY (status / flagged / tier / evidence / located / kind / mode-block).
|
||||
// The check-runner is injected via options.runCheck so no real subprocess is spawned.
|
||||
describe('prohibition-probe verify-tier: test-tier ENFORCEMENT (REQ-PROHIB-07 / ADR-550 D5d)', () => {
|
||||
const testTierProhibition = {
|
||||
requirement_id: 'R1',
|
||||
category: 'safety',
|
||||
status: 'resolved',
|
||||
verification: 'test',
|
||||
resolution: null,
|
||||
reason: null,
|
||||
statement: 'MUST NOT read source files and text-search them in tests',
|
||||
};
|
||||
|
||||
// Test A — node --test negative test that PASSES -> green + non-empty evidence.
|
||||
test('A: wired node-test check that passes disposes green with non-empty enforcement evidence', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.ok(result && typeof result === 'object', 'result must be a structured object');
|
||||
assert.equal(result.status, 'green', 'a passing wired node-test check must dispose green');
|
||||
assert.equal(result.flagged, false, 'a green test-tier disposition must not be flagged');
|
||||
assert.equal(result.tier, 'test', 'tier must be preserved as test');
|
||||
assert.equal(result.located, true, 'the wired check was locatable');
|
||||
assert.ok(Array.isArray(result.evidence) && result.evidence.length >= 1,
|
||||
'a passing check must build non-empty enforcementEvidence (the array dispositionForProhibition reads)');
|
||||
});
|
||||
|
||||
// Test B — lint/AST-rule (no-source-grep) check that PASSES -> green (D4 dogfood anchor).
|
||||
test('B: wired lint-rule (no-source-grep) check that passes disposes green', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'lint-rule', rule: 'local/no-source-grep', target: 'tests/', failFirst: true },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.equal(result.status, 'green', 'a passing wired lint-rule check must dispose green');
|
||||
assert.equal(result.flagged, false, 'a green disposition must not be flagged');
|
||||
assert.equal(result.kind, 'lint-rule', 'the located check kind must be the lint-rule kind');
|
||||
assert.ok(Array.isArray(result.evidence) && result.evidence.length >= 1,
|
||||
'a passing lint-rule check must build non-empty enforcementEvidence');
|
||||
});
|
||||
|
||||
// Test C — MISSING check (no locatable wired check) -> hard-gate (non-green, flagged).
|
||||
test('C: a test-tier prohibition with NO locatable wired check hard-gates (non-green, flagged)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
null,
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green', 'a missing wired check must NEVER be green (fail-closed)');
|
||||
assert.equal(result.flagged, true, 'a missing wired check must be flagged unverified');
|
||||
assert.equal(result.located, false, 'no check was locatable');
|
||||
assert.ok(Array.isArray(result.evidence) && result.evidence.length === 0,
|
||||
'a missing check builds no enforcement evidence');
|
||||
});
|
||||
|
||||
// Test D — FAILING check -> hard-gate in BOTH modes (interactive + autonomous).
|
||||
test('D: a wired check that FAILS hard-gates (non-green, flagged) in both modes', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
for (const mode of ['interactive', 'autonomous']) {
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: true },
|
||||
{ runCheck: () => ({ passed: false }), mode },
|
||||
);
|
||||
assert.notEqual(result.status, 'green',
|
||||
`a failing wired check must NEVER be green (mode=${mode})`);
|
||||
assert.equal(result.flagged, true,
|
||||
`a failing wired check must be flagged in both modes (mode=${mode})`);
|
||||
assert.equal(result.located, true, 'the check was located even though it failed');
|
||||
}
|
||||
});
|
||||
|
||||
// D (fail-first not satisfied) — a check that is NOT fail-first is not a valid regression proof.
|
||||
test('D2: a wired check that is not fail-first hard-gates (non-green, flagged)', () => {
|
||||
const enforce = require(ENFORCEMENT_LIB);
|
||||
const result = enforce.runProhibitionEnforcement(
|
||||
testTierProhibition,
|
||||
{ kind: 'node-test', target: 'tests/some-negative.test.cjs', failFirst: false },
|
||||
{ runCheck: () => ({ passed: true }) },
|
||||
);
|
||||
assert.notEqual(result.status, 'green',
|
||||
'a check that is not fail-first is not a valid regression-must-fail-first proof — never green');
|
||||
assert.equal(result.flagged, true, 'a non-fail-first check must be flagged');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -85,6 +85,6 @@
|
||||
"undo.md": 10431,
|
||||
"update.md": 21053,
|
||||
"validate-phase.md": 10745,
|
||||
"verify-phase.md": 33161,
|
||||
"verify-phase.md": 35362,
|
||||
"verify-work.md": 31157
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user