Files
msd-core/tests/probe-core.property.test.cjs
Rezolv 3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00

168 lines
7.0 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
'use strict';
/**
* Property-based tests for probe-core.cjs (ADR-550 Decision 7).
*
* Module: gsd-core/bin/lib/probe-core.cjs (generated from src/probe-core.cts)
* Exercised: analyzeCoverage(items, resolutions?, validators) — the generic
* merge/rollup/orphan-reject engine shared by the edge probe and the #644
* prohibition probe.
*
* trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
* transformation/rollup module (items × resolutions → CoverageReport) — exactly the
* class the predicate covers. The example-based suite (tests/probe-core.test.cjs)
* pins specific scenarios; these properties pin the algebraic invariants that must
* hold for EVERY valid scenario.
*
* Properties tested:
* (a) closed-set identity: applicable === resolved + unresolved (and === items.length)
* (b) byVerification sums ≤ resolved (dismissed is closed but unverified)
* (c) per-tier byVerification ≤ resolved, and only `resolved`-status items are counted
* (d) determinism: same input → identical CoverageReport (stable rollup)
* (e) orphan rejection is stable: a resolution matching no proposed item always throws
*/
const { describe, test } = require('node:test');
const path = require('node:path');
const fc = require('./helpers/fast-check-setup.cjs');
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
const pc = require(BUILT_SCRIPT);
// The same representative validators bundle the edge adapter injects (see
// tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe.
const VALIDATORS = {
categories: ['adjacency', 'empty', 'ordering'],
verification: ['explicit', 'backstop'],
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
};
function bareItem(requirement_id, category) {
return {
requirement_id,
category,
status: 'unresolved',
verification: null,
resolution: null,
reason: null,
probe: `probe-for-${category}`,
};
}
const catArb = fc.constantFrom(...VALIDATORS.categories);
const idArb = fc.constantFrom('R1', 'R2', 'R3', 'R4', 'R5');
const keyArb = fc.record({ requirement_id: idArb, category: catArb });
// Unique (requirement_id, category) keys — the merge keys analyzeCoverage maps on.
const keyOf = (k) => `${k.requirement_id}::${k.category}`;
const uniqueKeysArb = fc.uniqueArray(keyArb, { selector: keyOf, minLength: 0, maxLength: 12 });
// Each unique item key gets one resolution disposition. Resolution text/reason use fixed
// non-empty literals — the counting invariants are independent of their content, and this
// keeps the generator off the validateResolution rejection paths (which the example suite
// already covers exhaustively).
const DISPOSITIONS = ['none', 'resolved-explicit', 'resolved-backstop', 'dismissed', 'unresolved'];
function resolutionFor(k, disposition) {
const base = { requirement_id: k.requirement_id, category: k.category };
switch (disposition) {
case 'resolved-explicit':
return { ...base, status: 'resolved', verification: 'explicit', resolution: 'AC#1' };
case 'resolved-backstop':
return { ...base, status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' };
case 'dismissed':
return { ...base, status: 'dismissed', reason: 'bounded enum — not applicable' };
case 'unresolved':
return { ...base, status: 'unresolved' };
default: // 'none' — author left no resolution; item rolls up verbatim (bare unresolved)
return null;
}
}
// A fully valid scenario: unique items (all bare-unresolved) + a per-item resolution choice.
const scenarioArb = uniqueKeysArb.chain((keys) =>
fc.tuple(...keys.map(() => fc.constantFrom(...DISPOSITIONS))).map((choices) => {
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
const resolutions = [];
keys.forEach((k, i) => {
const r = resolutionFor(k, choices[i]);
if (r) resolutions.push(r);
});
return { items, resolutions };
}),
);
describe('probe-core property: analyzeCoverage algebraic invariants', () => {
test('(a) closed-set identity: applicable === resolved + unresolved === items.length', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
return (
coverage.applicable === coverage.resolved + coverage.unresolved &&
coverage.applicable === items.length
);
}),
);
});
test('(b) sum(byVerification) ≤ resolved — dismissed counts closed but unverified', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
const verifiedTotal = Object.values(coverage.byVerification).reduce((a, b) => a + b, 0);
return verifiedTotal <= coverage.resolved && verifiedTotal >= 0;
}),
);
});
test('(c) byVerification only counts resolved-status items, and matches a direct recount', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const rep = pc.analyzeCoverage(items, resolutions, VALIDATORS);
for (const tier of VALIDATORS.verification) {
const recount = rep.items.filter((i) => i.status === 'resolved' && i.verification === tier).length;
if (rep.coverage.byVerification[tier] !== recount) return false;
if (rep.coverage.byVerification[tier] > rep.coverage.resolved) return false;
}
return true;
}),
);
});
test('(d) determinism: identical inputs produce an identical CoverageReport', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const a = pc.analyzeCoverage(items, resolutions, VALIDATORS);
const b = pc.analyzeCoverage(items, resolutions, VALIDATORS);
return JSON.stringify(a) === JSON.stringify(b);
}),
);
});
});
describe('probe-core property: orphan rejection is stable', () => {
// An orphan resolution carries an id ('Z9') that no generated item ever uses, so its
// (requirement_id, category) key never matches a proposed item. The resolution is itself
// structurally VALID (a bare unresolved), so it clears validateResolution and reaches the
// orphan-reject guard — isolating that guard from input-validation throws.
const orphanScenarioArb = fc.record({
keys: uniqueKeysArb,
orphanCategory: catArb,
});
test('(e) a resolution matching no proposed item always throws', () => {
fc.assert(
fc.property(orphanScenarioArb, ({ keys, orphanCategory }) => {
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
const orphan = { requirement_id: 'Z9', category: orphanCategory, status: 'unresolved' };
let threwForOrphan = false;
try {
pc.analyzeCoverage(items, [orphan], VALIDATORS);
} catch (e) {
threwForOrphan = /unknown resolution|no matching proposed item/i.test(e.message);
}
return threwForOrphan;
}),
);
});
});