Files
msd-core/tests/edge-probe.test.cjs
Rezolv 3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00

442 lines
21 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* Edge-probe reference core unit tests.
*
* Asserts the LOCKED export surface of the spec-completeness edge-probe against
* the BUILT artifact (`gsd-core/bin/lib/edge-probe.cjs`), which
* `npm run build:lib` (run by pretest) emits from `src/edge-probe.cts`.
*
* Post ADR-550 Decision 7: the generic resolution model lives in `probe-core`;
* edge-probe is its first adapter (shapes/TAXONOMY/proposeEdges + the
* `{explicit, backstop}` verification validators). The resolution model is the
* status×verification re-cut: `status: resolved | dismissed | unresolved` ×
* `verification: explicit | backstop`. `covered`/`backstop` are no longer status
* values — `covered → {resolved, explicit}`, `backstop → {resolved, backstop}`.
*/
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const path = require('node:path');
const fs = require('node:fs');
const os = require('node:os');
const { execFileSync, spawnSync } = require('node:child_process');
const { cleanup } = require('./helpers.cjs');
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs');
const ep = require(BUILT_SCRIPT);
describe('edge-probe: classifyShape', () => {
test('detects numeric-range from rounding/threshold cues', () => {
assert.deepEqual(ep.classifyShape('Round a number to N decimal places').sort(),
['numeric-range']);
});
test('detects collection from interval/merge cues', () => {
const shapes = ep.classifyShape('Merge a list of overlapping intervals');
assert.ok(shapes.includes('collection'));
});
test('detects text from truncate/string cues', () => {
const shapes = ep.classifyShape('Truncate a string to a maximum length');
assert.ok(shapes.includes('text'));
});
test('returns [] when no cue matches', () => {
assert.deepEqual(ep.classifyShape('Display the company logo'), []);
});
});
describe('edge-probe: TAXONOMY + applicableCategories', () => {
test('TAXONOMY has the 8 documented categories in order', () => {
assert.deepEqual(ep.TAXONOMY.map((c) => c.id),
['boundary', 'adjacency', 'empty', 'encoding', 'ordering', 'precision', 'idempotency', 'concurrency']);
});
test('every category has name, shapes[], probe', () => {
for (const c of ep.TAXONOMY) {
assert.equal(typeof c.name, 'string');
assert.ok(Array.isArray(c.shapes) && c.shapes.length >= 1);
assert.equal(typeof c.probe, 'string');
}
});
test('numeric-range raises boundary + precision only', () => {
assert.deepEqual(ep.applicableCategories(['numeric-range']).sort(),
['boundary', 'precision']);
});
test('collection raises adjacency, empty, ordering', () => {
assert.deepEqual(ep.applicableCategories(['collection']).sort(),
['adjacency', 'empty', 'ordering']);
});
test('text raises empty + encoding', () => {
assert.deepEqual(ep.applicableCategories(['text']).sort(),
['empty', 'encoding']);
});
test('no shapes raises nothing', () => {
assert.deepEqual(ep.applicableCategories([]), []);
});
});
describe('edge-probe: proposeEdges', () => {
test('rounding requirement proposes boundary + precision, all unresolved (verification null)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places' });
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
for (const e of edges) {
assert.equal(e.requirement_id, 'R1');
assert.equal(e.status, 'unresolved');
assert.equal(e.verification, null);
assert.equal(e.resolution, null);
assert.equal(e.reason, null);
assert.equal(typeof e.probe, 'string');
}
});
test('authored shapes override prose classification', () => {
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque label', shapes: ['collection'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
});
});
describe('edge-probe: validateResolution', () => {
test('rejects an unknown status', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'maybe' }),
/invalid status/i);
});
test('rejects a former covered status (re-cut: covered is no longer a status)', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'covered', resolution: 'x' }),
/invalid status/i);
});
test('rejects dismissed without a reason', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: '' }),
/dismissed requires a reason/i);
});
test('accepts dismissed with a reason', () => {
assert.equal(ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: 'bounded enum' }), true);
});
test('rejects resolved with a missing verification tier', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', resolution: 'AC' }),
/verification/i);
});
test('rejects resolved with a verification tier outside {explicit, backstop}', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'judgment', resolution: 'AC' }),
/invalid verification/i);
});
});
describe('edge-probe: analyzeCoverage', () => {
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
test('with no resolutions, every applicable edge is unresolved (byVerification zeroed)', () => {
const rep = ep.analyzeCoverage(reqs, []);
assert.deepEqual(rep.coverage, { applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 } });
});
test('merges a resolved/explicit resolution and counts it resolved', () => {
const rep = ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
]);
const adj = rep.items.find((i) => i.category === 'adjacency');
assert.equal(adj.status, 'resolved');
assert.equal(adj.verification, 'explicit');
assert.equal(adj.resolution, 'AC#6: touching intervals merge');
assert.equal(rep.coverage.resolved, 1);
assert.equal(rep.coverage.unresolved, 2);
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 });
});
test('throws if a resolution is invalid (dismissed w/o reason)', () => {
assert.throws(() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'empty', status: 'dismissed' },
]), /dismissed requires a reason/i);
});
});
describe('edge-probe: CLI (built artifact)', () => {
test('reads a requirements file and prints a coverage report as JSON', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-'));
const reqPath = path.join(dir, 'requirements.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
const out = execFileSync('node', [BUILT_SCRIPT, reqPath], { encoding: 'utf8' });
const rep = JSON.parse(out);
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
});
test('with no args exits with status 2 (assert on exit code, not stderr prose)', () => {
let status;
try {
execFileSync('node', [BUILT_SCRIPT], { stdio: 'pipe' });
status = 0;
} catch (error) {
status = error.status;
}
assert.equal(status, 2);
});
});
describe('edge-probe: CLI JSON.parse error handling (RR-10)', () => {
test('invalid requirements JSON exits with status 2 (handled error, not uncaught throw)', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const badJson = path.join(dir, 'bad-req.json');
fs.writeFileSync(badJson, 'not valid json {{{');
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, badJson], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('invalid resolutions JSON exits with status 2 (handled error, not uncaught throw)', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const goodReq = path.join(dir, 'req.json');
const badRes = path.join(dir, 'bad-res.json');
fs.writeFileSync(goodReq, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
fs.writeFileSync(badRes, 'not valid json {{{');
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, goodReq, badRes], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('valid requirements file exits 0 and stdout is parseable JSON', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const reqPath = path.join(dir, 'req.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe', encoding: 'utf8' });
assert.equal(r.status, 0);
const rep = JSON.parse(r.stdout);
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
} finally {
cleanup(dir);
}
});
});
describe('edge-probe: proposeEdges — empty-shapes override (RR-06)', () => {
test('shapes: [] returns zero edges (explicit empty-shapes override)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] });
assert.deepEqual(edges, []);
});
test('absent shapes key classifies from prose (no override)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals' });
assert.ok(edges.length > 0, 'should classify collection edges from prose');
});
test('shapes: [collection] overrides prose and proposes collection categories', () => {
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque text with no cues', shapes: ['collection'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
});
});
describe('edge-probe: proposeEdges — invalid authored shapes fail closed (re-review #3 High)', () => {
// A non-empty but INVALID shapes array must NOT silently suppress every probe.
// shapes:['numeric'] (typo for the locked 'numeric-range') previously passed
// Array.isArray, matched no category, and returned applicable:0 — failing OPEN.
test('rejects an unknown shape value (typo for a locked shape)', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric'] }),
/invalid shape/i,
);
});
test('rejects a mixed array where one entry is invalid', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric-range', 'bogus'] }),
/invalid shape/i,
);
});
test('rejects a non-string shape entry', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: [42] }),
/invalid shape/i,
);
});
test('analyzeCoverage propagates the invalid-shape throw', () => {
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]),
/invalid shape/i,
);
});
test('CLI exits 2 (handled) on an invalid authored shape, not an uncaught trace', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-shape-'));
const reqPath = path.join(dir, 'req.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]));
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('a valid locked shape still proposes its categories (no false rejection)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'opaque', shapes: ['numeric-range'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
});
test('shapes: [] remains a valid zero-edge override (RR-06 intact)', () => {
assert.deepEqual(ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] }), []);
});
});
describe('edge-probe: input validation & orphan-resolution rejection (adversarial review)', () => {
// HIGH: a resolution whose (requirement_id, category) matches no proposed edge — a typo'd
// category or a non-applicable one — was silently DROPPED, so an author who typos `precison`
// sees the precision edge as still-unresolved with no error (a confirmed money-rounding exploit).
test('rejects an orphan resolution (typo category — no matching proposed edge)', () => {
assert.throws(
() => ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'precison', status: 'resolved', verification: 'explicit', resolution: 'AC: precision handled' }],
),
/unknown resolution|no matching proposed edge/i,
);
});
test('rejects a resolution for a valid-but-non-applicable category', () => {
// 'encoding' is a real taxonomy id but applies to text, not the numeric-range requirement.
assert.throws(
() => ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'encoding', status: 'resolved', verification: 'explicit', resolution: 'AC' }],
),
/unknown resolution|no matching proposed edge/i,
);
});
test('a matching resolution still resolves (no false orphan rejection)', () => {
const rep = ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'precision', status: 'resolved', verification: 'explicit', resolution: 'AC: precision tested' }],
);
assert.equal(rep.coverage.resolved, 1);
});
test('rejects requirements that is not an array', () => {
assert.throws(() => ep.analyzeCoverage('nope'), /requirements must be an array/i);
});
test('rejects a duplicate requirement id', () => {
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1', text: 'a' }, { id: 'R1', text: 'b' }]),
/duplicate requirement/i,
);
});
test('rejects a truthy non-array shapes (string instead of array)', () => {
// A bare string `shapes: "numeric-range"` previously fell through to prose classification,
// silently ignoring the authored override instead of honoring or rejecting it.
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'x', shapes: 'numeric-range' }),
/shapes must be an array/i,
);
});
test('rejects a missing requirement id', () => {
assert.throws(() => ep.proposeEdges({ text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects an empty requirement id', () => {
assert.throws(() => ep.proposeEdges({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects a non-string requirement text', () => {
assert.throws(() => ep.proposeEdges({ id: 'R1', text: 42 }), /text must be a string/i);
});
test('rejects a missing requirement text when no shapes override (M2 fail-open)', () => {
// Without text or an authored shape, prose classification yields zero shapes → zero edges →
// the requirement is silently DROPPED from coverage with no signal — the exact fail-open this
// feature exists to eliminate. The edge adapter's `text` is required, so reject it.
assert.throws(
() => ep.proposeEdges({ id: 'R1' }),
/text must be a non-empty string when no shapes override/i,
);
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1' }]),
/text must be a non-empty string when no shapes override/i,
);
});
test('rejects an empty/whitespace requirement text when no shapes override (M2)', () => {
assert.throws(() => ep.proposeEdges({ id: 'R1', text: '' }), /text must be a non-empty string when no shapes override/i);
assert.throws(() => ep.proposeEdges({ id: 'R1', text: ' ' }), /text must be a non-empty string when no shapes override/i);
});
test('allows missing/empty text WHEN an explicit shapes override is provided (M2 legitimate path)', () => {
// An authored `shapes` array (including `[]` for "no applicable categories") opts out of prose
// classification, so `text` is not required — this must remain valid.
assert.deepEqual(ep.proposeEdges({ id: 'R1', shapes: [] }), []);
const edges = ep.proposeEdges({ id: 'R1', shapes: ['numeric-range'] });
assert.ok(edges.length > 0, 'an explicit shape override must still propose edges without text');
});
});
describe('edge-probe: validateResolution — explicit-needs-resolution (RR-07, re-cut)', () => {
test('rejects resolved/explicit with empty resolution string', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: '' }),
/explicit requires a resolution/i,
);
});
test('rejects resolved/explicit with whitespace-only resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: ' ' }),
/explicit requires a resolution/i,
);
});
test('rejects resolved/explicit with missing resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit' }),
/explicit requires a resolution/i,
);
});
test('accepts resolved/explicit with a non-empty resolution', () => {
assert.equal(
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC#3: boundary tested in suite' }),
true,
);
});
});
describe('edge-probe: validateResolution — backstop-needs-resolution (RR-07 follow-up, re-cut)', () => {
test('rejects resolved/backstop with empty resolution string', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: '' }),
/backstop requires a resolution/i,
);
});
test('rejects resolved/backstop with whitespace-only resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: ' ' }),
/backstop requires a resolution/i,
);
});
test('rejects resolved/backstop with missing resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop' }),
/backstop requires a resolution/i,
);
});
test('accepts resolved/backstop with a non-empty resolution note', () => {
assert.equal(
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: 'held-out: covered by integration fuzz suite' }),
true,
);
});
});
describe('edge-probe: analyzeCoverage — duplicate rejection (RR-09)', () => {
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
test('rejects duplicate (requirement_id, category) resolution', () => {
assert.throws(
() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' },
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' },
]),
/duplicate resolution/i,
);
});
test('distinct pairs still analyze without throwing', () => {
// mirrors fixture 06-resolved-mixed
assert.doesNotThrow(() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', resolution: null, reason: 'output is canonically sorted; no tie possible' },
]));
});
});
describe('edge-probe: golden fixtures', () => {
const root = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures');
const fixtures = fs.readdirSync(root).filter((d) =>
fs.statSync(path.join(root, d)).isDirectory());
assert.ok(fixtures.length >= 6, 'expected at least 6 fixtures');
for (const name of fixtures) {
test(`fixture ${name} matches its golden coverage`, () => {
const dir = path.join(root, name);
const reqs = JSON.parse(fs.readFileSync(path.join(dir, 'requirements.json'), 'utf8'));
const resPath = path.join(dir, 'resolutions.json');
const res = fs.existsSync(resPath) ? JSON.parse(fs.readFileSync(resPath, 'utf8')) : [];
const expected = JSON.parse(fs.readFileSync(path.join(dir, 'expected-coverage.json'), 'utf8'));
assert.deepEqual(ep.analyzeCoverage(reqs, res), expected);
});
}
});