* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:
- spec-phase.md 15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
- plan-phase.md 93135 -> 94253 (+1118): covered/backstop edge lift into
must_haves.truths (the live <downstream_consumer> block)
Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.
* chore(#550): reconcile INVENTORY headline counts after rebase onto next
Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
- References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
- CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs
Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
168 lines
7.0 KiB
JavaScript
168 lines
7.0 KiB
JavaScript
'use strict';
|
||
|
||
/**
|
||
* Property-based tests for probe-core.cjs (ADR-550 Decision 7).
|
||
*
|
||
* Module: gsd-core/bin/lib/probe-core.cjs (generated from src/probe-core.cts)
|
||
* Exercised: analyzeCoverage(items, resolutions?, validators) — the generic
|
||
* merge/rollup/orphan-reject engine shared by the edge probe and the #644
|
||
* prohibition probe.
|
||
*
|
||
* trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
|
||
* transformation/rollup module (items × resolutions → CoverageReport) — exactly the
|
||
* class the predicate covers. The example-based suite (tests/probe-core.test.cjs)
|
||
* pins specific scenarios; these properties pin the algebraic invariants that must
|
||
* hold for EVERY valid scenario.
|
||
*
|
||
* Properties tested:
|
||
* (a) closed-set identity: applicable === resolved + unresolved (and === items.length)
|
||
* (b) byVerification sums ≤ resolved (dismissed is closed but unverified)
|
||
* (c) per-tier byVerification ≤ resolved, and only `resolved`-status items are counted
|
||
* (d) determinism: same input → identical CoverageReport (stable rollup)
|
||
* (e) orphan rejection is stable: a resolution matching no proposed item always throws
|
||
*/
|
||
|
||
const { describe, test } = require('node:test');
|
||
const path = require('node:path');
|
||
const fc = require('./helpers/fast-check-setup.cjs');
|
||
|
||
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
|
||
const pc = require(BUILT_SCRIPT);
|
||
|
||
// The same representative validators bundle the edge adapter injects (see
|
||
// tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe.
|
||
const VALIDATORS = {
|
||
categories: ['adjacency', 'empty', 'ordering'],
|
||
verification: ['explicit', 'backstop'],
|
||
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
|
||
};
|
||
|
||
function bareItem(requirement_id, category) {
|
||
return {
|
||
requirement_id,
|
||
category,
|
||
status: 'unresolved',
|
||
verification: null,
|
||
resolution: null,
|
||
reason: null,
|
||
probe: `probe-for-${category}`,
|
||
};
|
||
}
|
||
|
||
const catArb = fc.constantFrom(...VALIDATORS.categories);
|
||
const idArb = fc.constantFrom('R1', 'R2', 'R3', 'R4', 'R5');
|
||
const keyArb = fc.record({ requirement_id: idArb, category: catArb });
|
||
// Unique (requirement_id, category) keys — the merge keys analyzeCoverage maps on.
|
||
const keyOf = (k) => `${k.requirement_id}::${k.category}`;
|
||
const uniqueKeysArb = fc.uniqueArray(keyArb, { selector: keyOf, minLength: 0, maxLength: 12 });
|
||
|
||
// Each unique item key gets one resolution disposition. Resolution text/reason use fixed
|
||
// non-empty literals — the counting invariants are independent of their content, and this
|
||
// keeps the generator off the validateResolution rejection paths (which the example suite
|
||
// already covers exhaustively).
|
||
const DISPOSITIONS = ['none', 'resolved-explicit', 'resolved-backstop', 'dismissed', 'unresolved'];
|
||
|
||
function resolutionFor(k, disposition) {
|
||
const base = { requirement_id: k.requirement_id, category: k.category };
|
||
switch (disposition) {
|
||
case 'resolved-explicit':
|
||
return { ...base, status: 'resolved', verification: 'explicit', resolution: 'AC#1' };
|
||
case 'resolved-backstop':
|
||
return { ...base, status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' };
|
||
case 'dismissed':
|
||
return { ...base, status: 'dismissed', reason: 'bounded enum — not applicable' };
|
||
case 'unresolved':
|
||
return { ...base, status: 'unresolved' };
|
||
default: // 'none' — author left no resolution; item rolls up verbatim (bare unresolved)
|
||
return null;
|
||
}
|
||
}
|
||
|
||
// A fully valid scenario: unique items (all bare-unresolved) + a per-item resolution choice.
|
||
const scenarioArb = uniqueKeysArb.chain((keys) =>
|
||
fc.tuple(...keys.map(() => fc.constantFrom(...DISPOSITIONS))).map((choices) => {
|
||
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
|
||
const resolutions = [];
|
||
keys.forEach((k, i) => {
|
||
const r = resolutionFor(k, choices[i]);
|
||
if (r) resolutions.push(r);
|
||
});
|
||
return { items, resolutions };
|
||
}),
|
||
);
|
||
|
||
describe('probe-core property: analyzeCoverage algebraic invariants', () => {
|
||
test('(a) closed-set identity: applicable === resolved + unresolved === items.length', () => {
|
||
fc.assert(
|
||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||
return (
|
||
coverage.applicable === coverage.resolved + coverage.unresolved &&
|
||
coverage.applicable === items.length
|
||
);
|
||
}),
|
||
);
|
||
});
|
||
|
||
test('(b) sum(byVerification) ≤ resolved — dismissed counts closed but unverified', () => {
|
||
fc.assert(
|
||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||
const verifiedTotal = Object.values(coverage.byVerification).reduce((a, b) => a + b, 0);
|
||
return verifiedTotal <= coverage.resolved && verifiedTotal >= 0;
|
||
}),
|
||
);
|
||
});
|
||
|
||
test('(c) byVerification only counts resolved-status items, and matches a direct recount', () => {
|
||
fc.assert(
|
||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||
const rep = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||
for (const tier of VALIDATORS.verification) {
|
||
const recount = rep.items.filter((i) => i.status === 'resolved' && i.verification === tier).length;
|
||
if (rep.coverage.byVerification[tier] !== recount) return false;
|
||
if (rep.coverage.byVerification[tier] > rep.coverage.resolved) return false;
|
||
}
|
||
return true;
|
||
}),
|
||
);
|
||
});
|
||
|
||
test('(d) determinism: identical inputs produce an identical CoverageReport', () => {
|
||
fc.assert(
|
||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||
const a = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||
const b = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||
return JSON.stringify(a) === JSON.stringify(b);
|
||
}),
|
||
);
|
||
});
|
||
});
|
||
|
||
describe('probe-core property: orphan rejection is stable', () => {
|
||
// An orphan resolution carries an id ('Z9') that no generated item ever uses, so its
|
||
// (requirement_id, category) key never matches a proposed item. The resolution is itself
|
||
// structurally VALID (a bare unresolved), so it clears validateResolution and reaches the
|
||
// orphan-reject guard — isolating that guard from input-validation throws.
|
||
const orphanScenarioArb = fc.record({
|
||
keys: uniqueKeysArb,
|
||
orphanCategory: catArb,
|
||
});
|
||
|
||
test('(e) a resolution matching no proposed item always throws', () => {
|
||
fc.assert(
|
||
fc.property(orphanScenarioArb, ({ keys, orphanCategory }) => {
|
||
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
|
||
const orphan = { requirement_id: 'Z9', category: orphanCategory, status: 'unresolved' };
|
||
let threwForOrphan = false;
|
||
try {
|
||
pc.analyzeCoverage(items, [orphan], VALIDATORS);
|
||
} catch (e) {
|
||
threwForOrphan = /unknown resolution|no matching proposed item/i.test(e.message);
|
||
}
|
||
return threwForOrphan;
|
||
}),
|
||
);
|
||
});
|
||
});
|