From 3e836fef0d421949ac9b1be87fb0990740970558 Mon Sep 17 00:00:00 2001 From: Rezolv Date: Fri, 12 Jun 2026 11:05:31 -0400 Subject: [PATCH] feat(spec-phase): spec-completeness edge-probe (#550) (#584) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/ Rebased onto current next and relocated the whole feature from get-shit-done/ to gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow @-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated. Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved. Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted by plan-phase into must_haves.truths, extending the goal-backward verifier's reach to boundary edges no requirement was written for. Engine authored as strict TS (src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs. Folds in every prior review round on PR #584: - RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test. - RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned, source-checkout-gated build fallback) instead of LLM re-derivation; the engine capture is exit-checked and the report JSON-validated before use (fail closed). - RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2); per-artifact build sentinel. - Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES). - Adversarial-review hardening: orphan/typo resolution rejection, requirement input validation (id/text/shapes, duplicate id, non-array), zero-applicable guard. Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72. * test(#550): RED — status×verification re-cut + probe-core engine specs Re-cut the edge-probe resolution model onto two orthogonal axes per ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44): status: resolved | dismissed | unresolved (lifecycle, shared) verification: explicit | backstop | null (only when resolved) - tests/probe-core.test.cjs (new): behavioral specs for the generic engine to be extracted — validateResolution(r, validators), validateRequirement, analyzeCoverage(items, resolutions?, validators), byVerification rollup, runProbeCli I/O scaffold (injected-io unit tests). - tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→ {resolved,backstop}; coverage gains byVerification.{explicit,backstop}; proposeEdges items gain verification:null. - 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved COUNT preserved on every fixture (closed set = resolved+dismissed; doc line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose rewritten to the two-axis model. Fails as expected: probe-core.cjs has no source yet; edge-probe still emits the old covered/backstop enum (27/61 edge specs red). * feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7) Extract the generic resolution model into src/probe-core.cts (the shared seam the prohibition probe #644 is born on) and refactor edge-probe.cts into its first adapter. probe-core owns (probe-agnostic): - the status×verification re-cut: status: resolved|dismissed|unresolved × verification: |null - validateResolution(r, validators) / validateRequirement (generic id+text) - analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED items[] (core never assumes propose is deterministic — edge resolves via LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject - byVerification rollup; coverage.resolved = closed set (resolved+dismissed), count-preserved from the pre-re-cut engine - runProbeCli I/O scaffold (injected io; one bin per probe) - hybrid typing: generic params + injected runtime validators {categories, verification, requiredFieldsByVerification} (ADR-550 #5) edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/ classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS {explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped #584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green. * chore(#550): register probe-core.cjs artifact in ledgers + inventory New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled from src/probe-core.cts) needs registering in every artifact ledger: - .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source, never the emitted .cjs. - scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs is missing on a clean checkout. - docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the edge-probe.cjs row updated to reflect it is now the first probe-core adapter. - docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write). probe-core.test.cjs is a single test file (under the 2-file cap), so no lint-test-file-count allowlist entry is needed. * docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted] trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment 2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on PR #584 alongside the probe-core extraction (Decision 7) it governs, so the contract and its first implementation arrive together. Decisions: probe packaging (3 layers); recall->precision protocol; prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions: (truths untouched, no polarity); tiered verification test|judgment (judgment = mode-dependent soft-gate-with-flags, never silent pass / never hard-halt); CI tests the contract not the classifier; secure-phase ownership seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e), which this PR implements. * feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits: - validateResolution now enforces the 'verification is null unless resolved' invariant for EVERY status (not just resolved): a dismissed/unresolved resolution carrying a verification tier is rejected instead of merging verbatim. - An unresolved resolution carrying a resolution/reason payload is rejected (was silently dropped into the unresolved count). - runProbeCli structurally validates the report an adapter returns before writing it (was: any malformed object stringified as green output) — fails closed → exit 2. - coverage.resolved kept count-preserved (closed set) per the blessed migration contract; a new test locks that an all-dismissed run is NOT affirmatively covered (byVerification is the honest gate). Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0. * docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics Re-review #5 (trek-e) clarity edits: - Decision 5: annotate that only contract item (a) ships on #584 (the edge adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences. - Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)' parenthetical — the blessed/implemented semantics are count-preserved = the CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification carrying the per-tier resolved-status breakdown. The old parenthetical contradicted the shipped count. * test(#550): cover runProbeCli structural-guard numeric-count branch Second-pass coverage audit found the 'coverage object present but counts non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised — the {nope:true} malformed test fails earlier at the items[] check. Add a report with well-formed items[] + a coverage object carrying non-numeric counts so the numeric branch is hit. No source change; closes the line gap. * fix(#550): reject edge requirement with missing/empty text when no shapes override (M2) The edge adapter's `text` is the classification signal and a required field, but core `validateRequirement` left it optional, so a `{ id }` requirement classified to zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the exact fail-open this feature exists to eliminate. Reject missing/empty text unless an authored `shapes` override (incl. `[]`) opts out of prose classification. * fix(#550): validate verbatim items in analyzeCoverage shared seam (m1) A proposed item with no matching author resolution is rolled up VERBATIM, but its own status/fields were never validated -- an item carrying an out-of-enum status (the dropped "covered") or `dismissed` without a reason would be counted closed. The edge adapter only proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated items that arrive populated. An Item is structurally a superset of a Resolution, so reuse validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam. * fix(#550): move edge-coverage lift instruction to runtime planner surface (M1) templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline ), so the precise covered/backstop -> must_haves.truths lift instruction this PR added there never reached the planner -- and the RR-02 contract test asserted it in that dead file, giving false green. Move the instruction into plan-phase.md's runtime block (where the rest of the wire already lives), revert the dead-template edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning. * test(#550): lock machine<->SPEC vocabulary mapping against drift (m2) The machine contract uses orthogonal status x verification; the SPEC table renders a flat covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping table) -- a rename or remap now fails the suite. * docs(#550): add how-to for resolving edge-coverage findings (B1) Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts linked out to the reference. Register it in the docs/how-to index. * docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1) trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and '### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam contract, placed beside the Research Module feature-seam entries. * test(#550): add fast-check property suite for probe-core analyzeCoverage (N2) trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a transformation/rollup module, the class the property-testing predicate covers, and fast-check is already a dependency with an established *.property.test.cjs pattern. Adds 5 properties over the algebraic invariants: closed-set identity (applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier recount + resolved-status-only counting, rollup determinism, and stable orphan rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs. * test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3) trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product, docs-parity) differed from CONTEXT.md's canonical exemption category 'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240). All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical category fits; each now carries a one-line justification per the ruleset format. Free-text reason — lint behavior unchanged. * test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe Rebased onto next (e4f0910d), which replaced the line-based tier-max size ratchet (#597) with the byte-based per-file baseline guard (#1074). The edge-probe feature legitimately grows two workflows: - spec-phase.md 15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe - plan-phase.md 93135 -> 94253 (+1118): covered/backstop edge lift into must_haves.truths (the live block) Both remain under their tier hard caps (plan-phase XL 98304, ~4KB headroom; spec-phase DEFAULT 40960). Growth is real inline workflow content the feature requires at that step — not eager @-import proxy gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump (superseded by the byte baseline) via rebase. * chore(#550): reconcile INVENTORY headline counts after rebase onto next Rebase onto current next dropped the prior reconcile commit (stale counts refs 68 / CLI 107). Current next + the edge-probe additions yield: - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs Rows for all three already present; only the headline counts were stale. Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg). --- .changeset/clever-eagles-jump.md | 5 + .gitignore | 2 + CONTEXT.md | 6 + docs/COMMANDS.md | 28 ++ docs/FEATURES.md | 32 ++ docs/INVENTORY-MANIFEST.json | 3 + docs/INVENTORY.md | 7 +- docs/README.md | 1 + docs/adr/550-spec-phase-probe-contract.md | 77 +++ docs/how-to/resolve-edge-coverage-findings.md | 107 +++++ eslint.config.mjs | 2 + .../01-round-half-even/expected-coverage.json | 7 + .../01-round-half-even/requirements.json | 1 + .../02-merge-intervals/expected-coverage.json | 8 + .../02-merge-intervals/requirements.json | 1 + .../expected-coverage.json | 7 + .../03-truncate-graphemes/requirements.json | 1 + .../04-money-rounding/expected-coverage.json | 7 + .../04-money-rounding/requirements.json | 1 + .../05-list-dedupe/expected-coverage.json | 8 + .../05-list-dedupe/requirements.json | 1 + .../06-resolved-mixed/expected-coverage.json | 8 + .../06-resolved-mixed/requirements.json | 1 + .../06-resolved-mixed/resolutions.json | 4 + gsd-core/references/edge-probe.md | 261 +++++++++++ gsd-core/templates/spec.md | 12 + gsd-core/workflows/plan-phase.md | 11 + gsd-core/workflows/spec-phase.md | 129 +++++ scripts/lint-test-file-count.allowlist.json | 9 + src/edge-probe.cts | 217 +++++++++ src/probe-core.cts | 336 +++++++++++++ tests/edge-probe-docs-fixtures.test.cjs | 115 +++++ tests/edge-probe-planner-contract.test.cjs | 167 +++++++ tests/edge-probe-spec-phase-contract.test.cjs | 193 ++++++++ tests/edge-probe.test.cjs | 441 ++++++++++++++++++ tests/probe-core.property.test.cjs | 167 +++++++ tests/probe-core.test.cjs | 335 +++++++++++++ tests/workflow-size-baseline.json | 4 +- 38 files changed, 2718 insertions(+), 4 deletions(-) create mode 100644 .changeset/clever-eagles-jump.md create mode 100644 docs/adr/550-spec-phase-probe-contract.md create mode 100644 docs/how-to/resolve-edge-coverage-findings.md create mode 100644 gsd-core/references/edge-probe-fixtures/01-round-half-even/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/01-round-half-even/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/02-merge-intervals/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/02-merge-intervals/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/04-money-rounding/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/04-money-rounding/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/05-list-dedupe/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/05-list-dedupe/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/06-resolved-mixed/expected-coverage.json create mode 100644 gsd-core/references/edge-probe-fixtures/06-resolved-mixed/requirements.json create mode 100644 gsd-core/references/edge-probe-fixtures/06-resolved-mixed/resolutions.json create mode 100644 gsd-core/references/edge-probe.md create mode 100644 src/edge-probe.cts create mode 100644 src/probe-core.cts create mode 100644 tests/edge-probe-docs-fixtures.test.cjs create mode 100644 tests/edge-probe-planner-contract.test.cjs create mode 100644 tests/edge-probe-spec-phase-contract.test.cjs create mode 100644 tests/edge-probe.test.cjs create mode 100644 tests/probe-core.property.test.cjs create mode 100644 tests/probe-core.test.cjs diff --git a/.changeset/clever-eagles-jump.md b/.changeset/clever-eagles-jump.md new file mode 100644 index 000000000..07b9f0f20 --- /dev/null +++ b/.changeset/clever-eagles-jump.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 584 +--- +spec-phase: spec-completeness edge-probe — a taxonomy-driven Step 5.5 that walks each SPEC requirement against a closed 8-category edge taxonomy (boundary, adjacency, empty/degenerate, encoding, ordering, precision, idempotency, concurrency), proposes concrete candidate edges, and resolves each to covered/dismissed/backstop/unresolved. covered edges add acceptance criteria the planner lifts into must_haves.truths; a soft gate flags unresolved edges. Additive and optional: existing SPECs without an Edge Coverage section remain valid. diff --git a/.gitignore b/.gitignore index 3cca0f941..63f31d212 100644 --- a/.gitignore +++ b/.gitignore @@ -71,6 +71,8 @@ build/ /gsd-core/bin/lib/research-provider.cjs /gsd-core/bin/lib/package-legitimacy.cjs /gsd-core/bin/lib/semver-compare.cjs +/gsd-core/bin/lib/edge-probe.cjs +/gsd-core/bin/lib/probe-core.cjs /gsd-core/bin/lib/config-types.cjs /gsd-core/bin/lib/cli-exit.cjs /gsd-core/bin/lib/code-review-flags.cjs diff --git a/CONTEXT.md b/CONTEXT.md index 330972b03..63a45ea66 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -204,6 +204,12 @@ The GSD-RESEARCH capability behind an L2-hybrid seam: code owns cache + provider ### UAT-Passed Predicate Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fields with markdown-aware parsing that ignores false-positive contexts (frontmatter body, fenced code, HTML comments, blockquotes). Returns `passed: true` only when all required checks pass; supports `--require-verification` to demand at least one VERIFICATION.md file alongside UAT results. Output envelope: `{ passed, uat_files[], verification_files[], checks[], blockers[], policy }`. Source: `gsd-core/bin/lib/uat-predicate.cjs` (generated from `src/uat-predicate.cts`). Wired via `phase uat-passed` alias → `phase-command-router` → `cmdPhaseUatPassed`. +### Probe Core Module +Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the prohibition probe (#644) next. Exports: `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli`. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. + +### Edge Probe Module +First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the 8 edge categories), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. + ### MVP Mode Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`. diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index 514949128..3ba6e9dac 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -90,6 +90,34 @@ Manage GSD workspaces — create, list, or remove isolated workspace environment --- +### `/gsd-spec-phase` + +Clarify WHAT a phase delivers through Socratic questioning with quantitative ambiguity scoring, then probe for omitted edges. Produces `SPEC.md` before discuss-phase. + +| Argument | Required | Description | +|----------|----------|-------------| +| `N` | Yes | Phase number | + +| Flag | Description | +|------|-------------| +| `--auto` | Skip interactive questions; Claude selects recommended defaults and writes SPEC.md | +| `--text` | Use plain-text numbered lists instead of TUI menus (required for `/rc` remote sessions) | + +**Position in workflow:** `spec-phase → discuss-phase → plan-phase → execute-phase → verify` + +**Edge Coverage (Step 5.5):** After the ambiguity gate passes, spec-phase runs an edge-completeness probe over each requirement. It raises only applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency), proposes one concrete candidate edge per category, and records each as `covered` / `dismissed` (reason required) / `backstop` / `unresolved` in a `## Edge Coverage` SPEC section. Unresolved applicable edges soft-gate the spec (Resolve / Write-anyway-flagged / Keep-probing); `covered` and `backstop` edges are later lifted into plan-phase `must_haves`. Under `--auto` the probe **never auto-dismisses** — it auto-covers where a defensible acceptance criterion exists, otherwise auto-backstops. + +**Prerequisites:** `.planning/ROADMAP.md` exists +**Produces:** `{phase}-SPEC.md` (with a `## Edge Coverage` section) + +```bash +/gsd-spec-phase 1 # Interactive spec + edge probe for phase 1 +/gsd-spec-phase 3 --auto # Auto-select defaults; never auto-dismisses an edge +/gsd-spec-phase 2 --text # Plain-text menus for remote sessions +``` + +--- + ### `/gsd-discuss-phase` Gather phase context through adaptive questioning before planning. diff --git a/docs/FEATURES.md b/docs/FEATURES.md index f86733150..5e8120e99 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -165,6 +165,7 @@ - [Milestone Tag Creation Toggle](#141-milestone-tag-creation-toggle) - [Structured JSON Error Mode](#142-structured-json-error-mode) - [UAT-Passed Predicate](#143-uat-passed-predicate) + - [Spec-Phase Edge-Completeness Probe](#144-spec-phase-edge-completeness-probe) --- @@ -3086,3 +3087,34 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev - [docs index](README.md) **Reference:** [JSON Error Mode](json-errors.md) + +--- + +### 144. Spec-Phase Edge-Completeness Probe + +**Command:** `/gsd-spec-phase` + +**Purpose:** Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as `Step 5.5` of spec-phase, after the ambiguity gate. + +**Behavior:** For each SPEC requirement the probe classifies its data/behavior shape, then raises only the *applicable* categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states: + +| State | Meaning | Downstream effect | +|-------|---------|-------------------| +| `covered` | An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into `plan-phase` `must_haves.truths` | +| `dismissed` | The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected | +| `backstop` | Intent recorded, needs a held-out/property-based test | Lifted into `must_haves.truths` as a non-inferable check | +| `unresolved` | Deferred | Soft-gates the spec; row stamped `⚠ Edge unresolved — planner must treat as assumption` | + +The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. + +The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. + +**Requirements:** +- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a `## Edge Coverage` SPEC section. +- REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved. +- REQ-EDGE-03: A `dismissed` resolution MUST require a non-empty reason. +- REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption. +- REQ-EDGE-05: `--auto` MUST never auto-dismiss — auto-cover or auto-backstop only. +- REQ-EDGE-06: `plan-phase` MUST lift `covered` criteria and `backstop` notes into `must_haves.truths`. + +**Reference:** [Edge Probe](../gsd-core/references/edge-probe.md) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 60757e653..8091c1fb9 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -209,6 +209,7 @@ "decimal-phase-calculation.md", "doc-conflict-engine.md", "domain-probes.md", + "edge-probe.md", "execute-mvp-tdd.md", "executor-examples.md", "gate-prompts.md", @@ -295,6 +296,7 @@ "decisions.cjs", "docs.cjs", "drift.cjs", + "edge-probe.cjs", "fallow-runner.cjs", "federated-config.cjs", "frontmatter.cjs", @@ -329,6 +331,7 @@ "phases-command-router.cjs", "plan-scan.cjs", "planning-workspace.cjs", + "probe-core.cjs", "profile-output.cjs", "profile-pipeline.cjs", "project-root.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index ad483376c..38ed3a3cd 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -264,7 +264,7 @@ Full roster at `gsd-core/workflows/*.md`. Workflows are thin orchestrators that --- -## References (68 shipped) +## References (69 shipped) Full roster at `gsd-core/references/*.md`. References are shared knowledge documents that workflows and agents `@-reference`. The groupings below match [`docs/ARCHITECTURE.md`](ARCHITECTURE.md#references-gsd-corereferencesmd) — core, workflow, thinking-model clusters, and the modular planner decomposition. @@ -300,6 +300,7 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum | `context-budget.md` | Context window budget allocation rules. | | `continuation-format.md` | Session continuation/resume format. | | `domain-probes.md` | Domain-specific probing questions for discuss-phase. | +| `edge-probe.md` | Spec-phase edge-completeness probe — 8-category edge taxonomy, shape classification, and the `requirements → checks → verifier` resolution model (Step 5.5). | | `gate-prompts.md` | Gate/checkpoint prompt templates. | | `scout-codebase.md` | Phase-type→codebase-map selection table for discuss-phase scout step (extracted via #2551). | | `revision-loop.md` | Plan revision iteration patterns. | @@ -371,7 +372,7 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t --- -## CLI Modules (107 shipped) +## CLI Modules (109 shipped) Full listing: `gsd-core/bin/lib/*.cjs`. @@ -406,6 +407,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `decisions.cjs` | Parses CONTEXT.md `` blocks; accepts numeric (D-42) and alphanumeric (D-INFRA-01) IDs; returns `{id, text, category, tags, trackable}` | | `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection | | `drift.cjs` | Post-execute codebase structural drift detector (#2003): classifies file changes into new-dir/barrel/migration/route categories and round-trips `last_mapped_commit` frontmatter | +| `edge-probe.cjs` | Spec-completeness edge probe (compiled from `src/edge-probe.cts`, gitignored) — the first adapter of the `probe-core` resolution model (ADR-550 Decision 7): shape classification, applicable-category relevance filter, edge proposal, and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`; exports `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `TAXONOMY` (#550) | | `fallow-runner.cjs` | Fallow audit adapter for `/gsd-code-review`: binary resolution (`PATH` then `node_modules/.bin`), actionable missing-binary errors, and structural findings normalization | | `federated-config.cjs` | Defensive merge of capability-declared config slices into the loadConfig return value — ADR-857 phase 3b; exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig })` → `{ values, validKeys, warnings }`; no-op until a key is atomically removed from the central config-schema (the cutover step) | | `frontmatter.cjs` | YAML frontmatter CRUD operations | @@ -446,6 +448,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `prompt-budget.cjs` | Pure token-budget accounting for review prompts — estimates tokens, applies deterministic trim priority (head-shrink PROJECT.md, proportional plan truncation, drop context/research/requirements, hard-fail guard), returns structured metadata for `review.max_prompt_tokens` (#3081) | | `research-provider.cjs` | Research provider waterfall, confidence tiers, and planResearch (cache-hits + fetch plan) | | `research-store.cjs` | Content-addressed research cache: sha256 keys, per-source TTL staleness, two-tier (user ~/.gsd / project .planning) store | +| `probe-core.cjs` | Generic spec-phase probe resolution model (compiled from `src/probe-core.cts`, gitignored; ADR-550 Decision 7) — the status×verification re-cut (`status: resolved/dismissed/unresolved` × per-probe `verification`), `validateResolution`/`validateRequirement`, `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject, the `byVerification` rollup, and the `runProbeCli` I/O scaffold; the shared seam consumed by `edge-probe` (and the prohibition probe #644); exports `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` (#550) | | `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence | | `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` | | `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) | diff --git a/docs/README.md b/docs/README.md index aebdd7d81..7d169e842 100644 --- a/docs/README.md +++ b/docs/README.md @@ -18,6 +18,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md) - [Install on your runtime](how-to/install-on-your-runtime.md) — runtime-specific install steps for all 16 supported runtimes - [Install a minimal GSD and add skills later](how-to/install-minimal-and-add-skills.md) — install only the core skills, then grow the surface with profiles and `/gsd:surface` - [Discuss a phase](how-to/discuss-a-phase.md) — capture implementation decisions before planning begins +- [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions - [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality - [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents - [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md new file mode 100644 index 000000000..422cda73c --- /dev/null +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -0,0 +1,77 @@ +# ADR 550: spec-phase probe pattern and prohibition contract [Accepted] + +- **Status:** Accepted +- **Date:** 2026-06-03 + +> **Provenance.** Drafted during triage of #644 (prohibition probe), hardened in a `/grill-with-docs` session, and extended with a `probe-core` seam designed in an `/improve-codebase-architecture` grill (Decision 7). Endorsed by the maintainer and by the #644 author (who also authored the edge-probe, #550 / PR #584). This ADR lands **on PR #584** alongside the `probe-core` extraction that implements Decision 7; #644 (the prohibition probe itself) remains blocked until #584 merges and then implements the prohibition adapter. "What exists today" is verified against `next` + the PR-#584 head as of 2026-06-03. +> +> **Lineage.** The judgment-tier contract (Decision 4) was reached independently from two directions: from the domain model (truths are unenforced positive-observables, so a values prohibition cannot be a silent pass) and from the author's "verifier-abstention" (N17) calibration experiment (on irreducible values rules the verifier is confidently wrong — ~0.93 confidence, ECE 0.81 — so confidence-gating cannot help; the only honest move is abstain-and-flag). That parked experiment folds into this ADR as the verify-time half of judgment-tier rather than needing a separate feature. + +## Context + +`spec-phase` produces `SPEC.md` from an interview. Two features add *probes* — soft-gate steps that surface spec gaps before code is written: **#550 / PR #584 (edge-probe)** for data-shape edges, and **#644 (prohibition probe)** for values/safety "must-NOT" constraints. Both use identical three-layer packaging and a recall→precision protocol, and both want confirmed findings to reach `SPEC.md` and the plan contract. + +Grounding the design against the codebase surfaced the decisive facts: + +- **`truths` are positive, observable assertions.** `gsd-planner` derives them as *"observable truths (3-7, user perspective)"*; `verify-phase` checks each by *"determine if the codebase enables it"* — a presence/wiring check. A prohibition is the inverse and cannot be confirmed that way. +- **`truths` are not mechanically verified.** `verify.cjs` iterates `artifacts` and `key_links` but has no `truths` handler. A prohibition parked in `truths` would inherit that non-enforcement — a negative wearing a positive's clothing. +- **`must_haves` has no schema/type** — a convention parsed by `parseMustHavesBlock` (blocks `truths`/`artifacts`/`key_links`), read in ≥6 places. +- **The SPEC→plan lift is LLM reasoning** in `gsd-planner`'s `derive_must_haves`, not mechanical extraction. +- **Prohibitions already exist in GSD only as narrative prose** (``, `"MUST NOT loop"`), never as structured, verifiable data. +- **The edge-probe's load-bearing logic is a deep module on the live path.** `src/edge-probe.cts` (~297 lines) is compiled to `bin/lib/edge-probe.cjs` and invoked by `spec-phase` Step 5.5 at runtime. ~120 of its lines are a *generic resolution model* (status lifecycle, validation, merge/rollup, CLI); only the `Shape`/`SHAPE_CUES`/`TAXONOMY`/`classifyShape`/`proposeEdges` cluster is edge-specific. The prohibition probe is the **second adapter** of that model — making the seam real (one adapter is hypothetical; two is real). +- **No prior ADR** governs probes, soft gates, or `must_haves` evolution. + +The values/safety class #644 targets (*"must not become a guilt mechanic"*) is frequently **not reducible to a deterministic test** — which is why it needs an explicit verification contract rather than truth-style limbo. + +## Decision + +1. **Probe packaging (three layers).** Every spec-phase probe ships as: (a) a portable, dependency-free reference core under `references/.md` (taxonomy + protocol + output schema); (b) a `spec-phase.md` step run as a **soft gate** (write-anyway-with-flags), with explicit `--auto` and text-mode (non-Claude, no `AskUserQuestion`) handling, mirroring the ambiguity gate; (c) tests + docs. + +2. **Two-stage protocol.** **Recall** (adversarial over-generation) → **precision** (classifier dropping routine engineering items), surfacing a short confirmable list. Dismissals require a non-empty reason. + +3. **Prohibition home & representation.** Confirmed prohibitions live primarily as **`SPEC.md` acceptance criteria** (negative criteria). They MAY be projected into an optional **`must_haves.prohibitions:` sibling block** (alongside `truths`/`artifacts`/`key_links`) when they must survive into the plan contract. **`truths` is left untouched** — no `polarity` field is added. Each prohibition item carries `statement`, the orthogonal `status` + `verification` of Decision 7, and (when dismissed) a non-empty `reason`. + +4. **Tiered verification.** Each prohibition declares `verification: test | judgment`. + - **`test`-tier** (reducible to a deterministic assertion) → a **negative test** following `regression-must-fail-first` + negative-proof discipline. **Hard gate in both interactive and autonomous modes.** + - **`judgment`-tier** (irreducible values/safety rule) → **mode-dependent soft-gate-with-flags**: *interactive* verify requires explicit human resolution of each item; *autonomous* verify records an LLM-judge verdict **marked non-authoritative** and emits a prominent *"unverified prohibition — human review recommended"* flag in SUMMARY/verdict. **Never a silent pass; never a hard halt of AFK runs.** Autonomous completion reads *"complete with N flagged prohibitions."* + - *Optional future input:* cross-tier (or cross-model) disagreement MAY feed the judgment-tier flag as a cheap blind-spot signal. This is a breadcrumb, **not a dependency** — the flag stands without it. + +5. **The CI-testable surface is the contract, not the classifier.** The recall/precision/tier-assignment stages are LLM behavior and are **not** claimed as deterministic CI coverage (validated by offline batteries; in CI only by source-grep-under-`// allow-test-rule:`-exemption over reference/workflow prose, which *is* the product). CI deterministically tests the **contract**: (a) parse + validate an item (`statement` present; `status ∈ {resolved, dismissed, unresolved}`; `verification` in the probe's allowed set; dismissed items carry a non-empty `reason`); (b) round-trip / projection between `SPEC.md` and `must_haves.prohibitions:`; (c) a `DEFECT.GENERATIVE-FIX` **parity assertion** across template ↔ parser ↔ planner; (d) `test`-tier items are provably wired into verify-phase and cannot be silently skipped. A test that asserts the LLM's *judgment* is vacuous and is rejected per `RULESET.TESTS.delete-bad-tests`. *(Scope: on #584 only **(a)** ships — it is the edge adapter's parse+validate contract test (`tests/probe-core.test.cjs`). **(b)–(d)** require the `must_haves.prohibitions:` block, projection, and verify-phase tiering and are **#644 scope**, as the Consequences section records.)* + +6. **Ownership seam with security tooling.** The probe owns **bespoke product/values** prohibitions only. When precision classifies an item as **canon security/compliance** (OWASP/GDPR/fairness/prototype-pollution/path-traversal), it does **not** mint a SPEC prohibition; it emits a one-line breadcrumb (*"possible canon-security concern X — owned by `/gsd:secure-phase` / eslint"*) and stops. Canon checks are **referred, not duplicated** — keeping the surfaced list short (#644's ~2–3-item goal). + +7. **`probe-core` is the seam; probes are adapters.** The shared resolution model is extracted into `src/probe-core.cts`; `edge-probe.cts` refactors onto it as the first adapter and the prohibition probe is born on it as the second. Sub-decisions, each chosen against both probes: + - **7a — Orthogonal model.** The item carries two orthogonal dimensions: `status: resolved | dismissed | unresolved` (resolution lifecycle, shared) and `verification: | null` (edge: `explicit | backstop`; prohibition: `test | judgment`). This replaces edge-probe's shipped `status: covered | dismissed | backstop | unresolved`, which smuggled a verification fact (`backstop`) into a lifecycle enum. Migration is mechanical (`covered → {resolved, explicit}`, `backstop → {resolved, backstop}`, others straight across); `coverage.resolved` is preserved **count-for-count** — it remains the *closed* set (`resolved` + `dismissed` status = `applicable − unresolved`), so every fixture's count is unchanged, while `coverage.byVerification` carries the per-tier `resolved`-status breakdown; the 6 edge fixtures are re-generated on #584. Done now (one adapter) rather than against a fixture-locked enum later. + - **7b — Core ingests already-proposed items.** The seam is `analyzeCoverage(items, resolutions?, validators)`, **not** `(requirements, proposeFn, …)`. The probes have different deterministic surfaces — edge = deterministic propose (`proposeEdges`) + LLM resolve; prohibition = **LLM propose** (adversarial recall) + deterministic validate/merge/rollup — so core MUST NOT assume propose is deterministic. `proposeEdges` stays in the edge adapter; candidate-parsing stays in the prohibition adapter. + - **7c — Hybrid typing.** Generic type params on the exported interfaces for adapter-authoring DX, but the load-bearing enforcement is **injected runtime validators** (`{ categories, verification, requiredFieldsByVerification }`), because the probe runs as a CLI over JSON where TS types are erased. The contract test (Decision 5) pins the validators, not the types. + - **7d — Rollup carries a verification tally.** `CoverageReport.coverage` gains `byVerification: { : count }`, computed generically in core over items that have a verification set. Verify-phase reads `byVerification.judgment` as Decision 4's denominator without re-scanning. (Unresolved items carry no tier and are counted by `status`.) + - **7e — One bin per probe.** Each probe ships its own compiled bin (`edge-probe.cjs`, `prohibition-probe.cjs`) calling a shared `runProbeCli(...)`. A single dispatcher CLI is **deferred** as an independent follow-on: unlike the enum, it is pure invocation plumbing with no migration debt, so it does not justify enlarging the #584 blast radius. + + Indicative `probe-core` surface: + ```ts + export type Status = 'resolved' | 'dismissed' | 'unresolved'; + export interface Item { + requirement_id: string; category: TCat; status: Status; + verification: TVer | null; resolution: string | null; reason: string | null; + } + export interface CoverageReport { + items: Item[]; + coverage: { applicable: number; resolved: number; unresolved: number; byVerification: Record }; + } + export interface ProbeValidators { + categories: ReadonlySet; verification: ReadonlySet; + requiredFieldsByVerification?: Record>; + } + export function analyzeCoverage(items: Item[], resolutions: Resolution[] | undefined, v: ProbeValidators): CoverageReport; + export function validateResolution(r: Resolution, items: Item[], v: ProbeValidators): void; + export function validateItems(items: Item[], v: ProbeValidators): void; + export function runProbeCli(opts: { ingest: (argv: string[]) => { items: Item[]; resolutions?: Resolution[] }; validators: ProbeValidators }): never; + ``` + +## Consequences + +- **Positive:** `truths` keeps its positive-observable semantics; prohibitions get a first-class home with a real verification lifecycle; CI is honest (tests the contract, never fakes the model's judgment as green); no fake-green and no gutted autonomy; the secure-phase boundary is explicit and de-duplicated; the resolution model lives in one place, so the third probe is nearly free and `verification`/`status` cannot drift across the ≥6 `must_haves` read sites. +- **Costs / required work (on #584):** extract `probe-core.cts` and refactor `edge-probe.cts` onto it; re-cut the status enum into `status × verification` and re-generate the 6 edge fixtures. **On the #644 PR:** add the prohibition reference + the `must_haves.prohibitions:` block and its `parseMustHavesBlock` extension; the `DEFECT.GENERATIVE-FIX` parity assertion across template ↔ parser ↔ planner (**owned by the #644 author**); extend the SUMMARY/verdict schema with the `unverified-prohibition` flag; teach verify-phase the test-tier wiring and judgment-tier mode-dependent handling. A deterministic `projectProhibitions()` in `probe-core` is recommended so the parity assertion tests a function's round-trip rather than a prompt; every existing `must_haves` read site must tolerate the new block (absence = no prohibitions). +- **Docs debt:** `spec-phase` is undocumented in `FEATURES.md`/`COMMANDS.md` despite shipping in v1.38; the docs-required gate bites both probes. PR #584 establishes the spec-phase docs section. +- **Glossary:** `probe family`, `probe-core`, `verification tier`, and `bespoke vs canon prohibition` are added to `CONTEXT.md` when #584 merges (not while the contract is pre-merge); this ADR is their interim home. +- **Sequencing:** PR #584 lands ADR 550 + `probe-core` + the `edge-probe.cts` refactor + enum re-cut + fixture re-gen + the spec-phase docs section. #644 then adds the prohibition adapter (reference + `prohibitions:` block/parser/parity + tiered verification). #644 stays blocked until #584 merges. diff --git a/docs/how-to/resolve-edge-coverage-findings.md b/docs/how-to/resolve-edge-coverage-findings.md new file mode 100644 index 000000000..fed067e19 --- /dev/null +++ b/docs/how-to/resolve-edge-coverage-findings.md @@ -0,0 +1,107 @@ +# How to resolve edge-coverage findings while writing a spec + +**Goal:** Turn each domain-boundary edge the spec phase surfaces into an explicit, verifiable spec decision — so omitted boundaries (rounding ties, touching ranges, grapheme truncation) become checkable requirements *before* any code exists, instead of silent blind spots the verifier is confidently wrong about. + +**Prerequisites:** A phase whose `/gsd-spec-phase` run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically and presents its findings — you do not invoke it separately. + +For the category taxonomy and the reasoning behind front-of-pipeline edge analysis, see [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe). This guide covers only how to *act* on the findings. + +--- + +## Read a finding + +Each finding is one **applicable** edge for one requirement — a boundary the probe's relevance filter decided is in scope for that requirement's shape, and that you have not yet addressed. A finding is phrased as a probe question, for example: + +> **R3 · precision** — Where can precision loss or overflow occur, and what is the contract? + +You must resolve each finding into exactly one of four states. Claude presents them as a numbered choice (or an `AskUserQuestion` menu). + +--- + +## Specify it — write an acceptance criterion + +**Choose this when you can state the correct behaviour as a pass/fail check.** This is the strongest resolution: it produces a concrete assertion the planner and verifier can enforce. + +Claude writes a new line into the spec's **Acceptance Criteria** and marks the edge `covered`. For the `precision` finding above: + +> - [ ] Monetary amounts round half-to-even to 2 decimal places; `2.005 → 2.00` + +Prefer this state whenever a defensible criterion can be written. A `covered` edge is lifted into the plan's `must_haves.truths`, so the verifier checks it. + +--- + +## Dismiss it — record why it does not apply + +**Choose this when the edge genuinely cannot occur** — and say why. A dismissal **requires a non-empty reason**; silence is rejected. The reason is the audit trail. + +> ⛔ dismissed — input is a bounded enum; no boundary value exists + +A wrong dismissal is the exact silent failure this probe exists to prevent, so dismiss only when the reason is solid. + +--- + +## Backstop it — defer the contract to a held-out test + +**Choose this when you know the edge matters but cannot fully articulate the correct behaviour in prose yet.** Claude marks the edge `backstop` and notes that a held-out / property-based test stands in for the missing assertion. + +> 🧪 backstop — held-out property test: dedupe output order is stable under input permutation + +A `backstop` edge is also carried into the plan's `must_haves.truths` as a non-inferable check — the planner records the intent, and the test body is authored later during execution. + +--- + +## Defer it — leave it unresolved and flagged + +**Choose this only when you are not ready to decide.** The edge stays `unresolved` and is flagged. Unlike a dismissal, deferring makes no claim that the edge is safe — it is an explicit, visible assumption the planner must surface, not silently drop. + +--- + +## Clear the soft gate + +After you have worked through the findings, the probe runs a **soft gate**: + +- **All applicable edges resolved** → the spec proceeds to the next step. +- **One or more still `unresolved`** → Claude asks what to do: + - **Resolve now** — loop back and resolve the remaining edges. + - **Write the spec anyway** — the spec is written with those rows marked `⚠ Edge unresolved — planner must treat as assumption`. Use this deliberately; you are choosing to ship a known gap. + +The gate is *soft*: it never blocks you, but every unresolved edge remains visible in the spec's `## Edge Coverage` section. + +--- + +## Let Claude resolve them for you + +**If the edges are low-stakes or already implied by earlier phases**, run the spec phase in auto mode: + +```bash +/gsd-spec-phase 3 --auto +``` + +In `--auto`, Claude marks an edge `covered` where it can write a defensible acceptance criterion, and `backstop` otherwise. It **never auto-dismisses** — dismissing an edge requires a human reason, because a wrong auto-dismissal is precisely the silent failure being eliminated. Claude logs the tally, for example: + +``` +[auto] edge coverage: 4 covered, 2 backstop, 1 unresolved +``` + +Review the logged choices afterwards; auto mode is a fast first pass, not a substitute for judgement on edges that carry risk. + +--- + +## What happens to resolved findings downstream + +When you next run `/gsd-plan-phase`, the planner reads the spec's `## Edge Coverage` section and: + +- lifts every `covered` edge's acceptance criterion into `must_haves.truths`, +- carries every `backstop` edge into `must_haves.truths` as a non-inferable check (needing a held-out/property test), +- surfaces every `unresolved` edge as an explicit assumption. + +This is the payoff: a resolved edge becomes a unit the goal-backward verifier actually checks, extending its reach to boundaries the requirement prose never stated. + +--- + +## Related + +- [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe) — taxonomy, output schema, and the front-of-pipeline rationale +- [`/gsd-spec-phase`](../COMMANDS.md#gsd-spec-phase) — command reference and flags +- [Plan a phase](plan-a-phase.md) — where `covered`/`backstop` edges become `must_haves` +- [docs index](../README.md) diff --git a/eslint.config.mjs b/eslint.config.mjs index 65882ac86..579dfa5f1 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -36,6 +36,8 @@ export default tseslint.config( // ADR-457: tsc-generated runtime artifact — lint the src/*.cts source, not the emitted .cjs. 'gsd-core/bin/lib/semver-compare.cjs', 'gsd-core/bin/lib/cli-exit.cjs', + 'gsd-core/bin/lib/edge-probe.cjs', + 'gsd-core/bin/lib/probe-core.cjs', 'gsd-core/bin/lib/code-review-flags.cjs', 'gsd-core/bin/lib/context-utilization.cjs', 'gsd-core/bin/lib/artifacts.cjs', diff --git a/gsd-core/references/edge-probe-fixtures/01-round-half-even/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/01-round-half-even/expected-coverage.json new file mode 100644 index 000000000..87c90f32f --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/01-round-half-even/expected-coverage.json @@ -0,0 +1,7 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" }, + { "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/01-round-half-even/requirements.json b/gsd-core/references/edge-probe-fixtures/01-round-half-even/requirements.json new file mode 100644 index 000000000..f3ea92b83 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/01-round-half-even/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Round a number to N decimal places, rounding half to nearest" }] diff --git a/gsd-core/references/edge-probe-fixtures/02-merge-intervals/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/02-merge-intervals/expected-coverage.json new file mode 100644 index 000000000..b06b24b63 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/02-merge-intervals/expected-coverage.json @@ -0,0 +1,8 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/02-merge-intervals/requirements.json b/gsd-core/references/edge-probe-fixtures/02-merge-intervals/requirements.json new file mode 100644 index 000000000..4e013e47e --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/02-merge-intervals/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }] diff --git a/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/expected-coverage.json new file mode 100644 index 000000000..19f3e07b3 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/expected-coverage.json @@ -0,0 +1,7 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/requirements.json b/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/requirements.json new file mode 100644 index 000000000..4f3412080 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/03-truncate-graphemes/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Truncate a display string to its first N characters" }] diff --git a/gsd-core/references/edge-probe-fixtures/04-money-rounding/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/04-money-rounding/expected-coverage.json new file mode 100644 index 000000000..87c90f32f --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/04-money-rounding/expected-coverage.json @@ -0,0 +1,7 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" }, + { "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/04-money-rounding/requirements.json b/gsd-core/references/edge-probe-fixtures/04-money-rounding/requirements.json new file mode 100644 index 000000000..fa8326420 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/04-money-rounding/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Compute the total price as an amount rounded to two decimals" }] diff --git a/gsd-core/references/edge-probe-fixtures/05-list-dedupe/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/05-list-dedupe/expected-coverage.json new file mode 100644 index 000000000..b06b24b63 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/05-list-dedupe/expected-coverage.json @@ -0,0 +1,8 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/05-list-dedupe/requirements.json b/gsd-core/references/edge-probe-fixtures/05-list-dedupe/requirements.json new file mode 100644 index 000000000..559f96c9d --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/05-list-dedupe/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Return all items in the list with duplicates removed" }] diff --git a/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/expected-coverage.json b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/expected-coverage.json new file mode 100644 index 000000000..4323f87ad --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/expected-coverage.json @@ -0,0 +1,8 @@ +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } } +} diff --git a/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/requirements.json b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/requirements.json new file mode 100644 index 000000000..4e013e47e --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/requirements.json @@ -0,0 +1 @@ +[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }] diff --git a/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/resolutions.json b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/resolutions.json new file mode 100644 index 000000000..8336ef620 --- /dev/null +++ b/gsd-core/references/edge-probe-fixtures/06-resolved-mixed/resolutions.json @@ -0,0 +1,4 @@ +[ + { "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null }, + { "requirement_id": "R1", "category": "ordering", "status": "dismissed", "resolution": null, "reason": "output is canonically sorted; no tie possible" } +] diff --git a/gsd-core/references/edge-probe.md b/gsd-core/references/edge-probe.md new file mode 100644 index 000000000..acac6068e --- /dev/null +++ b/gsd-core/references/edge-probe.md @@ -0,0 +1,261 @@ +# Edge-Probe — Spec-Completeness Reference + +Shared reference for the spec/requirements phase. Companion to +`@~/.claude/gsd-core/references/domain-probes.md`: `domain-probes` covers the +**technology axis** (auth, search, caching, deployment); this covers the +**data/behavior-shape axis** (boundaries, adjacency, encoding, ordering…). Walk each +requirement against the closed taxonomy below, propose a concrete candidate edge for each +applicable category, and resolve each to exactly one state. Adopt the established QA names +verbatim — this is decades-old black-box test technique, moved upstream to the spec layer. + +This doc is written in generic `requirements → checks → verifier` terms with no +tool-specific vocabulary, so it is portable: copy it into any spec/requirements process. +A short mapping table at the end binds it to common host structures. + +## Why front-of-pipeline + +A goal-backward verifier only checks assertions that exist; an assertion only exists for a +requirement that was written down. A domain-boundary edge the author never surfaced is +invisible to the verifier — and worse, the verifier is *confidently wrong* about it +(measured: ~0.93 confidence while catching 0/12 omitted-edge defects, an expected +calibration error of 0.81 — worse than a coin flip — versus 100% / 0.03 on edges the spec +did state). The fix is not a better verifier or a confidence gate; confidence is +uninformative across these regimes. The fix is **spec completeness**: surface the omitted +edge into an explicit, checkable assertion *before* any code exists, after which the +verifier reliably catches it. + +The underlying techniques are classic and should be named as such — **Boundary Value +Analysis**, **Equivalence Partitioning**, the **Category-Partition method**, and +**Metamorphic Relations**; property-based testing (PBT) is the academic name for the +held-out backstop. The literature applies these at the *test* layer (back of pipeline, +generating checks). The differentiated move here is **placement, not technique**: apply the +same edge taxonomy at the *spec* layer (front of pipeline, generating requirements), with +the explicit goal of extending a verifier's reach. Recent LLM spec work corroborates the +core finding that the spec layer is the measured weak point: + +- SLD-Spec — program slicing + logical deletion (code→spec): arXiv 2509.09917 +- Specine / AutoReSpec / SpecMind — spec alignment & postcondition inference +- CodeSpecBench (2604.12268), OSVBench (2504.20964), VERINA (2505.23135) — benchmarks + showing models solve tasks far better than they generate precise behavioral specs +- LLM property-based tests for edge cases: arXiv 2510.25297 (PBT+EBT ≈ 81% edge detection) + +## Inputs + +A list of requirements, each a `{ id, text, shapes? }` record where `text` is a testable +statement and `shapes` is an optional author-supplied override of the data/behavior shape. +The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. When +`shapes` is absent, a heuristic classifier proposes them from the requirement prose +(propose-then-confirm) — the author may correct the shape. + +## Taxonomy (8 categories) + +Closed and small by design: a fixed eight the author must explicitly clear beats thirty +nobody finishes. The failure mode being eliminated was never "too few categories" — it was +that no taxonomy was *systematically applied at all*. Categories 1, 2, and 4 alone cover +the three canonical corpus defects (banker's-rounding ties, touching intervals, grapheme +truncation). Growth happens via optional domain packs, not by bloating the core. + +| id | name (QA term) | applies to shapes | probe question | +|----|----------------|-------------------|----------------| +| boundary | Boundary values | numeric-range | What happens exactly at each min/max/threshold — and one step either side? | +| adjacency | Adjacency / touching | collection | When two things are exactly equal or just touch, do they merge, collide, or separate? | +| empty | Empty / degenerate | collection, text | What is the result for empty, single-element, or null input? | +| encoding | Encoding / representation | text | Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form? | +| ordering | Ordering / stability | collection | When elements compare equal, is output order specified and stable? | +| precision | Precision / overflow | numeric-range | Where can precision loss or overflow occur, and what is the contract? | +| idempotency | Idempotency / repetition | stateful | What happens if this runs twice on the same input? | +| concurrency | Concurrency / effect ordering | stateful, io | If interrupted or run in parallel, what is guaranteed? | + +## Relevance filter + resolution states + +Two rules keep the probe honest and prevent an "everything is N/A" failure mode: + +1. **Relevance filter first.** Classify each requirement's shape, then raise only the + categories whose `applies to shapes` intersect that requirement's shapes. A pure-text + requirement is never asked about overflow. This is what makes an unresolved edge + meaningful: it is an edge that *applies* and was not addressed. +2. **Dismissal requires a reason string.** "N/A — input is a bounded enum, no boundary + exists" is valid; silence is not. The reason string is the audit trail. + +Each raised edge carries two orthogonal axes — a resolution **lifecycle** and, when +resolved, a **verification** tier (ADR-550 Decision 7, the shared probe-core model): + +- **status** — `resolved | dismissed | unresolved`: + - **resolved** — the edge is addressed; *how* it is addressed is the verification tier. + - **dismissed** — not applicable, accompanied by a required, non-empty reason string. + - **unresolved** — carried forward and flagged; the author chose not to resolve it yet. +- **verification** (only when `status` is `resolved`; `null` otherwise) — `explicit | backstop`: + - **explicit** — a checkable assertion for the edge is written (a SPEC acceptance criterion). + - **backstop** — a held-out / property-based test stands in for an edge the author knows + but cannot fully articulate in prose (records intent; the test body is authored later). + +Splitting these axes keeps the lifecycle enum free of a verification fact (the old single +enum smuggled `covered`/`backstop` — both *resolved* — into one flat list) and lets sibling +probes add their own verification tiers (e.g. `test | judgment`) without a parallel enum. + +`coverage.resolved` is the count of **closed** edges — `resolved` + `dismissed` (the +pre-re-cut "covered + dismissed + backstop" set, count-preserved). An `unresolved` +*applicable* edge is the precise signal a soft completeness gate raises. + +## Output schema + +The probe emits, per edge, an item of the form: + +``` +{ requirement_id, category, status, verification, resolution, reason, probe } +``` + +plus a coverage summary: + +``` +coverage: { applicable, resolved, unresolved, byVerification: { explicit, backstop } } +``` + +`applicable` is the number of raised edges, `resolved` = closed (`resolved` + `dismissed`) +status edges, `unresolved` is the remainder, and `byVerification` breaks the +`resolved`-status edges down by tier (probe-agnostic in core; the edge adapter declares +`{ explicit, backstop }`). This JSON is the stable contract both the reference +implementation and any third-party port emit. + +## Generic mapping (requirements → checks → verifier) + +| Host structure | "requirement" | a `resolved`/`explicit` edge becomes | a `resolved`/`backstop` edge becomes | +|----------------|---------------|--------------------------|---------------------------| +| GSD SPEC | a SPEC Requirement | an Acceptance Criterion that `plan-phase` lifts into `must_haves.truths` | a non-inferable check in `must_haves.truths` (needs a held-out/PBT test) | +| Gherkin feature | a Scenario | an additional `Then` assertion / Scenario Outline row | a tagged scenario routed to a property test | +| OpenAPI operation | an operation | a response/constraint example + schema rule | a contract/property test on the operation | +| Docstring contract | a documented behavior | an assertion in the contract test | a property-based test for the function | + +The portable invariant: a `resolved`/`explicit` edge produces **the unit your verifier +iterates over** (GSD: a `must_haves.truth`); a `resolved`/`backstop` edge produces a test +added to that same set as a non-inferable check. An `unresolved` edge is an explicit +assumption the downstream planner must surface, not silently drop. + +## Worked example (merge-intervals) + +Given a single requirement with no resolutions yet: + +``` +[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }] +``` + +the requirement classifies as a `collection`, which raises `adjacency`, `empty`, and +`ordering` (but not `boundary`/`precision`/`encoding`/`idempotency`/`concurrency`). With no +resolutions supplied, every applicable edge is `unresolved`: + +```json edge-probe:02-merge-intervals/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } } +} +``` + +The `adjacency` row is the one that catches the canonical defect: `[[1,2],[2,3]]` intervals +that only *touch* — does the spec say they merge? Resolving it `resolved`/`explicit` writes +that assertion, and the verifier can then enforce it. This worked-example block is kept +byte-for-byte (parsed-JSON) identical to its fixture by `edge-probe-docs-fixtures.test.cjs`, +so the doc and the reference implementation cannot silently drift. + +## Worked example (round-half-even) + +Given a requirement to round floating-point values using the banker's-rounding (half-even) +rule, the requirement classifies as `numeric-range`, which raises `boundary` and `precision` +(but not `adjacency`/`empty`/`ordering`/`encoding`/`idempotency`/`concurrency`): + +```json edge-probe:01-round-half-even/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" }, + { "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} +``` + +The `precision` row surfaces the canonical defect for this requirement: IEEE 754 floating-point +arithmetic rounds 2.5 to 2 (not 3) under half-even — the probe asks whether the spec +states that contract explicitly, so the verifier can enforce it. + +## Worked example (truncate-graphemes) + +Given a requirement to truncate a string to N grapheme clusters (not bytes or code points), +the requirement classifies as `text`, which raises `empty` and `encoding`: + +```json edge-probe:03-truncate-graphemes/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} +``` + +The `encoding` row is the load-bearing one: a requirement that says "truncate to 10 +characters" is ambiguous — the spec must state whether "character" means bytes, UTF-16 +code units, Unicode code points, or grapheme clusters (emoji sequences are 1 grapheme, +multiple code points). + +## Worked example (money-rounding) + +Given a requirement to round monetary amounts to two decimal places, the requirement +classifies as `numeric-range`, which raises `boundary` and `precision`: + +```json edge-probe:04-money-rounding/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" }, + { "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" } + ], + "coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } } +} +``` + +The `boundary` row asks what the minimum and maximum representable values are (negative +amounts? fractional cents?). The `precision` row asks whether IEEE 754 binary rounding +can produce `$0.30000000000000004` — the spec must commit to a representation. + +## Worked example (list-dedupe) + +Given a requirement to deduplicate a list of items, the requirement classifies as +`collection`, which raises `adjacency`, `empty`, and `ordering`: + +```json edge-probe:05-list-dedupe/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } } +} +``` + +The `ordering` row is the non-obvious one: dedupe removes duplicates, but which copy is +kept — first occurrence, last, or implementation-defined? The spec must commit. + +## Worked example (resolved-mixed) + +The same merge-intervals requirement after a resolution session where `adjacency` was +resolved with an explicit acceptance criterion, `ordering` was dismissed, and `empty` was +left unresolved: + +```json edge-probe:06-resolved-mixed/expected-coverage.json +{ + "items": [ + { "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" }, + { "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" }, + { "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" } + ], + "coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } } +} +``` + +`coverage.resolved` is 2 — the closed set (adjacency=`resolved`/`explicit` + ordering=`dismissed`); +`unresolved` is 1 (empty); `byVerification.explicit` is 1 (the single explicitly-verified +edge), `backstop` 0. The soft gate raises on this example because one applicable edge remains +unresolved — the author must either specify, dismiss, or backstop it before writing the SPEC. diff --git a/gsd-core/templates/spec.md b/gsd-core/templates/spec.md index 18b8ce6e1..952eb3503 100644 --- a/gsd-core/templates/spec.md +++ b/gsd-core/templates/spec.md @@ -68,6 +68,18 @@ If none: "No additional constraints beyond standard project conventions."] [Every acceptance criterion must be a checkbox that resolves to PASS or FAIL. No "should feel good", "looks reasonable", or "generally works" — those are not checkboxes.] +## Edge Coverage + +**Coverage:** [resolved]/[applicable] applicable edges resolved · [unresolved] unresolved + +| Category | Requirement | Status | Resolution / Reason | +|----------|-------------|--------|---------------------| +| [category] | [Rn] | [✅ covered / ⛔ dismissed / 🧪 backstop / ⚠ UNRESOLVED] | [acceptance criterion ref, dismissal reason, or backstop test note] | + +[Generated by the edge-completeness probe (Step 5.5). `covered` rows correspond to +Acceptance Criteria above; `backstop` rows must be carried into plan-phase `must_haves`. +`⚠ UNRESOLVED` rows are flagged: planner must treat as assumption.] + ## Ambiguity Report | Dimension | Score | Min | Status | Notes | diff --git a/gsd-core/workflows/plan-phase.md b/gsd-core/workflows/plan-phase.md index dd981d74b..52937b19a 100644 --- a/gsd-core/workflows/plan-phase.md +++ b/gsd-core/workflows/plan-phase.md @@ -810,6 +810,14 @@ PATTERNS_PATH=$(_gsd_field "$INIT" patterns_path) # Detect spike/sketch findings skills (project-local) SPIKE_FINDINGS_PATH=$(ls ./.claude/skills/spike-findings-*/SKILL.md 2>/dev/null | head -1 || true) SKETCH_FINDINGS_PATH=$(ls ./.claude/skills/sketch-findings-*/SKILL.md 2>/dev/null | head -1 || true) + +# Resolve the phase SPEC (carries the ## Edge Coverage section the planner lifts covered/ +# backstop edges from). UNCONDITIONAL — must NOT live in §4.5 Check AI-SPEC, which is skipped +# on non-AI phases; gating it there silently starves the planner of the SPEC (#550 review). +# Glob the plain phase SPEC, excluding the -AI-SPEC.md / -UI-SPEC.md variants. +PHASE_DIR_FOR_SPEC=$(_gsd_field "$INIT" phase_dir) +SPEC_FILE=$(ls "${PHASE_DIR_FOR_SPEC}"/*-SPEC.md 2>/dev/null | grep -Ev -- '-(AI|UI)-SPEC\.md$' | head -1) +SPEC_PATH="${SPEC_FILE}" ``` ## 7.5. Verify Nyquist Artifacts @@ -937,6 +945,7 @@ Planner prompt: - {uat_path} (UAT Gaps - if --gaps) - {reviews_path} (Cross-AI Review Feedback - if --reviews; actionable findings must be incorporated or explicitly deferred/rejected in PLAN.md) - {UI_SPEC_PATH} (UI Design Contract — visual/interaction specs, if exists) +- {SPEC_PATH} (Phase SPEC — carries the ## Edge Coverage section to lift covered/backstop edges from, if exists) - {SPIKE_FINDINGS_PATH} (Spike Findings — validated patterns, constraints, landmines from experiments, if exists) - {SKETCH_FINDINGS_PATH} (Sketch Findings — validated design decisions, CSS patterns, visual direction, if exists) - {API_SURFACE_PATH} (API Surface — HINT ONLY, if intel.enabled; see below) @@ -999,6 +1008,7 @@ Output consumed by /gsd:execute-phase. Plans need: - Tasks in XML format with read_first and acceptance_criteria fields (MANDATORY on every task) - Verification criteria - must_haves for goal-backward verification +- If the SPEC has an `## Edge Coverage` section, lift every `covered` edge's acceptance criterion into `must_haves.truths`, and every `backstop` edge into `must_haves.truths` as a non-inferable check (note it needs a held-out/property-based test). `unresolved` edges are explicit assumptions — surface them in the plan, do not silently drop them. - **"Artifacts this phase produces" section (MANDATORY)** — list every symbol this phase creates: decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths. The plan-review-convergence source-grounding pass reads this section to exclude newly-created symbols from drift verification; omitting it causes new symbols to be flagged for acknowledgement. @@ -1044,6 +1054,7 @@ Every task MUST include these fields — they are NOT optional: - [ ] Waves assigned for parallel execution - [ ] must_haves derived from phase goal - [ ] Every PLAN.md includes an "Artifacts this phase produces" section listing symbols created by this phase (decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths) +- [ ] Every SPEC ## Edge Coverage covered/backstop edge is represented in a plan's must_haves (no silent drops) ``` diff --git a/gsd-core/workflows/spec-phase.md b/gsd-core/workflows/spec-phase.md index 84bb21e6c..ce004c51d 100644 --- a/gsd-core/workflows/spec-phase.md +++ b/gsd-core/workflows/spec-phase.md @@ -187,10 +187,137 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met): ## Step 5: (covered inline — ambiguity scoring is per-round) +## Step 5.5: Edge-Completeness Probe + +Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague +ones). Reference: @~/.claude/gsd-core/references/edge-probe.md. + +**Runtime coverage compute — resolve and invoke edge-probe.cjs:** + +```bash +# Resolve the compiled edge-probe.cjs against the GSD install dir via RUNTIME_DIR (#448) +# — NOT the consuming project's git root — falling back to git toplevel / $HOME/.claude. +# Mirrors the ui-safety-gate.cjs resolution idiom at autonomous.md:290 / plan-phase.md:631. +_GSD_RT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}" +EDGE_PROBE_JS=$(for _c in \ + "$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \ + "$_GSD_RT/bin/lib/edge-probe.cjs" \ + "$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \ + "$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \ + "$HOME/.claude/bin/lib/edge-probe.cjs"; do + [ -f "$_c" ] && { echo "$_c"; break; } +done) + +# Graceful degradation — never silent skip (RR-04). Build ONLY when $_GSD_RT is a verified +# GSD source checkout (has tsconfig.build.json + src/edge-probe.cts), and pin npm to it with +# --prefix so we never trigger the CONSUMING project's own build:lib (its cwd package scripts: +# codegen/migrations/writes) during a spec workflow. Real installs ship the compiled .cjs via +# prepublishOnly, so this build path only matters in a GSD dev checkout (review High). +if [ -z "$EDGE_PROBE_JS" ]; then + if [ -f "$_GSD_RT/tsconfig.build.json" ] && [ -f "$_GSD_RT/src/edge-probe.cts" ]; then + npm --prefix "$_GSD_RT" run build:lib 2>/dev/null || true + EDGE_PROBE_JS=$(for _c in \ + "$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \ + "$_GSD_RT/bin/lib/edge-probe.cjs" \ + "$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \ + "$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \ + "$HOME/.claude/bin/lib/edge-probe.cjs"; do + [ -f "$_c" ] && { echo "$_c"; break; } + done) + fi + if [ -z "$EDGE_PROBE_JS" ]; then + echo "ERROR: edge-probe.cjs not found — reinstall GSD or run \`npm run build:lib\` in your GSD checkout." >&2 + exit 1 + fi +fi + +# Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the +# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object +# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes +# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04). +REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX.json") +cat > "$REQS_JSON" <<'JSON' +[ + { "id": "R1", "text": "" } +] +JSON +# Guard — never invoke on an empty/invalid array, OR one still holding the heredoc +# `` placeholder (a forgotten substitution would otherwise yield a +# meaningful-looking but bogus coverage report). Fail loud, not silent no-op. +if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("/dev/null; then + echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2 + exit 1 +fi +# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per +# requirement. The covered/backstop/dismissed/unresolved rows in $COVERAGE drive the +# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose). +# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture +# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE +# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary +# the engine validation exists to protect. Make the run fatal, then validate the captured +# report is well-formed JSON before the resolution loop consumes it. +if ! COVERAGE=$(node "$EDGE_PROBE_JS" "$REQS_JSON"); then + rm -f "$REQS_JSON" + echo "ERROR: edge-probe engine failed (invalid shapes or bad input) — fix the requirement(s) and re-run; never proceed with empty coverage." >&2 + exit 1 +fi +rm -f "$REQS_JSON" +# Exit-0-but-garbage guard: the report must parse as JSON with the expected { items[], coverage{} } shape. +if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let r;try{r=JSON.parse(s)}catch{process.exit(1)}if(!r||!Array.isArray(r.items)||typeof r.coverage!=="object"||r.coverage===null)process.exit(1)})'; then + echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2 + exit 1 +fi +# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY +# requirement is far more likely a shape-classification miss (or malformed requirements) than +# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding +# applicable:0. Surface it loudly; the author must explicitly confirm "no applicable edges" +# below rather than silently emitting a green empty ## Edge Coverage section. +APPLICABLE=$(printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let n=0;try{n=JSON.parse(s).coverage.applicable}catch{n=0}process.stdout.write(String(n))})') +if [ "$APPLICABLE" = "0" ]; then + echo "WARNING: edge-probe proposed ZERO applicable edges across all requirements — likely a classification miss or malformed requirements, not a genuinely edge-free spec. Do NOT silently write an empty Edge Coverage section." >&2 +fi +``` + +If `$APPLICABLE` is `0`, do NOT proceed silently: ask the author to confirm via AskUserQuestion +("The edge probe found no applicable edges for any requirement — is this genuinely an +edge-free spec, or should we revisit the requirement wording / authored shapes?"). Only write +an empty `## Edge Coverage` section after explicit confirmation. + +For each Requirement gathered so far: +1. Classify its shape and raise only applicable edge categories (relevance filter — see + the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already + surfaced as pre-`covered`. +2. For each raised category, propose a CONCRETE candidate edge (not "consider + boundaries" — e.g. "R2 merges intervals; what about `[[1,2],[2,3]]` that only touch?"). +3. Resolve each with the user (AskUserQuestion; text mode → numbered list): + - **Specify it** → write a new pass/fail line into Acceptance Criteria AND mark the + edge `covered`. + - **Dismiss (reason)** → mark `dismissed` with a required non-empty reason. + - **Backstop with a test** → mark `backstop`; note "held-out edge test" for plan-phase. + - **Defer** → leave `unresolved`. + +**Soft gate (after resolving):** +- All applicable edges resolved → proceed to Step 6. +- Any `unresolved` → AskUserQuestion: + - header: "Edge Coverage" + - question: "[N] edge(s) are unresolved: [list]. What do you want to do?" + - options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" / + "Keep probing" + - On "anyway": write SPEC.md with those rows marked `⚠ Edge unresolved — planner must + treat as assumption`. + +**`--auto` mode:** auto-`covered` where a defensible acceptance criterion can be written; +otherwise auto-`backstop` (never auto-dismiss — a wrong dismissal is the exact silent +failure being eliminated). Log: `[auto] edge coverage: C covered, B backstop, U unresolved`. + +Populate the `## Edge Coverage` section of SPEC.md from the resolved edges. + ## Step 6: Generate SPEC.md Use the SPEC.md template from @~/.claude/gsd-core/templates/spec.md. +- Populate the **Edge Coverage** section from Step 5.5 (covered/dismissed/backstop/unresolved rows). + **Requirements for every requirement entry:** - One specific, testable statement - Current state (what exists now) @@ -249,6 +376,7 @@ Next: /gsd:discuss-phase {X} - Do NOT ask about HOW to implement — that is discuss-phase territory - Scout the codebase BEFORE the first question — grounded questions only - Max 2–3 questions per round — do not frontload all questions at once +- Step 5.5 edge probe runs after the ambiguity gate; dismissals require a reason; --auto never auto-dismisses @@ -260,4 +388,5 @@ Next: /gsd:discuss-phase {X} - Acceptance criteria are pass/fail checkboxes - SPEC.md committed atomically (when commit_docs is true) - User directed to /gsd:discuss-phase as next step +- Edge-completeness probe run; Edge Coverage section populated; unresolved edges flagged as assumptions diff --git a/scripts/lint-test-file-count.allowlist.json b/scripts/lint-test-file-count.allowlist.json index bc7ed604c..4ca20eb7e 100644 --- a/scripts/lint-test-file-count.allowlist.json +++ b/scripts/lint-test-file-count.allowlist.json @@ -151,6 +151,15 @@ "validate-context.test.cjs" ], "issue": "892" + }, + "edge-probe": { + "files": [ + "edge-probe-docs-fixtures.test.cjs", + "edge-probe-planner-contract.test.cjs", + "edge-probe-spec-phase-contract.test.cjs", + "edge-probe.test.cjs" + ], + "issue": "550" } } } diff --git a/src/edge-probe.cts b/src/edge-probe.cts new file mode 100644 index 000000000..8d7f7a7ef --- /dev/null +++ b/src/edge-probe.cts @@ -0,0 +1,217 @@ +/** + * Spec-completeness edge-probe — the FIRST adapter of the probe-core resolution model + * (ADR-457 build model; ADR-550 Decision 7 seam). + * + * The generic resolution lifecycle, the status×verification re-cut, `validateResolution`, + * `validateRequirement`, the `analyzeCoverage` merge/rollup/orphan-reject engine, and the + * `runProbeCli` scaffold all live in `src/probe-core.cts`. This module keeps ONLY the + * edge-specific cluster: the five data/behavior shapes, the closed 8-category edge taxonomy, + * shape classification, edge proposal, and the `{ explicit, backstop }` verification validators. + * + * Authored as strict TypeScript (`src/edge-probe.cts`) and compiled by + * `tsc -p tsconfig.build.json` to the gitignored runtime artifact + * `gsd-core/bin/lib/edge-probe.cjs`. Do NOT hand-write the `.cjs`; it is emitted. Tests + * `require()` the built artifact; `pretest` runs `build:lib` first. + * + * Pure and dependency-free: it classifies each requirement's data/behavior shape, filters + * the closed 8-category edge taxonomy to applicable categories, proposes concrete candidate + * edges, and (via probe-core) merges author resolutions into a coverage report. + */ + +import { + type Item, + type Resolution, + type CoverageReport, + type Validators, + validateRequirement as coreValidateRequirement, + validateResolution as coreValidateResolution, + analyzeCoverage as coreAnalyzeCoverage, + runProbeCli, +} from './probe-core.cjs'; + +/** The five data/behavior shapes a requirement can exhibit. */ +export type Shape = 'numeric-range' | 'collection' | 'text' | 'stateful' | 'io'; + +/** The edge probe's verification tiers (the `verification` axis values for a resolved edge). */ +export type EdgeVerification = 'explicit' | 'backstop'; + +/** A single edge taxonomy category. */ +export interface TaxonomyEntry { + id: string; + name: string; + shapes: Shape[]; + probe: string; +} + +/** A SPEC requirement; `shapes` is an optional authored override of classification. */ +export interface Requirement { + id: string; + text: string; + shapes?: Shape[]; +} + +/** An edge item — a probe-core `Item` specialized to the edge verification vocabulary. */ +export type Edge = Item; + +/** + * Word-boundary cues mapping requirement prose -> data/behavior shape. + * Heuristic and intentionally lossy; an authored `shapes` array overrides it. + */ +export const SHAPE_CUES: Record = { + 'numeric-range': /\b(round(ing|ed)?|threshold|max(imum)?|min(imum)?|limit|bound(ary)?|between|cap|percent|amount|price|count|number|score|rate|decimal)\b/i, + 'collection': /\b(lists?|arrays?|sets?|items?|collections?|each|every|all|sort(ed|ing)?|merge|dedupe|group|ranges?|intervals?|overlap(ping)?)\b/i, + 'text': /\b(string|text|names?|labels?|truncate|substring|char(acter)?s?|length|slug|message|unicode)\b/i, + 'stateful': /\b(save|persist|store|update|toggle|create|delete|remove|submit|retry|apply|register|insert)\b/i, + 'io': /\b(files?|requests?|fetch|upload|download|network|api|endpoints?|connections?|sockets?)\b/i, +}; + +/** The locked shape vocabulary — exactly the keys of SHAPE_CUES (single source of truth). */ +export const VALID_SHAPES: ReadonlySet = new Set(Object.keys(SHAPE_CUES)); + +/** Detect which shapes a requirement's prose matches (heuristic). */ +export function classifyShape(text: string): Shape[] { + const shapes: Shape[] = []; + const subject = String(text == null ? '' : text); + for (const shape of Object.keys(SHAPE_CUES) as Shape[]) { + if (SHAPE_CUES[shape].test(subject)) shapes.push(shape); + } + return shapes; +} + +/** + * Closed taxonomy of 8 domain-boundary edge categories (established QA names). + * `shapes` lists which requirement shapes make the category relevant. + */ +export const TAXONOMY: TaxonomyEntry[] = [ + { id: 'boundary', name: 'Boundary values', shapes: ['numeric-range'], probe: 'What happens exactly at each min/max/threshold — and one step either side?' }, + { id: 'adjacency', name: 'Adjacency / touching', shapes: ['collection'], probe: 'When two things are exactly equal or just touch, do they merge, collide, or separate?' }, + { id: 'empty', name: 'Empty / degenerate', shapes: ['collection', 'text'], probe: 'What is the result for empty, single-element, or null input?' }, + { id: 'encoding', name: 'Encoding / representation', shapes: ['text'], probe: 'Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?' }, + { id: 'ordering', name: 'Ordering / stability', shapes: ['collection'], probe: 'When elements compare equal, is output order specified and stable?' }, + { id: 'precision', name: 'Precision / overflow', shapes: ['numeric-range'], probe: 'Where can precision loss or overflow occur, and what is the contract?' }, + { id: 'idempotency', name: 'Idempotency / repetition', shapes: ['stateful'], probe: 'What happens if this runs twice on the same input?' }, + { id: 'concurrency', name: 'Concurrency / effect ordering', shapes: ['stateful', 'io'], probe: 'If interrupted or run in parallel, what is guaranteed?' }, +]; + +/** Return taxonomy category ids whose applicable shapes intersect the input set. */ +export function applicableCategories(shapes: Shape[]): string[] { + const set = new Set(shapes); + return TAXONOMY.filter((c) => c.shapes.some((s) => set.has(s))).map((c) => c.id); +} + +/** + * The edge adapter's injected runtime validators (ADR-550 #5). `categories` is the closed + * taxonomy; both verification tiers require a non-empty `resolution` (an explicit AC's text + * or a backstop note) so plan-phase has a criterion to lift. + */ +export const EDGE_VALIDATORS: Validators = { + categories: TAXONOMY.map((c) => c.id), + verification: ['explicit', 'backstop'], + requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] }, +}; + +/** + * Validate a single requirement — the generic id/text checks (probe-core) plus the edge's + * `shapes`-must-be-an-array check. A bare string like `shapes:"numeric-range"` would otherwise + * fall through to prose classification, silently ignoring the authored override. + * + * The edge adapter's `text` is REQUIRED (the prose is the classification signal), so reject a + * missing/empty `text` when no authored `shapes` override is present. Without this, a `{ id }` + * requirement classifies to zero shapes → zero edges → it is silently DROPPED from coverage + * with no signal — the exact fail-open this feature exists to eliminate. An explicit `shapes` + * array (including `[]` for "no applicable categories") is the legitimate way to opt out of + * prose classification, so `text` is only required when `shapes` is absent. + */ +export function validateRequirement(requirement: Requirement): void { + coreValidateRequirement(requirement); + const r = requirement as unknown as { shapes?: unknown; text?: unknown }; + if (r.shapes != null && !Array.isArray(r.shapes)) { + throw new Error(`requirement ${requirement.id} shapes must be an array when present`); + } + if (r.shapes == null && !(typeof r.text === 'string' && r.text.trim())) { + throw new Error( + `requirement ${requirement.id} text must be a non-empty string when no shapes override is provided`, + ); + } +} + +/** Validate an edge resolution against the edge verification vocabulary. */ +export function validateResolution(resolution: Resolution): true { + return coreValidateResolution(resolution, EDGE_VALIDATORS); +} + +/** + * Propose candidate edges for a requirement. Uses authored `shapes` when present, else + * classifies from prose. Every proposed edge starts unresolved (verification null). + */ +export function proposeEdges(requirement: Requirement): Edge[] { + validateRequirement(requirement); + let shapes: Shape[]; + if (Array.isArray(requirement.shapes)) { + // Fail closed: an authored array must contain only locked shape values. A non-empty + // but invalid array (e.g. ['numeric'], a typo for 'numeric-range') would otherwise + // intersect no category and silently suppress every probe — the gate reads green while + // nothing was checked. An empty array stays a valid "no applicable categories" override. + for (const s of requirement.shapes) { + if (typeof s !== 'string' || !VALID_SHAPES.has(s)) { + throw new Error( + `invalid shape ${JSON.stringify(s)} for requirement ${requirement.id} — must be one of: ${[...VALID_SHAPES].join(', ')}`, + ); + } + } + shapes = requirement.shapes; + } else { + shapes = classifyShape(requirement.text); + } + return applicableCategories(shapes).map((catId): Edge => { + const cat = TAXONOMY.find((c) => c.id === catId); + return { + requirement_id: requirement.id, + category: catId, + status: 'unresolved', + verification: null, + resolution: null, + reason: null, + probe: cat ? cat.probe : '', + }; + }); +} + +/** + * Propose edges for every requirement (deterministic propose), then delegate the + * merge/rollup/orphan-reject to probe-core. Edge-specific pre-checks: requirements must be an + * array, requirement ids must be unique. Throws on any invalid resolution. + */ +export function analyzeCoverage( + requirements: Requirement[], + resolutions: Resolution[] = [], +): CoverageReport { + if (!Array.isArray(requirements)) { + throw new Error('requirements must be an array'); + } + const items: Edge[] = []; + const seenReqIds = new Set(); + for (const req of requirements) { + validateRequirement(req); + if (seenReqIds.has(req.id)) { + throw new Error(`duplicate requirement id ${JSON.stringify(req.id)}`); + } + seenReqIds.add(req.id); + for (const edge of proposeEdges(req)) items.push(edge); + } + return coreAnalyzeCoverage(items, resolutions, EDGE_VALIDATORS); +} + +/* + * CLI entry (EP-06 invokable surface): `edge-probe.cjs [resolutions.json]`. + * The generic I/O plumbing (parse, fail-closed exit 2, pretty-JSON out) lives in probe-core's + * `runProbeCli`; the edge adapter supplies its `analyzeCoverage`. Guarded by + * `require.main === module` so it runs only when the compiled `.cjs` is executed directly. + */ +if (require.main === module) { + runProbeCli( + (requirements, resolutions) => + analyzeCoverage(requirements as Requirement[], resolutions as Resolution[]), + { usage: 'edge-probe.cjs [resolutions.json]' }, + ); +} diff --git a/src/probe-core.cts b/src/probe-core.cts new file mode 100644 index 000000000..64d782b47 --- /dev/null +++ b/src/probe-core.cts @@ -0,0 +1,336 @@ +/** + * probe-core — generic spec-phase probe resolution model (ADR-550 Decision 7). + * + * Extracted from the edge-probe (the first adapter) once the prohibition probe (#644) + * proved it the *second* adapter of the same model: one adapter is a hypothetical seam, + * two is a real one. This module owns everything generic — the resolution lifecycle, + * the status×verification re-cut, `validateResolution`/`validateRequirement`, the + * `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject engine, + * the `byVerification` rollup, and the `runProbeCli` I/O scaffold. Each probe is a thin + * adapter: it supplies the proposal logic (deterministic for edge, LLM-recall for + * prohibition) and its closed vocabularies via injected validators. + * + * Authored as strict TypeScript (`src/probe-core.cts`) and compiled by + * `tsc -p tsconfig.build.json` to the gitignored runtime artifact + * `gsd-core/bin/lib/probe-core.cjs`. Do NOT hand-write the `.cjs`; it is emitted. + * + * Two orthogonal axes (the re-cut): + * - status: resolved | dismissed | unresolved — the resolution LIFECYCLE (shared) + * - verification: | null — HOW a resolved item is verified + * The edge adapter declares `verification: explicit | backstop`; the prohibition adapter + * (#644) will declare `test | judgment`. Splitting the axes keeps the lifecycle enum free + * of a verification fact and lets a sibling probe add its own tiers without a parallel enum. + * + * Typing is hybrid (ADR-550 #5): generic type params for adapter DX, but enforcement runs + * through injected runtime validators, because the CLI executes over JSON where TS types are + * erased. The contract test pins the validators, not the types. + */ + +import fs from 'node:fs'; + +/** Resolution lifecycle — shared across every probe adapter. */ +export type Status = 'resolved' | 'dismissed' | 'unresolved'; + +/** The LOCKED set of valid lifecycle statuses (the re-cut: no covered/backstop). */ +export const VALID_STATUS: Status[] = ['resolved', 'dismissed', 'unresolved']; + +/** + * A proposed or resolved item for a requirement/category pair. Generic over the probe's + * verification-tier vocabulary `V` (edge: `'explicit' | 'backstop'`). A freshly proposed + * item is `{ status: 'unresolved', verification: null, resolution: null, reason: null }`. + */ +export interface Item { + requirement_id: string; + category: string; + status: Status; + verification: V | null; + resolution: string | null; + reason: string | null; + probe: string; +} + +/** An author resolution merged onto a proposed item. */ +export interface Resolution { + requirement_id: string; + category: string; + status: Status; + verification?: V | null; + resolution?: string | null; + reason?: string | null; +} + +/** A coverage report: the merged items plus rollup counts (incl. the per-tier breakdown). */ +export interface CoverageReport { + items: Item[]; + coverage: { + applicable: number; + resolved: number; + unresolved: number; + byVerification: Record; + }; +} + +/** + * The injected runtime enforcement contract (ADR-550 #5). The probe declares its closed + * vocabularies so `analyzeCoverage`/`validateResolution` can enforce them at the JSON + * boundary where TS types no longer exist. + * - `categories` — the probe's valid category ids (a proposed item outside this set is an + * adapter bug, caught rather than silently rolled up). + * - `verification` — the valid verification tiers for a `resolved` item. + * - `requiredFieldsByVerification` — for each tier, which resolution fields MUST be a + * non-empty string (edge: every tier needs `resolution` text for plan-phase to lift). + */ +export interface Validators { + categories: string[]; + verification: string[]; + requiredFieldsByVerification: Record>; +} + +/** A generic requirement — every probe ingests at least `{ id, text }`. */ +export interface Requirement { + id: string; + text?: string; +} + +function errMessage(e: unknown): string { + return e instanceof Error ? e.message : String(e); +} + +/** + * Structural guard for the report an adapter's `analyze` returns. The scaffold types `analyze` + * loosely (it runs over JSON-parsed input the adapter `as`-casts), so a future adapter (#644) + * that forgets to validate inside its closure could hand back a malformed object. Rather than + * stringify garbage as green output, `runProbeCli` checks the report shape and fails closed. + */ +function isValidReport(report: unknown): report is CoverageReport { + if (report == null || typeof report !== 'object') return false; + const r = report as { items?: unknown; coverage?: unknown }; + if (!Array.isArray(r.items)) return false; + const c = r.coverage as + | { applicable?: unknown; resolved?: unknown; unresolved?: unknown; byVerification?: unknown } + | undefined; + if (c == null || typeof c !== 'object') return false; + if (typeof c.applicable !== 'number' || typeof c.resolved !== 'number' || typeof c.unresolved !== 'number') { + return false; + } + if (c.byVerification == null || typeof c.byVerification !== 'object') return false; + return true; +} + +/** + * Validate a requirement's generic structural fields — fail closed on malformed input rather + * than coercing it. Probe-specific fields (e.g. the edge adapter's `shapes`) are validated by + * the adapter. Typed loosely because the CLI casts arbitrary parsed JSON to `Requirement`. + */ +export function validateRequirement(requirement: Requirement): void { + const r = requirement as unknown as { id?: unknown; text?: unknown }; + if (typeof r.id !== 'string' || !r.id.trim()) { + throw new Error(`requirement id must be a non-empty string (got ${JSON.stringify(r.id)})`); + } + if (r.text != null && typeof r.text !== 'string') { + throw new Error(`requirement ${r.id} text must be a string when present`); + } +} + +/** + * Validate a resolution against the probe's injected validators. Rejects an unknown status, + * a dismissal without a non-empty reason, a `resolved` item with a missing/unknown + * verification tier, and a `resolved` item missing any field its tier requires (per + * `requiredFieldsByVerification`). Returns true on success. + */ +export function validateResolution(r: Resolution, validators: Validators): true { + const key = `${r.requirement_id}::${r.category}`; + if (!VALID_STATUS.includes(r.status)) { + throw new Error(`invalid status "${r.status}" for ${key}`); + } + // Invariant (this module's header): `verification` is null unless `status` is `resolved`. + // Enforce it for EVERY status — a dismissed/unresolved resolution carrying a verification + // tier would otherwise merge verbatim (`analyzeCoverage` below) and silently break the + // model for the second adapter (#644) that inherits this seam. Fail closed across the full + // status×verification space, not just `resolved`. + if (r.status !== 'resolved' && r.verification != null) { + throw new Error(`verification must be null unless status is "resolved" (got "${r.verification}") for ${key}`); + } + // An `unresolved` resolution is an UNACTED item: it must carry no resolution/reason payload. + // A populated payload is an authoring mistake (the author meant resolved/dismissed) that + // would otherwise be silently dropped into the unresolved count with no error pointing at + // it. Reject it so the mistake surfaces. + if (r.status === 'unresolved') { + if (r.resolution != null && String(r.resolution).trim()) { + throw new Error(`unresolved must not carry a resolution (${key})`); + } + if (r.reason != null && String(r.reason).trim()) { + throw new Error(`unresolved must not carry a reason (${key})`); + } + } + if (r.status === 'dismissed' && !(r.reason && String(r.reason).trim())) { + throw new Error(`dismissed requires a reason (${key})`); + } + if (r.status === 'resolved') { + const tier = r.verification; + if (tier == null) { + throw new Error(`resolved requires a verification tier (one of: ${validators.verification.join(', ')}) for ${key}`); + } + if (!validators.verification.includes(tier)) { + throw new Error(`invalid verification "${tier}" for ${key} — must be one of: ${validators.verification.join(', ')}`); + } + const required = validators.requiredFieldsByVerification[tier] ?? []; + for (const field of required) { + // field is 'resolution' | 'reason'; both are `string | null | undefined` on Resolution, + // so the indexed access is string-typed (no unknown-to-string coercion). + const value = r[field]; + if (!(value != null && String(value).trim())) { + throw new Error(`${tier} requires a ${field} (${key})`); + } + } + } + return true; +} + +/** + * Merge author resolutions onto ALREADY-PROPOSED items and roll up coverage counts. + * + * Core operates on `items[]`, never a `proposeFn`: probes have different deterministic + * surfaces (edge = deterministic propose + LLM resolve; prohibition = LLM propose + deterministic + * validate/merge), so proposal stays in each adapter and core must not assume it is deterministic. + * + * `coverage.resolved` is the COUNT of CLOSED items (`resolved` + `dismissed` status) = + * `applicable - unresolved` — the pre-re-cut "covered + dismissed + backstop" set, + * count-preserved. `byVerification` breaks the `resolved`-status items down by tier (each tier + * initialized to 0). Throws on any invalid resolution, a duplicate, an orphan (a resolution + * matching no proposed item), or a proposed item whose category is outside `validators.categories`. + */ +export function analyzeCoverage( + items: Item[], + resolutions: Resolution[] = [], + validators: Validators, +): CoverageReport { + if (!Array.isArray(items)) { + throw new Error('items must be an array'); + } + const key = (r: { requirement_id: string; category: string }): string => `${r.requirement_id}::${r.category}`; + const resMap = new Map>(); + for (const r of resolutions) { + validateResolution(r, validators); + if (resMap.has(key(r))) { + throw new Error(`duplicate resolution for ${key(r)}`); + } + resMap.set(key(r), r); + } + const validCategories = new Set(validators.categories); + const merged: Item[] = []; + const itemKeys = new Set(); + for (const item of items) { + if (!validCategories.has(item.category)) { + throw new Error(`item ${key(item)} has unknown category "${item.category}" — not one of: ${validators.categories.join(', ')}`); + } + itemKeys.add(key(item)); + const o = resMap.get(key(item)); + if (o) { + merged.push({ ...item, status: o.status, verification: o.verification ?? null, resolution: o.resolution ?? null, reason: o.reason ?? null }); + } else { + // No author resolution: the item is rolled up VERBATIM, so its own status/fields must be + // valid too. The edge adapter only proposes `unresolved` items, but the prohibition adapter + // (#644) proposes LLM-generated items that arrive already populated — one carrying an + // out-of-enum status (e.g. the dropped "covered") or `dismissed` with no reason would + // otherwise be counted closed without validation. An Item is structurally a superset of a + // Resolution, so the same fail-closed check guards both. (ADR-550 Decision 5 hardens this + // shared seam for the second adapter; m1.) + validateResolution(item as unknown as Resolution, validators); + merged.push(item); + } + } + // Reject orphan resolutions — a resolution whose (requirement_id, category) matches no + // proposed item (typo'd category or a non-applicable one) would otherwise be silently + // dropped, leaving the author believing an item is resolved while the report shows it + // unresolved (adversarial-review HIGH; preserved from the edge-probe's original engine). + for (const k of resMap.keys()) { + if (!itemKeys.has(k)) { + throw new Error(`unknown resolution for ${k} — no matching proposed item (typo'd category or non-applicable shape?)`); + } + } + const unresolved = merged.filter((i) => i.status === 'unresolved').length; + const applicable = merged.length; + const resolved = applicable - unresolved; // closed set: resolved-status + dismissed + const byVerification: Record = {}; + for (const tier of validators.verification) byVerification[tier] = 0; + for (const i of merged) { + if (i.status === 'resolved' && i.verification != null) { + byVerification[i.verification] = (byVerification[i.verification] ?? 0) + 1; + } + } + return { items: merged, coverage: { applicable, resolved, unresolved, byVerification } }; +} + +/* + * CLI scaffold (the EP-06 invokable surface, generalized). Each probe ships one bin that + * calls `runProbeCli` with its own `analyze` (closing over the adapter's propose + validators) + * and usage string; a single dispatcher CLI is a deferred follow-on. The I/O dependencies are + * injectable so the generic plumbing is unit-testable without spawning a process. + * + * `tsconfig.build.json` sets `"types": ["node"]`, so `process` and `node:fs` are typed. + */ + +/** Injectable I/O for `runProbeCli` (defaults wire to the real process). */ +export interface ProbeCliOptions { + usage: string; + argv?: string[]; + readFile?: (path: string) => string; + write?: (s: string) => void; + writeErr?: (s: string) => void; + exit?: (code: number) => void; +} + +/** + * Read the requirements file (and optional resolutions file), run the adapter's `analyze`, + * and write the report as pretty JSON + newline. With no requirements path, writes the usage + * line to stderr and exits 2. A JSON-parse failure or any `analyze` throw is a handled error: + * stderr + exit 2, never an uncaught stack trace — so the engine's fail-closed validation + * surfaces at the workflow boundary rather than failing open. + */ +export function runProbeCli( + analyze: (requirements: unknown, resolutions: unknown) => CoverageReport, + options: ProbeCliOptions, +): void { + const argv = options.argv ?? process.argv; + const readFile = options.readFile ?? ((p: string) => fs.readFileSync(p, 'utf8')); + const write = options.write ?? ((s: string) => { process.stdout.write(s); }); + const writeErr = options.writeErr ?? ((s: string) => { process.stderr.write(s); }); + const exit = options.exit ?? ((code: number) => { process.exit(code); }); + + const reqPath: string | undefined = argv[2]; + const resPath: string | undefined = argv[3]; + if (!reqPath) { + writeErr(`usage: ${options.usage}\n`); + exit(2); + return; + } + let requirements: unknown; + try { + requirements = JSON.parse(readFile(reqPath)); + } catch (e: unknown) { + writeErr(`error: cannot parse JSON from ${reqPath}: ${errMessage(e)}\n`); + exit(2); + return; + } + let resolutions: unknown = []; + if (resPath) { + try { + resolutions = JSON.parse(readFile(resPath)); + } catch (e: unknown) { + writeErr(`error: cannot parse JSON from ${resPath}: ${errMessage(e)}\n`); + exit(2); + return; + } + } + try { + const report = analyze(requirements, resolutions); + if (!isValidReport(report)) { + throw new Error('adapter returned a structurally-invalid coverage report (expected { items[], coverage{ applicable, resolved, unresolved, byVerification } })'); + } + write(`${JSON.stringify(report, null, 2)}\n`); + } catch (e: unknown) { + writeErr(`error: ${errMessage(e)}\n`); + exit(2); + } +} diff --git a/tests/edge-probe-docs-fixtures.test.cjs b/tests/edge-probe-docs-fixtures.test.cjs new file mode 100644 index 000000000..3b848bca6 --- /dev/null +++ b/tests/edge-probe-docs-fixtures.test.cjs @@ -0,0 +1,115 @@ +// allow-test-rule: runtime-contract-is-the-product — the rendered reference/SPEC/ADR vocab surfaces are the runtime contract; this pins their bijection to the code (docs-parity) +// Asserts the portable reference doc (gsd-core/references/edge-probe.md) keeps its +// worked-example JSON blocks in sync with the source-of-truth fixture files under +// gsd-core/references/edge-probe-fixtures/. The fixtures are the canonical data; the +// doc embeds copies. Per the CONTRIBUTING.md exception matrix this is `docs-parity`: a +// reference doc must mirror source-defined data and there is no runtime enumeration API. +// The comparison is PARSED JSON (deepEqual of JSON.parse on both sides), never a raw-text +// substring match — so a reformat that preserves the data does not fail, and any semantic +// drift between doc and fixture does. +'use strict'; +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const docPath = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe.md'); +const fixturesRoot = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures'); +const adrPath = path.join(__dirname, '..', 'docs', 'adr', '550-spec-phase-probe-contract.md'); +const specTemplatePath = path.join(__dirname, '..', 'gsd-core', 'templates', 'spec.md'); + +// Extract fenced blocks tagged ```json edge-probe:/ from the doc, keyed by ref. +// The \n? before the closing fence allows blocks whose closing fence has no preceding newline +// (fixes the silent-skip bug where a trailing-fence-with-no-newline was not matched). +function taggedJsonBlocks(md) { + const re = /```json edge-probe:([^\n]+)\n([\s\S]*?)\n?```/g; + const out = {}; + let m; + while ((m = re.exec(md))) out[m[1].trim()] = m[2]; + return out; +} + +describe('edge-probe doc/fixture sync', () => { + test('reference doc exists', () => { + assert.ok(fs.existsSync(docPath), `${docPath} must exist`); + }); + + test('doc embeds tagged fixture blocks for every expected-coverage.json fixture (count-equality)', () => { + const md = fs.readFileSync(docPath, 'utf8'); + const blocks = taggedJsonBlocks(md); + // Count the expected-coverage.json files under the fixtures root (one per fixture dir). + const expectedCount = fs.readdirSync(fixturesRoot, { withFileTypes: true }) + .filter(e => e.isDirectory()) + .filter(dir => fs.existsSync(path.join(fixturesRoot, dir.name, 'expected-coverage.json'))) + .length; + assert.strictEqual( + Object.keys(blocks).length, + expectedCount, + `edge-probe.md must embed exactly ${expectedCount} tagged blocks (one per fixture expected-coverage.json)` + ); + }); + + test('every tagged doc block parses and deepEquals its fixture file', () => { + const md = fs.readFileSync(docPath, 'utf8'); + const blocks = taggedJsonBlocks(md); + for (const [ref, body] of Object.entries(blocks)) { + const fixtureFile = path.join(fixturesRoot, ref); + const onDisk = fs.readFileSync(fixtureFile, 'utf8'); + assert.deepEqual(JSON.parse(body), JSON.parse(onDisk), + `doc block edge-probe:${ref} must deepEqual ${fixtureFile}`); + } + }); +}); + +// m2: lock the machine↔SPEC vocabulary mapping so the two layers cannot silently drift. The +// machine contract uses orthogonal `status` × `verification`; the SPEC table and planner prose +// render a flat `covered/dismissed/backstop/unresolved`. The migration is documented in ADR-550 +// Decision 7a and the reference's "Generic mapping" table, but the SPEC table is rendered by the +// LLM workflow (no JS renderer to round-trip against). So this pins the canonical map as code AND +// grounds it in every doc surface that renders the vocabulary — a renderer/parser drift fails here. +describe('edge-probe machine↔SPEC vocabulary mapping (m2 drift lock)', () => { + // The canonical migration map (ADR-550 Decision 7a): machine state → SPEC display label. + // `verification` is null for the lifecycle-only states (dismissed/unresolved). + const MACHINE_TO_DISPLAY = [ + { status: 'resolved', verification: 'explicit', display: 'covered' }, + { status: 'resolved', verification: 'backstop', display: 'backstop' }, + { status: 'dismissed', verification: null, display: 'dismissed' }, + { status: 'unresolved', verification: null, display: 'unresolved' }, + ]; + const machineKey = (s, v) => `${s}|${v ?? '∅'}`; + + test('the mapping is a bijection — machine→display→machine is identity, no shared labels', () => { + const toDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [machineKey(m.status, m.verification), m.display])); + const fromDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [m.display, machineKey(m.status, m.verification)])); + // No two machine states collapse onto the same display label (the silent-drift failure mode). + assert.equal(toDisplay.size, MACHINE_TO_DISPLAY.length, 'each machine state must have a distinct key'); + assert.equal(fromDisplay.size, MACHINE_TO_DISPLAY.length, 'each display label must map back to exactly one machine state'); + // Round-trip identity. + for (const m of MACHINE_TO_DISPLAY) { + const key = machineKey(m.status, m.verification); + assert.equal(fromDisplay.get(toDisplay.get(key)), key, `${key} must round-trip through its display label`); + } + }); + + test('ADR-550 Decision 7a documents the resolved/explicit↔covered and resolved/backstop↔backstop migration', () => { + const adr = fs.readFileSync(adrPath, 'utf8'); + assert.match(adr, /covered\b[^.]*resolved[^.]*explicit/i, 'ADR must document covered → {resolved, explicit}'); + assert.match(adr, /backstop\b[^.]*resolved[^.]*backstop/i, 'ADR must document backstop → {resolved, backstop}'); + assert.match(adr, /count-for-count|count-preserved/i, 'ADR must state coverage.resolved is count-preserved across the migration'); + }); + + test('the SPEC template legend renders exactly the four canonical display labels', () => { + const spec = fs.readFileSync(specTemplatePath, 'utf8'); + for (const { display } of MACHINE_TO_DISPLAY) { + assert.match(spec, new RegExp(`\\b${display}\\b`, 'i'), `spec.md Edge Coverage legend must render the "${display}" label`); + } + }); + + test('the reference Generic-mapping table distinguishes the resolved/explicit and resolved/backstop tiers', () => { + const md = fs.readFileSync(docPath, 'utf8'); + assert.match(md, /resolved`?\/`?explicit/i, 'reference must name the resolved/explicit tier in the mapping table'); + assert.match(md, /resolved`?\/`?backstop/i, 'reference must name the resolved/backstop tier in the mapping table'); + }); +}); diff --git a/tests/edge-probe-planner-contract.test.cjs b/tests/edge-probe-planner-contract.test.cjs new file mode 100644 index 000000000..592ee9c22 --- /dev/null +++ b/tests/edge-probe-planner-contract.test.cjs @@ -0,0 +1,167 @@ +// allow-test-rule: runtime-contract-is-the-product — plan-phase.md's planner prompt is the deployed runtime contract under assertion +// plan-phase.md is the deployed planning workflow contract; these checks lock +// the SPEC path wiring and quality-gate that the edge-probe review (RR-01/02/03) +// requires — assertions scope to extracted sub-blocks to avoid false positives. + +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const PLAN_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-phase.md'); +const PROMPT_PATH = path.join(__dirname, '..', 'gsd-core', 'templates', 'planner-subagent-prompt.md'); + +function readPlanPhase() { + return fs.readFileSync(PLAN_PHASE_PATH, 'utf8'); +} + +// Extract the planner block from plan-phase.md. The runtime planner is +// spawned from plan-phase.md's own inline /, so the +// load-bearing lift instruction must live here — NOT in templates/planner-subagent-prompt.md, +// which nothing loads at runtime (no @-import in agents/gsd-planner.md, no read in plan-phase.md). +function extractDownstreamConsumerBlock(content) { + const start = content.indexOf(''); + if (start === -1) return ''; + const end = content.indexOf('', start); + if (end === -1) return ''; + return content.slice(start, end + ''.length); +} + +// Extract the planner block that contains {UI_SPEC_PATH} +// There are multiple blocks in plan-phase.md; we need the one +// at ~line 890-912 inside the planning_context markdown block. +function extractPlannerFilesBlock(content) { + let pos = 0; + while (true) { + const start = content.indexOf('', pos); + if (start === -1) return ''; + const end = content.indexOf('', start); + if (end === -1) return ''; + const block = content.slice(start, end + ''.length); + if (block.includes('{UI_SPEC_PATH}')) { + return block; + } + pos = end + 1; + } +} + +// Extract the planner block (the last one in plan-phase.md, +// inside the planner prompt template) +function extractQualityGateBlock(content) { + const start = content.lastIndexOf(''); + if (start === -1) return ''; + const end = content.indexOf('', start); + if (end === -1) return ''; + return content.slice(start, end + ''.length); +} + +// Test A (RR-01): plan-phase.md resolves a phase *-SPEC.md into SPEC_FILE/SPEC_PATH +// Uses new RegExp to correctly match literal $ and ( characters in bash snippets. +// This MUST FAIL before the RR-01 fix (no SPEC_PATH resolution exists today) +test('RR-01: plan-phase.md resolves phase *-SPEC.md (excluding AI/UI variants) into SPEC_FILE/SPEC_PATH', () => { + const content = readPlanPhase(); + + // Assert the canonical SPEC_FILE resolution form is present. + // new RegExp used so that \$ and \( are treated as literal dollar-sign and open-paren + // (JS regex literals interpret \$ as end-anchor and \( as group open). + assert.match( + content, + new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md'), + 'plan-phase.md must resolve SPEC_FILE using ls "${...PHASE_DIR...}"/*-SPEC.md pattern' + ); + + // Assert {SPEC_PATH} token appears in the planner files_to_read block + const filesBlock = extractPlannerFilesBlock(content); + assert.match( + filesBlock, + /[{]SPEC_PATH[}]/, + 'The planner block (containing {UI_SPEC_PATH}) must also contain {SPEC_PATH}' + ); +}); + +// Test B (RR-01): The {SPEC_PATH} entry is labelled as carrying the ## Edge Coverage section +// This MUST FAIL before the RR-01 fix (no {SPEC_PATH} entry exists today) +test('RR-01: {SPEC_PATH} entry in files_to_read is labelled with Edge Coverage', () => { + const content = readPlanPhase(); + const filesBlock = extractPlannerFilesBlock(content); + + assert.match( + filesBlock, + /[{]SPEC_PATH[}][^\n]*Edge Coverage/, + '{SPEC_PATH} entry in planner files_to_read must be labelled as carrying the ## Edge Coverage section' + ); +}); + +// Extract a "## " section from its heading until the next "## " heading (or EOF). +function extractSection(content, heading) { + const start = content.indexOf(heading); + if (start === -1) return ''; + const next = content.indexOf('\n## ', start + heading.length); + return content.slice(start, next === -1 ? content.length : next); +} + +// Test E (RR-01 REACHABILITY — the assertion that catches the original no-op): +// Token presence is not enough. The SPEC resolution must live on an UN-GATED path. §4.5 +// "Check AI-SPEC" is skipped on every non-AI phase (ai_integration_phase_enabled false / +// --skip-ai-spec), so a resolution placed there leaves SPEC_PATH unbound and the planner +// never receives the SPEC — exactly the #550 silent no-op. Assert it is NOT in §4.5. +test('RR-01 reachability: SPEC_FILE resolution is NOT gated inside the skippable "## 4.5 Check AI-SPEC" section', () => { + const content = readPlanPhase(); + const aiSpecSection = extractSection(content, '## 4.5. Check AI-SPEC'); + const specResolution = new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md'); + assert.ok(aiSpecSection.length > 0, 'sanity: the §4.5 Check AI-SPEC section must exist to scope this test'); + assert.doesNotMatch( + aiSpecSection, + specResolution, + 'SPEC_FILE resolution must NOT live inside the skippable §4.5 "Check AI-SPEC" block — gating it there silently starves the planner of the SPEC on non-AI phases (the original #550 no-op)' + ); + assert.match(content, specResolution, 'SPEC_FILE resolution must still exist on an un-gated path elsewhere in plan-phase.md'); +}); + +// Test C (RR-02 consumer end): the lift instruction lives in the RUNTIME planner surface. +// plan-phase.md spawns the planner from its own inline ; the +// templates/planner-subagent-prompt.md file is orphaned (loaded by nothing), so asserting the +// contract there is false assurance — the test stays green even if the runtime never consumes it. +// Pin the contract to the block plan-phase.md actually sends the planner. +test('RR-02 consumer: plan-phase.md downstream_consumer instructs lifting covered/backstop edges into must_haves.truths', () => { + const block = extractDownstreamConsumerBlock(readPlanPhase()); + + assert.ok(block.length > 0, 'sanity: plan-phase.md must contain a block to scope this test'); + assert.match( + block, + /##\s*Edge Coverage/, + 'plan-phase.md must reference ## Edge Coverage' + ); + assert.match( + block, + /must_haves\.truths/, + 'plan-phase.md must reference must_haves.truths as the lift target' + ); + + // Regression guard against re-orphaning: the lift instruction must NOT be relocated back into + // templates/planner-subagent-prompt.md (which nothing loads) and presented as the consumer + // contract — that is exactly the false-assurance this retarget fixes. + const prompt = fs.readFileSync(PROMPT_PATH, 'utf8'); + assert.doesNotMatch( + prompt, + /Edge Coverage/, + 'planner-subagent-prompt.md is not loaded at runtime — the Edge Coverage lift instruction must not live there' + ); +}); + +// Test D (RR-03): plan-phase.md contains a covered/backstop ↔ must_haves item +// This MUST FAIL before the RR-03 fix (no such quality_gate item exists today) +test('RR-03: planner quality_gate requires covered/backstop edges represented in must_haves', () => { + const content = readPlanPhase(); + const qgBlock = extractQualityGateBlock(content); + + assert.match( + qgBlock, + /covered.*backstop.*must_haves|backstop.*covered.*must_haves/i, + 'planner quality_gate must contain a checklist item tying covered/backstop edges to must_haves' + ); +}); diff --git a/tests/edge-probe-spec-phase-contract.test.cjs b/tests/edge-probe-spec-phase-contract.test.cjs new file mode 100644 index 000000000..4933cbe71 --- /dev/null +++ b/tests/edge-probe-spec-phase-contract.test.cjs @@ -0,0 +1,193 @@ +// allow-test-rule: runtime-contract-is-the-product — spec-phase.md Step 5.5 is the deployed workflow runtime contract under assertion +// spec-phase.md is the deployed spec workflow contract; these checks lock +// the Step 5.5 wiring so the edge-probe.cjs runtime invocation cannot +// silently rot the way the original plan-phase no-op did (reviewer finding RR-11). +// Assertions scope to the extracted Step 5.5 block to avoid false positives +// from incidental mentions elsewhere in the file. + +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const SPEC_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'spec-phase.md'); + +function readSpecPhase() { + return fs.readFileSync(SPEC_PHASE_PATH, 'utf8'); +} + +// Slice the Step 5.5 block: from the "Step 5.5" heading to the next "## " or "Step " heading. +// This scopes assertions to Step 5.5 only, preventing false positives from mentions elsewhere. +function extractStep55Block(content) { + const startIdx = content.indexOf('## Step 5.5'); + if (startIdx === -1) { + // Also try without the ## prefix + const altIdx = content.indexOf('Step 5.5'); + if (altIdx === -1) return ''; + // Find end: next heading starting with ## or Step N (not Step 5.5) + const rest = content.slice(altIdx + 'Step 5.5'.length); + const nextHeading = rest.search(/\n## |\nStep \d/); + if (nextHeading === -1) return content.slice(altIdx); + return content.slice(altIdx, altIdx + 'Step 5.5'.length + nextHeading); + } + const rest = content.slice(startIdx + '## Step 5.5'.length); + const nextHeading = rest.search(/\n## /); + if (nextHeading === -1) return content.slice(startIdx); + return content.slice(startIdx, startIdx + '## Step 5.5'.length + nextHeading); +} + +// Test A (RR-11): Step 5.5 resolves and invokes edge-probe.cjs via node. +// MUST FAIL before the RR-04 wire (Step 5.5 is prose-only today — no CLI invocation). +test('RR-11: spec-phase Step 5.5 resolves edge-probe.cjs via path-fallback loop', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // Assert the path-fallback resolution loop for edge-probe.cjs is present in Step 5.5. + // The token "edge-probe.cjs" must appear inside the block (the artifact being resolved). + assert.match( + block, + /edge-probe\.cjs/, + 'Step 5.5 must reference edge-probe.cjs as the artifact being resolved' + ); + + // Assert node invocation of edge-probe.cjs in Step 5.5. + // Matches: node "$EDGE_PROBE_JS" or node ... edge-probe.cjs + assert.match( + block, + /node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/, + 'Step 5.5 must invoke edge-probe.cjs via node (e.g. node "$EDGE_PROBE_JS" ...)' + ); +}); + +// Test C (RR-11 FUNCTION — the assertion that catches "decorative bash"): +// Token presence is not enough. The invocation is a no-op unless $REQS_JSON is actually +// POPULATED before the engine runs (the original block only mktemp'd it + left a comment, +// so the CLI parsed an empty file). Assert the block (a) writes $REQS_JSON via a redirect, +// (b) does so BEFORE the node invocation, and (c) guards against an empty/invalid file. +test('RR-11 function: Step 5.5 writes $REQS_JSON before invoking, and guards against empty input', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // (a) A redirect that writes the requirements into $REQS_JSON (e.g. `cat > "$REQS_JSON"`). + const writeIdx = block.search(/>\s*"\$REQS_JSON"/); + assert.ok( + writeIdx !== -1, + 'Step 5.5 must WRITE requirements into $REQS_JSON (a redirect like `cat > "$REQS_JSON"`), not just mktemp it — an empty file makes the probe a silent no-op' + ); + + // (b) The write must precede the node invocation of the engine. + const invokeIdx = block.search(/node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/); + assert.ok(invokeIdx !== -1, 'Step 5.5 must invoke the engine via node'); + assert.ok( + writeIdx < invokeIdx, + 'Step 5.5 must populate $REQS_JSON BEFORE invoking edge-probe.cjs (write precedes the node call)' + ); + + // (c) A guard that refuses to run on an empty/invalid requirements array. + assert.match( + block, + /Array\.isArray|empty\/invalid|empty or invalid|REQS_JSON[^\n]*empty/i, + 'Step 5.5 must guard against an empty/invalid $REQS_JSON before invoking (fail loud, not silent no-op)' + ); +}); + +// Test B (RR-11): Step 5.5 has an explicit not-found branch — build:lib or error token. +// MUST FAIL before the RR-04 wire (no not-found handling today). +test('RR-11: spec-phase Step 5.5 has an explicit not-found branch (build:lib or blocking error)', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // Assert either a build:lib invocation or an explicit "not found" / error message exists. + // This prevents the wire from being added as a silent-skip with no fallback. + assert.match( + block, + /build:lib|not found|ERROR.*edge-probe|edge-probe.*not found/i, + 'Step 5.5 must have an explicit not-found branch (build:lib attempt or clear blocking error)' + ); +}); + +// Test D (review High): the build fallback must NEVER run the CONSUMING project's package +// scripts. Every executable `build:lib` invocation must be pinned to the GSD dir with +// `npm --prefix`, and the build must be gated behind a verified GSD source checkout. A bare +// `npm run build:lib` (no --prefix) uses cwd — which, under the git-toplevel fallback, is the +// consumer repo — and would execute its codegen/migrations during a spec workflow. +test('review High: Step 5.5 build:lib is --prefix-pinned to the GSD dir and gated on a source checkout', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // Collect lines that actually INVOKE npm (trimmed start === "npm"), excluding echo/comment + // mentions (e.g. the error message that quotes `npm run build:lib` for the user). + const buildInvocations = block + .split('\n') + .filter((l) => l.trim().startsWith('npm') && l.includes('build:lib')); + + assert.ok( + buildInvocations.length > 0, + 'Step 5.5 must contain at least one npm build:lib invocation (the dev-checkout fallback)' + ); + for (const line of buildInvocations) { + assert.match( + line, + /npm\s+--prefix\s+"?\$?\{?_?GSD_RT/, + `build:lib must be pinned with \`npm --prefix "$_GSD_RT"\` so it never runs the consuming project's scripts — offending line: ${line.trim()}` + ); + } + + // The build must be gated behind a verified GSD source checkout (tsconfig.build.json present), + // so it cannot fire inside a plain consumer repo where the artifact merely happens to be absent. + assert.match( + block, + /tsconfig\.build\.json/, + 'Step 5.5 must gate the build behind a GSD source checkout (e.g. test -f "$_GSD_RT/tsconfig.build.json")' + ); +}); + +// Test E (review #4 High): the engine's fail-closed exit(2) must NOT be swallowed by command +// substitution. A bare `COVERAGE=$(node "$EDGE_PROBE_JS" ...)` discards the exit status, leaves +// $COVERAGE empty on an invalid-shapes failure, and lets the workflow proceed into prose +// re-derivation — fail-OPEN at the very boundary the engine validation exists to protect. +test('review #4 High: Step 5.5 exit-checks the engine capture and validates the report (no fail-open)', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // The engine capture must be FATAL — guarded by `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)` + // (or an explicit exit-status check) that exits non-zero on failure. + assert.match( + block, + /if\s+!\s+COVERAGE=\$\(node\s+"\$EDGE_PROBE_JS"/, + 'Step 5.5 must exit-check the engine invocation (e.g. `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)`) — a bare command substitution swallows the engine exit code and fails open' + ); + + // And the captured report must be validated as JSON before the resolution loop consumes it + // (guards against an exit-0-but-garbage capture). + assert.match( + block, + /(COVERAGE[\s\S]{0,500}JSON\.parse)|(JSON\.parse[\s\S]{0,500}\$COVERAGE)/, + 'Step 5.5 must validate $COVERAGE parses as JSON before use (guard against exit-0-but-malformed output)' + ); +}); + +// Test F (adversarial review): a report with ZERO applicable edges across all requirements is +// the likely-classification-miss fail-open (same shape as an invalid shape yielding applicable:0). +// Step 5.5 must surface it, not silently emit a green empty ## Edge Coverage section. +test('adversarial review: Step 5.5 guards a zero-applicable coverage report', () => { + const content = readSpecPhase(); + const block = extractStep55Block(content); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + assert.match( + block, + /coverage\.applicable/, + 'Step 5.5 must read coverage.applicable and guard the zero-applicable case (warn/confirm, not silently proceed)' + ); +}); diff --git a/tests/edge-probe.test.cjs b/tests/edge-probe.test.cjs new file mode 100644 index 000000000..2436bc1c0 --- /dev/null +++ b/tests/edge-probe.test.cjs @@ -0,0 +1,441 @@ +/** + * Edge-probe reference core unit tests. + * + * Asserts the LOCKED export surface of the spec-completeness edge-probe against + * the BUILT artifact (`gsd-core/bin/lib/edge-probe.cjs`), which + * `npm run build:lib` (run by pretest) emits from `src/edge-probe.cts`. + * + * Post ADR-550 Decision 7: the generic resolution model lives in `probe-core`; + * edge-probe is its first adapter (shapes/TAXONOMY/proposeEdges + the + * `{explicit, backstop}` verification validators). The resolution model is the + * status×verification re-cut: `status: resolved | dismissed | unresolved` × + * `verification: explicit | backstop`. `covered`/`backstop` are no longer status + * values — `covered → {resolved, explicit}`, `backstop → {resolved, backstop}`. + */ +'use strict'; +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const path = require('node:path'); +const fs = require('node:fs'); +const os = require('node:os'); +const { execFileSync, spawnSync } = require('node:child_process'); +const { cleanup } = require('./helpers.cjs'); + +const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs'); +const ep = require(BUILT_SCRIPT); + +describe('edge-probe: classifyShape', () => { + test('detects numeric-range from rounding/threshold cues', () => { + assert.deepEqual(ep.classifyShape('Round a number to N decimal places').sort(), + ['numeric-range']); + }); + test('detects collection from interval/merge cues', () => { + const shapes = ep.classifyShape('Merge a list of overlapping intervals'); + assert.ok(shapes.includes('collection')); + }); + test('detects text from truncate/string cues', () => { + const shapes = ep.classifyShape('Truncate a string to a maximum length'); + assert.ok(shapes.includes('text')); + }); + test('returns [] when no cue matches', () => { + assert.deepEqual(ep.classifyShape('Display the company logo'), []); + }); +}); + +describe('edge-probe: TAXONOMY + applicableCategories', () => { + test('TAXONOMY has the 8 documented categories in order', () => { + assert.deepEqual(ep.TAXONOMY.map((c) => c.id), + ['boundary', 'adjacency', 'empty', 'encoding', 'ordering', 'precision', 'idempotency', 'concurrency']); + }); + test('every category has name, shapes[], probe', () => { + for (const c of ep.TAXONOMY) { + assert.equal(typeof c.name, 'string'); + assert.ok(Array.isArray(c.shapes) && c.shapes.length >= 1); + assert.equal(typeof c.probe, 'string'); + } + }); + test('numeric-range raises boundary + precision only', () => { + assert.deepEqual(ep.applicableCategories(['numeric-range']).sort(), + ['boundary', 'precision']); + }); + test('collection raises adjacency, empty, ordering', () => { + assert.deepEqual(ep.applicableCategories(['collection']).sort(), + ['adjacency', 'empty', 'ordering']); + }); + test('text raises empty + encoding', () => { + assert.deepEqual(ep.applicableCategories(['text']).sort(), + ['empty', 'encoding']); + }); + test('no shapes raises nothing', () => { + assert.deepEqual(ep.applicableCategories([]), []); + }); +}); + +describe('edge-probe: proposeEdges', () => { + test('rounding requirement proposes boundary + precision, all unresolved (verification null)', () => { + const edges = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places' }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']); + for (const e of edges) { + assert.equal(e.requirement_id, 'R1'); + assert.equal(e.status, 'unresolved'); + assert.equal(e.verification, null); + assert.equal(e.resolution, null); + assert.equal(e.reason, null); + assert.equal(typeof e.probe, 'string'); + } + }); + test('authored shapes override prose classification', () => { + const edges = ep.proposeEdges({ id: 'R9', text: 'opaque label', shapes: ['collection'] }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']); + }); +}); + +describe('edge-probe: validateResolution', () => { + test('rejects an unknown status', () => { + assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'maybe' }), + /invalid status/i); + }); + test('rejects a former covered status (re-cut: covered is no longer a status)', () => { + assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'covered', resolution: 'x' }), + /invalid status/i); + }); + test('rejects dismissed without a reason', () => { + assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: '' }), + /dismissed requires a reason/i); + }); + test('accepts dismissed with a reason', () => { + assert.equal(ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: 'bounded enum' }), true); + }); + test('rejects resolved with a missing verification tier', () => { + assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', resolution: 'AC' }), + /verification/i); + }); + test('rejects resolved with a verification tier outside {explicit, backstop}', () => { + assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'judgment', resolution: 'AC' }), + /invalid verification/i); + }); +}); + +describe('edge-probe: analyzeCoverage', () => { + const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }]; + test('with no resolutions, every applicable edge is unresolved (byVerification zeroed)', () => { + const rep = ep.analyzeCoverage(reqs, []); + assert.deepEqual(rep.coverage, { applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 } }); + }); + test('merges a resolved/explicit resolution and counts it resolved', () => { + const rep = ep.analyzeCoverage(reqs, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' }, + ]); + const adj = rep.items.find((i) => i.category === 'adjacency'); + assert.equal(adj.status, 'resolved'); + assert.equal(adj.verification, 'explicit'); + assert.equal(adj.resolution, 'AC#6: touching intervals merge'); + assert.equal(rep.coverage.resolved, 1); + assert.equal(rep.coverage.unresolved, 2); + assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 }); + }); + test('throws if a resolution is invalid (dismissed w/o reason)', () => { + assert.throws(() => ep.analyzeCoverage(reqs, [ + { requirement_id: 'R1', category: 'empty', status: 'dismissed' }, + ]), /dismissed requires a reason/i); + }); +}); + +describe('edge-probe: CLI (built artifact)', () => { + test('reads a requirements file and prints a coverage report as JSON', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-')); + const reqPath = path.join(dir, 'requirements.json'); + fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }])); + const out = execFileSync('node', [BUILT_SCRIPT, reqPath], { encoding: 'utf8' }); + const rep = JSON.parse(out); + assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } }); + }); + test('with no args exits with status 2 (assert on exit code, not stderr prose)', () => { + let status; + try { + execFileSync('node', [BUILT_SCRIPT], { stdio: 'pipe' }); + status = 0; + } catch (error) { + status = error.status; + } + assert.equal(status, 2); + }); +}); + +describe('edge-probe: CLI JSON.parse error handling (RR-10)', () => { + test('invalid requirements JSON exits with status 2 (handled error, not uncaught throw)', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-')); + const badJson = path.join(dir, 'bad-req.json'); + fs.writeFileSync(badJson, 'not valid json {{{'); + try { + const r = spawnSync(process.execPath, [BUILT_SCRIPT, badJson], { stdio: 'pipe' }); + assert.equal(r.status, 2); + } finally { + cleanup(dir); + } + }); + test('invalid resolutions JSON exits with status 2 (handled error, not uncaught throw)', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-')); + const goodReq = path.join(dir, 'req.json'); + const badRes = path.join(dir, 'bad-res.json'); + fs.writeFileSync(goodReq, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }])); + fs.writeFileSync(badRes, 'not valid json {{{'); + try { + const r = spawnSync(process.execPath, [BUILT_SCRIPT, goodReq, badRes], { stdio: 'pipe' }); + assert.equal(r.status, 2); + } finally { + cleanup(dir); + } + }); + test('valid requirements file exits 0 and stdout is parseable JSON', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-')); + const reqPath = path.join(dir, 'req.json'); + fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }])); + try { + const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe', encoding: 'utf8' }); + assert.equal(r.status, 0); + const rep = JSON.parse(r.stdout); + assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } }); + } finally { + cleanup(dir); + } + }); +}); + +describe('edge-probe: proposeEdges — empty-shapes override (RR-06)', () => { + test('shapes: [] returns zero edges (explicit empty-shapes override)', () => { + const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] }); + assert.deepEqual(edges, []); + }); + test('absent shapes key classifies from prose (no override)', () => { + const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals' }); + assert.ok(edges.length > 0, 'should classify collection edges from prose'); + }); + test('shapes: [collection] overrides prose and proposes collection categories', () => { + const edges = ep.proposeEdges({ id: 'R9', text: 'opaque text with no cues', shapes: ['collection'] }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']); + }); +}); + +describe('edge-probe: proposeEdges — invalid authored shapes fail closed (re-review #3 High)', () => { + // A non-empty but INVALID shapes array must NOT silently suppress every probe. + // shapes:['numeric'] (typo for the locked 'numeric-range') previously passed + // Array.isArray, matched no category, and returned applicable:0 — failing OPEN. + test('rejects an unknown shape value (typo for a locked shape)', () => { + assert.throws( + () => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric'] }), + /invalid shape/i, + ); + }); + test('rejects a mixed array where one entry is invalid', () => { + assert.throws( + () => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric-range', 'bogus'] }), + /invalid shape/i, + ); + }); + test('rejects a non-string shape entry', () => { + assert.throws( + () => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: [42] }), + /invalid shape/i, + ); + }); + test('analyzeCoverage propagates the invalid-shape throw', () => { + assert.throws( + () => ep.analyzeCoverage([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]), + /invalid shape/i, + ); + }); + test('CLI exits 2 (handled) on an invalid authored shape, not an uncaught trace', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-shape-')); + const reqPath = path.join(dir, 'req.json'); + fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }])); + try { + const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe' }); + assert.equal(r.status, 2); + } finally { + cleanup(dir); + } + }); + test('a valid locked shape still proposes its categories (no false rejection)', () => { + const edges = ep.proposeEdges({ id: 'R1', text: 'opaque', shapes: ['numeric-range'] }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']); + }); + test('shapes: [] remains a valid zero-edge override (RR-06 intact)', () => { + assert.deepEqual(ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] }), []); + }); +}); + +describe('edge-probe: input validation & orphan-resolution rejection (adversarial review)', () => { + // HIGH: a resolution whose (requirement_id, category) matches no proposed edge — a typo'd + // category or a non-applicable one — was silently DROPPED, so an author who typos `precison` + // sees the precision edge as still-unresolved with no error (a confirmed money-rounding exploit). + test('rejects an orphan resolution (typo category — no matching proposed edge)', () => { + assert.throws( + () => ep.analyzeCoverage( + [{ id: 'R1', text: 'Round a number to N decimal places' }], + [{ requirement_id: 'R1', category: 'precison', status: 'resolved', verification: 'explicit', resolution: 'AC: precision handled' }], + ), + /unknown resolution|no matching proposed edge/i, + ); + }); + test('rejects a resolution for a valid-but-non-applicable category', () => { + // 'encoding' is a real taxonomy id but applies to text, not the numeric-range requirement. + assert.throws( + () => ep.analyzeCoverage( + [{ id: 'R1', text: 'Round a number to N decimal places' }], + [{ requirement_id: 'R1', category: 'encoding', status: 'resolved', verification: 'explicit', resolution: 'AC' }], + ), + /unknown resolution|no matching proposed edge/i, + ); + }); + test('a matching resolution still resolves (no false orphan rejection)', () => { + const rep = ep.analyzeCoverage( + [{ id: 'R1', text: 'Round a number to N decimal places' }], + [{ requirement_id: 'R1', category: 'precision', status: 'resolved', verification: 'explicit', resolution: 'AC: precision tested' }], + ); + assert.equal(rep.coverage.resolved, 1); + }); + test('rejects requirements that is not an array', () => { + assert.throws(() => ep.analyzeCoverage('nope'), /requirements must be an array/i); + }); + test('rejects a duplicate requirement id', () => { + assert.throws( + () => ep.analyzeCoverage([{ id: 'R1', text: 'a' }, { id: 'R1', text: 'b' }]), + /duplicate requirement/i, + ); + }); + test('rejects a truthy non-array shapes (string instead of array)', () => { + // A bare string `shapes: "numeric-range"` previously fell through to prose classification, + // silently ignoring the authored override instead of honoring or rejecting it. + assert.throws( + () => ep.proposeEdges({ id: 'R1', text: 'x', shapes: 'numeric-range' }), + /shapes must be an array/i, + ); + }); + test('rejects a missing requirement id', () => { + assert.throws(() => ep.proposeEdges({ text: 'x' }), /requirement id must be a non-empty string/i); + }); + test('rejects an empty requirement id', () => { + assert.throws(() => ep.proposeEdges({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i); + }); + test('rejects a non-string requirement text', () => { + assert.throws(() => ep.proposeEdges({ id: 'R1', text: 42 }), /text must be a string/i); + }); + test('rejects a missing requirement text when no shapes override (M2 fail-open)', () => { + // Without text or an authored shape, prose classification yields zero shapes → zero edges → + // the requirement is silently DROPPED from coverage with no signal — the exact fail-open this + // feature exists to eliminate. The edge adapter's `text` is required, so reject it. + assert.throws( + () => ep.proposeEdges({ id: 'R1' }), + /text must be a non-empty string when no shapes override/i, + ); + assert.throws( + () => ep.analyzeCoverage([{ id: 'R1' }]), + /text must be a non-empty string when no shapes override/i, + ); + }); + test('rejects an empty/whitespace requirement text when no shapes override (M2)', () => { + assert.throws(() => ep.proposeEdges({ id: 'R1', text: '' }), /text must be a non-empty string when no shapes override/i); + assert.throws(() => ep.proposeEdges({ id: 'R1', text: ' ' }), /text must be a non-empty string when no shapes override/i); + }); + test('allows missing/empty text WHEN an explicit shapes override is provided (M2 legitimate path)', () => { + // An authored `shapes` array (including `[]` for "no applicable categories") opts out of prose + // classification, so `text` is not required — this must remain valid. + assert.deepEqual(ep.proposeEdges({ id: 'R1', shapes: [] }), []); + const edges = ep.proposeEdges({ id: 'R1', shapes: ['numeric-range'] }); + assert.ok(edges.length > 0, 'an explicit shape override must still propose edges without text'); + }); +}); + +describe('edge-probe: validateResolution — explicit-needs-resolution (RR-07, re-cut)', () => { + test('rejects resolved/explicit with empty resolution string', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: '' }), + /explicit requires a resolution/i, + ); + }); + test('rejects resolved/explicit with whitespace-only resolution', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: ' ' }), + /explicit requires a resolution/i, + ); + }); + test('rejects resolved/explicit with missing resolution', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit' }), + /explicit requires a resolution/i, + ); + }); + test('accepts resolved/explicit with a non-empty resolution', () => { + assert.equal( + ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC#3: boundary tested in suite' }), + true, + ); + }); +}); + +describe('edge-probe: validateResolution — backstop-needs-resolution (RR-07 follow-up, re-cut)', () => { + test('rejects resolved/backstop with empty resolution string', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: '' }), + /backstop requires a resolution/i, + ); + }); + test('rejects resolved/backstop with whitespace-only resolution', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: ' ' }), + /backstop requires a resolution/i, + ); + }); + test('rejects resolved/backstop with missing resolution', () => { + assert.throws( + () => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop' }), + /backstop requires a resolution/i, + ); + }); + test('accepts resolved/backstop with a non-empty resolution note', () => { + assert.equal( + ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: 'held-out: covered by integration fuzz suite' }), + true, + ); + }); +}); + +describe('edge-probe: analyzeCoverage — duplicate rejection (RR-09)', () => { + const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }]; + test('rejects duplicate (requirement_id, category) resolution', () => { + assert.throws( + () => ep.analyzeCoverage(reqs, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' }, + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' }, + ]), + /duplicate resolution/i, + ); + }); + test('distinct pairs still analyze without throwing', () => { + // mirrors fixture 06-resolved-mixed + assert.doesNotThrow(() => ep.analyzeCoverage(reqs, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' }, + { requirement_id: 'R1', category: 'ordering', status: 'dismissed', resolution: null, reason: 'output is canonically sorted; no tie possible' }, + ])); + }); +}); + +describe('edge-probe: golden fixtures', () => { + const root = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures'); + const fixtures = fs.readdirSync(root).filter((d) => + fs.statSync(path.join(root, d)).isDirectory()); + assert.ok(fixtures.length >= 6, 'expected at least 6 fixtures'); + for (const name of fixtures) { + test(`fixture ${name} matches its golden coverage`, () => { + const dir = path.join(root, name); + const reqs = JSON.parse(fs.readFileSync(path.join(dir, 'requirements.json'), 'utf8')); + const resPath = path.join(dir, 'resolutions.json'); + const res = fs.existsSync(resPath) ? JSON.parse(fs.readFileSync(resPath, 'utf8')) : []; + const expected = JSON.parse(fs.readFileSync(path.join(dir, 'expected-coverage.json'), 'utf8')); + assert.deepEqual(ep.analyzeCoverage(reqs, res), expected); + }); + } +}); diff --git a/tests/probe-core.property.test.cjs b/tests/probe-core.property.test.cjs new file mode 100644 index 000000000..a7a7de064 --- /dev/null +++ b/tests/probe-core.property.test.cjs @@ -0,0 +1,167 @@ +'use strict'; + +/** + * Property-based tests for probe-core.cjs (ADR-550 Decision 7). + * + * Module: gsd-core/bin/lib/probe-core.cjs (generated from src/probe-core.cts) + * Exercised: analyzeCoverage(items, resolutions?, validators) — the generic + * merge/rollup/orphan-reject engine shared by the edge probe and the #644 + * prohibition probe. + * + * trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a + * transformation/rollup module (items × resolutions → CoverageReport) — exactly the + * class the predicate covers. The example-based suite (tests/probe-core.test.cjs) + * pins specific scenarios; these properties pin the algebraic invariants that must + * hold for EVERY valid scenario. + * + * Properties tested: + * (a) closed-set identity: applicable === resolved + unresolved (and === items.length) + * (b) byVerification sums ≤ resolved (dismissed is closed but unverified) + * (c) per-tier byVerification ≤ resolved, and only `resolved`-status items are counted + * (d) determinism: same input → identical CoverageReport (stable rollup) + * (e) orphan rejection is stable: a resolution matching no proposed item always throws + */ + +const { describe, test } = require('node:test'); +const path = require('node:path'); +const fc = require('./helpers/fast-check-setup.cjs'); + +const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs'); +const pc = require(BUILT_SCRIPT); + +// The same representative validators bundle the edge adapter injects (see +// tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe. +const VALIDATORS = { + categories: ['adjacency', 'empty', 'ordering'], + verification: ['explicit', 'backstop'], + requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] }, +}; + +function bareItem(requirement_id, category) { + return { + requirement_id, + category, + status: 'unresolved', + verification: null, + resolution: null, + reason: null, + probe: `probe-for-${category}`, + }; +} + +const catArb = fc.constantFrom(...VALIDATORS.categories); +const idArb = fc.constantFrom('R1', 'R2', 'R3', 'R4', 'R5'); +const keyArb = fc.record({ requirement_id: idArb, category: catArb }); +// Unique (requirement_id, category) keys — the merge keys analyzeCoverage maps on. +const keyOf = (k) => `${k.requirement_id}::${k.category}`; +const uniqueKeysArb = fc.uniqueArray(keyArb, { selector: keyOf, minLength: 0, maxLength: 12 }); + +// Each unique item key gets one resolution disposition. Resolution text/reason use fixed +// non-empty literals — the counting invariants are independent of their content, and this +// keeps the generator off the validateResolution rejection paths (which the example suite +// already covers exhaustively). +const DISPOSITIONS = ['none', 'resolved-explicit', 'resolved-backstop', 'dismissed', 'unresolved']; + +function resolutionFor(k, disposition) { + const base = { requirement_id: k.requirement_id, category: k.category }; + switch (disposition) { + case 'resolved-explicit': + return { ...base, status: 'resolved', verification: 'explicit', resolution: 'AC#1' }; + case 'resolved-backstop': + return { ...base, status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' }; + case 'dismissed': + return { ...base, status: 'dismissed', reason: 'bounded enum — not applicable' }; + case 'unresolved': + return { ...base, status: 'unresolved' }; + default: // 'none' — author left no resolution; item rolls up verbatim (bare unresolved) + return null; + } +} + +// A fully valid scenario: unique items (all bare-unresolved) + a per-item resolution choice. +const scenarioArb = uniqueKeysArb.chain((keys) => + fc.tuple(...keys.map(() => fc.constantFrom(...DISPOSITIONS))).map((choices) => { + const items = keys.map((k) => bareItem(k.requirement_id, k.category)); + const resolutions = []; + keys.forEach((k, i) => { + const r = resolutionFor(k, choices[i]); + if (r) resolutions.push(r); + }); + return { items, resolutions }; + }), +); + +describe('probe-core property: analyzeCoverage algebraic invariants', () => { + test('(a) closed-set identity: applicable === resolved + unresolved === items.length', () => { + fc.assert( + fc.property(scenarioArb, ({ items, resolutions }) => { + const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS); + return ( + coverage.applicable === coverage.resolved + coverage.unresolved && + coverage.applicable === items.length + ); + }), + ); + }); + + test('(b) sum(byVerification) ≤ resolved — dismissed counts closed but unverified', () => { + fc.assert( + fc.property(scenarioArb, ({ items, resolutions }) => { + const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS); + const verifiedTotal = Object.values(coverage.byVerification).reduce((a, b) => a + b, 0); + return verifiedTotal <= coverage.resolved && verifiedTotal >= 0; + }), + ); + }); + + test('(c) byVerification only counts resolved-status items, and matches a direct recount', () => { + fc.assert( + fc.property(scenarioArb, ({ items, resolutions }) => { + const rep = pc.analyzeCoverage(items, resolutions, VALIDATORS); + for (const tier of VALIDATORS.verification) { + const recount = rep.items.filter((i) => i.status === 'resolved' && i.verification === tier).length; + if (rep.coverage.byVerification[tier] !== recount) return false; + if (rep.coverage.byVerification[tier] > rep.coverage.resolved) return false; + } + return true; + }), + ); + }); + + test('(d) determinism: identical inputs produce an identical CoverageReport', () => { + fc.assert( + fc.property(scenarioArb, ({ items, resolutions }) => { + const a = pc.analyzeCoverage(items, resolutions, VALIDATORS); + const b = pc.analyzeCoverage(items, resolutions, VALIDATORS); + return JSON.stringify(a) === JSON.stringify(b); + }), + ); + }); +}); + +describe('probe-core property: orphan rejection is stable', () => { + // An orphan resolution carries an id ('Z9') that no generated item ever uses, so its + // (requirement_id, category) key never matches a proposed item. The resolution is itself + // structurally VALID (a bare unresolved), so it clears validateResolution and reaches the + // orphan-reject guard — isolating that guard from input-validation throws. + const orphanScenarioArb = fc.record({ + keys: uniqueKeysArb, + orphanCategory: catArb, + }); + + test('(e) a resolution matching no proposed item always throws', () => { + fc.assert( + fc.property(orphanScenarioArb, ({ keys, orphanCategory }) => { + const items = keys.map((k) => bareItem(k.requirement_id, k.category)); + const orphan = { requirement_id: 'Z9', category: orphanCategory, status: 'unresolved' }; + let threwForOrphan = false; + try { + pc.analyzeCoverage(items, [orphan], VALIDATORS); + } catch (e) { + threwForOrphan = /unknown resolution|no matching proposed item/i.test(e.message); + } + return threwForOrphan; + }), + ); + }); +}); diff --git a/tests/probe-core.test.cjs b/tests/probe-core.test.cjs new file mode 100644 index 000000000..3d7300d2a --- /dev/null +++ b/tests/probe-core.test.cjs @@ -0,0 +1,335 @@ +/** + * probe-core reference-model unit tests (ADR-550 Decision 7). + * + * probe-core is the GENERIC spec-phase probe resolution model extracted from the + * edge-probe (the first adapter): the resolution lifecycle, the two-axis + * status×verification re-cut, `validateResolution`/`validateRequirement`, the + * `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject + * engine, the `byVerification` rollup, and the `runProbeCli` I/O scaffold. + * + * Asserts the LOCKED export surface against the BUILT artifact + * (`gsd-core/bin/lib/probe-core.cjs`), which `npm run build:lib` (run by pretest / + * the run-tests sentinel) emits from `src/probe-core.cts`. + * + * The injected runtime validators are the enforcement contract (ADR-550 #5): the + * CLI runs over JSON where TS types are erased, so `analyzeCoverage` is told its + * probe's closed vocabularies — `{ categories, verification, requiredFieldsByVerification }` + * — rather than relying on the type system. These tests pin the validators' behavior. + */ +'use strict'; +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const path = require('node:path'); + +const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs'); +const pc = require(BUILT_SCRIPT); + +// A representative validators bundle — the shape the edge adapter injects, used here +// to exercise the generic engine independent of any one probe. +const VALIDATORS = { + categories: ['adjacency', 'empty', 'ordering'], + verification: ['explicit', 'backstop'], + requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] }, +}; + +function item(category, overrides = {}) { + return { + requirement_id: 'R1', + category, + status: 'unresolved', + verification: null, + resolution: null, + reason: null, + probe: `probe-for-${category}`, + ...overrides, + }; +} + +const UNRESOLVED_ITEMS = [item('adjacency'), item('empty'), item('ordering')]; + +describe('probe-core: VALID_STATUS is the re-cut lifecycle enum', () => { + test('exposes exactly resolved | dismissed | unresolved (no covered/backstop)', () => { + assert.deepEqual([...ep_sorted(pc.VALID_STATUS)], ['dismissed', 'resolved', 'unresolved']); + assert.ok(!pc.VALID_STATUS.includes('covered'), 'covered must not survive the re-cut as a status'); + assert.ok(!pc.VALID_STATUS.includes('backstop'), 'backstop must not survive the re-cut as a status'); + }); +}); + +function ep_sorted(arr) { + return [...arr].sort(); +} + +describe('probe-core: validateResolution (status×verification)', () => { + const v = (r) => pc.validateResolution(r, VALIDATORS); + test('rejects an unknown status', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'maybe' }), /invalid status/i); + }); + test('rejects a former covered status (re-cut: no longer a valid status)', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'covered', resolution: 'x' }), /invalid status/i); + }); + test('rejects dismissed without a reason', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: '' }), /dismissed requires a reason/i); + }); + test('accepts dismissed with a reason', () => { + assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' }), true); + }); + test('rejects resolved with a missing verification tier', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', resolution: 'AC' }), /verification/i); + }); + test('rejects resolved with an unknown verification tier', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'judgment', resolution: 'AC' }), /invalid verification/i); + }); + test('rejects resolved/explicit with empty resolution text', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: ' ' }), /explicit requires a resolution/i); + }); + test('rejects resolved/backstop with missing resolution note', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop' }), /backstop requires a resolution/i); + }); + test('accepts resolved/explicit with a resolution', () => { + assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' }), true); + }); + test('accepts resolved/backstop with a resolution note', () => { + assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' }), true); + }); + // Re-review #5 Medium — fail closed across the FULL status×verification model, not just + // `resolved`. The invariant (probe-core header): verification is null unless status is + // resolved. A dismissed/unresolved resolution carrying a tier otherwise merges verbatim + // (analyzeCoverage line ~189), breaking the model for the second adapter (#644) that + // inherits this seam. + test('rejects a dismissed resolution carrying a verification tier (null unless resolved)', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'n/a', verification: 'explicit' }), /verification must be null/i); + }); + test('rejects an unresolved resolution carrying a verification tier', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', verification: 'backstop' }), /verification must be null/i); + }); + // An unresolved resolution is an UNACTED item — a populated resolution/reason payload is an + // authoring mistake (the author meant resolved/dismissed) that today is silently dropped into + // the unresolved count with no error pointing at it. Reject the payload. + test('rejects an unresolved resolution carrying a resolution payload', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', resolution: 'AC#1' }), /unresolved must not carry/i); + }); + test('rejects an unresolved resolution carrying a reason payload', () => { + assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', reason: 'because' }), /unresolved must not carry/i); + }); + test('accepts a bare unresolved resolution (no payload — a harmless no-op merge)', () => { + assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved' }), true); + }); +}); + +describe('probe-core: validateRequirement (generic id/text only)', () => { + test('rejects a missing id', () => { + assert.throws(() => pc.validateRequirement({ text: 'x' }), /requirement id must be a non-empty string/i); + }); + test('rejects an empty id', () => { + assert.throws(() => pc.validateRequirement({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i); + }); + test('rejects a non-string text', () => { + assert.throws(() => pc.validateRequirement({ id: 'R1', text: 42 }), /text must be a string/i); + }); + test('accepts a valid requirement', () => { + assert.doesNotThrow(() => pc.validateRequirement({ id: 'R1', text: 'a testable statement' })); + }); +}); + +describe('probe-core: analyzeCoverage (merge · rollup · byVerification)', () => { + test('no resolutions → every item unresolved; resolved 0; byVerification zeroed per tier', () => { + const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [], VALIDATORS); + assert.deepEqual(rep.coverage, { + applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 }, + }); + }); + test('merges a resolved/explicit resolution and counts byVerification.explicit', () => { + const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' }, + ], VALIDATORS); + const adj = rep.items.find((i) => i.category === 'adjacency'); + assert.equal(adj.status, 'resolved'); + assert.equal(adj.verification, 'explicit'); + assert.equal(adj.resolution, 'AC#6: touching intervals merge'); + assert.equal(rep.coverage.resolved, 1); + assert.equal(rep.coverage.unresolved, 2); + assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 }); + }); + test('a dismissed item counts toward coverage.resolved (closed set) but NOT byVerification', () => { + // coverage.resolved preserves the pre-re-cut "closed" semantic = applicable - unresolved + // (covered + dismissed + backstop), per edge-probe.md and the migration contract. + const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted; no tie' }, + ], VALIDATORS); + assert.equal(rep.coverage.resolved, 1); + assert.equal(rep.coverage.unresolved, 2); + assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 }); + const ord = rep.items.find((i) => i.category === 'ordering'); + assert.equal(ord.status, 'dismissed'); + assert.equal(ord.verification, null); + assert.equal(ord.reason, 'canonically sorted; no tie'); + }); + test('mixed resolved/explicit + backstop + dismissed: resolved = closed = applicable - unresolved', () => { + const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' }, + { requirement_id: 'R1', category: 'empty', status: 'resolved', verification: 'backstop', resolution: 'held-out empty-input PBT' }, + { requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' }, + ], VALIDATORS); + assert.equal(rep.coverage.applicable, 3); + assert.equal(rep.coverage.unresolved, 0); + assert.equal(rep.coverage.resolved, 3); // 2 resolved-status + 1 dismissed + assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 1 }); + }); + test('an all-dismissed run is CLOSED but NOT affirmatively covered (byVerification is the honest gate)', () => { + // Re-review #5 Medium — coverage.resolved is the closed set (resolved + dismissed), kept + // count-preserved per the blessed migration contract. So an all-dismissed spec has + // resolved === applicable while NOTHING was affirmatively resolved/backstopped. A CI gate + // keying on `resolved === applicable` would read green here; the honest signal is + // byVerification (all tiers zero). This test locks that distinction. + const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' }, + { requirement_id: 'R1', category: 'empty', status: 'dismissed', reason: 'guaranteed non-empty' }, + { requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' }, + ], VALIDATORS); + assert.equal(rep.coverage.unresolved, 0); + assert.equal(rep.coverage.resolved, 3); // closed set counts the dismissals + assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 }); // nothing affirmatively verified + }); + test('rejects a duplicate (requirement_id, category) resolution', () => { + assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' }, + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' }, + ], VALIDATORS), /duplicate resolution/i); + }); + test('rejects an orphan resolution (no matching proposed item)', () => { + assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC' }, + ], VALIDATORS), /unknown resolution|no matching proposed/i); + }); + test('propagates an invalid-resolution throw from validateResolution', () => { + assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [ + { requirement_id: 'R1', category: 'empty', status: 'dismissed' }, + ], VALIDATORS), /dismissed requires a reason/i); + }); + test('rejects items that is not an array', () => { + assert.throws(() => pc.analyzeCoverage('nope', [], VALIDATORS), /items must be an array/i); + }); + test('rejects a proposed item whose category is not in validators.categories', () => { + // Adapter self-consistency: an item carrying a category outside the probe's closed + // vocabulary is an adapter bug, caught here rather than silently rolled up. + assert.throws(() => pc.analyzeCoverage([item('bogus-category')], [], VALIDATORS), /unknown category/i); + }); + + // m1: a verbatim item (no matching author resolution) is rolled up as-is, so its OWN + // status/fields must also be validated. The edge adapter only proposes `unresolved` items, but + // the prohibition adapter (#644) proposes LLM-generated items that arrive already populated — + // an out-of-enum status or a `dismissed` with no reason must fail closed, not count as closed. + test('m1: rejects a verbatim item carrying an out-of-enum status (e.g. the dropped "covered")', () => { + assert.throws( + () => pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [], VALIDATORS), + /invalid status/i, + ); + }); + test('m1: rejects a verbatim item dismissed without a reason', () => { + assert.throws( + () => pc.analyzeCoverage([item('adjacency', { status: 'dismissed' })], [], VALIDATORS), + /dismissed requires a reason/i, + ); + }); + test('m1: rejects a verbatim resolved item missing its verification tier', () => { + assert.throws( + () => pc.analyzeCoverage([item('adjacency', { status: 'resolved', resolution: 'AC' })], [], VALIDATORS), + /requires a verification tier/i, + ); + }); + test('m1: a matching resolution still governs — item validation targets VERBATIM items only', () => { + // When a resolution matches, the item is rebuilt from the (already-validated) resolution, so a + // pre-populated item status is irrelevant and must NOT cause a throw. Guards against the m1 + // fix over-reaching into the merge path. + const rep = pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [ + { requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' }, + ], VALIDATORS); + const adj = rep.items.find((i) => i.category === 'adjacency'); + assert.equal(adj.status, 'resolved'); + assert.equal(adj.verification, 'explicit'); + assert.equal(rep.coverage.resolved, 1); + }); +}); + +describe('probe-core: runProbeCli (generic I/O scaffold, injected io)', () => { + const report = { items: [], coverage: { applicable: 0, resolved: 0, unresolved: 0, byVerification: {} } }; + test('no requirements path → writes usage to stderr and exits 2', () => { + let code; let err = ''; + pc.runProbeCli(() => report, { + usage: 'demo-probe.cjs [resolutions.json]', + argv: ['node', 'demo'], writeErr: (s) => { err += s; }, exit: (c) => { code = c; }, + }); + assert.equal(code, 2); + assert.match(err, /usage: demo-probe\.cjs/); + }); + test('valid requirements path → calls analyze and prints the report as JSON', () => { + let out = ''; + pc.runProbeCli((reqs, res) => { + assert.deepEqual(reqs, [{ id: 'R1', text: 'x' }]); + assert.deepEqual(res, []); + return report; + }, { + usage: 'demo', argv: ['node', 'demo', '/fake/req.json'], + readFile: () => '[{"id":"R1","text":"x"}]', write: (s) => { out += s; }, exit: () => {}, + }); + assert.deepEqual(JSON.parse(out), report); + }); + test('reads the optional resolutions file when a second path is given', () => { + let seenRes; + pc.runProbeCli((reqs, res) => { seenRes = res; return report; }, { + usage: 'demo', argv: ['node', 'demo', '/req.json', '/res.json'], + readFile: (p) => (p === '/res.json' ? '[{"requirement_id":"R1"}]' : '[{"id":"R1"}]'), + write: () => {}, exit: () => {}, + }); + assert.deepEqual(seenRes, [{ requirement_id: 'R1' }]); + }); + test('invalid requirements JSON → exits 2 (handled, not an uncaught throw)', () => { + let code; + pc.runProbeCli(() => report, { + usage: 'demo', argv: ['node', 'demo', '/req.json'], + readFile: () => 'not json {{{', writeErr: () => {}, exit: (c) => { code = c; }, + }); + assert.equal(code, 2); + }); + test('an analyze throw → exits 2 (the engine fail-closed surfaces, never silently passes)', () => { + let code; + pc.runProbeCli(() => { throw new Error('boom'); }, { + usage: 'demo', argv: ['node', 'demo', '/req.json'], + readFile: () => '[]', writeErr: () => {}, exit: (c) => { code = c; }, + }); + assert.equal(code, 2); + }); + // Re-review #5 Low — the scaffold trusts each adapter's `as` cast for the returned report. + // A future adapter (#644) that forgets to validate inside its closure would otherwise have a + // structurally-broken report written as green output. Guard the report shape structurally. + test('a structurally-invalid report from analyze → exits 2 and writes nothing (no silent malformed output)', () => { + let code; let out = ''; + pc.runProbeCli(() => ({ nope: true }), { + usage: 'demo', argv: ['node', 'demo', '/req.json'], + readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; }, + }); + assert.equal(code, 2); + assert.equal(out, ''); + }); + test('a report whose coverage object carries non-numeric counts → exits 2 (partial-guard case)', () => { + // items[] is well-formed and `coverage` IS an object, so the cheaper checks pass — this + // exercises the numeric-count branch of the structural guard specifically. + let code; let out = ''; + pc.runProbeCli(() => ({ items: [], coverage: { applicable: 'x', resolved: null, unresolved: undefined, byVerification: {} } }), { + usage: 'demo', argv: ['node', 'demo', '/req.json'], + readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; }, + }); + assert.equal(code, 2); + assert.equal(out, ''); + }); + test('a well-formed report still writes and does not trip the structural guard', () => { + let out = ''; + pc.runProbeCli(() => report, { + usage: 'demo', argv: ['node', 'demo', '/req.json'], + readFile: () => '[]', write: (s) => { out += s; }, exit: () => {}, + }); + assert.deepEqual(JSON.parse(out), report); + }); +}); diff --git a/tests/workflow-size-baseline.json b/tests/workflow-size-baseline.json index e8b235b3a..3acfb3825 100644 --- a/tests/workflow-size-baseline.json +++ b/tests/workflow-size-baseline.json @@ -51,7 +51,7 @@ "note.md": 6563, "pause-work.md": 13654, "plan-milestone-gaps.md": 11765, - "plan-phase.md": 93135, + "plan-phase.md": 94253, "plan-review-convergence.md": 22949, "plant-seed.md": 11741, "pr-branch.md": 4994, @@ -72,7 +72,7 @@ "ship.md": 20896, "sketch-wrap-up.md": 14223, "sketch.md": 19960, - "spec-phase.md": 15131, + "spec-phase.md": 23094, "spike-wrap-up.md": 15092, "spike.md": 24517, "stats.md": 6718,