feat(spec-phase): spec-completeness edge-probe (#550) (#584)

* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
This commit is contained in:
Rezolv
2026-06-12 11:05:31 -04:00
committed by GitHub
parent e4f0910d62
commit 3e836fef0d
38 changed files with 2718 additions and 4 deletions

View File

@@ -0,0 +1,5 @@
---
type: Added
pr: 584
---
spec-phase: spec-completeness edge-probe — a taxonomy-driven Step 5.5 that walks each SPEC requirement against a closed 8-category edge taxonomy (boundary, adjacency, empty/degenerate, encoding, ordering, precision, idempotency, concurrency), proposes concrete candidate edges, and resolves each to covered/dismissed/backstop/unresolved. covered edges add acceptance criteria the planner lifts into must_haves.truths; a soft gate flags unresolved edges. Additive and optional: existing SPECs without an Edge Coverage section remain valid.

2
.gitignore vendored
View File

@@ -71,6 +71,8 @@ build/
/gsd-core/bin/lib/research-provider.cjs
/gsd-core/bin/lib/package-legitimacy.cjs
/gsd-core/bin/lib/semver-compare.cjs
/gsd-core/bin/lib/edge-probe.cjs
/gsd-core/bin/lib/probe-core.cjs
/gsd-core/bin/lib/config-types.cjs
/gsd-core/bin/lib/cli-exit.cjs
/gsd-core/bin/lib/code-review-flags.cjs

View File

@@ -204,6 +204,12 @@ The GSD-RESEARCH capability behind an L2-hybrid seam: code owns cache + provider
### UAT-Passed Predicate
Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fields with markdown-aware parsing that ignores false-positive contexts (frontmatter body, fenced code, HTML comments, blockquotes). Returns `passed: true` only when all required checks pass; supports `--require-verification` to demand at least one VERIFICATION.md file alongside UAT results. Output envelope: `{ passed, uat_files[], verification_files[], checks[], blockers[], policy }`. Source: `gsd-core/bin/lib/uat-predicate.cjs` (generated from `src/uat-predicate.cts`). Wired via `phase uat-passed` alias → `phase-command-router` → `cmdPhaseUatPassed`.
### Probe Core Module
Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the prohibition probe (#644) next. Exports: `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli`. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module.
### Edge Probe Module
First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the 8 edge categories), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module.
### MVP Mode
Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`.

View File

@@ -90,6 +90,34 @@ Manage GSD workspaces — create, list, or remove isolated workspace environment
---
### `/gsd-spec-phase`
Clarify WHAT a phase delivers through Socratic questioning with quantitative ambiguity scoring, then probe for omitted edges. Produces `SPEC.md` before discuss-phase.
| Argument | Required | Description |
|----------|----------|-------------|
| `N` | Yes | Phase number |
| Flag | Description |
|------|-------------|
| `--auto` | Skip interactive questions; Claude selects recommended defaults and writes SPEC.md |
| `--text` | Use plain-text numbered lists instead of TUI menus (required for `/rc` remote sessions) |
**Position in workflow:** `spec-phase → discuss-phase → plan-phase → execute-phase → verify`
**Edge Coverage (Step 5.5):** After the ambiguity gate passes, spec-phase runs an edge-completeness probe over each requirement. It raises only applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency), proposes one concrete candidate edge per category, and records each as `covered` / `dismissed` (reason required) / `backstop` / `unresolved` in a `## Edge Coverage` SPEC section. Unresolved applicable edges soft-gate the spec (Resolve / Write-anyway-flagged / Keep-probing); `covered` and `backstop` edges are later lifted into plan-phase `must_haves`. Under `--auto` the probe **never auto-dismisses** — it auto-covers where a defensible acceptance criterion exists, otherwise auto-backstops.
**Prerequisites:** `.planning/ROADMAP.md` exists
**Produces:** `{phase}-SPEC.md` (with a `## Edge Coverage` section)
```bash
/gsd-spec-phase 1 # Interactive spec + edge probe for phase 1
/gsd-spec-phase 3 --auto # Auto-select defaults; never auto-dismisses an edge
/gsd-spec-phase 2 --text # Plain-text menus for remote sessions
```
---
### `/gsd-discuss-phase`
Gather phase context through adaptive questioning before planning.

View File

@@ -165,6 +165,7 @@
- [Milestone Tag Creation Toggle](#141-milestone-tag-creation-toggle)
- [Structured JSON Error Mode](#142-structured-json-error-mode)
- [UAT-Passed Predicate](#143-uat-passed-predicate)
- [Spec-Phase Edge-Completeness Probe](#144-spec-phase-edge-completeness-probe)
---
@@ -3086,3 +3087,34 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev
- [docs index](README.md)
**Reference:** [JSON Error Mode](json-errors.md)
---
### 144. Spec-Phase Edge-Completeness Probe
**Command:** `/gsd-spec-phase`
**Purpose:** Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as `Step 5.5` of spec-phase, after the ambiguity gate.
**Behavior:** For each SPEC requirement the probe classifies its data/behavior shape, then raises only the *applicable* categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:
| State | Meaning | Downstream effect |
|-------|---------|-------------------|
| `covered` | An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into `plan-phase` `must_haves.truths` |
| `dismissed` | The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
| `backstop` | Intent recorded, needs a held-out/property-based test | Lifted into `must_haves.truths` as a non-inferable check |
| `unresolved` | Deferred | Soft-gates the spec; row stamped `⚠ Edge unresolved — planner must treat as assumption` |
The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`.
The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation.
**Requirements:**
- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a `## Edge Coverage` SPEC section.
- REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
- REQ-EDGE-03: A `dismissed` resolution MUST require a non-empty reason.
- REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
- REQ-EDGE-05: `--auto` MUST never auto-dismiss — auto-cover or auto-backstop only.
- REQ-EDGE-06: `plan-phase` MUST lift `covered` criteria and `backstop` notes into `must_haves.truths`.
**Reference:** [Edge Probe](../gsd-core/references/edge-probe.md)

View File

@@ -209,6 +209,7 @@
"decimal-phase-calculation.md",
"doc-conflict-engine.md",
"domain-probes.md",
"edge-probe.md",
"execute-mvp-tdd.md",
"executor-examples.md",
"gate-prompts.md",
@@ -295,6 +296,7 @@
"decisions.cjs",
"docs.cjs",
"drift.cjs",
"edge-probe.cjs",
"fallow-runner.cjs",
"federated-config.cjs",
"frontmatter.cjs",
@@ -329,6 +331,7 @@
"phases-command-router.cjs",
"plan-scan.cjs",
"planning-workspace.cjs",
"probe-core.cjs",
"profile-output.cjs",
"profile-pipeline.cjs",
"project-root.cjs",

View File

@@ -264,7 +264,7 @@ Full roster at `gsd-core/workflows/*.md`. Workflows are thin orchestrators that
---
## References (68 shipped)
## References (69 shipped)
Full roster at `gsd-core/references/*.md`. References are shared knowledge documents that workflows and agents `@-reference`. The groupings below match [`docs/ARCHITECTURE.md`](ARCHITECTURE.md#references-gsd-corereferencesmd) — core, workflow, thinking-model clusters, and the modular planner decomposition.
@@ -300,6 +300,7 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum
| `context-budget.md` | Context window budget allocation rules. |
| `continuation-format.md` | Session continuation/resume format. |
| `domain-probes.md` | Domain-specific probing questions for discuss-phase. |
| `edge-probe.md` | Spec-phase edge-completeness probe — 8-category edge taxonomy, shape classification, and the `requirements → checks → verifier` resolution model (Step 5.5). |
| `gate-prompts.md` | Gate/checkpoint prompt templates. |
| `scout-codebase.md` | Phase-type→codebase-map selection table for discuss-phase scout step (extracted via #2551). |
| `revision-loop.md` | Plan revision iteration patterns. |
@@ -371,7 +372,7 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
---
## CLI Modules (107 shipped)
## CLI Modules (109 shipped)
Full listing: `gsd-core/bin/lib/*.cjs`.
@@ -406,6 +407,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `decisions.cjs` | Parses CONTEXT.md `<decisions>` blocks; accepts numeric (D-42) and alphanumeric (D-INFRA-01) IDs; returns `{id, text, category, tags, trackable}` |
| `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection |
| `drift.cjs` | Post-execute codebase structural drift detector (#2003): classifies file changes into new-dir/barrel/migration/route categories and round-trips `last_mapped_commit` frontmatter |
| `edge-probe.cjs` | Spec-completeness edge probe (compiled from `src/edge-probe.cts`, gitignored) — the first adapter of the `probe-core` resolution model (ADR-550 Decision 7): shape classification, applicable-category relevance filter, edge proposal, and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`; exports `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `TAXONOMY` (#550) |
| `fallow-runner.cjs` | Fallow audit adapter for `/gsd-code-review`: binary resolution (`PATH` then `node_modules/.bin`), actionable missing-binary errors, and structural findings normalization |
| `federated-config.cjs` | Defensive merge of capability-declared config slices into the loadConfig return value — ADR-857 phase 3b; exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig })` → `{ values, validKeys, warnings }`; no-op until a key is atomically removed from the central config-schema (the cutover step) |
| `frontmatter.cjs` | YAML frontmatter CRUD operations |
@@ -446,6 +448,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `prompt-budget.cjs` | Pure token-budget accounting for review prompts — estimates tokens, applies deterministic trim priority (head-shrink PROJECT.md, proportional plan truncation, drop context/research/requirements, hard-fail guard), returns structured metadata for `review.max_prompt_tokens` (#3081) |
| `research-provider.cjs` | Research provider waterfall, confidence tiers, and planResearch (cache-hits + fetch plan) |
| `research-store.cjs` | Content-addressed research cache: sha256 keys, per-source TTL staleness, two-tier (user ~/.gsd / project .planning) store |
| `probe-core.cjs` | Generic spec-phase probe resolution model (compiled from `src/probe-core.cts`, gitignored; ADR-550 Decision 7) — the status×verification re-cut (`status: resolved/dismissed/unresolved` × per-probe `verification`), `validateResolution`/`validateRequirement`, `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject, the `byVerification` rollup, and the `runProbeCli` I/O scaffold; the shared seam consumed by `edge-probe` (and the prohibition probe #644); exports `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` (#550) |
| `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence |
| `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` |
| `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) |

View File

@@ -18,6 +18,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [Install on your runtime](how-to/install-on-your-runtime.md) — runtime-specific install steps for all 16 supported runtimes
- [Install a minimal GSD and add skills later](how-to/install-minimal-and-add-skills.md) — install only the core skills, then grow the surface with profiles and `/gsd:surface`
- [Discuss a phase](how-to/discuss-a-phase.md) — capture implementation decisions before planning begins
- [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality
- [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents
- [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR

View File

@@ -0,0 +1,77 @@
# ADR 550: spec-phase probe pattern and prohibition contract [Accepted]
- **Status:** Accepted
- **Date:** 2026-06-03
> **Provenance.** Drafted during triage of #644 (prohibition probe), hardened in a `/grill-with-docs` session, and extended with a `probe-core` seam designed in an `/improve-codebase-architecture` grill (Decision 7). Endorsed by the maintainer and by the #644 author (who also authored the edge-probe, #550 / PR #584). This ADR lands **on PR #584** alongside the `probe-core` extraction that implements Decision 7; #644 (the prohibition probe itself) remains blocked until #584 merges and then implements the prohibition adapter. "What exists today" is verified against `next` + the PR-#584 head as of 2026-06-03.
>
> **Lineage.** The judgment-tier contract (Decision 4) was reached independently from two directions: from the domain model (truths are unenforced positive-observables, so a values prohibition cannot be a silent pass) and from the author's "verifier-abstention" (N17) calibration experiment (on irreducible values rules the verifier is confidently wrong — ~0.93 confidence, ECE 0.81 — so confidence-gating cannot help; the only honest move is abstain-and-flag). That parked experiment folds into this ADR as the verify-time half of judgment-tier rather than needing a separate feature.
## Context
`spec-phase` produces `SPEC.md` from an interview. Two features add *probes* — soft-gate steps that surface spec gaps before code is written: **#550 / PR #584 (edge-probe)** for data-shape edges, and **#644 (prohibition probe)** for values/safety "must-NOT" constraints. Both use identical three-layer packaging and a recall→precision protocol, and both want confirmed findings to reach `SPEC.md` and the plan contract.
Grounding the design against the codebase surfaced the decisive facts:
- **`truths` are positive, observable assertions.** `gsd-planner` derives them as *"observable truths (3-7, user perspective)"*; `verify-phase` checks each by *"determine if the codebase enables it"* — a presence/wiring check. A prohibition is the inverse and cannot be confirmed that way.
- **`truths` are not mechanically verified.** `verify.cjs` iterates `artifacts` and `key_links` but has no `truths` handler. A prohibition parked in `truths` would inherit that non-enforcement — a negative wearing a positive's clothing.
- **`must_haves` has no schema/type** — a convention parsed by `parseMustHavesBlock` (blocks `truths`/`artifacts`/`key_links`), read in ≥6 places.
- **The SPEC→plan lift is LLM reasoning** in `gsd-planner`'s `derive_must_haves`, not mechanical extraction.
- **Prohibitions already exist in GSD only as narrative prose** (`<scope_reduction_prohibition>`, `"MUST NOT loop"`), never as structured, verifiable data.
- **The edge-probe's load-bearing logic is a deep module on the live path.** `src/edge-probe.cts` (~297 lines) is compiled to `bin/lib/edge-probe.cjs` and invoked by `spec-phase` Step 5.5 at runtime. ~120 of its lines are a *generic resolution model* (status lifecycle, validation, merge/rollup, CLI); only the `Shape`/`SHAPE_CUES`/`TAXONOMY`/`classifyShape`/`proposeEdges` cluster is edge-specific. The prohibition probe is the **second adapter** of that model — making the seam real (one adapter is hypothetical; two is real).
- **No prior ADR** governs probes, soft gates, or `must_haves` evolution.
The values/safety class #644 targets (*"must not become a guilt mechanic"*) is frequently **not reducible to a deterministic test** — which is why it needs an explicit verification contract rather than truth-style limbo.
## Decision
1. **Probe packaging (three layers).** Every spec-phase probe ships as: (a) a portable, dependency-free reference core under `references/<probe>.md` (taxonomy + protocol + output schema); (b) a `spec-phase.md` step run as a **soft gate** (write-anyway-with-flags), with explicit `--auto` and text-mode (non-Claude, no `AskUserQuestion`) handling, mirroring the ambiguity gate; (c) tests + docs.
2. **Two-stage protocol.** **Recall** (adversarial over-generation) → **precision** (classifier dropping routine engineering items), surfacing a short confirmable list. Dismissals require a non-empty reason.
3. **Prohibition home & representation.** Confirmed prohibitions live primarily as **`SPEC.md` acceptance criteria** (negative criteria). They MAY be projected into an optional **`must_haves.prohibitions:` sibling block** (alongside `truths`/`artifacts`/`key_links`) when they must survive into the plan contract. **`truths` is left untouched** — no `polarity` field is added. Each prohibition item carries `statement`, the orthogonal `status` + `verification` of Decision 7, and (when dismissed) a non-empty `reason`.
4. **Tiered verification.** Each prohibition declares `verification: test | judgment`.
- **`test`-tier** (reducible to a deterministic assertion) → a **negative test** following `regression-must-fail-first` + negative-proof discipline. **Hard gate in both interactive and autonomous modes.**
- **`judgment`-tier** (irreducible values/safety rule) → **mode-dependent soft-gate-with-flags**: *interactive* verify requires explicit human resolution of each item; *autonomous* verify records an LLM-judge verdict **marked non-authoritative** and emits a prominent *"unverified prohibition — human review recommended"* flag in SUMMARY/verdict. **Never a silent pass; never a hard halt of AFK runs.** Autonomous completion reads *"complete with N flagged prohibitions."*
- *Optional future input:* cross-tier (or cross-model) disagreement MAY feed the judgment-tier flag as a cheap blind-spot signal. This is a breadcrumb, **not a dependency** — the flag stands without it.
5. **The CI-testable surface is the contract, not the classifier.** The recall/precision/tier-assignment stages are LLM behavior and are **not** claimed as deterministic CI coverage (validated by offline batteries; in CI only by source-grep-under-`// allow-test-rule:`-exemption over reference/workflow prose, which *is* the product). CI deterministically tests the **contract**: (a) parse + validate an item (`statement` present; `status ∈ {resolved, dismissed, unresolved}`; `verification` in the probe's allowed set; dismissed items carry a non-empty `reason`); (b) round-trip / projection between `SPEC.md` and `must_haves.prohibitions:`; (c) a `DEFECT.GENERATIVE-FIX` **parity assertion** across template ↔ parser ↔ planner; (d) `test`-tier items are provably wired into verify-phase and cannot be silently skipped. A test that asserts the LLM's *judgment* is vacuous and is rejected per `RULESET.TESTS.delete-bad-tests`. *(Scope: on #584 only **(a)** ships — it is the edge adapter's parse+validate contract test (`tests/probe-core.test.cjs`). **(b)–(d)** require the `must_haves.prohibitions:` block, projection, and verify-phase tiering and are **#644 scope**, as the Consequences section records.)*
6. **Ownership seam with security tooling.** The probe owns **bespoke product/values** prohibitions only. When precision classifies an item as **canon security/compliance** (OWASP/GDPR/fairness/prototype-pollution/path-traversal), it does **not** mint a SPEC prohibition; it emits a one-line breadcrumb (*"possible canon-security concern X — owned by `/gsd:secure-phase` / eslint"*) and stops. Canon checks are **referred, not duplicated** — keeping the surfaced list short (#644's ~2–3-item goal).
7. **`probe-core` is the seam; probes are adapters.** The shared resolution model is extracted into `src/probe-core.cts`; `edge-probe.cts` refactors onto it as the first adapter and the prohibition probe is born on it as the second. Sub-decisions, each chosen against both probes:
- **7a — Orthogonal model.** The item carries two orthogonal dimensions: `status: resolved | dismissed | unresolved` (resolution lifecycle, shared) and `verification: <probe-defined> | null` (edge: `explicit | backstop`; prohibition: `test | judgment`). This replaces edge-probe's shipped `status: covered | dismissed | backstop | unresolved`, which smuggled a verification fact (`backstop`) into a lifecycle enum. Migration is mechanical (`covered → {resolved, explicit}`, `backstop → {resolved, backstop}`, others straight across); `coverage.resolved` is preserved **count-for-count** — it remains the *closed* set (`resolved` + `dismissed` status = `applicable − unresolved`), so every fixture's count is unchanged, while `coverage.byVerification` carries the per-tier `resolved`-status breakdown; the 6 edge fixtures are re-generated on #584. Done now (one adapter) rather than against a fixture-locked enum later.
- **7b — Core ingests already-proposed items.** The seam is `analyzeCoverage(items, resolutions?, validators)`, **not** `(requirements, proposeFn, …)`. The probes have different deterministic surfaces — edge = deterministic propose (`proposeEdges`) + LLM resolve; prohibition = **LLM propose** (adversarial recall) + deterministic validate/merge/rollup — so core MUST NOT assume propose is deterministic. `proposeEdges` stays in the edge adapter; candidate-parsing stays in the prohibition adapter.
- **7c — Hybrid typing.** Generic type params on the exported interfaces for adapter-authoring DX, but the load-bearing enforcement is **injected runtime validators** (`{ categories, verification, requiredFieldsByVerification }`), because the probe runs as a CLI over JSON where TS types are erased. The contract test (Decision 5) pins the validators, not the types.
- **7d — Rollup carries a verification tally.** `CoverageReport.coverage` gains `byVerification: { <tier>: count }`, computed generically in core over items that have a verification set. Verify-phase reads `byVerification.judgment` as Decision 4's denominator without re-scanning. (Unresolved items carry no tier and are counted by `status`.)
- **7e — One bin per probe.** Each probe ships its own compiled bin (`edge-probe.cjs`, `prohibition-probe.cjs`) calling a shared `runProbeCli(...)`. A single dispatcher CLI is **deferred** as an independent follow-on: unlike the enum, it is pure invocation plumbing with no migration debt, so it does not justify enlarging the #584 blast radius.
Indicative `probe-core` surface:
```ts
export type Status = 'resolved' | 'dismissed' | 'unresolved';
export interface Item<TCat extends string = string, TVer extends string = string> {
requirement_id: string; category: TCat; status: Status;
verification: TVer | null; resolution: string | null; reason: string | null;
}
export interface CoverageReport {
items: Item[];
coverage: { applicable: number; resolved: number; unresolved: number; byVerification: Record<string, number> };
}
export interface ProbeValidators {
categories: ReadonlySet<string>; verification: ReadonlySet<string>;
requiredFieldsByVerification?: Record<string, ReadonlyArray<keyof Item>>;
}
export function analyzeCoverage(items: Item[], resolutions: Resolution[] | undefined, v: ProbeValidators): CoverageReport;
export function validateResolution(r: Resolution, items: Item[], v: ProbeValidators): void;
export function validateItems(items: Item[], v: ProbeValidators): void;
export function runProbeCli(opts: { ingest: (argv: string[]) => { items: Item[]; resolutions?: Resolution[] }; validators: ProbeValidators }): never;
```
## Consequences
- **Positive:** `truths` keeps its positive-observable semantics; prohibitions get a first-class home with a real verification lifecycle; CI is honest (tests the contract, never fakes the model's judgment as green); no fake-green and no gutted autonomy; the secure-phase boundary is explicit and de-duplicated; the resolution model lives in one place, so the third probe is nearly free and `verification`/`status` cannot drift across the ≥6 `must_haves` read sites.
- **Costs / required work (on #584):** extract `probe-core.cts` and refactor `edge-probe.cts` onto it; re-cut the status enum into `status × verification` and re-generate the 6 edge fixtures. **On the #644 PR:** add the prohibition reference + the `must_haves.prohibitions:` block and its `parseMustHavesBlock` extension; the `DEFECT.GENERATIVE-FIX` parity assertion across template ↔ parser ↔ planner (**owned by the #644 author**); extend the SUMMARY/verdict schema with the `unverified-prohibition` flag; teach verify-phase the test-tier wiring and judgment-tier mode-dependent handling. A deterministic `projectProhibitions()` in `probe-core` is recommended so the parity assertion tests a function's round-trip rather than a prompt; every existing `must_haves` read site must tolerate the new block (absence = no prohibitions).
- **Docs debt:** `spec-phase` is undocumented in `FEATURES.md`/`COMMANDS.md` despite shipping in v1.38; the docs-required gate bites both probes. PR #584 establishes the spec-phase docs section.
- **Glossary:** `probe family`, `probe-core`, `verification tier`, and `bespoke vs canon prohibition` are added to `CONTEXT.md` when #584 merges (not while the contract is pre-merge); this ADR is their interim home.
- **Sequencing:** PR #584 lands ADR 550 + `probe-core` + the `edge-probe.cts` refactor + enum re-cut + fixture re-gen + the spec-phase docs section. #644 then adds the prohibition adapter (reference + `prohibitions:` block/parser/parity + tiered verification). #644 stays blocked until #584 merges.

View File

@@ -0,0 +1,107 @@
# How to resolve edge-coverage findings while writing a spec
**Goal:** Turn each domain-boundary edge the spec phase surfaces into an explicit, verifiable spec decision — so omitted boundaries (rounding ties, touching ranges, grapheme truncation) become checkable requirements *before* any code exists, instead of silent blind spots the verifier is confidently wrong about.
**Prerequisites:** A phase whose `/gsd-spec-phase` run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically and presents its findings — you do not invoke it separately.
For the category taxonomy and the reasoning behind front-of-pipeline edge analysis, see [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe). This guide covers only how to *act* on the findings.
---
## Read a finding
Each finding is one **applicable** edge for one requirement — a boundary the probe's relevance filter decided is in scope for that requirement's shape, and that you have not yet addressed. A finding is phrased as a probe question, for example:
> **R3 · precision** — Where can precision loss or overflow occur, and what is the contract?
You must resolve each finding into exactly one of four states. Claude presents them as a numbered choice (or an `AskUserQuestion` menu).
---
## Specify it — write an acceptance criterion
**Choose this when you can state the correct behaviour as a pass/fail check.** This is the strongest resolution: it produces a concrete assertion the planner and verifier can enforce.
Claude writes a new line into the spec's **Acceptance Criteria** and marks the edge `covered`. For the `precision` finding above:
> - [ ] Monetary amounts round half-to-even to 2 decimal places; `2.005 → 2.00`
Prefer this state whenever a defensible criterion can be written. A `covered` edge is lifted into the plan's `must_haves.truths`, so the verifier checks it.
---
## Dismiss it — record why it does not apply
**Choose this when the edge genuinely cannot occur** — and say why. A dismissal **requires a non-empty reason**; silence is rejected. The reason is the audit trail.
> ⛔ dismissed — input is a bounded enum; no boundary value exists
A wrong dismissal is the exact silent failure this probe exists to prevent, so dismiss only when the reason is solid.
---
## Backstop it — defer the contract to a held-out test
**Choose this when you know the edge matters but cannot fully articulate the correct behaviour in prose yet.** Claude marks the edge `backstop` and notes that a held-out / property-based test stands in for the missing assertion.
> 🧪 backstop — held-out property test: dedupe output order is stable under input permutation
A `backstop` edge is also carried into the plan's `must_haves.truths` as a non-inferable check — the planner records the intent, and the test body is authored later during execution.
---
## Defer it — leave it unresolved and flagged
**Choose this only when you are not ready to decide.** The edge stays `unresolved` and is flagged. Unlike a dismissal, deferring makes no claim that the edge is safe — it is an explicit, visible assumption the planner must surface, not silently drop.
---
## Clear the soft gate
After you have worked through the findings, the probe runs a **soft gate**:
- **All applicable edges resolved** → the spec proceeds to the next step.
- **One or more still `unresolved`** → Claude asks what to do:
- **Resolve now** — loop back and resolve the remaining edges.
- **Write the spec anyway** — the spec is written with those rows marked `⚠ Edge unresolved — planner must treat as assumption`. Use this deliberately; you are choosing to ship a known gap.
The gate is *soft*: it never blocks you, but every unresolved edge remains visible in the spec's `## Edge Coverage` section.
---
## Let Claude resolve them for you
**If the edges are low-stakes or already implied by earlier phases**, run the spec phase in auto mode:
```bash
/gsd-spec-phase 3 --auto
```
In `--auto`, Claude marks an edge `covered` where it can write a defensible acceptance criterion, and `backstop` otherwise. It **never auto-dismisses** — dismissing an edge requires a human reason, because a wrong auto-dismissal is precisely the silent failure being eliminated. Claude logs the tally, for example:
```
[auto] edge coverage: 4 covered, 2 backstop, 1 unresolved
```
Review the logged choices afterwards; auto mode is a fast first pass, not a substitute for judgement on edges that carry risk.
---
## What happens to resolved findings downstream
When you next run `/gsd-plan-phase`, the planner reads the spec's `## Edge Coverage` section and:
- lifts every `covered` edge's acceptance criterion into `must_haves.truths`,
- carries every `backstop` edge into `must_haves.truths` as a non-inferable check (needing a held-out/property test),
- surfaces every `unresolved` edge as an explicit assumption.
This is the payoff: a resolved edge becomes a unit the goal-backward verifier actually checks, extending its reach to boundaries the requirement prose never stated.
---
## Related
- [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe) — taxonomy, output schema, and the front-of-pipeline rationale
- [`/gsd-spec-phase`](../COMMANDS.md#gsd-spec-phase) — command reference and flags
- [Plan a phase](plan-a-phase.md) — where `covered`/`backstop` edges become `must_haves`
- [docs index](../README.md)

View File

@@ -36,6 +36,8 @@ export default tseslint.config(
// ADR-457: tsc-generated runtime artifact — lint the src/*.cts source, not the emitted .cjs.
'gsd-core/bin/lib/semver-compare.cjs',
'gsd-core/bin/lib/cli-exit.cjs',
'gsd-core/bin/lib/edge-probe.cjs',
'gsd-core/bin/lib/probe-core.cjs',
'gsd-core/bin/lib/code-review-flags.cjs',
'gsd-core/bin/lib/context-utilization.cjs',
'gsd-core/bin/lib/artifacts.cjs',

View File

@@ -0,0 +1,7 @@
{
"items": [
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Round a number to N decimal places, rounding half to nearest" }]

View File

@@ -0,0 +1,8 @@
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]

View File

@@ -0,0 +1,7 @@
{
"items": [
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Truncate a display string to its first N characters" }]

View File

@@ -0,0 +1,7 @@
{
"items": [
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Compute the total price as an amount rounded to two decimals" }]

View File

@@ -0,0 +1,8 @@
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Return all items in the list with duplicates removed" }]

View File

@@ -0,0 +1,8 @@
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } }
}

View File

@@ -0,0 +1 @@
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]

View File

@@ -0,0 +1,4 @@
[
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null },
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "resolution": null, "reason": "output is canonically sorted; no tie possible" }
]

View File

@@ -0,0 +1,261 @@
# Edge-Probe — Spec-Completeness Reference
Shared reference for the spec/requirements phase. Companion to
`@~/.claude/gsd-core/references/domain-probes.md`: `domain-probes` covers the
**technology axis** (auth, search, caching, deployment); this covers the
**data/behavior-shape axis** (boundaries, adjacency, encoding, ordering…). Walk each
requirement against the closed taxonomy below, propose a concrete candidate edge for each
applicable category, and resolve each to exactly one state. Adopt the established QA names
verbatim — this is decades-old black-box test technique, moved upstream to the spec layer.
This doc is written in generic `requirements → checks → verifier` terms with no
tool-specific vocabulary, so it is portable: copy it into any spec/requirements process.
A short mapping table at the end binds it to common host structures.
## Why front-of-pipeline
A goal-backward verifier only checks assertions that exist; an assertion only exists for a
requirement that was written down. A domain-boundary edge the author never surfaced is
invisible to the verifier — and worse, the verifier is *confidently wrong* about it
(measured: ~0.93 confidence while catching 0/12 omitted-edge defects, an expected
calibration error of 0.81 — worse than a coin flip — versus 100% / 0.03 on edges the spec
did state). The fix is not a better verifier or a confidence gate; confidence is
uninformative across these regimes. The fix is **spec completeness**: surface the omitted
edge into an explicit, checkable assertion *before* any code exists, after which the
verifier reliably catches it.
The underlying techniques are classic and should be named as such — **Boundary Value
Analysis**, **Equivalence Partitioning**, the **Category-Partition method**, and
**Metamorphic Relations**; property-based testing (PBT) is the academic name for the
held-out backstop. The literature applies these at the *test* layer (back of pipeline,
generating checks). The differentiated move here is **placement, not technique**: apply the
same edge taxonomy at the *spec* layer (front of pipeline, generating requirements), with
the explicit goal of extending a verifier's reach. Recent LLM spec work corroborates the
core finding that the spec layer is the measured weak point:
- SLD-Spec — program slicing + logical deletion (code→spec): arXiv 2509.09917
- Specine / AutoReSpec / SpecMind — spec alignment & postcondition inference
- CodeSpecBench (2604.12268), OSVBench (2504.20964), VERINA (2505.23135) — benchmarks
showing models solve tasks far better than they generate precise behavioral specs
- LLM property-based tests for edge cases: arXiv 2510.25297 (PBT+EBT ≈ 81% edge detection)
## Inputs
A list of requirements, each a `{ id, text, shapes? }` record where `text` is a testable
statement and `shapes` is an optional author-supplied override of the data/behavior shape.
The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. When
`shapes` is absent, a heuristic classifier proposes them from the requirement prose
(propose-then-confirm) — the author may correct the shape.
## Taxonomy (8 categories)
Closed and small by design: a fixed eight the author must explicitly clear beats thirty
nobody finishes. The failure mode being eliminated was never "too few categories" — it was
that no taxonomy was *systematically applied at all*. Categories 1, 2, and 4 alone cover
the three canonical corpus defects (banker's-rounding ties, touching intervals, grapheme
truncation). Growth happens via optional domain packs, not by bloating the core.
| id | name (QA term) | applies to shapes | probe question |
|----|----------------|-------------------|----------------|
| boundary | Boundary values | numeric-range | What happens exactly at each min/max/threshold — and one step either side? |
| adjacency | Adjacency / touching | collection | When two things are exactly equal or just touch, do they merge, collide, or separate? |
| empty | Empty / degenerate | collection, text | What is the result for empty, single-element, or null input? |
| encoding | Encoding / representation | text | Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form? |
| ordering | Ordering / stability | collection | When elements compare equal, is output order specified and stable? |
| precision | Precision / overflow | numeric-range | Where can precision loss or overflow occur, and what is the contract? |
| idempotency | Idempotency / repetition | stateful | What happens if this runs twice on the same input? |
| concurrency | Concurrency / effect ordering | stateful, io | If interrupted or run in parallel, what is guaranteed? |
## Relevance filter + resolution states
Two rules keep the probe honest and prevent an "everything is N/A" failure mode:
1. **Relevance filter first.** Classify each requirement's shape, then raise only the
categories whose `applies to shapes` intersect that requirement's shapes. A pure-text
requirement is never asked about overflow. This is what makes an unresolved edge
meaningful: it is an edge that *applies* and was not addressed.
2. **Dismissal requires a reason string.** "N/A — input is a bounded enum, no boundary
exists" is valid; silence is not. The reason string is the audit trail.
Each raised edge carries two orthogonal axes — a resolution **lifecycle** and, when
resolved, a **verification** tier (ADR-550 Decision 7, the shared probe-core model):
- **status** — `resolved | dismissed | unresolved`:
- **resolved** — the edge is addressed; *how* it is addressed is the verification tier.
- **dismissed** — not applicable, accompanied by a required, non-empty reason string.
- **unresolved** — carried forward and flagged; the author chose not to resolve it yet.
- **verification** (only when `status` is `resolved`; `null` otherwise) — `explicit | backstop`:
- **explicit** — a checkable assertion for the edge is written (a SPEC acceptance criterion).
- **backstop** — a held-out / property-based test stands in for an edge the author knows
but cannot fully articulate in prose (records intent; the test body is authored later).
Splitting these axes keeps the lifecycle enum free of a verification fact (the old single
enum smuggled `covered`/`backstop` — both *resolved* — into one flat list) and lets sibling
probes add their own verification tiers (e.g. `test | judgment`) without a parallel enum.
`coverage.resolved` is the count of **closed** edges — `resolved` + `dismissed` (the
pre-re-cut "covered + dismissed + backstop" set, count-preserved). An `unresolved`
*applicable* edge is the precise signal a soft completeness gate raises.
## Output schema
The probe emits, per edge, an item of the form:
```
{ requirement_id, category, status, verification, resolution, reason, probe }
```
plus a coverage summary:
```
coverage: { applicable, resolved, unresolved, byVerification: { explicit, backstop } }
```
`applicable` is the number of raised edges, `resolved` = closed (`resolved` + `dismissed`)
status edges, `unresolved` is the remainder, and `byVerification` breaks the
`resolved`-status edges down by tier (probe-agnostic in core; the edge adapter declares
`{ explicit, backstop }`). This JSON is the stable contract both the reference
implementation and any third-party port emit.
## Generic mapping (requirements → checks → verifier)
| Host structure | "requirement" | a `resolved`/`explicit` edge becomes | a `resolved`/`backstop` edge becomes |
|----------------|---------------|--------------------------|---------------------------|
| GSD SPEC | a SPEC Requirement | an Acceptance Criterion that `plan-phase` lifts into `must_haves.truths` | a non-inferable check in `must_haves.truths` (needs a held-out/PBT test) |
| Gherkin feature | a Scenario | an additional `Then` assertion / Scenario Outline row | a tagged scenario routed to a property test |
| OpenAPI operation | an operation | a response/constraint example + schema rule | a contract/property test on the operation |
| Docstring contract | a documented behavior | an assertion in the contract test | a property-based test for the function |
The portable invariant: a `resolved`/`explicit` edge produces **the unit your verifier
iterates over** (GSD: a `must_haves.truth`); a `resolved`/`backstop` edge produces a test
added to that same set as a non-inferable check. An `unresolved` edge is an explicit
assumption the downstream planner must surface, not silently drop.
## Worked example (merge-intervals)
Given a single requirement with no resolutions yet:
```
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]
```
the requirement classifies as a `collection`, which raises `adjacency`, `empty`, and
`ordering` (but not `boundary`/`precision`/`encoding`/`idempotency`/`concurrency`). With no
resolutions supplied, every applicable edge is `unresolved`:
```json edge-probe:02-merge-intervals/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
}
```
The `adjacency` row is the one that catches the canonical defect: `[[1,2],[2,3]]` intervals
that only *touch* — does the spec say they merge? Resolving it `resolved`/`explicit` writes
that assertion, and the verifier can then enforce it. This worked-example block is kept
byte-for-byte (parsed-JSON) identical to its fixture by `edge-probe-docs-fixtures.test.cjs`,
so the doc and the reference implementation cannot silently drift.
## Worked example (round-half-even)
Given a requirement to round floating-point values using the banker's-rounding (half-even)
rule, the requirement classifies as `numeric-range`, which raises `boundary` and `precision`
(but not `adjacency`/`empty`/`ordering`/`encoding`/`idempotency`/`concurrency`):
```json edge-probe:01-round-half-even/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}
```
The `precision` row surfaces the canonical defect for this requirement: IEEE 754 floating-point
arithmetic rounds 2.5 to 2 (not 3) under half-even — the probe asks whether the spec
states that contract explicitly, so the verifier can enforce it.
## Worked example (truncate-graphemes)
Given a requirement to truncate a string to N grapheme clusters (not bytes or code points),
the requirement classifies as `text`, which raises `empty` and `encoding`:
```json edge-probe:03-truncate-graphemes/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}
```
The `encoding` row is the load-bearing one: a requirement that says "truncate to 10
characters" is ambiguous — the spec must state whether "character" means bytes, UTF-16
code units, Unicode code points, or grapheme clusters (emoji sequences are 1 grapheme,
multiple code points).
## Worked example (money-rounding)
Given a requirement to round monetary amounts to two decimal places, the requirement
classifies as `numeric-range`, which raises `boundary` and `precision`:
```json edge-probe:04-money-rounding/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
],
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
}
```
The `boundary` row asks what the minimum and maximum representable values are (negative
amounts? fractional cents?). The `precision` row asks whether IEEE 754 binary rounding
can produce `$0.30000000000000004` — the spec must commit to a representation.
## Worked example (list-dedupe)
Given a requirement to deduplicate a list of items, the requirement classifies as
`collection`, which raises `adjacency`, `empty`, and `ordering`:
```json edge-probe:05-list-dedupe/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
}
```
The `ordering` row is the non-obvious one: dedupe removes duplicates, but which copy is
kept — first occurrence, last, or implementation-defined? The spec must commit.
## Worked example (resolved-mixed)
The same merge-intervals requirement after a resolution session where `adjacency` was
resolved with an explicit acceptance criterion, `ordering` was dismissed, and `empty` was
left unresolved:
```json edge-probe:06-resolved-mixed/expected-coverage.json
{
"items": [
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" }
],
"coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } }
}
```
`coverage.resolved` is 2 — the closed set (adjacency=`resolved`/`explicit` + ordering=`dismissed`);
`unresolved` is 1 (empty); `byVerification.explicit` is 1 (the single explicitly-verified
edge), `backstop` 0. The soft gate raises on this example because one applicable edge remains
unresolved — the author must either specify, dismiss, or backstop it before writing the SPEC.

View File

@@ -68,6 +68,18 @@ If none: "No additional constraints beyond standard project conventions."]
[Every acceptance criterion must be a checkbox that resolves to PASS or FAIL.
No "should feel good", "looks reasonable", or "generally works" — those are not checkboxes.]
## Edge Coverage
**Coverage:** [resolved]/[applicable] applicable edges resolved · [unresolved] unresolved
| Category | Requirement | Status | Resolution / Reason |
|----------|-------------|--------|---------------------|
| [category] | [Rn] | [✅ covered / ⛔ dismissed / 🧪 backstop / ⚠ UNRESOLVED] | [acceptance criterion ref, dismissal reason, or backstop test note] |
[Generated by the edge-completeness probe (Step 5.5). `covered` rows correspond to
Acceptance Criteria above; `backstop` rows must be carried into plan-phase `must_haves`.
`⚠ UNRESOLVED` rows are flagged: planner must treat as assumption.]
## Ambiguity Report
| Dimension | Score | Min | Status | Notes |

View File

@@ -810,6 +810,14 @@ PATTERNS_PATH=$(_gsd_field "$INIT" patterns_path)
# Detect spike/sketch findings skills (project-local)
SPIKE_FINDINGS_PATH=$(ls ./.claude/skills/spike-findings-*/SKILL.md 2>/dev/null | head -1 || true)
SKETCH_FINDINGS_PATH=$(ls ./.claude/skills/sketch-findings-*/SKILL.md 2>/dev/null | head -1 || true)
# Resolve the phase SPEC (carries the ## Edge Coverage section the planner lifts covered/
# backstop edges from). UNCONDITIONAL — must NOT live in §4.5 Check AI-SPEC, which is skipped
# on non-AI phases; gating it there silently starves the planner of the SPEC (#550 review).
# Glob the plain phase SPEC, excluding the -AI-SPEC.md / -UI-SPEC.md variants.
PHASE_DIR_FOR_SPEC=$(_gsd_field "$INIT" phase_dir)
SPEC_FILE=$(ls "${PHASE_DIR_FOR_SPEC}"/*-SPEC.md 2>/dev/null | grep -Ev -- '-(AI|UI)-SPEC\.md$' | head -1)
SPEC_PATH="${SPEC_FILE}"
```
## 7.5. Verify Nyquist Artifacts
@@ -937,6 +945,7 @@ Planner prompt:
- {uat_path} (UAT Gaps - if --gaps)
- {reviews_path} (Cross-AI Review Feedback - if --reviews; actionable findings must be incorporated or explicitly deferred/rejected in PLAN.md)
- {UI_SPEC_PATH} (UI Design Contract — visual/interaction specs, if exists)
- {SPEC_PATH} (Phase SPEC — carries the ## Edge Coverage section to lift covered/backstop edges from, if exists)
- {SPIKE_FINDINGS_PATH} (Spike Findings — validated patterns, constraints, landmines from experiments, if exists)
- {SKETCH_FINDINGS_PATH} (Sketch Findings — validated design decisions, CSS patterns, visual direction, if exists)
- {API_SURFACE_PATH} (API Surface — HINT ONLY, if intel.enabled; see <intel_surface_hint> below)
@@ -999,6 +1008,7 @@ Output consumed by /gsd:execute-phase. Plans need:
- Tasks in XML format with read_first and acceptance_criteria fields (MANDATORY on every task)
- Verification criteria
- must_haves for goal-backward verification
- If the SPEC has an `## Edge Coverage` section, lift every `covered` edge's acceptance criterion into `must_haves.truths`, and every `backstop` edge into `must_haves.truths` as a non-inferable check (note it needs a held-out/property-based test). `unresolved` edges are explicit assumptions — surface them in the plan, do not silently drop them.
- **"Artifacts this phase produces" section (MANDATORY)** — list every symbol this phase creates: decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths. The plan-review-convergence source-grounding pass reads this section to exclude newly-created symbols from drift verification; omitting it causes new symbols to be flagged for acknowledgement.
</downstream_consumer>
@@ -1044,6 +1054,7 @@ Every task MUST include these fields — they are NOT optional:
- [ ] Waves assigned for parallel execution
- [ ] must_haves derived from phase goal
- [ ] Every PLAN.md includes an "Artifacts this phase produces" section listing symbols created by this phase (decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths)
- [ ] Every SPEC ## Edge Coverage covered/backstop edge is represented in a plan's must_haves (no silent drops)
</quality_gate>
```

View File

@@ -187,10 +187,137 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met):
## Step 5: (covered inline — ambiguity scoring is per-round)
## Step 5.5: Edge-Completeness Probe
Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague
ones). Reference: @~/.claude/gsd-core/references/edge-probe.md.
**Runtime coverage compute — resolve and invoke edge-probe.cjs:**
```bash
# Resolve the compiled edge-probe.cjs against the GSD install dir via RUNTIME_DIR (#448)
# — NOT the consuming project's git root — falling back to git toplevel / $HOME/.claude.
# Mirrors the ui-safety-gate.cjs resolution idiom at autonomous.md:290 / plan-phase.md:631.
_GSD_RT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"
EDGE_PROBE_JS=$(for _c in \
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
"$_GSD_RT/bin/lib/edge-probe.cjs" \
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
[ -f "$_c" ] && { echo "$_c"; break; }
done)
# Graceful degradation — never silent skip (RR-04). Build ONLY when $_GSD_RT is a verified
# GSD source checkout (has tsconfig.build.json + src/edge-probe.cts), and pin npm to it with
# --prefix so we never trigger the CONSUMING project's own build:lib (its cwd package scripts:
# codegen/migrations/writes) during a spec workflow. Real installs ship the compiled .cjs via
# prepublishOnly, so this build path only matters in a GSD dev checkout (review High).
if [ -z "$EDGE_PROBE_JS" ]; then
if [ -f "$_GSD_RT/tsconfig.build.json" ] && [ -f "$_GSD_RT/src/edge-probe.cts" ]; then
npm --prefix "$_GSD_RT" run build:lib 2>/dev/null || true
EDGE_PROBE_JS=$(for _c in \
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
"$_GSD_RT/bin/lib/edge-probe.cjs" \
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
[ -f "$_c" ] && { echo "$_c"; break; }
done)
fi
if [ -z "$EDGE_PROBE_JS" ]; then
echo "ERROR: edge-probe.cjs not found — reinstall GSD or run \`npm run build:lib\` in your GSD checkout." >&2
exit 1
fi
fi
# Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the
# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object
# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes
# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04).
REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX.json")
cat > "$REQS_JSON" <<'JSON'
[
{ "id": "R1", "text": "<replace: requirement text from the SPEC>" }
]
JSON
# Guard — never invoke on an empty/invalid array, OR one still holding the heredoc
# `<replace: …>` placeholder (a forgotten substitution would otherwise yield a
# meaningful-looking but bogus coverage report). Fail loud, not silent no-op.
if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("<replace:")))process.exit(1)' "$REQS_JSON" 2>/dev/null; then
echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the <replace: …> placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2
exit 1
fi
# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per
# requirement. The covered/backstop/dismissed/unresolved rows in $COVERAGE drive the
# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose).
# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture
# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE
# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary
# the engine validation exists to protect. Make the run fatal, then validate the captured
# report is well-formed JSON before the resolution loop consumes it.
if ! COVERAGE=$(node "$EDGE_PROBE_JS" "$REQS_JSON"); then
rm -f "$REQS_JSON"
echo "ERROR: edge-probe engine failed (invalid shapes or bad input) — fix the requirement(s) and re-run; never proceed with empty coverage." >&2
exit 1
fi
rm -f "$REQS_JSON"
# Exit-0-but-garbage guard: the report must parse as JSON with the expected { items[], coverage{} } shape.
if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let r;try{r=JSON.parse(s)}catch{process.exit(1)}if(!r||!Array.isArray(r.items)||typeof r.coverage!=="object"||r.coverage===null)process.exit(1)})'; then
echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2
exit 1
fi
# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY
# requirement is far more likely a shape-classification miss (or malformed requirements) than
# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding
# applicable:0. Surface it loudly; the author must explicitly confirm "no applicable edges"
# below rather than silently emitting a green empty ## Edge Coverage section.
APPLICABLE=$(printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let n=0;try{n=JSON.parse(s).coverage.applicable}catch{n=0}process.stdout.write(String(n))})')
if [ "$APPLICABLE" = "0" ]; then
echo "WARNING: edge-probe proposed ZERO applicable edges across all requirements — likely a classification miss or malformed requirements, not a genuinely edge-free spec. Do NOT silently write an empty Edge Coverage section." >&2
fi
```
If `$APPLICABLE` is `0`, do NOT proceed silently: ask the author to confirm via AskUserQuestion
("The edge probe found no applicable edges for any requirement — is this genuinely an
edge-free spec, or should we revisit the requirement wording / authored shapes?"). Only write
an empty `## Edge Coverage` section after explicit confirmation.
For each Requirement gathered so far:
1. Classify its shape and raise only applicable edge categories (relevance filter — see
the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already
surfaced as pre-`covered`.
2. For each raised category, propose a CONCRETE candidate edge (not "consider
boundaries" — e.g. "R2 merges intervals; what about `[[1,2],[2,3]]` that only touch?").
3. Resolve each with the user (AskUserQuestion; text mode → numbered list):
- **Specify it** → write a new pass/fail line into Acceptance Criteria AND mark the
edge `covered`.
- **Dismiss (reason)** → mark `dismissed` with a required non-empty reason.
- **Backstop with a test** → mark `backstop`; note "held-out edge test" for plan-phase.
- **Defer** → leave `unresolved`.
**Soft gate (after resolving):**
- All applicable edges resolved → proceed to Step 6.
- Any `unresolved` → AskUserQuestion:
- header: "Edge Coverage"
- question: "[N] edge(s) are unresolved: [list]. What do you want to do?"
- options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" /
"Keep probing"
- On "anyway": write SPEC.md with those rows marked `⚠ Edge unresolved — planner must
treat as assumption`.
**`--auto` mode:** auto-`covered` where a defensible acceptance criterion can be written;
otherwise auto-`backstop` (never auto-dismiss — a wrong dismissal is the exact silent
failure being eliminated). Log: `[auto] edge coverage: C covered, B backstop, U unresolved`.
Populate the `## Edge Coverage` section of SPEC.md from the resolved edges.
## Step 6: Generate SPEC.md
Use the SPEC.md template from @~/.claude/gsd-core/templates/spec.md.
- Populate the **Edge Coverage** section from Step 5.5 (covered/dismissed/backstop/unresolved rows).
**Requirements for every requirement entry:**
- One specific, testable statement
- Current state (what exists now)
@@ -249,6 +376,7 @@ Next: /gsd:discuss-phase {X}
- Do NOT ask about HOW to implement — that is discuss-phase territory
- Scout the codebase BEFORE the first question — grounded questions only
- Max 2–3 questions per round — do not frontload all questions at once
- Step 5.5 edge probe runs after the ambiguity gate; dismissals require a reason; --auto never auto-dismisses
</critical_rules>
<success_criteria>
@@ -260,4 +388,5 @@ Next: /gsd:discuss-phase {X}
- Acceptance criteria are pass/fail checkboxes
- SPEC.md committed atomically (when commit_docs is true)
- User directed to /gsd:discuss-phase as next step
- Edge-completeness probe run; Edge Coverage section populated; unresolved edges flagged as assumptions
</success_criteria>

View File

@@ -151,6 +151,15 @@
"validate-context.test.cjs"
],
"issue": "892"
},
"edge-probe": {
"files": [
"edge-probe-docs-fixtures.test.cjs",
"edge-probe-planner-contract.test.cjs",
"edge-probe-spec-phase-contract.test.cjs",
"edge-probe.test.cjs"
],
"issue": "550"
}
}
}

217
src/edge-probe.cts Normal file
View File

@@ -0,0 +1,217 @@
/**
* Spec-completeness edge-probe — the FIRST adapter of the probe-core resolution model
* (ADR-457 build model; ADR-550 Decision 7 seam).
*
* The generic resolution lifecycle, the status×verification re-cut, `validateResolution`,
* `validateRequirement`, the `analyzeCoverage` merge/rollup/orphan-reject engine, and the
* `runProbeCli` scaffold all live in `src/probe-core.cts`. This module keeps ONLY the
* edge-specific cluster: the five data/behavior shapes, the closed 8-category edge taxonomy,
* shape classification, edge proposal, and the `{ explicit, backstop }` verification validators.
*
* Authored as strict TypeScript (`src/edge-probe.cts`) and compiled by
* `tsc -p tsconfig.build.json` to the gitignored runtime artifact
* `gsd-core/bin/lib/edge-probe.cjs`. Do NOT hand-write the `.cjs`; it is emitted. Tests
* `require()` the built artifact; `pretest` runs `build:lib` first.
*
* Pure and dependency-free: it classifies each requirement's data/behavior shape, filters
* the closed 8-category edge taxonomy to applicable categories, proposes concrete candidate
* edges, and (via probe-core) merges author resolutions into a coverage report.
*/
import {
type Item,
type Resolution,
type CoverageReport,
type Validators,
validateRequirement as coreValidateRequirement,
validateResolution as coreValidateResolution,
analyzeCoverage as coreAnalyzeCoverage,
runProbeCli,
} from './probe-core.cjs';
/** The five data/behavior shapes a requirement can exhibit. */
export type Shape = 'numeric-range' | 'collection' | 'text' | 'stateful' | 'io';
/** The edge probe's verification tiers (the `verification` axis values for a resolved edge). */
export type EdgeVerification = 'explicit' | 'backstop';
/** A single edge taxonomy category. */
export interface TaxonomyEntry {
id: string;
name: string;
shapes: Shape[];
probe: string;
}
/** A SPEC requirement; `shapes` is an optional authored override of classification. */
export interface Requirement {
id: string;
text: string;
shapes?: Shape[];
}
/** An edge item — a probe-core `Item` specialized to the edge verification vocabulary. */
export type Edge = Item<EdgeVerification>;
/**
* Word-boundary cues mapping requirement prose -> data/behavior shape.
* Heuristic and intentionally lossy; an authored `shapes` array overrides it.
*/
export const SHAPE_CUES: Record<Shape, RegExp> = {
'numeric-range': /\b(round(ing|ed)?|threshold|max(imum)?|min(imum)?|limit|bound(ary)?|between|cap|percent|amount|price|count|number|score|rate|decimal)\b/i,
'collection': /\b(lists?|arrays?|sets?|items?|collections?|each|every|all|sort(ed|ing)?|merge|dedupe|group|ranges?|intervals?|overlap(ping)?)\b/i,
'text': /\b(string|text|names?|labels?|truncate|substring|char(acter)?s?|length|slug|message|unicode)\b/i,
'stateful': /\b(save|persist|store|update|toggle|create|delete|remove|submit|retry|apply|register|insert)\b/i,
'io': /\b(files?|requests?|fetch|upload|download|network|api|endpoints?|connections?|sockets?)\b/i,
};
/** The locked shape vocabulary — exactly the keys of SHAPE_CUES (single source of truth). */
export const VALID_SHAPES: ReadonlySet<string> = new Set(Object.keys(SHAPE_CUES));
/** Detect which shapes a requirement's prose matches (heuristic). */
export function classifyShape(text: string): Shape[] {
const shapes: Shape[] = [];
const subject = String(text == null ? '' : text);
for (const shape of Object.keys(SHAPE_CUES) as Shape[]) {
if (SHAPE_CUES[shape].test(subject)) shapes.push(shape);
}
return shapes;
}
/**
* Closed taxonomy of 8 domain-boundary edge categories (established QA names).
* `shapes` lists which requirement shapes make the category relevant.
*/
export const TAXONOMY: TaxonomyEntry[] = [
{ id: 'boundary', name: 'Boundary values', shapes: ['numeric-range'], probe: 'What happens exactly at each min/max/threshold — and one step either side?' },
{ id: 'adjacency', name: 'Adjacency / touching', shapes: ['collection'], probe: 'When two things are exactly equal or just touch, do they merge, collide, or separate?' },
{ id: 'empty', name: 'Empty / degenerate', shapes: ['collection', 'text'], probe: 'What is the result for empty, single-element, or null input?' },
{ id: 'encoding', name: 'Encoding / representation', shapes: ['text'], probe: 'Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?' },
{ id: 'ordering', name: 'Ordering / stability', shapes: ['collection'], probe: 'When elements compare equal, is output order specified and stable?' },
{ id: 'precision', name: 'Precision / overflow', shapes: ['numeric-range'], probe: 'Where can precision loss or overflow occur, and what is the contract?' },
{ id: 'idempotency', name: 'Idempotency / repetition', shapes: ['stateful'], probe: 'What happens if this runs twice on the same input?' },
{ id: 'concurrency', name: 'Concurrency / effect ordering', shapes: ['stateful', 'io'], probe: 'If interrupted or run in parallel, what is guaranteed?' },
];
/** Return taxonomy category ids whose applicable shapes intersect the input set. */
export function applicableCategories(shapes: Shape[]): string[] {
const set = new Set<Shape>(shapes);
return TAXONOMY.filter((c) => c.shapes.some((s) => set.has(s))).map((c) => c.id);
}
/**
* The edge adapter's injected runtime validators (ADR-550 #5). `categories` is the closed
* taxonomy; both verification tiers require a non-empty `resolution` (an explicit AC's text
* or a backstop note) so plan-phase has a criterion to lift.
*/
export const EDGE_VALIDATORS: Validators = {
categories: TAXONOMY.map((c) => c.id),
verification: ['explicit', 'backstop'],
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
};
/**
* Validate a single requirement — the generic id/text checks (probe-core) plus the edge's
* `shapes`-must-be-an-array check. A bare string like `shapes:"numeric-range"` would otherwise
* fall through to prose classification, silently ignoring the authored override.
*
* The edge adapter's `text` is REQUIRED (the prose is the classification signal), so reject a
* missing/empty `text` when no authored `shapes` override is present. Without this, a `{ id }`
* requirement classifies to zero shapes → zero edges → it is silently DROPPED from coverage
* with no signal — the exact fail-open this feature exists to eliminate. An explicit `shapes`
* array (including `[]` for "no applicable categories") is the legitimate way to opt out of
* prose classification, so `text` is only required when `shapes` is absent.
*/
export function validateRequirement(requirement: Requirement): void {
coreValidateRequirement(requirement);
const r = requirement as unknown as { shapes?: unknown; text?: unknown };
if (r.shapes != null && !Array.isArray(r.shapes)) {
throw new Error(`requirement ${requirement.id} shapes must be an array when present`);
}
if (r.shapes == null && !(typeof r.text === 'string' && r.text.trim())) {
throw new Error(
`requirement ${requirement.id} text must be a non-empty string when no shapes override is provided`,
);
}
}
/** Validate an edge resolution against the edge verification vocabulary. */
export function validateResolution(resolution: Resolution<EdgeVerification>): true {
return coreValidateResolution(resolution, EDGE_VALIDATORS);
}
/**
* Propose candidate edges for a requirement. Uses authored `shapes` when present, else
* classifies from prose. Every proposed edge starts unresolved (verification null).
*/
export function proposeEdges(requirement: Requirement): Edge[] {
validateRequirement(requirement);
let shapes: Shape[];
if (Array.isArray(requirement.shapes)) {
// Fail closed: an authored array must contain only locked shape values. A non-empty
// but invalid array (e.g. ['numeric'], a typo for 'numeric-range') would otherwise
// intersect no category and silently suppress every probe — the gate reads green while
// nothing was checked. An empty array stays a valid "no applicable categories" override.
for (const s of requirement.shapes) {
if (typeof s !== 'string' || !VALID_SHAPES.has(s)) {
throw new Error(
`invalid shape ${JSON.stringify(s)} for requirement ${requirement.id} — must be one of: ${[...VALID_SHAPES].join(', ')}`,
);
}
}
shapes = requirement.shapes;
} else {
shapes = classifyShape(requirement.text);
}
return applicableCategories(shapes).map((catId): Edge => {
const cat = TAXONOMY.find((c) => c.id === catId);
return {
requirement_id: requirement.id,
category: catId,
status: 'unresolved',
verification: null,
resolution: null,
reason: null,
probe: cat ? cat.probe : '',
};
});
}
/**
* Propose edges for every requirement (deterministic propose), then delegate the
* merge/rollup/orphan-reject to probe-core. Edge-specific pre-checks: requirements must be an
* array, requirement ids must be unique. Throws on any invalid resolution.
*/
export function analyzeCoverage(
requirements: Requirement[],
resolutions: Resolution<EdgeVerification>[] = [],
): CoverageReport<EdgeVerification> {
if (!Array.isArray(requirements)) {
throw new Error('requirements must be an array');
}
const items: Edge[] = [];
const seenReqIds = new Set<string>();
for (const req of requirements) {
validateRequirement(req);
if (seenReqIds.has(req.id)) {
throw new Error(`duplicate requirement id ${JSON.stringify(req.id)}`);
}
seenReqIds.add(req.id);
for (const edge of proposeEdges(req)) items.push(edge);
}
return coreAnalyzeCoverage(items, resolutions, EDGE_VALIDATORS);
}
/*
* CLI entry (EP-06 invokable surface): `edge-probe.cjs <requirements.json> [resolutions.json]`.
* The generic I/O plumbing (parse, fail-closed exit 2, pretty-JSON out) lives in probe-core's
* `runProbeCli`; the edge adapter supplies its `analyzeCoverage`. Guarded by
* `require.main === module` so it runs only when the compiled `.cjs` is executed directly.
*/
if (require.main === module) {
runProbeCli(
(requirements, resolutions) =>
analyzeCoverage(requirements as Requirement[], resolutions as Resolution<EdgeVerification>[]),
{ usage: 'edge-probe.cjs <requirements.json> [resolutions.json]' },
);
}

336
src/probe-core.cts Normal file
View File

@@ -0,0 +1,336 @@
/**
* probe-core — generic spec-phase probe resolution model (ADR-550 Decision 7).
*
* Extracted from the edge-probe (the first adapter) once the prohibition probe (#644)
* proved it the *second* adapter of the same model: one adapter is a hypothetical seam,
* two is a real one. This module owns everything generic — the resolution lifecycle,
* the status×verification re-cut, `validateResolution`/`validateRequirement`, the
* `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject engine,
* the `byVerification` rollup, and the `runProbeCli` I/O scaffold. Each probe is a thin
* adapter: it supplies the proposal logic (deterministic for edge, LLM-recall for
* prohibition) and its closed vocabularies via injected validators.
*
* Authored as strict TypeScript (`src/probe-core.cts`) and compiled by
* `tsc -p tsconfig.build.json` to the gitignored runtime artifact
* `gsd-core/bin/lib/probe-core.cjs`. Do NOT hand-write the `.cjs`; it is emitted.
*
* Two orthogonal axes (the re-cut):
* - status: resolved | dismissed | unresolved — the resolution LIFECYCLE (shared)
* - verification: <probe-defined> | null — HOW a resolved item is verified
* The edge adapter declares `verification: explicit | backstop`; the prohibition adapter
* (#644) will declare `test | judgment`. Splitting the axes keeps the lifecycle enum free
* of a verification fact and lets a sibling probe add its own tiers without a parallel enum.
*
* Typing is hybrid (ADR-550 #5): generic type params for adapter DX, but enforcement runs
* through injected runtime validators, because the CLI executes over JSON where TS types are
* erased. The contract test pins the validators, not the types.
*/
import fs from 'node:fs';
/** Resolution lifecycle — shared across every probe adapter. */
export type Status = 'resolved' | 'dismissed' | 'unresolved';
/** The LOCKED set of valid lifecycle statuses (the re-cut: no covered/backstop). */
export const VALID_STATUS: Status[] = ['resolved', 'dismissed', 'unresolved'];
/**
* A proposed or resolved item for a requirement/category pair. Generic over the probe's
* verification-tier vocabulary `V` (edge: `'explicit' | 'backstop'`). A freshly proposed
* item is `{ status: 'unresolved', verification: null, resolution: null, reason: null }`.
*/
export interface Item<V extends string = string> {
requirement_id: string;
category: string;
status: Status;
verification: V | null;
resolution: string | null;
reason: string | null;
probe: string;
}
/** An author resolution merged onto a proposed item. */
export interface Resolution<V extends string = string> {
requirement_id: string;
category: string;
status: Status;
verification?: V | null;
resolution?: string | null;
reason?: string | null;
}
/** A coverage report: the merged items plus rollup counts (incl. the per-tier breakdown). */
export interface CoverageReport<V extends string = string> {
items: Item<V>[];
coverage: {
applicable: number;
resolved: number;
unresolved: number;
byVerification: Record<string, number>;
};
}
/**
* The injected runtime enforcement contract (ADR-550 #5). The probe declares its closed
* vocabularies so `analyzeCoverage`/`validateResolution` can enforce them at the JSON
* boundary where TS types no longer exist.
* - `categories` — the probe's valid category ids (a proposed item outside this set is an
* adapter bug, caught rather than silently rolled up).
* - `verification` — the valid verification tiers for a `resolved` item.
* - `requiredFieldsByVerification` — for each tier, which resolution fields MUST be a
* non-empty string (edge: every tier needs `resolution` text for plan-phase to lift).
*/
export interface Validators {
categories: string[];
verification: string[];
requiredFieldsByVerification: Record<string, Array<'resolution' | 'reason'>>;
}
/** A generic requirement — every probe ingests at least `{ id, text }`. */
export interface Requirement {
id: string;
text?: string;
}
function errMessage(e: unknown): string {
return e instanceof Error ? e.message : String(e);
}
/**
* Structural guard for the report an adapter's `analyze` returns. The scaffold types `analyze`
* loosely (it runs over JSON-parsed input the adapter `as`-casts), so a future adapter (#644)
* that forgets to validate inside its closure could hand back a malformed object. Rather than
* stringify garbage as green output, `runProbeCli` checks the report shape and fails closed.
*/
function isValidReport(report: unknown): report is CoverageReport {
if (report == null || typeof report !== 'object') return false;
const r = report as { items?: unknown; coverage?: unknown };
if (!Array.isArray(r.items)) return false;
const c = r.coverage as
| { applicable?: unknown; resolved?: unknown; unresolved?: unknown; byVerification?: unknown }
| undefined;
if (c == null || typeof c !== 'object') return false;
if (typeof c.applicable !== 'number' || typeof c.resolved !== 'number' || typeof c.unresolved !== 'number') {
return false;
}
if (c.byVerification == null || typeof c.byVerification !== 'object') return false;
return true;
}
/**
* Validate a requirement's generic structural fields — fail closed on malformed input rather
* than coercing it. Probe-specific fields (e.g. the edge adapter's `shapes`) are validated by
* the adapter. Typed loosely because the CLI casts arbitrary parsed JSON to `Requirement`.
*/
export function validateRequirement(requirement: Requirement): void {
const r = requirement as unknown as { id?: unknown; text?: unknown };
if (typeof r.id !== 'string' || !r.id.trim()) {
throw new Error(`requirement id must be a non-empty string (got ${JSON.stringify(r.id)})`);
}
if (r.text != null && typeof r.text !== 'string') {
throw new Error(`requirement ${r.id} text must be a string when present`);
}
}
/**
* Validate a resolution against the probe's injected validators. Rejects an unknown status,
* a dismissal without a non-empty reason, a `resolved` item with a missing/unknown
* verification tier, and a `resolved` item missing any field its tier requires (per
* `requiredFieldsByVerification`). Returns true on success.
*/
export function validateResolution<V extends string>(r: Resolution<V>, validators: Validators): true {
const key = `${r.requirement_id}::${r.category}`;
if (!VALID_STATUS.includes(r.status)) {
throw new Error(`invalid status "${r.status}" for ${key}`);
}
// Invariant (this module's header): `verification` is null unless `status` is `resolved`.
// Enforce it for EVERY status — a dismissed/unresolved resolution carrying a verification
// tier would otherwise merge verbatim (`analyzeCoverage` below) and silently break the
// model for the second adapter (#644) that inherits this seam. Fail closed across the full
// status×verification space, not just `resolved`.
if (r.status !== 'resolved' && r.verification != null) {
throw new Error(`verification must be null unless status is "resolved" (got "${r.verification}") for ${key}`);
}
// An `unresolved` resolution is an UNACTED item: it must carry no resolution/reason payload.
// A populated payload is an authoring mistake (the author meant resolved/dismissed) that
// would otherwise be silently dropped into the unresolved count with no error pointing at
// it. Reject it so the mistake surfaces.
if (r.status === 'unresolved') {
if (r.resolution != null && String(r.resolution).trim()) {
throw new Error(`unresolved must not carry a resolution (${key})`);
}
if (r.reason != null && String(r.reason).trim()) {
throw new Error(`unresolved must not carry a reason (${key})`);
}
}
if (r.status === 'dismissed' && !(r.reason && String(r.reason).trim())) {
throw new Error(`dismissed requires a reason (${key})`);
}
if (r.status === 'resolved') {
const tier = r.verification;
if (tier == null) {
throw new Error(`resolved requires a verification tier (one of: ${validators.verification.join(', ')}) for ${key}`);
}
if (!validators.verification.includes(tier)) {
throw new Error(`invalid verification "${tier}" for ${key} — must be one of: ${validators.verification.join(', ')}`);
}
const required = validators.requiredFieldsByVerification[tier] ?? [];
for (const field of required) {
// field is 'resolution' | 'reason'; both are `string | null | undefined` on Resolution,
// so the indexed access is string-typed (no unknown-to-string coercion).
const value = r[field];
if (!(value != null && String(value).trim())) {
throw new Error(`${tier} requires a ${field} (${key})`);
}
}
}
return true;
}
/**
* Merge author resolutions onto ALREADY-PROPOSED items and roll up coverage counts.
*
* Core operates on `items[]`, never a `proposeFn`: probes have different deterministic
* surfaces (edge = deterministic propose + LLM resolve; prohibition = LLM propose + deterministic
* validate/merge), so proposal stays in each adapter and core must not assume it is deterministic.
*
* `coverage.resolved` is the COUNT of CLOSED items (`resolved` + `dismissed` status) =
* `applicable - unresolved` — the pre-re-cut "covered + dismissed + backstop" set,
* count-preserved. `byVerification` breaks the `resolved`-status items down by tier (each tier
* initialized to 0). Throws on any invalid resolution, a duplicate, an orphan (a resolution
* matching no proposed item), or a proposed item whose category is outside `validators.categories`.
*/
export function analyzeCoverage<V extends string>(
items: Item<V>[],
resolutions: Resolution<V>[] = [],
validators: Validators,
): CoverageReport<V> {
if (!Array.isArray(items)) {
throw new Error('items must be an array');
}
const key = (r: { requirement_id: string; category: string }): string => `${r.requirement_id}::${r.category}`;
const resMap = new Map<string, Resolution<V>>();
for (const r of resolutions) {
validateResolution(r, validators);
if (resMap.has(key(r))) {
throw new Error(`duplicate resolution for ${key(r)}`);
}
resMap.set(key(r), r);
}
const validCategories = new Set(validators.categories);
const merged: Item<V>[] = [];
const itemKeys = new Set<string>();
for (const item of items) {
if (!validCategories.has(item.category)) {
throw new Error(`item ${key(item)} has unknown category "${item.category}" — not one of: ${validators.categories.join(', ')}`);
}
itemKeys.add(key(item));
const o = resMap.get(key(item));
if (o) {
merged.push({ ...item, status: o.status, verification: o.verification ?? null, resolution: o.resolution ?? null, reason: o.reason ?? null });
} else {
// No author resolution: the item is rolled up VERBATIM, so its own status/fields must be
// valid too. The edge adapter only proposes `unresolved` items, but the prohibition adapter
// (#644) proposes LLM-generated items that arrive already populated — one carrying an
// out-of-enum status (e.g. the dropped "covered") or `dismissed` with no reason would
// otherwise be counted closed without validation. An Item is structurally a superset of a
// Resolution, so the same fail-closed check guards both. (ADR-550 Decision 5 hardens this
// shared seam for the second adapter; m1.)
validateResolution(item as unknown as Resolution<V>, validators);
merged.push(item);
}
}
// Reject orphan resolutions — a resolution whose (requirement_id, category) matches no
// proposed item (typo'd category or a non-applicable one) would otherwise be silently
// dropped, leaving the author believing an item is resolved while the report shows it
// unresolved (adversarial-review HIGH; preserved from the edge-probe's original engine).
for (const k of resMap.keys()) {
if (!itemKeys.has(k)) {
throw new Error(`unknown resolution for ${k} — no matching proposed item (typo'd category or non-applicable shape?)`);
}
}
const unresolved = merged.filter((i) => i.status === 'unresolved').length;
const applicable = merged.length;
const resolved = applicable - unresolved; // closed set: resolved-status + dismissed
const byVerification: Record<string, number> = {};
for (const tier of validators.verification) byVerification[tier] = 0;
for (const i of merged) {
if (i.status === 'resolved' && i.verification != null) {
byVerification[i.verification] = (byVerification[i.verification] ?? 0) + 1;
}
}
return { items: merged, coverage: { applicable, resolved, unresolved, byVerification } };
}
/*
* CLI scaffold (the EP-06 invokable surface, generalized). Each probe ships one bin that
* calls `runProbeCli` with its own `analyze` (closing over the adapter's propose + validators)
* and usage string; a single dispatcher CLI is a deferred follow-on. The I/O dependencies are
* injectable so the generic plumbing is unit-testable without spawning a process.
*
* `tsconfig.build.json` sets `"types": ["node"]`, so `process` and `node:fs` are typed.
*/
/** Injectable I/O for `runProbeCli` (defaults wire to the real process). */
export interface ProbeCliOptions {
usage: string;
argv?: string[];
readFile?: (path: string) => string;
write?: (s: string) => void;
writeErr?: (s: string) => void;
exit?: (code: number) => void;
}
/**
* Read the requirements file (and optional resolutions file), run the adapter's `analyze`,
* and write the report as pretty JSON + newline. With no requirements path, writes the usage
* line to stderr and exits 2. A JSON-parse failure or any `analyze` throw is a handled error:
* stderr + exit 2, never an uncaught stack trace — so the engine's fail-closed validation
* surfaces at the workflow boundary rather than failing open.
*/
export function runProbeCli(
analyze: (requirements: unknown, resolutions: unknown) => CoverageReport,
options: ProbeCliOptions,
): void {
const argv = options.argv ?? process.argv;
const readFile = options.readFile ?? ((p: string) => fs.readFileSync(p, 'utf8'));
const write = options.write ?? ((s: string) => { process.stdout.write(s); });
const writeErr = options.writeErr ?? ((s: string) => { process.stderr.write(s); });
const exit = options.exit ?? ((code: number) => { process.exit(code); });
const reqPath: string | undefined = argv[2];
const resPath: string | undefined = argv[3];
if (!reqPath) {
writeErr(`usage: ${options.usage}\n`);
exit(2);
return;
}
let requirements: unknown;
try {
requirements = JSON.parse(readFile(reqPath));
} catch (e: unknown) {
writeErr(`error: cannot parse JSON from ${reqPath}: ${errMessage(e)}\n`);
exit(2);
return;
}
let resolutions: unknown = [];
if (resPath) {
try {
resolutions = JSON.parse(readFile(resPath));
} catch (e: unknown) {
writeErr(`error: cannot parse JSON from ${resPath}: ${errMessage(e)}\n`);
exit(2);
return;
}
}
try {
const report = analyze(requirements, resolutions);
if (!isValidReport(report)) {
throw new Error('adapter returned a structurally-invalid coverage report (expected { items[], coverage{ applicable, resolved, unresolved, byVerification } })');
}
write(`${JSON.stringify(report, null, 2)}\n`);
} catch (e: unknown) {
writeErr(`error: ${errMessage(e)}\n`);
exit(2);
}
}

View File

@@ -0,0 +1,115 @@
// allow-test-rule: runtime-contract-is-the-product — the rendered reference/SPEC/ADR vocab surfaces are the runtime contract; this pins their bijection to the code (docs-parity)
// Asserts the portable reference doc (gsd-core/references/edge-probe.md) keeps its
// worked-example JSON blocks in sync with the source-of-truth fixture files under
// gsd-core/references/edge-probe-fixtures/. The fixtures are the canonical data; the
// doc embeds copies. Per the CONTRIBUTING.md exception matrix this is `docs-parity`: a
// reference doc must mirror source-defined data and there is no runtime enumeration API.
// The comparison is PARSED JSON (deepEqual of JSON.parse on both sides), never a raw-text
// substring match — so a reformat that preserves the data does not fail, and any semantic
// drift between doc and fixture does.
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const docPath = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe.md');
const fixturesRoot = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures');
const adrPath = path.join(__dirname, '..', 'docs', 'adr', '550-spec-phase-probe-contract.md');
const specTemplatePath = path.join(__dirname, '..', 'gsd-core', 'templates', 'spec.md');
// Extract fenced blocks tagged ```json edge-probe:<dir>/<file> from the doc, keyed by ref.
// The \n? before the closing fence allows blocks whose closing fence has no preceding newline
// (fixes the silent-skip bug where a trailing-fence-with-no-newline was not matched).
function taggedJsonBlocks(md) {
const re = /```json edge-probe:([^\n]+)\n([\s\S]*?)\n?```/g;
const out = {};
let m;
while ((m = re.exec(md))) out[m[1].trim()] = m[2];
return out;
}
describe('edge-probe doc/fixture sync', () => {
test('reference doc exists', () => {
assert.ok(fs.existsSync(docPath), `${docPath} must exist`);
});
test('doc embeds tagged fixture blocks for every expected-coverage.json fixture (count-equality)', () => {
const md = fs.readFileSync(docPath, 'utf8');
const blocks = taggedJsonBlocks(md);
// Count the expected-coverage.json files under the fixtures root (one per fixture dir).
const expectedCount = fs.readdirSync(fixturesRoot, { withFileTypes: true })
.filter(e => e.isDirectory())
.filter(dir => fs.existsSync(path.join(fixturesRoot, dir.name, 'expected-coverage.json')))
.length;
assert.strictEqual(
Object.keys(blocks).length,
expectedCount,
`edge-probe.md must embed exactly ${expectedCount} tagged blocks (one per fixture expected-coverage.json)`
);
});
test('every tagged doc block parses and deepEquals its fixture file', () => {
const md = fs.readFileSync(docPath, 'utf8');
const blocks = taggedJsonBlocks(md);
for (const [ref, body] of Object.entries(blocks)) {
const fixtureFile = path.join(fixturesRoot, ref);
const onDisk = fs.readFileSync(fixtureFile, 'utf8');
assert.deepEqual(JSON.parse(body), JSON.parse(onDisk),
`doc block edge-probe:${ref} must deepEqual ${fixtureFile}`);
}
});
});
// m2: lock the machine↔SPEC vocabulary mapping so the two layers cannot silently drift. The
// machine contract uses orthogonal `status` × `verification`; the SPEC table and planner prose
// render a flat `covered/dismissed/backstop/unresolved`. The migration is documented in ADR-550
// Decision 7a and the reference's "Generic mapping" table, but the SPEC table is rendered by the
// LLM workflow (no JS renderer to round-trip against). So this pins the canonical map as code AND
// grounds it in every doc surface that renders the vocabulary — a renderer/parser drift fails here.
describe('edge-probe machine↔SPEC vocabulary mapping (m2 drift lock)', () => {
// The canonical migration map (ADR-550 Decision 7a): machine state → SPEC display label.
// `verification` is null for the lifecycle-only states (dismissed/unresolved).
const MACHINE_TO_DISPLAY = [
{ status: 'resolved', verification: 'explicit', display: 'covered' },
{ status: 'resolved', verification: 'backstop', display: 'backstop' },
{ status: 'dismissed', verification: null, display: 'dismissed' },
{ status: 'unresolved', verification: null, display: 'unresolved' },
];
const machineKey = (s, v) => `${s}|${v ?? '∅'}`;
test('the mapping is a bijection — machine→display→machine is identity, no shared labels', () => {
const toDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [machineKey(m.status, m.verification), m.display]));
const fromDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [m.display, machineKey(m.status, m.verification)]));
// No two machine states collapse onto the same display label (the silent-drift failure mode).
assert.equal(toDisplay.size, MACHINE_TO_DISPLAY.length, 'each machine state must have a distinct key');
assert.equal(fromDisplay.size, MACHINE_TO_DISPLAY.length, 'each display label must map back to exactly one machine state');
// Round-trip identity.
for (const m of MACHINE_TO_DISPLAY) {
const key = machineKey(m.status, m.verification);
assert.equal(fromDisplay.get(toDisplay.get(key)), key, `${key} must round-trip through its display label`);
}
});
test('ADR-550 Decision 7a documents the resolved/explicit↔covered and resolved/backstop↔backstop migration', () => {
const adr = fs.readFileSync(adrPath, 'utf8');
assert.match(adr, /covered\b[^.]*resolved[^.]*explicit/i, 'ADR must document covered → {resolved, explicit}');
assert.match(adr, /backstop\b[^.]*resolved[^.]*backstop/i, 'ADR must document backstop → {resolved, backstop}');
assert.match(adr, /count-for-count|count-preserved/i, 'ADR must state coverage.resolved is count-preserved across the migration');
});
test('the SPEC template legend renders exactly the four canonical display labels', () => {
const spec = fs.readFileSync(specTemplatePath, 'utf8');
for (const { display } of MACHINE_TO_DISPLAY) {
assert.match(spec, new RegExp(`\\b${display}\\b`, 'i'), `spec.md Edge Coverage legend must render the "${display}" label`);
}
});
test('the reference Generic-mapping table distinguishes the resolved/explicit and resolved/backstop tiers', () => {
const md = fs.readFileSync(docPath, 'utf8');
assert.match(md, /resolved`?\/`?explicit/i, 'reference must name the resolved/explicit tier in the mapping table');
assert.match(md, /resolved`?\/`?backstop/i, 'reference must name the resolved/backstop tier in the mapping table');
});
});

View File

@@ -0,0 +1,167 @@
// allow-test-rule: runtime-contract-is-the-product — plan-phase.md's planner prompt is the deployed runtime contract under assertion
// plan-phase.md is the deployed planning workflow contract; these checks lock
// the SPEC path wiring and quality-gate that the edge-probe review (RR-01/02/03)
// requires — assertions scope to extracted sub-blocks to avoid false positives.
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const PLAN_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-phase.md');
const PROMPT_PATH = path.join(__dirname, '..', 'gsd-core', 'templates', 'planner-subagent-prompt.md');
function readPlanPhase() {
return fs.readFileSync(PLAN_PHASE_PATH, 'utf8');
}
// Extract the planner <downstream_consumer> block from plan-phase.md. The runtime planner is
// spawned from plan-phase.md's own inline <planning_context>/<downstream_consumer>, so the
// load-bearing lift instruction must live here — NOT in templates/planner-subagent-prompt.md,
// which nothing loads at runtime (no @-import in agents/gsd-planner.md, no read in plan-phase.md).
function extractDownstreamConsumerBlock(content) {
const start = content.indexOf('<downstream_consumer>');
if (start === -1) return '';
const end = content.indexOf('</downstream_consumer>', start);
if (end === -1) return '';
return content.slice(start, end + '</downstream_consumer>'.length);
}
// Extract the planner <files_to_read> block that contains {UI_SPEC_PATH}
// There are multiple <files_to_read> blocks in plan-phase.md; we need the one
// at ~line 890-912 inside the planning_context markdown block.
function extractPlannerFilesBlock(content) {
let pos = 0;
while (true) {
const start = content.indexOf('<files_to_read>', pos);
if (start === -1) return '';
const end = content.indexOf('</files_to_read>', start);
if (end === -1) return '';
const block = content.slice(start, end + '</files_to_read>'.length);
if (block.includes('{UI_SPEC_PATH}')) {
return block;
}
pos = end + 1;
}
}
// Extract the planner <quality_gate> block (the last one in plan-phase.md,
// inside the planner prompt template)
function extractQualityGateBlock(content) {
const start = content.lastIndexOf('<quality_gate>');
if (start === -1) return '';
const end = content.indexOf('</quality_gate>', start);
if (end === -1) return '';
return content.slice(start, end + '</quality_gate>'.length);
}
// Test A (RR-01): plan-phase.md resolves a phase *-SPEC.md into SPEC_FILE/SPEC_PATH
// Uses new RegExp to correctly match literal $ and ( characters in bash snippets.
// This MUST FAIL before the RR-01 fix (no SPEC_PATH resolution exists today)
test('RR-01: plan-phase.md resolves phase *-SPEC.md (excluding AI/UI variants) into SPEC_FILE/SPEC_PATH', () => {
const content = readPlanPhase();
// Assert the canonical SPEC_FILE resolution form is present.
// new RegExp used so that \$ and \( are treated as literal dollar-sign and open-paren
// (JS regex literals interpret \$ as end-anchor and \( as group open).
assert.match(
content,
new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md'),
'plan-phase.md must resolve SPEC_FILE using ls "${...PHASE_DIR...}"/*-SPEC.md pattern'
);
// Assert {SPEC_PATH} token appears in the planner files_to_read block
const filesBlock = extractPlannerFilesBlock(content);
assert.match(
filesBlock,
/[{]SPEC_PATH[}]/,
'The planner <files_to_read> block (containing {UI_SPEC_PATH}) must also contain {SPEC_PATH}'
);
});
// Test B (RR-01): The {SPEC_PATH} entry is labelled as carrying the ## Edge Coverage section
// This MUST FAIL before the RR-01 fix (no {SPEC_PATH} entry exists today)
test('RR-01: {SPEC_PATH} entry in files_to_read is labelled with Edge Coverage', () => {
const content = readPlanPhase();
const filesBlock = extractPlannerFilesBlock(content);
assert.match(
filesBlock,
/[{]SPEC_PATH[}][^\n]*Edge Coverage/,
'{SPEC_PATH} entry in planner files_to_read must be labelled as carrying the ## Edge Coverage section'
);
});
// Extract a "## " section from its heading until the next "## " heading (or EOF).
function extractSection(content, heading) {
const start = content.indexOf(heading);
if (start === -1) return '';
const next = content.indexOf('\n## ', start + heading.length);
return content.slice(start, next === -1 ? content.length : next);
}
// Test E (RR-01 REACHABILITY — the assertion that catches the original no-op):
// Token presence is not enough. The SPEC resolution must live on an UN-GATED path. §4.5
// "Check AI-SPEC" is skipped on every non-AI phase (ai_integration_phase_enabled false /
// --skip-ai-spec), so a resolution placed there leaves SPEC_PATH unbound and the planner
// never receives the SPEC — exactly the #550 silent no-op. Assert it is NOT in §4.5.
test('RR-01 reachability: SPEC_FILE resolution is NOT gated inside the skippable "## 4.5 Check AI-SPEC" section', () => {
const content = readPlanPhase();
const aiSpecSection = extractSection(content, '## 4.5. Check AI-SPEC');
const specResolution = new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md');
assert.ok(aiSpecSection.length > 0, 'sanity: the §4.5 Check AI-SPEC section must exist to scope this test');
assert.doesNotMatch(
aiSpecSection,
specResolution,
'SPEC_FILE resolution must NOT live inside the skippable §4.5 "Check AI-SPEC" block — gating it there silently starves the planner of the SPEC on non-AI phases (the original #550 no-op)'
);
assert.match(content, specResolution, 'SPEC_FILE resolution must still exist on an un-gated path elsewhere in plan-phase.md');
});
// Test C (RR-02 consumer end): the lift instruction lives in the RUNTIME planner surface.
// plan-phase.md spawns the planner from its own inline <planning_context>; the
// templates/planner-subagent-prompt.md file is orphaned (loaded by nothing), so asserting the
// contract there is false assurance — the test stays green even if the runtime never consumes it.
// Pin the contract to the <downstream_consumer> block plan-phase.md actually sends the planner.
test('RR-02 consumer: plan-phase.md downstream_consumer instructs lifting covered/backstop edges into must_haves.truths', () => {
const block = extractDownstreamConsumerBlock(readPlanPhase());
assert.ok(block.length > 0, 'sanity: plan-phase.md must contain a <downstream_consumer> block to scope this test');
assert.match(
block,
/##\s*Edge Coverage/,
'plan-phase.md <downstream_consumer> must reference ## Edge Coverage'
);
assert.match(
block,
/must_haves\.truths/,
'plan-phase.md <downstream_consumer> must reference must_haves.truths as the lift target'
);
// Regression guard against re-orphaning: the lift instruction must NOT be relocated back into
// templates/planner-subagent-prompt.md (which nothing loads) and presented as the consumer
// contract — that is exactly the false-assurance this retarget fixes.
const prompt = fs.readFileSync(PROMPT_PATH, 'utf8');
assert.doesNotMatch(
prompt,
/Edge Coverage/,
'planner-subagent-prompt.md is not loaded at runtime — the Edge Coverage lift instruction must not live there'
);
});
// Test D (RR-03): plan-phase.md <quality_gate> contains a covered/backstop ↔ must_haves item
// This MUST FAIL before the RR-03 fix (no such quality_gate item exists today)
test('RR-03: planner quality_gate requires covered/backstop edges represented in must_haves', () => {
const content = readPlanPhase();
const qgBlock = extractQualityGateBlock(content);
assert.match(
qgBlock,
/covered.*backstop.*must_haves|backstop.*covered.*must_haves/i,
'planner quality_gate must contain a checklist item tying covered/backstop edges to must_haves'
);
});

View File

@@ -0,0 +1,193 @@
// allow-test-rule: runtime-contract-is-the-product — spec-phase.md Step 5.5 is the deployed workflow runtime contract under assertion
// spec-phase.md is the deployed spec workflow contract; these checks lock
// the Step 5.5 wiring so the edge-probe.cjs runtime invocation cannot
// silently rot the way the original plan-phase no-op did (reviewer finding RR-11).
// Assertions scope to the extracted Step 5.5 block to avoid false positives
// from incidental mentions elsewhere in the file.
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const SPEC_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'spec-phase.md');
function readSpecPhase() {
return fs.readFileSync(SPEC_PHASE_PATH, 'utf8');
}
// Slice the Step 5.5 block: from the "Step 5.5" heading to the next "## " or "Step " heading.
// This scopes assertions to Step 5.5 only, preventing false positives from mentions elsewhere.
function extractStep55Block(content) {
const startIdx = content.indexOf('## Step 5.5');
if (startIdx === -1) {
// Also try without the ## prefix
const altIdx = content.indexOf('Step 5.5');
if (altIdx === -1) return '';
// Find end: next heading starting with ## or Step N (not Step 5.5)
const rest = content.slice(altIdx + 'Step 5.5'.length);
const nextHeading = rest.search(/\n## |\nStep \d/);
if (nextHeading === -1) return content.slice(altIdx);
return content.slice(altIdx, altIdx + 'Step 5.5'.length + nextHeading);
}
const rest = content.slice(startIdx + '## Step 5.5'.length);
const nextHeading = rest.search(/\n## /);
if (nextHeading === -1) return content.slice(startIdx);
return content.slice(startIdx, startIdx + '## Step 5.5'.length + nextHeading);
}
// Test A (RR-11): Step 5.5 resolves and invokes edge-probe.cjs via node.
// MUST FAIL before the RR-04 wire (Step 5.5 is prose-only today — no CLI invocation).
test('RR-11: spec-phase Step 5.5 resolves edge-probe.cjs via path-fallback loop', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// Assert the path-fallback resolution loop for edge-probe.cjs is present in Step 5.5.
// The token "edge-probe.cjs" must appear inside the block (the artifact being resolved).
assert.match(
block,
/edge-probe\.cjs/,
'Step 5.5 must reference edge-probe.cjs as the artifact being resolved'
);
// Assert node invocation of edge-probe.cjs in Step 5.5.
// Matches: node "$EDGE_PROBE_JS" or node ... edge-probe.cjs
assert.match(
block,
/node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/,
'Step 5.5 must invoke edge-probe.cjs via node (e.g. node "$EDGE_PROBE_JS" ...)'
);
});
// Test C (RR-11 FUNCTION — the assertion that catches "decorative bash"):
// Token presence is not enough. The invocation is a no-op unless $REQS_JSON is actually
// POPULATED before the engine runs (the original block only mktemp'd it + left a comment,
// so the CLI parsed an empty file). Assert the block (a) writes $REQS_JSON via a redirect,
// (b) does so BEFORE the node invocation, and (c) guards against an empty/invalid file.
test('RR-11 function: Step 5.5 writes $REQS_JSON before invoking, and guards against empty input', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// (a) A redirect that writes the requirements into $REQS_JSON (e.g. `cat > "$REQS_JSON"`).
const writeIdx = block.search(/>\s*"\$REQS_JSON"/);
assert.ok(
writeIdx !== -1,
'Step 5.5 must WRITE requirements into $REQS_JSON (a redirect like `cat > "$REQS_JSON"`), not just mktemp it — an empty file makes the probe a silent no-op'
);
// (b) The write must precede the node invocation of the engine.
const invokeIdx = block.search(/node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/);
assert.ok(invokeIdx !== -1, 'Step 5.5 must invoke the engine via node');
assert.ok(
writeIdx < invokeIdx,
'Step 5.5 must populate $REQS_JSON BEFORE invoking edge-probe.cjs (write precedes the node call)'
);
// (c) A guard that refuses to run on an empty/invalid requirements array.
assert.match(
block,
/Array\.isArray|empty\/invalid|empty or invalid|REQS_JSON[^\n]*empty/i,
'Step 5.5 must guard against an empty/invalid $REQS_JSON before invoking (fail loud, not silent no-op)'
);
});
// Test B (RR-11): Step 5.5 has an explicit not-found branch — build:lib or error token.
// MUST FAIL before the RR-04 wire (no not-found handling today).
test('RR-11: spec-phase Step 5.5 has an explicit not-found branch (build:lib or blocking error)', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// Assert either a build:lib invocation or an explicit "not found" / error message exists.
// This prevents the wire from being added as a silent-skip with no fallback.
assert.match(
block,
/build:lib|not found|ERROR.*edge-probe|edge-probe.*not found/i,
'Step 5.5 must have an explicit not-found branch (build:lib attempt or clear blocking error)'
);
});
// Test D (review High): the build fallback must NEVER run the CONSUMING project's package
// scripts. Every executable `build:lib` invocation must be pinned to the GSD dir with
// `npm --prefix`, and the build must be gated behind a verified GSD source checkout. A bare
// `npm run build:lib` (no --prefix) uses cwd — which, under the git-toplevel fallback, is the
// consumer repo — and would execute its codegen/migrations during a spec workflow.
test('review High: Step 5.5 build:lib is --prefix-pinned to the GSD dir and gated on a source checkout', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// Collect lines that actually INVOKE npm (trimmed start === "npm"), excluding echo/comment
// mentions (e.g. the error message that quotes `npm run build:lib` for the user).
const buildInvocations = block
.split('\n')
.filter((l) => l.trim().startsWith('npm') && l.includes('build:lib'));
assert.ok(
buildInvocations.length > 0,
'Step 5.5 must contain at least one npm build:lib invocation (the dev-checkout fallback)'
);
for (const line of buildInvocations) {
assert.match(
line,
/npm\s+--prefix\s+"?\$?\{?_?GSD_RT/,
`build:lib must be pinned with \`npm --prefix "$_GSD_RT"\` so it never runs the consuming project's scripts — offending line: ${line.trim()}`
);
}
// The build must be gated behind a verified GSD source checkout (tsconfig.build.json present),
// so it cannot fire inside a plain consumer repo where the artifact merely happens to be absent.
assert.match(
block,
/tsconfig\.build\.json/,
'Step 5.5 must gate the build behind a GSD source checkout (e.g. test -f "$_GSD_RT/tsconfig.build.json")'
);
});
// Test E (review #4 High): the engine's fail-closed exit(2) must NOT be swallowed by command
// substitution. A bare `COVERAGE=$(node "$EDGE_PROBE_JS" ...)` discards the exit status, leaves
// $COVERAGE empty on an invalid-shapes failure, and lets the workflow proceed into prose
// re-derivation — fail-OPEN at the very boundary the engine validation exists to protect.
test('review #4 High: Step 5.5 exit-checks the engine capture and validates the report (no fail-open)', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// The engine capture must be FATAL — guarded by `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)`
// (or an explicit exit-status check) that exits non-zero on failure.
assert.match(
block,
/if\s+!\s+COVERAGE=\$\(node\s+"\$EDGE_PROBE_JS"/,
'Step 5.5 must exit-check the engine invocation (e.g. `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)`) — a bare command substitution swallows the engine exit code and fails open'
);
// And the captured report must be validated as JSON before the resolution loop consumes it
// (guards against an exit-0-but-garbage capture).
assert.match(
block,
/(COVERAGE[\s\S]{0,500}JSON\.parse)|(JSON\.parse[\s\S]{0,500}\$COVERAGE)/,
'Step 5.5 must validate $COVERAGE parses as JSON before use (guard against exit-0-but-malformed output)'
);
});
// Test F (adversarial review): a report with ZERO applicable edges across all requirements is
// the likely-classification-miss fail-open (same shape as an invalid shape yielding applicable:0).
// Step 5.5 must surface it, not silently emit a green empty ## Edge Coverage section.
test('adversarial review: Step 5.5 guards a zero-applicable coverage report', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
assert.match(
block,
/coverage\.applicable/,
'Step 5.5 must read coverage.applicable and guard the zero-applicable case (warn/confirm, not silently proceed)'
);
});

441
tests/edge-probe.test.cjs Normal file
View File

@@ -0,0 +1,441 @@
/**
* Edge-probe reference core unit tests.
*
* Asserts the LOCKED export surface of the spec-completeness edge-probe against
* the BUILT artifact (`gsd-core/bin/lib/edge-probe.cjs`), which
* `npm run build:lib` (run by pretest) emits from `src/edge-probe.cts`.
*
* Post ADR-550 Decision 7: the generic resolution model lives in `probe-core`;
* edge-probe is its first adapter (shapes/TAXONOMY/proposeEdges + the
* `{explicit, backstop}` verification validators). The resolution model is the
* status×verification re-cut: `status: resolved | dismissed | unresolved` ×
* `verification: explicit | backstop`. `covered`/`backstop` are no longer status
* values — `covered → {resolved, explicit}`, `backstop → {resolved, backstop}`.
*/
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const path = require('node:path');
const fs = require('node:fs');
const os = require('node:os');
const { execFileSync, spawnSync } = require('node:child_process');
const { cleanup } = require('./helpers.cjs');
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs');
const ep = require(BUILT_SCRIPT);
describe('edge-probe: classifyShape', () => {
test('detects numeric-range from rounding/threshold cues', () => {
assert.deepEqual(ep.classifyShape('Round a number to N decimal places').sort(),
['numeric-range']);
});
test('detects collection from interval/merge cues', () => {
const shapes = ep.classifyShape('Merge a list of overlapping intervals');
assert.ok(shapes.includes('collection'));
});
test('detects text from truncate/string cues', () => {
const shapes = ep.classifyShape('Truncate a string to a maximum length');
assert.ok(shapes.includes('text'));
});
test('returns [] when no cue matches', () => {
assert.deepEqual(ep.classifyShape('Display the company logo'), []);
});
});
describe('edge-probe: TAXONOMY + applicableCategories', () => {
test('TAXONOMY has the 8 documented categories in order', () => {
assert.deepEqual(ep.TAXONOMY.map((c) => c.id),
['boundary', 'adjacency', 'empty', 'encoding', 'ordering', 'precision', 'idempotency', 'concurrency']);
});
test('every category has name, shapes[], probe', () => {
for (const c of ep.TAXONOMY) {
assert.equal(typeof c.name, 'string');
assert.ok(Array.isArray(c.shapes) && c.shapes.length >= 1);
assert.equal(typeof c.probe, 'string');
}
});
test('numeric-range raises boundary + precision only', () => {
assert.deepEqual(ep.applicableCategories(['numeric-range']).sort(),
['boundary', 'precision']);
});
test('collection raises adjacency, empty, ordering', () => {
assert.deepEqual(ep.applicableCategories(['collection']).sort(),
['adjacency', 'empty', 'ordering']);
});
test('text raises empty + encoding', () => {
assert.deepEqual(ep.applicableCategories(['text']).sort(),
['empty', 'encoding']);
});
test('no shapes raises nothing', () => {
assert.deepEqual(ep.applicableCategories([]), []);
});
});
describe('edge-probe: proposeEdges', () => {
test('rounding requirement proposes boundary + precision, all unresolved (verification null)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places' });
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
for (const e of edges) {
assert.equal(e.requirement_id, 'R1');
assert.equal(e.status, 'unresolved');
assert.equal(e.verification, null);
assert.equal(e.resolution, null);
assert.equal(e.reason, null);
assert.equal(typeof e.probe, 'string');
}
});
test('authored shapes override prose classification', () => {
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque label', shapes: ['collection'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
});
});
describe('edge-probe: validateResolution', () => {
test('rejects an unknown status', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'maybe' }),
/invalid status/i);
});
test('rejects a former covered status (re-cut: covered is no longer a status)', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'covered', resolution: 'x' }),
/invalid status/i);
});
test('rejects dismissed without a reason', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: '' }),
/dismissed requires a reason/i);
});
test('accepts dismissed with a reason', () => {
assert.equal(ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: 'bounded enum' }), true);
});
test('rejects resolved with a missing verification tier', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', resolution: 'AC' }),
/verification/i);
});
test('rejects resolved with a verification tier outside {explicit, backstop}', () => {
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'judgment', resolution: 'AC' }),
/invalid verification/i);
});
});
describe('edge-probe: analyzeCoverage', () => {
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
test('with no resolutions, every applicable edge is unresolved (byVerification zeroed)', () => {
const rep = ep.analyzeCoverage(reqs, []);
assert.deepEqual(rep.coverage, { applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 } });
});
test('merges a resolved/explicit resolution and counts it resolved', () => {
const rep = ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
]);
const adj = rep.items.find((i) => i.category === 'adjacency');
assert.equal(adj.status, 'resolved');
assert.equal(adj.verification, 'explicit');
assert.equal(adj.resolution, 'AC#6: touching intervals merge');
assert.equal(rep.coverage.resolved, 1);
assert.equal(rep.coverage.unresolved, 2);
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 });
});
test('throws if a resolution is invalid (dismissed w/o reason)', () => {
assert.throws(() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'empty', status: 'dismissed' },
]), /dismissed requires a reason/i);
});
});
describe('edge-probe: CLI (built artifact)', () => {
test('reads a requirements file and prints a coverage report as JSON', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-'));
const reqPath = path.join(dir, 'requirements.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
const out = execFileSync('node', [BUILT_SCRIPT, reqPath], { encoding: 'utf8' });
const rep = JSON.parse(out);
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
});
test('with no args exits with status 2 (assert on exit code, not stderr prose)', () => {
let status;
try {
execFileSync('node', [BUILT_SCRIPT], { stdio: 'pipe' });
status = 0;
} catch (error) {
status = error.status;
}
assert.equal(status, 2);
});
});
describe('edge-probe: CLI JSON.parse error handling (RR-10)', () => {
test('invalid requirements JSON exits with status 2 (handled error, not uncaught throw)', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const badJson = path.join(dir, 'bad-req.json');
fs.writeFileSync(badJson, 'not valid json {{{');
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, badJson], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('invalid resolutions JSON exits with status 2 (handled error, not uncaught throw)', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const goodReq = path.join(dir, 'req.json');
const badRes = path.join(dir, 'bad-res.json');
fs.writeFileSync(goodReq, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
fs.writeFileSync(badRes, 'not valid json {{{');
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, goodReq, badRes], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('valid requirements file exits 0 and stdout is parseable JSON', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
const reqPath = path.join(dir, 'req.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe', encoding: 'utf8' });
assert.equal(r.status, 0);
const rep = JSON.parse(r.stdout);
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
} finally {
cleanup(dir);
}
});
});
describe('edge-probe: proposeEdges — empty-shapes override (RR-06)', () => {
test('shapes: [] returns zero edges (explicit empty-shapes override)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] });
assert.deepEqual(edges, []);
});
test('absent shapes key classifies from prose (no override)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals' });
assert.ok(edges.length > 0, 'should classify collection edges from prose');
});
test('shapes: [collection] overrides prose and proposes collection categories', () => {
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque text with no cues', shapes: ['collection'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
});
});
describe('edge-probe: proposeEdges — invalid authored shapes fail closed (re-review #3 High)', () => {
// A non-empty but INVALID shapes array must NOT silently suppress every probe.
// shapes:['numeric'] (typo for the locked 'numeric-range') previously passed
// Array.isArray, matched no category, and returned applicable:0 — failing OPEN.
test('rejects an unknown shape value (typo for a locked shape)', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric'] }),
/invalid shape/i,
);
});
test('rejects a mixed array where one entry is invalid', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric-range', 'bogus'] }),
/invalid shape/i,
);
});
test('rejects a non-string shape entry', () => {
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: [42] }),
/invalid shape/i,
);
});
test('analyzeCoverage propagates the invalid-shape throw', () => {
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]),
/invalid shape/i,
);
});
test('CLI exits 2 (handled) on an invalid authored shape, not an uncaught trace', () => {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-shape-'));
const reqPath = path.join(dir, 'req.json');
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]));
try {
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe' });
assert.equal(r.status, 2);
} finally {
cleanup(dir);
}
});
test('a valid locked shape still proposes its categories (no false rejection)', () => {
const edges = ep.proposeEdges({ id: 'R1', text: 'opaque', shapes: ['numeric-range'] });
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
});
test('shapes: [] remains a valid zero-edge override (RR-06 intact)', () => {
assert.deepEqual(ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] }), []);
});
});
describe('edge-probe: input validation & orphan-resolution rejection (adversarial review)', () => {
// HIGH: a resolution whose (requirement_id, category) matches no proposed edge — a typo'd
// category or a non-applicable one — was silently DROPPED, so an author who typos `precison`
// sees the precision edge as still-unresolved with no error (a confirmed money-rounding exploit).
test('rejects an orphan resolution (typo category — no matching proposed edge)', () => {
assert.throws(
() => ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'precison', status: 'resolved', verification: 'explicit', resolution: 'AC: precision handled' }],
),
/unknown resolution|no matching proposed edge/i,
);
});
test('rejects a resolution for a valid-but-non-applicable category', () => {
// 'encoding' is a real taxonomy id but applies to text, not the numeric-range requirement.
assert.throws(
() => ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'encoding', status: 'resolved', verification: 'explicit', resolution: 'AC' }],
),
/unknown resolution|no matching proposed edge/i,
);
});
test('a matching resolution still resolves (no false orphan rejection)', () => {
const rep = ep.analyzeCoverage(
[{ id: 'R1', text: 'Round a number to N decimal places' }],
[{ requirement_id: 'R1', category: 'precision', status: 'resolved', verification: 'explicit', resolution: 'AC: precision tested' }],
);
assert.equal(rep.coverage.resolved, 1);
});
test('rejects requirements that is not an array', () => {
assert.throws(() => ep.analyzeCoverage('nope'), /requirements must be an array/i);
});
test('rejects a duplicate requirement id', () => {
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1', text: 'a' }, { id: 'R1', text: 'b' }]),
/duplicate requirement/i,
);
});
test('rejects a truthy non-array shapes (string instead of array)', () => {
// A bare string `shapes: "numeric-range"` previously fell through to prose classification,
// silently ignoring the authored override instead of honoring or rejecting it.
assert.throws(
() => ep.proposeEdges({ id: 'R1', text: 'x', shapes: 'numeric-range' }),
/shapes must be an array/i,
);
});
test('rejects a missing requirement id', () => {
assert.throws(() => ep.proposeEdges({ text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects an empty requirement id', () => {
assert.throws(() => ep.proposeEdges({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects a non-string requirement text', () => {
assert.throws(() => ep.proposeEdges({ id: 'R1', text: 42 }), /text must be a string/i);
});
test('rejects a missing requirement text when no shapes override (M2 fail-open)', () => {
// Without text or an authored shape, prose classification yields zero shapes → zero edges →
// the requirement is silently DROPPED from coverage with no signal — the exact fail-open this
// feature exists to eliminate. The edge adapter's `text` is required, so reject it.
assert.throws(
() => ep.proposeEdges({ id: 'R1' }),
/text must be a non-empty string when no shapes override/i,
);
assert.throws(
() => ep.analyzeCoverage([{ id: 'R1' }]),
/text must be a non-empty string when no shapes override/i,
);
});
test('rejects an empty/whitespace requirement text when no shapes override (M2)', () => {
assert.throws(() => ep.proposeEdges({ id: 'R1', text: '' }), /text must be a non-empty string when no shapes override/i);
assert.throws(() => ep.proposeEdges({ id: 'R1', text: ' ' }), /text must be a non-empty string when no shapes override/i);
});
test('allows missing/empty text WHEN an explicit shapes override is provided (M2 legitimate path)', () => {
// An authored `shapes` array (including `[]` for "no applicable categories") opts out of prose
// classification, so `text` is not required — this must remain valid.
assert.deepEqual(ep.proposeEdges({ id: 'R1', shapes: [] }), []);
const edges = ep.proposeEdges({ id: 'R1', shapes: ['numeric-range'] });
assert.ok(edges.length > 0, 'an explicit shape override must still propose edges without text');
});
});
describe('edge-probe: validateResolution — explicit-needs-resolution (RR-07, re-cut)', () => {
test('rejects resolved/explicit with empty resolution string', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: '' }),
/explicit requires a resolution/i,
);
});
test('rejects resolved/explicit with whitespace-only resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: ' ' }),
/explicit requires a resolution/i,
);
});
test('rejects resolved/explicit with missing resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit' }),
/explicit requires a resolution/i,
);
});
test('accepts resolved/explicit with a non-empty resolution', () => {
assert.equal(
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC#3: boundary tested in suite' }),
true,
);
});
});
describe('edge-probe: validateResolution — backstop-needs-resolution (RR-07 follow-up, re-cut)', () => {
test('rejects resolved/backstop with empty resolution string', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: '' }),
/backstop requires a resolution/i,
);
});
test('rejects resolved/backstop with whitespace-only resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: ' ' }),
/backstop requires a resolution/i,
);
});
test('rejects resolved/backstop with missing resolution', () => {
assert.throws(
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop' }),
/backstop requires a resolution/i,
);
});
test('accepts resolved/backstop with a non-empty resolution note', () => {
assert.equal(
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: 'held-out: covered by integration fuzz suite' }),
true,
);
});
});
describe('edge-probe: analyzeCoverage — duplicate rejection (RR-09)', () => {
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
test('rejects duplicate (requirement_id, category) resolution', () => {
assert.throws(
() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' },
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' },
]),
/duplicate resolution/i,
);
});
test('distinct pairs still analyze without throwing', () => {
// mirrors fixture 06-resolved-mixed
assert.doesNotThrow(() => ep.analyzeCoverage(reqs, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', resolution: null, reason: 'output is canonically sorted; no tie possible' },
]));
});
});
describe('edge-probe: golden fixtures', () => {
const root = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures');
const fixtures = fs.readdirSync(root).filter((d) =>
fs.statSync(path.join(root, d)).isDirectory());
assert.ok(fixtures.length >= 6, 'expected at least 6 fixtures');
for (const name of fixtures) {
test(`fixture ${name} matches its golden coverage`, () => {
const dir = path.join(root, name);
const reqs = JSON.parse(fs.readFileSync(path.join(dir, 'requirements.json'), 'utf8'));
const resPath = path.join(dir, 'resolutions.json');
const res = fs.existsSync(resPath) ? JSON.parse(fs.readFileSync(resPath, 'utf8')) : [];
const expected = JSON.parse(fs.readFileSync(path.join(dir, 'expected-coverage.json'), 'utf8'));
assert.deepEqual(ep.analyzeCoverage(reqs, res), expected);
});
}
});

View File

@@ -0,0 +1,167 @@
'use strict';
/**
* Property-based tests for probe-core.cjs (ADR-550 Decision 7).
*
* Module: gsd-core/bin/lib/probe-core.cjs (generated from src/probe-core.cts)
* Exercised: analyzeCoverage(items, resolutions?, validators) — the generic
* merge/rollup/orphan-reject engine shared by the edge probe and the #644
* prohibition probe.
*
* trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
* transformation/rollup module (items × resolutions → CoverageReport) — exactly the
* class the predicate covers. The example-based suite (tests/probe-core.test.cjs)
* pins specific scenarios; these properties pin the algebraic invariants that must
* hold for EVERY valid scenario.
*
* Properties tested:
* (a) closed-set identity: applicable === resolved + unresolved (and === items.length)
* (b) byVerification sums ≤ resolved (dismissed is closed but unverified)
* (c) per-tier byVerification ≤ resolved, and only `resolved`-status items are counted
* (d) determinism: same input → identical CoverageReport (stable rollup)
* (e) orphan rejection is stable: a resolution matching no proposed item always throws
*/
const { describe, test } = require('node:test');
const path = require('node:path');
const fc = require('./helpers/fast-check-setup.cjs');
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
const pc = require(BUILT_SCRIPT);
// The same representative validators bundle the edge adapter injects (see
// tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe.
const VALIDATORS = {
categories: ['adjacency', 'empty', 'ordering'],
verification: ['explicit', 'backstop'],
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
};
function bareItem(requirement_id, category) {
return {
requirement_id,
category,
status: 'unresolved',
verification: null,
resolution: null,
reason: null,
probe: `probe-for-${category}`,
};
}
const catArb = fc.constantFrom(...VALIDATORS.categories);
const idArb = fc.constantFrom('R1', 'R2', 'R3', 'R4', 'R5');
const keyArb = fc.record({ requirement_id: idArb, category: catArb });
// Unique (requirement_id, category) keys — the merge keys analyzeCoverage maps on.
const keyOf = (k) => `${k.requirement_id}::${k.category}`;
const uniqueKeysArb = fc.uniqueArray(keyArb, { selector: keyOf, minLength: 0, maxLength: 12 });
// Each unique item key gets one resolution disposition. Resolution text/reason use fixed
// non-empty literals — the counting invariants are independent of their content, and this
// keeps the generator off the validateResolution rejection paths (which the example suite
// already covers exhaustively).
const DISPOSITIONS = ['none', 'resolved-explicit', 'resolved-backstop', 'dismissed', 'unresolved'];
function resolutionFor(k, disposition) {
const base = { requirement_id: k.requirement_id, category: k.category };
switch (disposition) {
case 'resolved-explicit':
return { ...base, status: 'resolved', verification: 'explicit', resolution: 'AC#1' };
case 'resolved-backstop':
return { ...base, status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' };
case 'dismissed':
return { ...base, status: 'dismissed', reason: 'bounded enum — not applicable' };
case 'unresolved':
return { ...base, status: 'unresolved' };
default: // 'none' — author left no resolution; item rolls up verbatim (bare unresolved)
return null;
}
}
// A fully valid scenario: unique items (all bare-unresolved) + a per-item resolution choice.
const scenarioArb = uniqueKeysArb.chain((keys) =>
fc.tuple(...keys.map(() => fc.constantFrom(...DISPOSITIONS))).map((choices) => {
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
const resolutions = [];
keys.forEach((k, i) => {
const r = resolutionFor(k, choices[i]);
if (r) resolutions.push(r);
});
return { items, resolutions };
}),
);
describe('probe-core property: analyzeCoverage algebraic invariants', () => {
test('(a) closed-set identity: applicable === resolved + unresolved === items.length', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
return (
coverage.applicable === coverage.resolved + coverage.unresolved &&
coverage.applicable === items.length
);
}),
);
});
test('(b) sum(byVerification) ≤ resolved — dismissed counts closed but unverified', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
const verifiedTotal = Object.values(coverage.byVerification).reduce((a, b) => a + b, 0);
return verifiedTotal <= coverage.resolved && verifiedTotal >= 0;
}),
);
});
test('(c) byVerification only counts resolved-status items, and matches a direct recount', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const rep = pc.analyzeCoverage(items, resolutions, VALIDATORS);
for (const tier of VALIDATORS.verification) {
const recount = rep.items.filter((i) => i.status === 'resolved' && i.verification === tier).length;
if (rep.coverage.byVerification[tier] !== recount) return false;
if (rep.coverage.byVerification[tier] > rep.coverage.resolved) return false;
}
return true;
}),
);
});
test('(d) determinism: identical inputs produce an identical CoverageReport', () => {
fc.assert(
fc.property(scenarioArb, ({ items, resolutions }) => {
const a = pc.analyzeCoverage(items, resolutions, VALIDATORS);
const b = pc.analyzeCoverage(items, resolutions, VALIDATORS);
return JSON.stringify(a) === JSON.stringify(b);
}),
);
});
});
describe('probe-core property: orphan rejection is stable', () => {
// An orphan resolution carries an id ('Z9') that no generated item ever uses, so its
// (requirement_id, category) key never matches a proposed item. The resolution is itself
// structurally VALID (a bare unresolved), so it clears validateResolution and reaches the
// orphan-reject guard — isolating that guard from input-validation throws.
const orphanScenarioArb = fc.record({
keys: uniqueKeysArb,
orphanCategory: catArb,
});
test('(e) a resolution matching no proposed item always throws', () => {
fc.assert(
fc.property(orphanScenarioArb, ({ keys, orphanCategory }) => {
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
const orphan = { requirement_id: 'Z9', category: orphanCategory, status: 'unresolved' };
let threwForOrphan = false;
try {
pc.analyzeCoverage(items, [orphan], VALIDATORS);
} catch (e) {
threwForOrphan = /unknown resolution|no matching proposed item/i.test(e.message);
}
return threwForOrphan;
}),
);
});
});

335
tests/probe-core.test.cjs Normal file
View File

@@ -0,0 +1,335 @@
/**
* probe-core reference-model unit tests (ADR-550 Decision 7).
*
* probe-core is the GENERIC spec-phase probe resolution model extracted from the
* edge-probe (the first adapter): the resolution lifecycle, the two-axis
* status×verification re-cut, `validateResolution`/`validateRequirement`, the
* `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject
* engine, the `byVerification` rollup, and the `runProbeCli` I/O scaffold.
*
* Asserts the LOCKED export surface against the BUILT artifact
* (`gsd-core/bin/lib/probe-core.cjs`), which `npm run build:lib` (run by pretest /
* the run-tests sentinel) emits from `src/probe-core.cts`.
*
* The injected runtime validators are the enforcement contract (ADR-550 #5): the
* CLI runs over JSON where TS types are erased, so `analyzeCoverage` is told its
* probe's closed vocabularies — `{ categories, verification, requiredFieldsByVerification }`
* — rather than relying on the type system. These tests pin the validators' behavior.
*/
'use strict';
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const path = require('node:path');
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
const pc = require(BUILT_SCRIPT);
// A representative validators bundle — the shape the edge adapter injects, used here
// to exercise the generic engine independent of any one probe.
const VALIDATORS = {
categories: ['adjacency', 'empty', 'ordering'],
verification: ['explicit', 'backstop'],
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
};
function item(category, overrides = {}) {
return {
requirement_id: 'R1',
category,
status: 'unresolved',
verification: null,
resolution: null,
reason: null,
probe: `probe-for-${category}`,
...overrides,
};
}
const UNRESOLVED_ITEMS = [item('adjacency'), item('empty'), item('ordering')];
describe('probe-core: VALID_STATUS is the re-cut lifecycle enum', () => {
test('exposes exactly resolved | dismissed | unresolved (no covered/backstop)', () => {
assert.deepEqual([...ep_sorted(pc.VALID_STATUS)], ['dismissed', 'resolved', 'unresolved']);
assert.ok(!pc.VALID_STATUS.includes('covered'), 'covered must not survive the re-cut as a status');
assert.ok(!pc.VALID_STATUS.includes('backstop'), 'backstop must not survive the re-cut as a status');
});
});
function ep_sorted(arr) {
return [...arr].sort();
}
describe('probe-core: validateResolution (status×verification)', () => {
const v = (r) => pc.validateResolution(r, VALIDATORS);
test('rejects an unknown status', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'maybe' }), /invalid status/i);
});
test('rejects a former covered status (re-cut: no longer a valid status)', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'covered', resolution: 'x' }), /invalid status/i);
});
test('rejects dismissed without a reason', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: '' }), /dismissed requires a reason/i);
});
test('accepts dismissed with a reason', () => {
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' }), true);
});
test('rejects resolved with a missing verification tier', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', resolution: 'AC' }), /verification/i);
});
test('rejects resolved with an unknown verification tier', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'judgment', resolution: 'AC' }), /invalid verification/i);
});
test('rejects resolved/explicit with empty resolution text', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: ' ' }), /explicit requires a resolution/i);
});
test('rejects resolved/backstop with missing resolution note', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop' }), /backstop requires a resolution/i);
});
test('accepts resolved/explicit with a resolution', () => {
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' }), true);
});
test('accepts resolved/backstop with a resolution note', () => {
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' }), true);
});
// Re-review #5 Medium — fail closed across the FULL status×verification model, not just
// `resolved`. The invariant (probe-core header): verification is null unless status is
// resolved. A dismissed/unresolved resolution carrying a tier otherwise merges verbatim
// (analyzeCoverage line ~189), breaking the model for the second adapter (#644) that
// inherits this seam.
test('rejects a dismissed resolution carrying a verification tier (null unless resolved)', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'n/a', verification: 'explicit' }), /verification must be null/i);
});
test('rejects an unresolved resolution carrying a verification tier', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', verification: 'backstop' }), /verification must be null/i);
});
// An unresolved resolution is an UNACTED item — a populated resolution/reason payload is an
// authoring mistake (the author meant resolved/dismissed) that today is silently dropped into
// the unresolved count with no error pointing at it. Reject the payload.
test('rejects an unresolved resolution carrying a resolution payload', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', resolution: 'AC#1' }), /unresolved must not carry/i);
});
test('rejects an unresolved resolution carrying a reason payload', () => {
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', reason: 'because' }), /unresolved must not carry/i);
});
test('accepts a bare unresolved resolution (no payload — a harmless no-op merge)', () => {
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved' }), true);
});
});
describe('probe-core: validateRequirement (generic id/text only)', () => {
test('rejects a missing id', () => {
assert.throws(() => pc.validateRequirement({ text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects an empty id', () => {
assert.throws(() => pc.validateRequirement({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i);
});
test('rejects a non-string text', () => {
assert.throws(() => pc.validateRequirement({ id: 'R1', text: 42 }), /text must be a string/i);
});
test('accepts a valid requirement', () => {
assert.doesNotThrow(() => pc.validateRequirement({ id: 'R1', text: 'a testable statement' }));
});
});
describe('probe-core: analyzeCoverage (merge · rollup · byVerification)', () => {
test('no resolutions → every item unresolved; resolved 0; byVerification zeroed per tier', () => {
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [], VALIDATORS);
assert.deepEqual(rep.coverage, {
applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 },
});
});
test('merges a resolved/explicit resolution and counts byVerification.explicit', () => {
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
], VALIDATORS);
const adj = rep.items.find((i) => i.category === 'adjacency');
assert.equal(adj.status, 'resolved');
assert.equal(adj.verification, 'explicit');
assert.equal(adj.resolution, 'AC#6: touching intervals merge');
assert.equal(rep.coverage.resolved, 1);
assert.equal(rep.coverage.unresolved, 2);
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 });
});
test('a dismissed item counts toward coverage.resolved (closed set) but NOT byVerification', () => {
// coverage.resolved preserves the pre-re-cut "closed" semantic = applicable - unresolved
// (covered + dismissed + backstop), per edge-probe.md and the migration contract.
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted; no tie' },
], VALIDATORS);
assert.equal(rep.coverage.resolved, 1);
assert.equal(rep.coverage.unresolved, 2);
assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 });
const ord = rep.items.find((i) => i.category === 'ordering');
assert.equal(ord.status, 'dismissed');
assert.equal(ord.verification, null);
assert.equal(ord.reason, 'canonically sorted; no tie');
});
test('mixed resolved/explicit + backstop + dismissed: resolved = closed = applicable - unresolved', () => {
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' },
{ requirement_id: 'R1', category: 'empty', status: 'resolved', verification: 'backstop', resolution: 'held-out empty-input PBT' },
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' },
], VALIDATORS);
assert.equal(rep.coverage.applicable, 3);
assert.equal(rep.coverage.unresolved, 0);
assert.equal(rep.coverage.resolved, 3); // 2 resolved-status + 1 dismissed
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 1 });
});
test('an all-dismissed run is CLOSED but NOT affirmatively covered (byVerification is the honest gate)', () => {
// Re-review #5 Medium — coverage.resolved is the closed set (resolved + dismissed), kept
// count-preserved per the blessed migration contract. So an all-dismissed spec has
// resolved === applicable while NOTHING was affirmatively resolved/backstopped. A CI gate
// keying on `resolved === applicable` would read green here; the honest signal is
// byVerification (all tiers zero). This test locks that distinction.
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' },
{ requirement_id: 'R1', category: 'empty', status: 'dismissed', reason: 'guaranteed non-empty' },
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' },
], VALIDATORS);
assert.equal(rep.coverage.unresolved, 0);
assert.equal(rep.coverage.resolved, 3); // closed set counts the dismissals
assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 }); // nothing affirmatively verified
});
test('rejects a duplicate (requirement_id, category) resolution', () => {
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' },
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' },
], VALIDATORS), /duplicate resolution/i);
});
test('rejects an orphan resolution (no matching proposed item)', () => {
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC' },
], VALIDATORS), /unknown resolution|no matching proposed/i);
});
test('propagates an invalid-resolution throw from validateResolution', () => {
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
{ requirement_id: 'R1', category: 'empty', status: 'dismissed' },
], VALIDATORS), /dismissed requires a reason/i);
});
test('rejects items that is not an array', () => {
assert.throws(() => pc.analyzeCoverage('nope', [], VALIDATORS), /items must be an array/i);
});
test('rejects a proposed item whose category is not in validators.categories', () => {
// Adapter self-consistency: an item carrying a category outside the probe's closed
// vocabulary is an adapter bug, caught here rather than silently rolled up.
assert.throws(() => pc.analyzeCoverage([item('bogus-category')], [], VALIDATORS), /unknown category/i);
});
// m1: a verbatim item (no matching author resolution) is rolled up as-is, so its OWN
// status/fields must also be validated. The edge adapter only proposes `unresolved` items, but
// the prohibition adapter (#644) proposes LLM-generated items that arrive already populated —
// an out-of-enum status or a `dismissed` with no reason must fail closed, not count as closed.
test('m1: rejects a verbatim item carrying an out-of-enum status (e.g. the dropped "covered")', () => {
assert.throws(
() => pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [], VALIDATORS),
/invalid status/i,
);
});
test('m1: rejects a verbatim item dismissed without a reason', () => {
assert.throws(
() => pc.analyzeCoverage([item('adjacency', { status: 'dismissed' })], [], VALIDATORS),
/dismissed requires a reason/i,
);
});
test('m1: rejects a verbatim resolved item missing its verification tier', () => {
assert.throws(
() => pc.analyzeCoverage([item('adjacency', { status: 'resolved', resolution: 'AC' })], [], VALIDATORS),
/requires a verification tier/i,
);
});
test('m1: a matching resolution still governs — item validation targets VERBATIM items only', () => {
// When a resolution matches, the item is rebuilt from the (already-validated) resolution, so a
// pre-populated item status is irrelevant and must NOT cause a throw. Guards against the m1
// fix over-reaching into the merge path.
const rep = pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' },
], VALIDATORS);
const adj = rep.items.find((i) => i.category === 'adjacency');
assert.equal(adj.status, 'resolved');
assert.equal(adj.verification, 'explicit');
assert.equal(rep.coverage.resolved, 1);
});
});
describe('probe-core: runProbeCli (generic I/O scaffold, injected io)', () => {
const report = { items: [], coverage: { applicable: 0, resolved: 0, unresolved: 0, byVerification: {} } };
test('no requirements path → writes usage to stderr and exits 2', () => {
let code; let err = '';
pc.runProbeCli(() => report, {
usage: 'demo-probe.cjs <requirements.json> [resolutions.json]',
argv: ['node', 'demo'], writeErr: (s) => { err += s; }, exit: (c) => { code = c; },
});
assert.equal(code, 2);
assert.match(err, /usage: demo-probe\.cjs/);
});
test('valid requirements path → calls analyze and prints the report as JSON', () => {
let out = '';
pc.runProbeCli((reqs, res) => {
assert.deepEqual(reqs, [{ id: 'R1', text: 'x' }]);
assert.deepEqual(res, []);
return report;
}, {
usage: 'demo', argv: ['node', 'demo', '/fake/req.json'],
readFile: () => '[{"id":"R1","text":"x"}]', write: (s) => { out += s; }, exit: () => {},
});
assert.deepEqual(JSON.parse(out), report);
});
test('reads the optional resolutions file when a second path is given', () => {
let seenRes;
pc.runProbeCli((reqs, res) => { seenRes = res; return report; }, {
usage: 'demo', argv: ['node', 'demo', '/req.json', '/res.json'],
readFile: (p) => (p === '/res.json' ? '[{"requirement_id":"R1"}]' : '[{"id":"R1"}]'),
write: () => {}, exit: () => {},
});
assert.deepEqual(seenRes, [{ requirement_id: 'R1' }]);
});
test('invalid requirements JSON → exits 2 (handled, not an uncaught throw)', () => {
let code;
pc.runProbeCli(() => report, {
usage: 'demo', argv: ['node', 'demo', '/req.json'],
readFile: () => 'not json {{{', writeErr: () => {}, exit: (c) => { code = c; },
});
assert.equal(code, 2);
});
test('an analyze throw → exits 2 (the engine fail-closed surfaces, never silently passes)', () => {
let code;
pc.runProbeCli(() => { throw new Error('boom'); }, {
usage: 'demo', argv: ['node', 'demo', '/req.json'],
readFile: () => '[]', writeErr: () => {}, exit: (c) => { code = c; },
});
assert.equal(code, 2);
});
// Re-review #5 Low — the scaffold trusts each adapter's `as` cast for the returned report.
// A future adapter (#644) that forgets to validate inside its closure would otherwise have a
// structurally-broken report written as green output. Guard the report shape structurally.
test('a structurally-invalid report from analyze → exits 2 and writes nothing (no silent malformed output)', () => {
let code; let out = '';
pc.runProbeCli(() => ({ nope: true }), {
usage: 'demo', argv: ['node', 'demo', '/req.json'],
readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; },
});
assert.equal(code, 2);
assert.equal(out, '');
});
test('a report whose coverage object carries non-numeric counts → exits 2 (partial-guard case)', () => {
// items[] is well-formed and `coverage` IS an object, so the cheaper checks pass — this
// exercises the numeric-count branch of the structural guard specifically.
let code; let out = '';
pc.runProbeCli(() => ({ items: [], coverage: { applicable: 'x', resolved: null, unresolved: undefined, byVerification: {} } }), {
usage: 'demo', argv: ['node', 'demo', '/req.json'],
readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; },
});
assert.equal(code, 2);
assert.equal(out, '');
});
test('a well-formed report still writes and does not trip the structural guard', () => {
let out = '';
pc.runProbeCli(() => report, {
usage: 'demo', argv: ['node', 'demo', '/req.json'],
readFile: () => '[]', write: (s) => { out += s; }, exit: () => {},
});
assert.deepEqual(JSON.parse(out), report);
});
});

View File

@@ -51,7 +51,7 @@
"note.md": 6563,
"pause-work.md": 13654,
"plan-milestone-gaps.md": 11765,
"plan-phase.md": 93135,
"plan-phase.md": 94253,
"plan-review-convergence.md": 22949,
"plant-seed.md": 11741,
"pr-branch.md": 4994,
@@ -72,7 +72,7 @@
"ship.md": 20896,
"sketch-wrap-up.md": 14223,
"sketch.md": 19960,
"spec-phase.md": 15131,
"spec-phase.md": 23094,
"spike-wrap-up.md": 15092,
"spike.md": 24517,
"stats.md": 6718,