* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:
- spec-phase.md 15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
- plan-phase.md 93135 -> 94253 (+1118): covered/backstop edge lift into
must_haves.truths (the live <downstream_consumer> block)
Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.
* chore(#550): reconcile INVENTORY headline counts after rebase onto next
Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
- References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
- CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs
Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
This commit is contained in:
5
.changeset/clever-eagles-jump.md
Normal file
5
.changeset/clever-eagles-jump.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Added
|
||||
pr: 584
|
||||
---
|
||||
spec-phase: spec-completeness edge-probe — a taxonomy-driven Step 5.5 that walks each SPEC requirement against a closed 8-category edge taxonomy (boundary, adjacency, empty/degenerate, encoding, ordering, precision, idempotency, concurrency), proposes concrete candidate edges, and resolves each to covered/dismissed/backstop/unresolved. covered edges add acceptance criteria the planner lifts into must_haves.truths; a soft gate flags unresolved edges. Additive and optional: existing SPECs without an Edge Coverage section remain valid.
|
||||
2
.gitignore
vendored
2
.gitignore
vendored
@@ -71,6 +71,8 @@ build/
|
||||
/gsd-core/bin/lib/research-provider.cjs
|
||||
/gsd-core/bin/lib/package-legitimacy.cjs
|
||||
/gsd-core/bin/lib/semver-compare.cjs
|
||||
/gsd-core/bin/lib/edge-probe.cjs
|
||||
/gsd-core/bin/lib/probe-core.cjs
|
||||
/gsd-core/bin/lib/config-types.cjs
|
||||
/gsd-core/bin/lib/cli-exit.cjs
|
||||
/gsd-core/bin/lib/code-review-flags.cjs
|
||||
|
||||
@@ -204,6 +204,12 @@ The GSD-RESEARCH capability behind an L2-hybrid seam: code owns cache + provider
|
||||
### UAT-Passed Predicate
|
||||
Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fields with markdown-aware parsing that ignores false-positive contexts (frontmatter body, fenced code, HTML comments, blockquotes). Returns `passed: true` only when all required checks pass; supports `--require-verification` to demand at least one VERIFICATION.md file alongside UAT results. Output envelope: `{ passed, uat_files[], verification_files[], checks[], blockers[], policy }`. Source: `gsd-core/bin/lib/uat-predicate.cjs` (generated from `src/uat-predicate.cts`). Wired via `phase uat-passed` alias → `phase-command-router` → `cmdPhaseUatPassed`.
|
||||
|
||||
### Probe Core Module
|
||||
Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the prohibition probe (#644) next. Exports: `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli`. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module.
|
||||
|
||||
### Edge Probe Module
|
||||
First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the 8 edge categories), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module.
|
||||
|
||||
### MVP Mode
|
||||
Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`.
|
||||
|
||||
|
||||
@@ -90,6 +90,34 @@ Manage GSD workspaces — create, list, or remove isolated workspace environment
|
||||
|
||||
---
|
||||
|
||||
### `/gsd-spec-phase`
|
||||
|
||||
Clarify WHAT a phase delivers through Socratic questioning with quantitative ambiguity scoring, then probe for omitted edges. Produces `SPEC.md` before discuss-phase.
|
||||
|
||||
| Argument | Required | Description |
|
||||
|----------|----------|-------------|
|
||||
| `N` | Yes | Phase number |
|
||||
|
||||
| Flag | Description |
|
||||
|------|-------------|
|
||||
| `--auto` | Skip interactive questions; Claude selects recommended defaults and writes SPEC.md |
|
||||
| `--text` | Use plain-text numbered lists instead of TUI menus (required for `/rc` remote sessions) |
|
||||
|
||||
**Position in workflow:** `spec-phase → discuss-phase → plan-phase → execute-phase → verify`
|
||||
|
||||
**Edge Coverage (Step 5.5):** After the ambiguity gate passes, spec-phase runs an edge-completeness probe over each requirement. It raises only applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency), proposes one concrete candidate edge per category, and records each as `covered` / `dismissed` (reason required) / `backstop` / `unresolved` in a `## Edge Coverage` SPEC section. Unresolved applicable edges soft-gate the spec (Resolve / Write-anyway-flagged / Keep-probing); `covered` and `backstop` edges are later lifted into plan-phase `must_haves`. Under `--auto` the probe **never auto-dismisses** — it auto-covers where a defensible acceptance criterion exists, otherwise auto-backstops.
|
||||
|
||||
**Prerequisites:** `.planning/ROADMAP.md` exists
|
||||
**Produces:** `{phase}-SPEC.md` (with a `## Edge Coverage` section)
|
||||
|
||||
```bash
|
||||
/gsd-spec-phase 1 # Interactive spec + edge probe for phase 1
|
||||
/gsd-spec-phase 3 --auto # Auto-select defaults; never auto-dismisses an edge
|
||||
/gsd-spec-phase 2 --text # Plain-text menus for remote sessions
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### `/gsd-discuss-phase`
|
||||
|
||||
Gather phase context through adaptive questioning before planning.
|
||||
|
||||
@@ -165,6 +165,7 @@
|
||||
- [Milestone Tag Creation Toggle](#141-milestone-tag-creation-toggle)
|
||||
- [Structured JSON Error Mode](#142-structured-json-error-mode)
|
||||
- [UAT-Passed Predicate](#143-uat-passed-predicate)
|
||||
- [Spec-Phase Edge-Completeness Probe](#144-spec-phase-edge-completeness-probe)
|
||||
|
||||
---
|
||||
|
||||
@@ -3086,3 +3087,34 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev
|
||||
- [docs index](README.md)
|
||||
|
||||
**Reference:** [JSON Error Mode](json-errors.md)
|
||||
|
||||
---
|
||||
|
||||
### 144. Spec-Phase Edge-Completeness Probe
|
||||
|
||||
**Command:** `/gsd-spec-phase`
|
||||
|
||||
**Purpose:** Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as `Step 5.5` of spec-phase, after the ambiguity gate.
|
||||
|
||||
**Behavior:** For each SPEC requirement the probe classifies its data/behavior shape, then raises only the *applicable* categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:
|
||||
|
||||
| State | Meaning | Downstream effect |
|
||||
|-------|---------|-------------------|
|
||||
| `covered` | An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into `plan-phase` `must_haves.truths` |
|
||||
| `dismissed` | The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
|
||||
| `backstop` | Intent recorded, needs a held-out/property-based test | Lifted into `must_haves.truths` as a non-inferable check |
|
||||
| `unresolved` | Deferred | Soft-gates the spec; row stamped `⚠ Edge unresolved — planner must treat as assumption` |
|
||||
|
||||
The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`.
|
||||
|
||||
The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation.
|
||||
|
||||
**Requirements:**
|
||||
- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a `## Edge Coverage` SPEC section.
|
||||
- REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
|
||||
- REQ-EDGE-03: A `dismissed` resolution MUST require a non-empty reason.
|
||||
- REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
|
||||
- REQ-EDGE-05: `--auto` MUST never auto-dismiss — auto-cover or auto-backstop only.
|
||||
- REQ-EDGE-06: `plan-phase` MUST lift `covered` criteria and `backstop` notes into `must_haves.truths`.
|
||||
|
||||
**Reference:** [Edge Probe](../gsd-core/references/edge-probe.md)
|
||||
|
||||
@@ -209,6 +209,7 @@
|
||||
"decimal-phase-calculation.md",
|
||||
"doc-conflict-engine.md",
|
||||
"domain-probes.md",
|
||||
"edge-probe.md",
|
||||
"execute-mvp-tdd.md",
|
||||
"executor-examples.md",
|
||||
"gate-prompts.md",
|
||||
@@ -295,6 +296,7 @@
|
||||
"decisions.cjs",
|
||||
"docs.cjs",
|
||||
"drift.cjs",
|
||||
"edge-probe.cjs",
|
||||
"fallow-runner.cjs",
|
||||
"federated-config.cjs",
|
||||
"frontmatter.cjs",
|
||||
@@ -329,6 +331,7 @@
|
||||
"phases-command-router.cjs",
|
||||
"plan-scan.cjs",
|
||||
"planning-workspace.cjs",
|
||||
"probe-core.cjs",
|
||||
"profile-output.cjs",
|
||||
"profile-pipeline.cjs",
|
||||
"project-root.cjs",
|
||||
|
||||
@@ -264,7 +264,7 @@ Full roster at `gsd-core/workflows/*.md`. Workflows are thin orchestrators that
|
||||
|
||||
---
|
||||
|
||||
## References (68 shipped)
|
||||
## References (69 shipped)
|
||||
|
||||
Full roster at `gsd-core/references/*.md`. References are shared knowledge documents that workflows and agents `@-reference`. The groupings below match [`docs/ARCHITECTURE.md`](ARCHITECTURE.md#references-gsd-corereferencesmd) — core, workflow, thinking-model clusters, and the modular planner decomposition.
|
||||
|
||||
@@ -300,6 +300,7 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum
|
||||
| `context-budget.md` | Context window budget allocation rules. |
|
||||
| `continuation-format.md` | Session continuation/resume format. |
|
||||
| `domain-probes.md` | Domain-specific probing questions for discuss-phase. |
|
||||
| `edge-probe.md` | Spec-phase edge-completeness probe — 8-category edge taxonomy, shape classification, and the `requirements → checks → verifier` resolution model (Step 5.5). |
|
||||
| `gate-prompts.md` | Gate/checkpoint prompt templates. |
|
||||
| `scout-codebase.md` | Phase-type→codebase-map selection table for discuss-phase scout step (extracted via #2551). |
|
||||
| `revision-loop.md` | Plan revision iteration patterns. |
|
||||
@@ -371,7 +372,7 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
|
||||
|
||||
---
|
||||
|
||||
## CLI Modules (107 shipped)
|
||||
## CLI Modules (109 shipped)
|
||||
|
||||
Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
|
||||
@@ -406,6 +407,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
| `decisions.cjs` | Parses CONTEXT.md `<decisions>` blocks; accepts numeric (D-42) and alphanumeric (D-INFRA-01) IDs; returns `{id, text, category, tags, trackable}` |
|
||||
| `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection |
|
||||
| `drift.cjs` | Post-execute codebase structural drift detector (#2003): classifies file changes into new-dir/barrel/migration/route categories and round-trips `last_mapped_commit` frontmatter |
|
||||
| `edge-probe.cjs` | Spec-completeness edge probe (compiled from `src/edge-probe.cts`, gitignored) — the first adapter of the `probe-core` resolution model (ADR-550 Decision 7): shape classification, applicable-category relevance filter, edge proposal, and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`; exports `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `TAXONOMY` (#550) |
|
||||
| `fallow-runner.cjs` | Fallow audit adapter for `/gsd-code-review`: binary resolution (`PATH` then `node_modules/.bin`), actionable missing-binary errors, and structural findings normalization |
|
||||
| `federated-config.cjs` | Defensive merge of capability-declared config slices into the loadConfig return value — ADR-857 phase 3b; exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig })` → `{ values, validKeys, warnings }`; no-op until a key is atomically removed from the central config-schema (the cutover step) |
|
||||
| `frontmatter.cjs` | YAML frontmatter CRUD operations |
|
||||
@@ -446,6 +448,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
| `prompt-budget.cjs` | Pure token-budget accounting for review prompts — estimates tokens, applies deterministic trim priority (head-shrink PROJECT.md, proportional plan truncation, drop context/research/requirements, hard-fail guard), returns structured metadata for `review.max_prompt_tokens` (#3081) |
|
||||
| `research-provider.cjs` | Research provider waterfall, confidence tiers, and planResearch (cache-hits + fetch plan) |
|
||||
| `research-store.cjs` | Content-addressed research cache: sha256 keys, per-source TTL staleness, two-tier (user ~/.gsd / project .planning) store |
|
||||
| `probe-core.cjs` | Generic spec-phase probe resolution model (compiled from `src/probe-core.cts`, gitignored; ADR-550 Decision 7) — the status×verification re-cut (`status: resolved/dismissed/unresolved` × per-probe `verification`), `validateResolution`/`validateRequirement`, `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject, the `byVerification` rollup, and the `runProbeCli` I/O scaffold; the shared seam consumed by `edge-probe` (and the prohibition probe #644); exports `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` (#550) |
|
||||
| `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence |
|
||||
| `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` |
|
||||
| `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) |
|
||||
|
||||
@@ -18,6 +18,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
|
||||
- [Install on your runtime](how-to/install-on-your-runtime.md) — runtime-specific install steps for all 16 supported runtimes
|
||||
- [Install a minimal GSD and add skills later](how-to/install-minimal-and-add-skills.md) — install only the core skills, then grow the surface with profiles and `/gsd:surface`
|
||||
- [Discuss a phase](how-to/discuss-a-phase.md) — capture implementation decisions before planning begins
|
||||
- [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
|
||||
- [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality
|
||||
- [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents
|
||||
- [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR
|
||||
|
||||
77
docs/adr/550-spec-phase-probe-contract.md
Normal file
77
docs/adr/550-spec-phase-probe-contract.md
Normal file
@@ -0,0 +1,77 @@
|
||||
# ADR 550: spec-phase probe pattern and prohibition contract [Accepted]
|
||||
|
||||
- **Status:** Accepted
|
||||
- **Date:** 2026-06-03
|
||||
|
||||
> **Provenance.** Drafted during triage of #644 (prohibition probe), hardened in a `/grill-with-docs` session, and extended with a `probe-core` seam designed in an `/improve-codebase-architecture` grill (Decision 7). Endorsed by the maintainer and by the #644 author (who also authored the edge-probe, #550 / PR #584). This ADR lands **on PR #584** alongside the `probe-core` extraction that implements Decision 7; #644 (the prohibition probe itself) remains blocked until #584 merges and then implements the prohibition adapter. "What exists today" is verified against `next` + the PR-#584 head as of 2026-06-03.
|
||||
>
|
||||
> **Lineage.** The judgment-tier contract (Decision 4) was reached independently from two directions: from the domain model (truths are unenforced positive-observables, so a values prohibition cannot be a silent pass) and from the author's "verifier-abstention" (N17) calibration experiment (on irreducible values rules the verifier is confidently wrong — ~0.93 confidence, ECE 0.81 — so confidence-gating cannot help; the only honest move is abstain-and-flag). That parked experiment folds into this ADR as the verify-time half of judgment-tier rather than needing a separate feature.
|
||||
|
||||
## Context
|
||||
|
||||
`spec-phase` produces `SPEC.md` from an interview. Two features add *probes* — soft-gate steps that surface spec gaps before code is written: **#550 / PR #584 (edge-probe)** for data-shape edges, and **#644 (prohibition probe)** for values/safety "must-NOT" constraints. Both use identical three-layer packaging and a recall→precision protocol, and both want confirmed findings to reach `SPEC.md` and the plan contract.
|
||||
|
||||
Grounding the design against the codebase surfaced the decisive facts:
|
||||
|
||||
- **`truths` are positive, observable assertions.** `gsd-planner` derives them as *"observable truths (3-7, user perspective)"*; `verify-phase` checks each by *"determine if the codebase enables it"* — a presence/wiring check. A prohibition is the inverse and cannot be confirmed that way.
|
||||
- **`truths` are not mechanically verified.** `verify.cjs` iterates `artifacts` and `key_links` but has no `truths` handler. A prohibition parked in `truths` would inherit that non-enforcement — a negative wearing a positive's clothing.
|
||||
- **`must_haves` has no schema/type** — a convention parsed by `parseMustHavesBlock` (blocks `truths`/`artifacts`/`key_links`), read in ≥6 places.
|
||||
- **The SPEC→plan lift is LLM reasoning** in `gsd-planner`'s `derive_must_haves`, not mechanical extraction.
|
||||
- **Prohibitions already exist in GSD only as narrative prose** (`<scope_reduction_prohibition>`, `"MUST NOT loop"`), never as structured, verifiable data.
|
||||
- **The edge-probe's load-bearing logic is a deep module on the live path.** `src/edge-probe.cts` (~297 lines) is compiled to `bin/lib/edge-probe.cjs` and invoked by `spec-phase` Step 5.5 at runtime. ~120 of its lines are a *generic resolution model* (status lifecycle, validation, merge/rollup, CLI); only the `Shape`/`SHAPE_CUES`/`TAXONOMY`/`classifyShape`/`proposeEdges` cluster is edge-specific. The prohibition probe is the **second adapter** of that model — making the seam real (one adapter is hypothetical; two is real).
|
||||
- **No prior ADR** governs probes, soft gates, or `must_haves` evolution.
|
||||
|
||||
The values/safety class #644 targets (*"must not become a guilt mechanic"*) is frequently **not reducible to a deterministic test** — which is why it needs an explicit verification contract rather than truth-style limbo.
|
||||
|
||||
## Decision
|
||||
|
||||
1. **Probe packaging (three layers).** Every spec-phase probe ships as: (a) a portable, dependency-free reference core under `references/<probe>.md` (taxonomy + protocol + output schema); (b) a `spec-phase.md` step run as a **soft gate** (write-anyway-with-flags), with explicit `--auto` and text-mode (non-Claude, no `AskUserQuestion`) handling, mirroring the ambiguity gate; (c) tests + docs.
|
||||
|
||||
2. **Two-stage protocol.** **Recall** (adversarial over-generation) → **precision** (classifier dropping routine engineering items), surfacing a short confirmable list. Dismissals require a non-empty reason.
|
||||
|
||||
3. **Prohibition home & representation.** Confirmed prohibitions live primarily as **`SPEC.md` acceptance criteria** (negative criteria). They MAY be projected into an optional **`must_haves.prohibitions:` sibling block** (alongside `truths`/`artifacts`/`key_links`) when they must survive into the plan contract. **`truths` is left untouched** — no `polarity` field is added. Each prohibition item carries `statement`, the orthogonal `status` + `verification` of Decision 7, and (when dismissed) a non-empty `reason`.
|
||||
|
||||
4. **Tiered verification.** Each prohibition declares `verification: test | judgment`.
|
||||
- **`test`-tier** (reducible to a deterministic assertion) → a **negative test** following `regression-must-fail-first` + negative-proof discipline. **Hard gate in both interactive and autonomous modes.**
|
||||
- **`judgment`-tier** (irreducible values/safety rule) → **mode-dependent soft-gate-with-flags**: *interactive* verify requires explicit human resolution of each item; *autonomous* verify records an LLM-judge verdict **marked non-authoritative** and emits a prominent *"unverified prohibition — human review recommended"* flag in SUMMARY/verdict. **Never a silent pass; never a hard halt of AFK runs.** Autonomous completion reads *"complete with N flagged prohibitions."*
|
||||
- *Optional future input:* cross-tier (or cross-model) disagreement MAY feed the judgment-tier flag as a cheap blind-spot signal. This is a breadcrumb, **not a dependency** — the flag stands without it.
|
||||
|
||||
5. **The CI-testable surface is the contract, not the classifier.** The recall/precision/tier-assignment stages are LLM behavior and are **not** claimed as deterministic CI coverage (validated by offline batteries; in CI only by source-grep-under-`// allow-test-rule:`-exemption over reference/workflow prose, which *is* the product). CI deterministically tests the **contract**: (a) parse + validate an item (`statement` present; `status ∈ {resolved, dismissed, unresolved}`; `verification` in the probe's allowed set; dismissed items carry a non-empty `reason`); (b) round-trip / projection between `SPEC.md` and `must_haves.prohibitions:`; (c) a `DEFECT.GENERATIVE-FIX` **parity assertion** across template ↔ parser ↔ planner; (d) `test`-tier items are provably wired into verify-phase and cannot be silently skipped. A test that asserts the LLM's *judgment* is vacuous and is rejected per `RULESET.TESTS.delete-bad-tests`. *(Scope: on #584 only **(a)** ships — it is the edge adapter's parse+validate contract test (`tests/probe-core.test.cjs`). **(b)–(d)** require the `must_haves.prohibitions:` block, projection, and verify-phase tiering and are **#644 scope**, as the Consequences section records.)*
|
||||
|
||||
6. **Ownership seam with security tooling.** The probe owns **bespoke product/values** prohibitions only. When precision classifies an item as **canon security/compliance** (OWASP/GDPR/fairness/prototype-pollution/path-traversal), it does **not** mint a SPEC prohibition; it emits a one-line breadcrumb (*"possible canon-security concern X — owned by `/gsd:secure-phase` / eslint"*) and stops. Canon checks are **referred, not duplicated** — keeping the surfaced list short (#644's ~2–3-item goal).
|
||||
|
||||
7. **`probe-core` is the seam; probes are adapters.** The shared resolution model is extracted into `src/probe-core.cts`; `edge-probe.cts` refactors onto it as the first adapter and the prohibition probe is born on it as the second. Sub-decisions, each chosen against both probes:
|
||||
- **7a — Orthogonal model.** The item carries two orthogonal dimensions: `status: resolved | dismissed | unresolved` (resolution lifecycle, shared) and `verification: <probe-defined> | null` (edge: `explicit | backstop`; prohibition: `test | judgment`). This replaces edge-probe's shipped `status: covered | dismissed | backstop | unresolved`, which smuggled a verification fact (`backstop`) into a lifecycle enum. Migration is mechanical (`covered → {resolved, explicit}`, `backstop → {resolved, backstop}`, others straight across); `coverage.resolved` is preserved **count-for-count** — it remains the *closed* set (`resolved` + `dismissed` status = `applicable − unresolved`), so every fixture's count is unchanged, while `coverage.byVerification` carries the per-tier `resolved`-status breakdown; the 6 edge fixtures are re-generated on #584. Done now (one adapter) rather than against a fixture-locked enum later.
|
||||
- **7b — Core ingests already-proposed items.** The seam is `analyzeCoverage(items, resolutions?, validators)`, **not** `(requirements, proposeFn, …)`. The probes have different deterministic surfaces — edge = deterministic propose (`proposeEdges`) + LLM resolve; prohibition = **LLM propose** (adversarial recall) + deterministic validate/merge/rollup — so core MUST NOT assume propose is deterministic. `proposeEdges` stays in the edge adapter; candidate-parsing stays in the prohibition adapter.
|
||||
- **7c — Hybrid typing.** Generic type params on the exported interfaces for adapter-authoring DX, but the load-bearing enforcement is **injected runtime validators** (`{ categories, verification, requiredFieldsByVerification }`), because the probe runs as a CLI over JSON where TS types are erased. The contract test (Decision 5) pins the validators, not the types.
|
||||
- **7d — Rollup carries a verification tally.** `CoverageReport.coverage` gains `byVerification: { <tier>: count }`, computed generically in core over items that have a verification set. Verify-phase reads `byVerification.judgment` as Decision 4's denominator without re-scanning. (Unresolved items carry no tier and are counted by `status`.)
|
||||
- **7e — One bin per probe.** Each probe ships its own compiled bin (`edge-probe.cjs`, `prohibition-probe.cjs`) calling a shared `runProbeCli(...)`. A single dispatcher CLI is **deferred** as an independent follow-on: unlike the enum, it is pure invocation plumbing with no migration debt, so it does not justify enlarging the #584 blast radius.
|
||||
|
||||
Indicative `probe-core` surface:
|
||||
```ts
|
||||
export type Status = 'resolved' | 'dismissed' | 'unresolved';
|
||||
export interface Item<TCat extends string = string, TVer extends string = string> {
|
||||
requirement_id: string; category: TCat; status: Status;
|
||||
verification: TVer | null; resolution: string | null; reason: string | null;
|
||||
}
|
||||
export interface CoverageReport {
|
||||
items: Item[];
|
||||
coverage: { applicable: number; resolved: number; unresolved: number; byVerification: Record<string, number> };
|
||||
}
|
||||
export interface ProbeValidators {
|
||||
categories: ReadonlySet<string>; verification: ReadonlySet<string>;
|
||||
requiredFieldsByVerification?: Record<string, ReadonlyArray<keyof Item>>;
|
||||
}
|
||||
export function analyzeCoverage(items: Item[], resolutions: Resolution[] | undefined, v: ProbeValidators): CoverageReport;
|
||||
export function validateResolution(r: Resolution, items: Item[], v: ProbeValidators): void;
|
||||
export function validateItems(items: Item[], v: ProbeValidators): void;
|
||||
export function runProbeCli(opts: { ingest: (argv: string[]) => { items: Item[]; resolutions?: Resolution[] }; validators: ProbeValidators }): never;
|
||||
```
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** `truths` keeps its positive-observable semantics; prohibitions get a first-class home with a real verification lifecycle; CI is honest (tests the contract, never fakes the model's judgment as green); no fake-green and no gutted autonomy; the secure-phase boundary is explicit and de-duplicated; the resolution model lives in one place, so the third probe is nearly free and `verification`/`status` cannot drift across the ≥6 `must_haves` read sites.
|
||||
- **Costs / required work (on #584):** extract `probe-core.cts` and refactor `edge-probe.cts` onto it; re-cut the status enum into `status × verification` and re-generate the 6 edge fixtures. **On the #644 PR:** add the prohibition reference + the `must_haves.prohibitions:` block and its `parseMustHavesBlock` extension; the `DEFECT.GENERATIVE-FIX` parity assertion across template ↔ parser ↔ planner (**owned by the #644 author**); extend the SUMMARY/verdict schema with the `unverified-prohibition` flag; teach verify-phase the test-tier wiring and judgment-tier mode-dependent handling. A deterministic `projectProhibitions()` in `probe-core` is recommended so the parity assertion tests a function's round-trip rather than a prompt; every existing `must_haves` read site must tolerate the new block (absence = no prohibitions).
|
||||
- **Docs debt:** `spec-phase` is undocumented in `FEATURES.md`/`COMMANDS.md` despite shipping in v1.38; the docs-required gate bites both probes. PR #584 establishes the spec-phase docs section.
|
||||
- **Glossary:** `probe family`, `probe-core`, `verification tier`, and `bespoke vs canon prohibition` are added to `CONTEXT.md` when #584 merges (not while the contract is pre-merge); this ADR is their interim home.
|
||||
- **Sequencing:** PR #584 lands ADR 550 + `probe-core` + the `edge-probe.cts` refactor + enum re-cut + fixture re-gen + the spec-phase docs section. #644 then adds the prohibition adapter (reference + `prohibitions:` block/parser/parity + tiered verification). #644 stays blocked until #584 merges.
|
||||
107
docs/how-to/resolve-edge-coverage-findings.md
Normal file
107
docs/how-to/resolve-edge-coverage-findings.md
Normal file
@@ -0,0 +1,107 @@
|
||||
# How to resolve edge-coverage findings while writing a spec
|
||||
|
||||
**Goal:** Turn each domain-boundary edge the spec phase surfaces into an explicit, verifiable spec decision — so omitted boundaries (rounding ties, touching ranges, grapheme truncation) become checkable requirements *before* any code exists, instead of silent blind spots the verifier is confidently wrong about.
|
||||
|
||||
**Prerequisites:** A phase whose `/gsd-spec-phase` run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically and presents its findings — you do not invoke it separately.
|
||||
|
||||
For the category taxonomy and the reasoning behind front-of-pipeline edge analysis, see [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe). This guide covers only how to *act* on the findings.
|
||||
|
||||
---
|
||||
|
||||
## Read a finding
|
||||
|
||||
Each finding is one **applicable** edge for one requirement — a boundary the probe's relevance filter decided is in scope for that requirement's shape, and that you have not yet addressed. A finding is phrased as a probe question, for example:
|
||||
|
||||
> **R3 · precision** — Where can precision loss or overflow occur, and what is the contract?
|
||||
|
||||
You must resolve each finding into exactly one of four states. Claude presents them as a numbered choice (or an `AskUserQuestion` menu).
|
||||
|
||||
---
|
||||
|
||||
## Specify it — write an acceptance criterion
|
||||
|
||||
**Choose this when you can state the correct behaviour as a pass/fail check.** This is the strongest resolution: it produces a concrete assertion the planner and verifier can enforce.
|
||||
|
||||
Claude writes a new line into the spec's **Acceptance Criteria** and marks the edge `covered`. For the `precision` finding above:
|
||||
|
||||
> - [ ] Monetary amounts round half-to-even to 2 decimal places; `2.005 → 2.00`
|
||||
|
||||
Prefer this state whenever a defensible criterion can be written. A `covered` edge is lifted into the plan's `must_haves.truths`, so the verifier checks it.
|
||||
|
||||
---
|
||||
|
||||
## Dismiss it — record why it does not apply
|
||||
|
||||
**Choose this when the edge genuinely cannot occur** — and say why. A dismissal **requires a non-empty reason**; silence is rejected. The reason is the audit trail.
|
||||
|
||||
> ⛔ dismissed — input is a bounded enum; no boundary value exists
|
||||
|
||||
A wrong dismissal is the exact silent failure this probe exists to prevent, so dismiss only when the reason is solid.
|
||||
|
||||
---
|
||||
|
||||
## Backstop it — defer the contract to a held-out test
|
||||
|
||||
**Choose this when you know the edge matters but cannot fully articulate the correct behaviour in prose yet.** Claude marks the edge `backstop` and notes that a held-out / property-based test stands in for the missing assertion.
|
||||
|
||||
> 🧪 backstop — held-out property test: dedupe output order is stable under input permutation
|
||||
|
||||
A `backstop` edge is also carried into the plan's `must_haves.truths` as a non-inferable check — the planner records the intent, and the test body is authored later during execution.
|
||||
|
||||
---
|
||||
|
||||
## Defer it — leave it unresolved and flagged
|
||||
|
||||
**Choose this only when you are not ready to decide.** The edge stays `unresolved` and is flagged. Unlike a dismissal, deferring makes no claim that the edge is safe — it is an explicit, visible assumption the planner must surface, not silently drop.
|
||||
|
||||
---
|
||||
|
||||
## Clear the soft gate
|
||||
|
||||
After you have worked through the findings, the probe runs a **soft gate**:
|
||||
|
||||
- **All applicable edges resolved** → the spec proceeds to the next step.
|
||||
- **One or more still `unresolved`** → Claude asks what to do:
|
||||
- **Resolve now** — loop back and resolve the remaining edges.
|
||||
- **Write the spec anyway** — the spec is written with those rows marked `⚠ Edge unresolved — planner must treat as assumption`. Use this deliberately; you are choosing to ship a known gap.
|
||||
|
||||
The gate is *soft*: it never blocks you, but every unresolved edge remains visible in the spec's `## Edge Coverage` section.
|
||||
|
||||
---
|
||||
|
||||
## Let Claude resolve them for you
|
||||
|
||||
**If the edges are low-stakes or already implied by earlier phases**, run the spec phase in auto mode:
|
||||
|
||||
```bash
|
||||
/gsd-spec-phase 3 --auto
|
||||
```
|
||||
|
||||
In `--auto`, Claude marks an edge `covered` where it can write a defensible acceptance criterion, and `backstop` otherwise. It **never auto-dismisses** — dismissing an edge requires a human reason, because a wrong auto-dismissal is precisely the silent failure being eliminated. Claude logs the tally, for example:
|
||||
|
||||
```
|
||||
[auto] edge coverage: 4 covered, 2 backstop, 1 unresolved
|
||||
```
|
||||
|
||||
Review the logged choices afterwards; auto mode is a fast first pass, not a substitute for judgement on edges that carry risk.
|
||||
|
||||
---
|
||||
|
||||
## What happens to resolved findings downstream
|
||||
|
||||
When you next run `/gsd-plan-phase`, the planner reads the spec's `## Edge Coverage` section and:
|
||||
|
||||
- lifts every `covered` edge's acceptance criterion into `must_haves.truths`,
|
||||
- carries every `backstop` edge into `must_haves.truths` as a non-inferable check (needing a held-out/property test),
|
||||
- surfaces every `unresolved` edge as an explicit assumption.
|
||||
|
||||
This is the payoff: a resolved edge becomes a unit the goal-backward verifier actually checks, extending its reach to boundaries the requirement prose never stated.
|
||||
|
||||
---
|
||||
|
||||
## Related
|
||||
|
||||
- [Spec-Phase Edge-Completeness Probe](../FEATURES.md#143-spec-phase-edge-completeness-probe) — taxonomy, output schema, and the front-of-pipeline rationale
|
||||
- [`/gsd-spec-phase`](../COMMANDS.md#gsd-spec-phase) — command reference and flags
|
||||
- [Plan a phase](plan-a-phase.md) — where `covered`/`backstop` edges become `must_haves`
|
||||
- [docs index](../README.md)
|
||||
@@ -36,6 +36,8 @@ export default tseslint.config(
|
||||
// ADR-457: tsc-generated runtime artifact — lint the src/*.cts source, not the emitted .cjs.
|
||||
'gsd-core/bin/lib/semver-compare.cjs',
|
||||
'gsd-core/bin/lib/cli-exit.cjs',
|
||||
'gsd-core/bin/lib/edge-probe.cjs',
|
||||
'gsd-core/bin/lib/probe-core.cjs',
|
||||
'gsd-core/bin/lib/code-review-flags.cjs',
|
||||
'gsd-core/bin/lib/context-utilization.cjs',
|
||||
'gsd-core/bin/lib/artifacts.cjs',
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
|
||||
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Round a number to N decimal places, rounding half to nearest" }]
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]
|
||||
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Truncate a display string to its first N characters" }]
|
||||
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
|
||||
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Compute the total price as an amount rounded to two decimals" }]
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Return all items in the list with duplicates removed" }]
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } }
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]
|
||||
@@ -0,0 +1,4 @@
|
||||
[
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "resolution": null, "reason": "output is canonically sorted; no tie possible" }
|
||||
]
|
||||
261
gsd-core/references/edge-probe.md
Normal file
261
gsd-core/references/edge-probe.md
Normal file
@@ -0,0 +1,261 @@
|
||||
# Edge-Probe — Spec-Completeness Reference
|
||||
|
||||
Shared reference for the spec/requirements phase. Companion to
|
||||
`@~/.claude/gsd-core/references/domain-probes.md`: `domain-probes` covers the
|
||||
**technology axis** (auth, search, caching, deployment); this covers the
|
||||
**data/behavior-shape axis** (boundaries, adjacency, encoding, ordering…). Walk each
|
||||
requirement against the closed taxonomy below, propose a concrete candidate edge for each
|
||||
applicable category, and resolve each to exactly one state. Adopt the established QA names
|
||||
verbatim — this is decades-old black-box test technique, moved upstream to the spec layer.
|
||||
|
||||
This doc is written in generic `requirements → checks → verifier` terms with no
|
||||
tool-specific vocabulary, so it is portable: copy it into any spec/requirements process.
|
||||
A short mapping table at the end binds it to common host structures.
|
||||
|
||||
## Why front-of-pipeline
|
||||
|
||||
A goal-backward verifier only checks assertions that exist; an assertion only exists for a
|
||||
requirement that was written down. A domain-boundary edge the author never surfaced is
|
||||
invisible to the verifier — and worse, the verifier is *confidently wrong* about it
|
||||
(measured: ~0.93 confidence while catching 0/12 omitted-edge defects, an expected
|
||||
calibration error of 0.81 — worse than a coin flip — versus 100% / 0.03 on edges the spec
|
||||
did state). The fix is not a better verifier or a confidence gate; confidence is
|
||||
uninformative across these regimes. The fix is **spec completeness**: surface the omitted
|
||||
edge into an explicit, checkable assertion *before* any code exists, after which the
|
||||
verifier reliably catches it.
|
||||
|
||||
The underlying techniques are classic and should be named as such — **Boundary Value
|
||||
Analysis**, **Equivalence Partitioning**, the **Category-Partition method**, and
|
||||
**Metamorphic Relations**; property-based testing (PBT) is the academic name for the
|
||||
held-out backstop. The literature applies these at the *test* layer (back of pipeline,
|
||||
generating checks). The differentiated move here is **placement, not technique**: apply the
|
||||
same edge taxonomy at the *spec* layer (front of pipeline, generating requirements), with
|
||||
the explicit goal of extending a verifier's reach. Recent LLM spec work corroborates the
|
||||
core finding that the spec layer is the measured weak point:
|
||||
|
||||
- SLD-Spec — program slicing + logical deletion (code→spec): arXiv 2509.09917
|
||||
- Specine / AutoReSpec / SpecMind — spec alignment & postcondition inference
|
||||
- CodeSpecBench (2604.12268), OSVBench (2504.20964), VERINA (2505.23135) — benchmarks
|
||||
showing models solve tasks far better than they generate precise behavioral specs
|
||||
- LLM property-based tests for edge cases: arXiv 2510.25297 (PBT+EBT ≈ 81% edge detection)
|
||||
|
||||
## Inputs
|
||||
|
||||
A list of requirements, each a `{ id, text, shapes? }` record where `text` is a testable
|
||||
statement and `shapes` is an optional author-supplied override of the data/behavior shape.
|
||||
The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. When
|
||||
`shapes` is absent, a heuristic classifier proposes them from the requirement prose
|
||||
(propose-then-confirm) — the author may correct the shape.
|
||||
|
||||
## Taxonomy (8 categories)
|
||||
|
||||
Closed and small by design: a fixed eight the author must explicitly clear beats thirty
|
||||
nobody finishes. The failure mode being eliminated was never "too few categories" — it was
|
||||
that no taxonomy was *systematically applied at all*. Categories 1, 2, and 4 alone cover
|
||||
the three canonical corpus defects (banker's-rounding ties, touching intervals, grapheme
|
||||
truncation). Growth happens via optional domain packs, not by bloating the core.
|
||||
|
||||
| id | name (QA term) | applies to shapes | probe question |
|
||||
|----|----------------|-------------------|----------------|
|
||||
| boundary | Boundary values | numeric-range | What happens exactly at each min/max/threshold — and one step either side? |
|
||||
| adjacency | Adjacency / touching | collection | When two things are exactly equal or just touch, do they merge, collide, or separate? |
|
||||
| empty | Empty / degenerate | collection, text | What is the result for empty, single-element, or null input? |
|
||||
| encoding | Encoding / representation | text | Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form? |
|
||||
| ordering | Ordering / stability | collection | When elements compare equal, is output order specified and stable? |
|
||||
| precision | Precision / overflow | numeric-range | Where can precision loss or overflow occur, and what is the contract? |
|
||||
| idempotency | Idempotency / repetition | stateful | What happens if this runs twice on the same input? |
|
||||
| concurrency | Concurrency / effect ordering | stateful, io | If interrupted or run in parallel, what is guaranteed? |
|
||||
|
||||
## Relevance filter + resolution states
|
||||
|
||||
Two rules keep the probe honest and prevent an "everything is N/A" failure mode:
|
||||
|
||||
1. **Relevance filter first.** Classify each requirement's shape, then raise only the
|
||||
categories whose `applies to shapes` intersect that requirement's shapes. A pure-text
|
||||
requirement is never asked about overflow. This is what makes an unresolved edge
|
||||
meaningful: it is an edge that *applies* and was not addressed.
|
||||
2. **Dismissal requires a reason string.** "N/A — input is a bounded enum, no boundary
|
||||
exists" is valid; silence is not. The reason string is the audit trail.
|
||||
|
||||
Each raised edge carries two orthogonal axes — a resolution **lifecycle** and, when
|
||||
resolved, a **verification** tier (ADR-550 Decision 7, the shared probe-core model):
|
||||
|
||||
- **status** — `resolved | dismissed | unresolved`:
|
||||
- **resolved** — the edge is addressed; *how* it is addressed is the verification tier.
|
||||
- **dismissed** — not applicable, accompanied by a required, non-empty reason string.
|
||||
- **unresolved** — carried forward and flagged; the author chose not to resolve it yet.
|
||||
- **verification** (only when `status` is `resolved`; `null` otherwise) — `explicit | backstop`:
|
||||
- **explicit** — a checkable assertion for the edge is written (a SPEC acceptance criterion).
|
||||
- **backstop** — a held-out / property-based test stands in for an edge the author knows
|
||||
but cannot fully articulate in prose (records intent; the test body is authored later).
|
||||
|
||||
Splitting these axes keeps the lifecycle enum free of a verification fact (the old single
|
||||
enum smuggled `covered`/`backstop` — both *resolved* — into one flat list) and lets sibling
|
||||
probes add their own verification tiers (e.g. `test | judgment`) without a parallel enum.
|
||||
|
||||
`coverage.resolved` is the count of **closed** edges — `resolved` + `dismissed` (the
|
||||
pre-re-cut "covered + dismissed + backstop" set, count-preserved). An `unresolved`
|
||||
*applicable* edge is the precise signal a soft completeness gate raises.
|
||||
|
||||
## Output schema
|
||||
|
||||
The probe emits, per edge, an item of the form:
|
||||
|
||||
```
|
||||
{ requirement_id, category, status, verification, resolution, reason, probe }
|
||||
```
|
||||
|
||||
plus a coverage summary:
|
||||
|
||||
```
|
||||
coverage: { applicable, resolved, unresolved, byVerification: { explicit, backstop } }
|
||||
```
|
||||
|
||||
`applicable` is the number of raised edges, `resolved` = closed (`resolved` + `dismissed`)
|
||||
status edges, `unresolved` is the remainder, and `byVerification` breaks the
|
||||
`resolved`-status edges down by tier (probe-agnostic in core; the edge adapter declares
|
||||
`{ explicit, backstop }`). This JSON is the stable contract both the reference
|
||||
implementation and any third-party port emit.
|
||||
|
||||
## Generic mapping (requirements → checks → verifier)
|
||||
|
||||
| Host structure | "requirement" | a `resolved`/`explicit` edge becomes | a `resolved`/`backstop` edge becomes |
|
||||
|----------------|---------------|--------------------------|---------------------------|
|
||||
| GSD SPEC | a SPEC Requirement | an Acceptance Criterion that `plan-phase` lifts into `must_haves.truths` | a non-inferable check in `must_haves.truths` (needs a held-out/PBT test) |
|
||||
| Gherkin feature | a Scenario | an additional `Then` assertion / Scenario Outline row | a tagged scenario routed to a property test |
|
||||
| OpenAPI operation | an operation | a response/constraint example + schema rule | a contract/property test on the operation |
|
||||
| Docstring contract | a documented behavior | an assertion in the contract test | a property-based test for the function |
|
||||
|
||||
The portable invariant: a `resolved`/`explicit` edge produces **the unit your verifier
|
||||
iterates over** (GSD: a `must_haves.truth`); a `resolved`/`backstop` edge produces a test
|
||||
added to that same set as a non-inferable check. An `unresolved` edge is an explicit
|
||||
assumption the downstream planner must surface, not silently drop.
|
||||
|
||||
## Worked example (merge-intervals)
|
||||
|
||||
Given a single requirement with no resolutions yet:
|
||||
|
||||
```
|
||||
[{ "id": "R1", "text": "Merge a list of overlapping intervals into the minimal set" }]
|
||||
```
|
||||
|
||||
the requirement classifies as a `collection`, which raises `adjacency`, `empty`, and
|
||||
`ordering` (but not `boundary`/`precision`/`encoding`/`idempotency`/`concurrency`). With no
|
||||
resolutions supplied, every applicable edge is `unresolved`:
|
||||
|
||||
```json edge-probe:02-merge-intervals/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
The `adjacency` row is the one that catches the canonical defect: `[[1,2],[2,3]]` intervals
|
||||
that only *touch* — does the spec say they merge? Resolving it `resolved`/`explicit` writes
|
||||
that assertion, and the verifier can then enforce it. This worked-example block is kept
|
||||
byte-for-byte (parsed-JSON) identical to its fixture by `edge-probe-docs-fixtures.test.cjs`,
|
||||
so the doc and the reference implementation cannot silently drift.
|
||||
|
||||
## Worked example (round-half-even)
|
||||
|
||||
Given a requirement to round floating-point values using the banker's-rounding (half-even)
|
||||
rule, the requirement classifies as `numeric-range`, which raises `boundary` and `precision`
|
||||
(but not `adjacency`/`empty`/`ordering`/`encoding`/`idempotency`/`concurrency`):
|
||||
|
||||
```json edge-probe:01-round-half-even/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
|
||||
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
The `precision` row surfaces the canonical defect for this requirement: IEEE 754 floating-point
|
||||
arithmetic rounds 2.5 to 2 (not 3) under half-even — the probe asks whether the spec
|
||||
states that contract explicitly, so the verifier can enforce it.
|
||||
|
||||
## Worked example (truncate-graphemes)
|
||||
|
||||
Given a requirement to truncate a string to N grapheme clusters (not bytes or code points),
|
||||
the requirement classifies as `text`, which raises `empty` and `encoding`:
|
||||
|
||||
```json edge-probe:03-truncate-graphemes/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "encoding", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
The `encoding` row is the load-bearing one: a requirement that says "truncate to 10
|
||||
characters" is ambiguous — the spec must state whether "character" means bytes, UTF-16
|
||||
code units, Unicode code points, or grapheme clusters (emoji sequences are 1 grapheme,
|
||||
multiple code points).
|
||||
|
||||
## Worked example (money-rounding)
|
||||
|
||||
Given a requirement to round monetary amounts to two decimal places, the requirement
|
||||
classifies as `numeric-range`, which raises `boundary` and `precision`:
|
||||
|
||||
```json edge-probe:04-money-rounding/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "boundary", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What happens exactly at each min/max/threshold — and one step either side?" },
|
||||
{ "requirement_id": "R1", "category": "precision", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "Where can precision loss or overflow occur, and what is the contract?" }
|
||||
],
|
||||
"coverage": { "applicable": 2, "resolved": 0, "unresolved": 2, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
The `boundary` row asks what the minimum and maximum representable values are (negative
|
||||
amounts? fractional cents?). The `precision` row asks whether IEEE 754 binary rounding
|
||||
can produce `$0.30000000000000004` — the spec must commit to a representation.
|
||||
|
||||
## Worked example (list-dedupe)
|
||||
|
||||
Given a requirement to deduplicate a list of items, the requirement classifies as
|
||||
`collection`, which raises `adjacency`, `empty`, and `ordering`:
|
||||
|
||||
```json edge-probe:05-list-dedupe/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 0, "unresolved": 3, "byVerification": { "explicit": 0, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
The `ordering` row is the non-obvious one: dedupe removes duplicates, but which copy is
|
||||
kept — first occurrence, last, or implementation-defined? The spec must commit.
|
||||
|
||||
## Worked example (resolved-mixed)
|
||||
|
||||
The same merge-intervals requirement after a resolution session where `adjacency` was
|
||||
resolved with an explicit acceptance criterion, `ordering` was dismissed, and `empty` was
|
||||
left unresolved:
|
||||
|
||||
```json edge-probe:06-resolved-mixed/expected-coverage.json
|
||||
{
|
||||
"items": [
|
||||
{ "requirement_id": "R1", "category": "adjacency", "status": "resolved", "verification": "explicit", "resolution": "AC#6: touching intervals merge", "reason": null, "probe": "When two things are exactly equal or just touch, do they merge, collide, or separate?" },
|
||||
{ "requirement_id": "R1", "category": "empty", "status": "unresolved", "verification": null, "resolution": null, "reason": null, "probe": "What is the result for empty, single-element, or null input?" },
|
||||
{ "requirement_id": "R1", "category": "ordering", "status": "dismissed", "verification": null, "resolution": null, "reason": "output is canonically sorted; no tie possible", "probe": "When elements compare equal, is output order specified and stable?" }
|
||||
],
|
||||
"coverage": { "applicable": 3, "resolved": 2, "unresolved": 1, "byVerification": { "explicit": 1, "backstop": 0 } }
|
||||
}
|
||||
```
|
||||
|
||||
`coverage.resolved` is 2 — the closed set (adjacency=`resolved`/`explicit` + ordering=`dismissed`);
|
||||
`unresolved` is 1 (empty); `byVerification.explicit` is 1 (the single explicitly-verified
|
||||
edge), `backstop` 0. The soft gate raises on this example because one applicable edge remains
|
||||
unresolved — the author must either specify, dismiss, or backstop it before writing the SPEC.
|
||||
@@ -68,6 +68,18 @@ If none: "No additional constraints beyond standard project conventions."]
|
||||
[Every acceptance criterion must be a checkbox that resolves to PASS or FAIL.
|
||||
No "should feel good", "looks reasonable", or "generally works" — those are not checkboxes.]
|
||||
|
||||
## Edge Coverage
|
||||
|
||||
**Coverage:** [resolved]/[applicable] applicable edges resolved · [unresolved] unresolved
|
||||
|
||||
| Category | Requirement | Status | Resolution / Reason |
|
||||
|----------|-------------|--------|---------------------|
|
||||
| [category] | [Rn] | [✅ covered / ⛔ dismissed / 🧪 backstop / ⚠ UNRESOLVED] | [acceptance criterion ref, dismissal reason, or backstop test note] |
|
||||
|
||||
[Generated by the edge-completeness probe (Step 5.5). `covered` rows correspond to
|
||||
Acceptance Criteria above; `backstop` rows must be carried into plan-phase `must_haves`.
|
||||
`⚠ UNRESOLVED` rows are flagged: planner must treat as assumption.]
|
||||
|
||||
## Ambiguity Report
|
||||
|
||||
| Dimension | Score | Min | Status | Notes |
|
||||
|
||||
@@ -810,6 +810,14 @@ PATTERNS_PATH=$(_gsd_field "$INIT" patterns_path)
|
||||
# Detect spike/sketch findings skills (project-local)
|
||||
SPIKE_FINDINGS_PATH=$(ls ./.claude/skills/spike-findings-*/SKILL.md 2>/dev/null | head -1 || true)
|
||||
SKETCH_FINDINGS_PATH=$(ls ./.claude/skills/sketch-findings-*/SKILL.md 2>/dev/null | head -1 || true)
|
||||
|
||||
# Resolve the phase SPEC (carries the ## Edge Coverage section the planner lifts covered/
|
||||
# backstop edges from). UNCONDITIONAL — must NOT live in §4.5 Check AI-SPEC, which is skipped
|
||||
# on non-AI phases; gating it there silently starves the planner of the SPEC (#550 review).
|
||||
# Glob the plain phase SPEC, excluding the -AI-SPEC.md / -UI-SPEC.md variants.
|
||||
PHASE_DIR_FOR_SPEC=$(_gsd_field "$INIT" phase_dir)
|
||||
SPEC_FILE=$(ls "${PHASE_DIR_FOR_SPEC}"/*-SPEC.md 2>/dev/null | grep -Ev -- '-(AI|UI)-SPEC\.md$' | head -1)
|
||||
SPEC_PATH="${SPEC_FILE}"
|
||||
```
|
||||
|
||||
## 7.5. Verify Nyquist Artifacts
|
||||
@@ -937,6 +945,7 @@ Planner prompt:
|
||||
- {uat_path} (UAT Gaps - if --gaps)
|
||||
- {reviews_path} (Cross-AI Review Feedback - if --reviews; actionable findings must be incorporated or explicitly deferred/rejected in PLAN.md)
|
||||
- {UI_SPEC_PATH} (UI Design Contract — visual/interaction specs, if exists)
|
||||
- {SPEC_PATH} (Phase SPEC — carries the ## Edge Coverage section to lift covered/backstop edges from, if exists)
|
||||
- {SPIKE_FINDINGS_PATH} (Spike Findings — validated patterns, constraints, landmines from experiments, if exists)
|
||||
- {SKETCH_FINDINGS_PATH} (Sketch Findings — validated design decisions, CSS patterns, visual direction, if exists)
|
||||
- {API_SURFACE_PATH} (API Surface — HINT ONLY, if intel.enabled; see <intel_surface_hint> below)
|
||||
@@ -999,6 +1008,7 @@ Output consumed by /gsd:execute-phase. Plans need:
|
||||
- Tasks in XML format with read_first and acceptance_criteria fields (MANDATORY on every task)
|
||||
- Verification criteria
|
||||
- must_haves for goal-backward verification
|
||||
- If the SPEC has an `## Edge Coverage` section, lift every `covered` edge's acceptance criterion into `must_haves.truths`, and every `backstop` edge into `must_haves.truths` as a non-inferable check (note it needs a held-out/property-based test). `unresolved` edges are explicit assumptions — surface them in the plan, do not silently drop them.
|
||||
- **"Artifacts this phase produces" section (MANDATORY)** — list every symbol this phase creates: decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths. The plan-review-convergence source-grounding pass reads this section to exclude newly-created symbols from drift verification; omitting it causes new symbols to be flagged for acknowledgement.
|
||||
</downstream_consumer>
|
||||
|
||||
@@ -1044,6 +1054,7 @@ Every task MUST include these fields — they are NOT optional:
|
||||
- [ ] Waves assigned for parallel execution
|
||||
- [ ] must_haves derived from phase goal
|
||||
- [ ] Every PLAN.md includes an "Artifacts this phase produces" section listing symbols created by this phase (decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths)
|
||||
- [ ] Every SPEC ## Edge Coverage covered/backstop edge is represented in a plan's must_haves (no silent drops)
|
||||
</quality_gate>
|
||||
```
|
||||
|
||||
|
||||
@@ -187,10 +187,137 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met):
|
||||
|
||||
## Step 5: (covered inline — ambiguity scoring is per-round)
|
||||
|
||||
## Step 5.5: Edge-Completeness Probe
|
||||
|
||||
Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague
|
||||
ones). Reference: @~/.claude/gsd-core/references/edge-probe.md.
|
||||
|
||||
**Runtime coverage compute — resolve and invoke edge-probe.cjs:**
|
||||
|
||||
```bash
|
||||
# Resolve the compiled edge-probe.cjs against the GSD install dir via RUNTIME_DIR (#448)
|
||||
# — NOT the consuming project's git root — falling back to git toplevel / $HOME/.claude.
|
||||
# Mirrors the ui-safety-gate.cjs resolution idiom at autonomous.md:290 / plan-phase.md:631.
|
||||
_GSD_RT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"
|
||||
EDGE_PROBE_JS=$(for _c in \
|
||||
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
|
||||
"$_GSD_RT/bin/lib/edge-probe.cjs" \
|
||||
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
|
||||
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
|
||||
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
|
||||
[ -f "$_c" ] && { echo "$_c"; break; }
|
||||
done)
|
||||
|
||||
# Graceful degradation — never silent skip (RR-04). Build ONLY when $_GSD_RT is a verified
|
||||
# GSD source checkout (has tsconfig.build.json + src/edge-probe.cts), and pin npm to it with
|
||||
# --prefix so we never trigger the CONSUMING project's own build:lib (its cwd package scripts:
|
||||
# codegen/migrations/writes) during a spec workflow. Real installs ship the compiled .cjs via
|
||||
# prepublishOnly, so this build path only matters in a GSD dev checkout (review High).
|
||||
if [ -z "$EDGE_PROBE_JS" ]; then
|
||||
if [ -f "$_GSD_RT/tsconfig.build.json" ] && [ -f "$_GSD_RT/src/edge-probe.cts" ]; then
|
||||
npm --prefix "$_GSD_RT" run build:lib 2>/dev/null || true
|
||||
EDGE_PROBE_JS=$(for _c in \
|
||||
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
|
||||
"$_GSD_RT/bin/lib/edge-probe.cjs" \
|
||||
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
|
||||
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
|
||||
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
|
||||
[ -f "$_c" ] && { echo "$_c"; break; }
|
||||
done)
|
||||
fi
|
||||
if [ -z "$EDGE_PROBE_JS" ]; then
|
||||
echo "ERROR: edge-probe.cjs not found — reinstall GSD or run \`npm run build:lib\` in your GSD checkout." >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the
|
||||
# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object
|
||||
# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes
|
||||
# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04).
|
||||
REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX.json")
|
||||
cat > "$REQS_JSON" <<'JSON'
|
||||
[
|
||||
{ "id": "R1", "text": "<replace: requirement text from the SPEC>" }
|
||||
]
|
||||
JSON
|
||||
# Guard — never invoke on an empty/invalid array, OR one still holding the heredoc
|
||||
# `<replace: …>` placeholder (a forgotten substitution would otherwise yield a
|
||||
# meaningful-looking but bogus coverage report). Fail loud, not silent no-op.
|
||||
if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("<replace:")))process.exit(1)' "$REQS_JSON" 2>/dev/null; then
|
||||
echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the <replace: …> placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2
|
||||
exit 1
|
||||
fi
|
||||
# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per
|
||||
# requirement. The covered/backstop/dismissed/unresolved rows in $COVERAGE drive the
|
||||
# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose).
|
||||
# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture
|
||||
# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE
|
||||
# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary
|
||||
# the engine validation exists to protect. Make the run fatal, then validate the captured
|
||||
# report is well-formed JSON before the resolution loop consumes it.
|
||||
if ! COVERAGE=$(node "$EDGE_PROBE_JS" "$REQS_JSON"); then
|
||||
rm -f "$REQS_JSON"
|
||||
echo "ERROR: edge-probe engine failed (invalid shapes or bad input) — fix the requirement(s) and re-run; never proceed with empty coverage." >&2
|
||||
exit 1
|
||||
fi
|
||||
rm -f "$REQS_JSON"
|
||||
# Exit-0-but-garbage guard: the report must parse as JSON with the expected { items[], coverage{} } shape.
|
||||
if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let r;try{r=JSON.parse(s)}catch{process.exit(1)}if(!r||!Array.isArray(r.items)||typeof r.coverage!=="object"||r.coverage===null)process.exit(1)})'; then
|
||||
echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2
|
||||
exit 1
|
||||
fi
|
||||
# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY
|
||||
# requirement is far more likely a shape-classification miss (or malformed requirements) than
|
||||
# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding
|
||||
# applicable:0. Surface it loudly; the author must explicitly confirm "no applicable edges"
|
||||
# below rather than silently emitting a green empty ## Edge Coverage section.
|
||||
APPLICABLE=$(printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let n=0;try{n=JSON.parse(s).coverage.applicable}catch{n=0}process.stdout.write(String(n))})')
|
||||
if [ "$APPLICABLE" = "0" ]; then
|
||||
echo "WARNING: edge-probe proposed ZERO applicable edges across all requirements — likely a classification miss or malformed requirements, not a genuinely edge-free spec. Do NOT silently write an empty Edge Coverage section." >&2
|
||||
fi
|
||||
```
|
||||
|
||||
If `$APPLICABLE` is `0`, do NOT proceed silently: ask the author to confirm via AskUserQuestion
|
||||
("The edge probe found no applicable edges for any requirement — is this genuinely an
|
||||
edge-free spec, or should we revisit the requirement wording / authored shapes?"). Only write
|
||||
an empty `## Edge Coverage` section after explicit confirmation.
|
||||
|
||||
For each Requirement gathered so far:
|
||||
1. Classify its shape and raise only applicable edge categories (relevance filter — see
|
||||
the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already
|
||||
surfaced as pre-`covered`.
|
||||
2. For each raised category, propose a CONCRETE candidate edge (not "consider
|
||||
boundaries" — e.g. "R2 merges intervals; what about `[[1,2],[2,3]]` that only touch?").
|
||||
3. Resolve each with the user (AskUserQuestion; text mode → numbered list):
|
||||
- **Specify it** → write a new pass/fail line into Acceptance Criteria AND mark the
|
||||
edge `covered`.
|
||||
- **Dismiss (reason)** → mark `dismissed` with a required non-empty reason.
|
||||
- **Backstop with a test** → mark `backstop`; note "held-out edge test" for plan-phase.
|
||||
- **Defer** → leave `unresolved`.
|
||||
|
||||
**Soft gate (after resolving):**
|
||||
- All applicable edges resolved → proceed to Step 6.
|
||||
- Any `unresolved` → AskUserQuestion:
|
||||
- header: "Edge Coverage"
|
||||
- question: "[N] edge(s) are unresolved: [list]. What do you want to do?"
|
||||
- options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" /
|
||||
"Keep probing"
|
||||
- On "anyway": write SPEC.md with those rows marked `⚠ Edge unresolved — planner must
|
||||
treat as assumption`.
|
||||
|
||||
**`--auto` mode:** auto-`covered` where a defensible acceptance criterion can be written;
|
||||
otherwise auto-`backstop` (never auto-dismiss — a wrong dismissal is the exact silent
|
||||
failure being eliminated). Log: `[auto] edge coverage: C covered, B backstop, U unresolved`.
|
||||
|
||||
Populate the `## Edge Coverage` section of SPEC.md from the resolved edges.
|
||||
|
||||
## Step 6: Generate SPEC.md
|
||||
|
||||
Use the SPEC.md template from @~/.claude/gsd-core/templates/spec.md.
|
||||
|
||||
- Populate the **Edge Coverage** section from Step 5.5 (covered/dismissed/backstop/unresolved rows).
|
||||
|
||||
**Requirements for every requirement entry:**
|
||||
- One specific, testable statement
|
||||
- Current state (what exists now)
|
||||
@@ -249,6 +376,7 @@ Next: /gsd:discuss-phase {X}
|
||||
- Do NOT ask about HOW to implement — that is discuss-phase territory
|
||||
- Scout the codebase BEFORE the first question — grounded questions only
|
||||
- Max 2–3 questions per round — do not frontload all questions at once
|
||||
- Step 5.5 edge probe runs after the ambiguity gate; dismissals require a reason; --auto never auto-dismisses
|
||||
</critical_rules>
|
||||
|
||||
<success_criteria>
|
||||
@@ -260,4 +388,5 @@ Next: /gsd:discuss-phase {X}
|
||||
- Acceptance criteria are pass/fail checkboxes
|
||||
- SPEC.md committed atomically (when commit_docs is true)
|
||||
- User directed to /gsd:discuss-phase as next step
|
||||
- Edge-completeness probe run; Edge Coverage section populated; unresolved edges flagged as assumptions
|
||||
</success_criteria>
|
||||
|
||||
@@ -151,6 +151,15 @@
|
||||
"validate-context.test.cjs"
|
||||
],
|
||||
"issue": "892"
|
||||
},
|
||||
"edge-probe": {
|
||||
"files": [
|
||||
"edge-probe-docs-fixtures.test.cjs",
|
||||
"edge-probe-planner-contract.test.cjs",
|
||||
"edge-probe-spec-phase-contract.test.cjs",
|
||||
"edge-probe.test.cjs"
|
||||
],
|
||||
"issue": "550"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
217
src/edge-probe.cts
Normal file
217
src/edge-probe.cts
Normal file
@@ -0,0 +1,217 @@
|
||||
/**
|
||||
* Spec-completeness edge-probe — the FIRST adapter of the probe-core resolution model
|
||||
* (ADR-457 build model; ADR-550 Decision 7 seam).
|
||||
*
|
||||
* The generic resolution lifecycle, the status×verification re-cut, `validateResolution`,
|
||||
* `validateRequirement`, the `analyzeCoverage` merge/rollup/orphan-reject engine, and the
|
||||
* `runProbeCli` scaffold all live in `src/probe-core.cts`. This module keeps ONLY the
|
||||
* edge-specific cluster: the five data/behavior shapes, the closed 8-category edge taxonomy,
|
||||
* shape classification, edge proposal, and the `{ explicit, backstop }` verification validators.
|
||||
*
|
||||
* Authored as strict TypeScript (`src/edge-probe.cts`) and compiled by
|
||||
* `tsc -p tsconfig.build.json` to the gitignored runtime artifact
|
||||
* `gsd-core/bin/lib/edge-probe.cjs`. Do NOT hand-write the `.cjs`; it is emitted. Tests
|
||||
* `require()` the built artifact; `pretest` runs `build:lib` first.
|
||||
*
|
||||
* Pure and dependency-free: it classifies each requirement's data/behavior shape, filters
|
||||
* the closed 8-category edge taxonomy to applicable categories, proposes concrete candidate
|
||||
* edges, and (via probe-core) merges author resolutions into a coverage report.
|
||||
*/
|
||||
|
||||
import {
|
||||
type Item,
|
||||
type Resolution,
|
||||
type CoverageReport,
|
||||
type Validators,
|
||||
validateRequirement as coreValidateRequirement,
|
||||
validateResolution as coreValidateResolution,
|
||||
analyzeCoverage as coreAnalyzeCoverage,
|
||||
runProbeCli,
|
||||
} from './probe-core.cjs';
|
||||
|
||||
/** The five data/behavior shapes a requirement can exhibit. */
|
||||
export type Shape = 'numeric-range' | 'collection' | 'text' | 'stateful' | 'io';
|
||||
|
||||
/** The edge probe's verification tiers (the `verification` axis values for a resolved edge). */
|
||||
export type EdgeVerification = 'explicit' | 'backstop';
|
||||
|
||||
/** A single edge taxonomy category. */
|
||||
export interface TaxonomyEntry {
|
||||
id: string;
|
||||
name: string;
|
||||
shapes: Shape[];
|
||||
probe: string;
|
||||
}
|
||||
|
||||
/** A SPEC requirement; `shapes` is an optional authored override of classification. */
|
||||
export interface Requirement {
|
||||
id: string;
|
||||
text: string;
|
||||
shapes?: Shape[];
|
||||
}
|
||||
|
||||
/** An edge item — a probe-core `Item` specialized to the edge verification vocabulary. */
|
||||
export type Edge = Item<EdgeVerification>;
|
||||
|
||||
/**
|
||||
* Word-boundary cues mapping requirement prose -> data/behavior shape.
|
||||
* Heuristic and intentionally lossy; an authored `shapes` array overrides it.
|
||||
*/
|
||||
export const SHAPE_CUES: Record<Shape, RegExp> = {
|
||||
'numeric-range': /\b(round(ing|ed)?|threshold|max(imum)?|min(imum)?|limit|bound(ary)?|between|cap|percent|amount|price|count|number|score|rate|decimal)\b/i,
|
||||
'collection': /\b(lists?|arrays?|sets?|items?|collections?|each|every|all|sort(ed|ing)?|merge|dedupe|group|ranges?|intervals?|overlap(ping)?)\b/i,
|
||||
'text': /\b(string|text|names?|labels?|truncate|substring|char(acter)?s?|length|slug|message|unicode)\b/i,
|
||||
'stateful': /\b(save|persist|store|update|toggle|create|delete|remove|submit|retry|apply|register|insert)\b/i,
|
||||
'io': /\b(files?|requests?|fetch|upload|download|network|api|endpoints?|connections?|sockets?)\b/i,
|
||||
};
|
||||
|
||||
/** The locked shape vocabulary — exactly the keys of SHAPE_CUES (single source of truth). */
|
||||
export const VALID_SHAPES: ReadonlySet<string> = new Set(Object.keys(SHAPE_CUES));
|
||||
|
||||
/** Detect which shapes a requirement's prose matches (heuristic). */
|
||||
export function classifyShape(text: string): Shape[] {
|
||||
const shapes: Shape[] = [];
|
||||
const subject = String(text == null ? '' : text);
|
||||
for (const shape of Object.keys(SHAPE_CUES) as Shape[]) {
|
||||
if (SHAPE_CUES[shape].test(subject)) shapes.push(shape);
|
||||
}
|
||||
return shapes;
|
||||
}
|
||||
|
||||
/**
|
||||
* Closed taxonomy of 8 domain-boundary edge categories (established QA names).
|
||||
* `shapes` lists which requirement shapes make the category relevant.
|
||||
*/
|
||||
export const TAXONOMY: TaxonomyEntry[] = [
|
||||
{ id: 'boundary', name: 'Boundary values', shapes: ['numeric-range'], probe: 'What happens exactly at each min/max/threshold — and one step either side?' },
|
||||
{ id: 'adjacency', name: 'Adjacency / touching', shapes: ['collection'], probe: 'When two things are exactly equal or just touch, do they merge, collide, or separate?' },
|
||||
{ id: 'empty', name: 'Empty / degenerate', shapes: ['collection', 'text'], probe: 'What is the result for empty, single-element, or null input?' },
|
||||
{ id: 'encoding', name: 'Encoding / representation', shapes: ['text'], probe: 'Whose definition of length/equality applies — bytes, code points, grapheme clusters, or normalized form?' },
|
||||
{ id: 'ordering', name: 'Ordering / stability', shapes: ['collection'], probe: 'When elements compare equal, is output order specified and stable?' },
|
||||
{ id: 'precision', name: 'Precision / overflow', shapes: ['numeric-range'], probe: 'Where can precision loss or overflow occur, and what is the contract?' },
|
||||
{ id: 'idempotency', name: 'Idempotency / repetition', shapes: ['stateful'], probe: 'What happens if this runs twice on the same input?' },
|
||||
{ id: 'concurrency', name: 'Concurrency / effect ordering', shapes: ['stateful', 'io'], probe: 'If interrupted or run in parallel, what is guaranteed?' },
|
||||
];
|
||||
|
||||
/** Return taxonomy category ids whose applicable shapes intersect the input set. */
|
||||
export function applicableCategories(shapes: Shape[]): string[] {
|
||||
const set = new Set<Shape>(shapes);
|
||||
return TAXONOMY.filter((c) => c.shapes.some((s) => set.has(s))).map((c) => c.id);
|
||||
}
|
||||
|
||||
/**
|
||||
* The edge adapter's injected runtime validators (ADR-550 #5). `categories` is the closed
|
||||
* taxonomy; both verification tiers require a non-empty `resolution` (an explicit AC's text
|
||||
* or a backstop note) so plan-phase has a criterion to lift.
|
||||
*/
|
||||
export const EDGE_VALIDATORS: Validators = {
|
||||
categories: TAXONOMY.map((c) => c.id),
|
||||
verification: ['explicit', 'backstop'],
|
||||
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
|
||||
};
|
||||
|
||||
/**
|
||||
* Validate a single requirement — the generic id/text checks (probe-core) plus the edge's
|
||||
* `shapes`-must-be-an-array check. A bare string like `shapes:"numeric-range"` would otherwise
|
||||
* fall through to prose classification, silently ignoring the authored override.
|
||||
*
|
||||
* The edge adapter's `text` is REQUIRED (the prose is the classification signal), so reject a
|
||||
* missing/empty `text` when no authored `shapes` override is present. Without this, a `{ id }`
|
||||
* requirement classifies to zero shapes → zero edges → it is silently DROPPED from coverage
|
||||
* with no signal — the exact fail-open this feature exists to eliminate. An explicit `shapes`
|
||||
* array (including `[]` for "no applicable categories") is the legitimate way to opt out of
|
||||
* prose classification, so `text` is only required when `shapes` is absent.
|
||||
*/
|
||||
export function validateRequirement(requirement: Requirement): void {
|
||||
coreValidateRequirement(requirement);
|
||||
const r = requirement as unknown as { shapes?: unknown; text?: unknown };
|
||||
if (r.shapes != null && !Array.isArray(r.shapes)) {
|
||||
throw new Error(`requirement ${requirement.id} shapes must be an array when present`);
|
||||
}
|
||||
if (r.shapes == null && !(typeof r.text === 'string' && r.text.trim())) {
|
||||
throw new Error(
|
||||
`requirement ${requirement.id} text must be a non-empty string when no shapes override is provided`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/** Validate an edge resolution against the edge verification vocabulary. */
|
||||
export function validateResolution(resolution: Resolution<EdgeVerification>): true {
|
||||
return coreValidateResolution(resolution, EDGE_VALIDATORS);
|
||||
}
|
||||
|
||||
/**
|
||||
* Propose candidate edges for a requirement. Uses authored `shapes` when present, else
|
||||
* classifies from prose. Every proposed edge starts unresolved (verification null).
|
||||
*/
|
||||
export function proposeEdges(requirement: Requirement): Edge[] {
|
||||
validateRequirement(requirement);
|
||||
let shapes: Shape[];
|
||||
if (Array.isArray(requirement.shapes)) {
|
||||
// Fail closed: an authored array must contain only locked shape values. A non-empty
|
||||
// but invalid array (e.g. ['numeric'], a typo for 'numeric-range') would otherwise
|
||||
// intersect no category and silently suppress every probe — the gate reads green while
|
||||
// nothing was checked. An empty array stays a valid "no applicable categories" override.
|
||||
for (const s of requirement.shapes) {
|
||||
if (typeof s !== 'string' || !VALID_SHAPES.has(s)) {
|
||||
throw new Error(
|
||||
`invalid shape ${JSON.stringify(s)} for requirement ${requirement.id} — must be one of: ${[...VALID_SHAPES].join(', ')}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
shapes = requirement.shapes;
|
||||
} else {
|
||||
shapes = classifyShape(requirement.text);
|
||||
}
|
||||
return applicableCategories(shapes).map((catId): Edge => {
|
||||
const cat = TAXONOMY.find((c) => c.id === catId);
|
||||
return {
|
||||
requirement_id: requirement.id,
|
||||
category: catId,
|
||||
status: 'unresolved',
|
||||
verification: null,
|
||||
resolution: null,
|
||||
reason: null,
|
||||
probe: cat ? cat.probe : '',
|
||||
};
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Propose edges for every requirement (deterministic propose), then delegate the
|
||||
* merge/rollup/orphan-reject to probe-core. Edge-specific pre-checks: requirements must be an
|
||||
* array, requirement ids must be unique. Throws on any invalid resolution.
|
||||
*/
|
||||
export function analyzeCoverage(
|
||||
requirements: Requirement[],
|
||||
resolutions: Resolution<EdgeVerification>[] = [],
|
||||
): CoverageReport<EdgeVerification> {
|
||||
if (!Array.isArray(requirements)) {
|
||||
throw new Error('requirements must be an array');
|
||||
}
|
||||
const items: Edge[] = [];
|
||||
const seenReqIds = new Set<string>();
|
||||
for (const req of requirements) {
|
||||
validateRequirement(req);
|
||||
if (seenReqIds.has(req.id)) {
|
||||
throw new Error(`duplicate requirement id ${JSON.stringify(req.id)}`);
|
||||
}
|
||||
seenReqIds.add(req.id);
|
||||
for (const edge of proposeEdges(req)) items.push(edge);
|
||||
}
|
||||
return coreAnalyzeCoverage(items, resolutions, EDGE_VALIDATORS);
|
||||
}
|
||||
|
||||
/*
|
||||
* CLI entry (EP-06 invokable surface): `edge-probe.cjs <requirements.json> [resolutions.json]`.
|
||||
* The generic I/O plumbing (parse, fail-closed exit 2, pretty-JSON out) lives in probe-core's
|
||||
* `runProbeCli`; the edge adapter supplies its `analyzeCoverage`. Guarded by
|
||||
* `require.main === module` so it runs only when the compiled `.cjs` is executed directly.
|
||||
*/
|
||||
if (require.main === module) {
|
||||
runProbeCli(
|
||||
(requirements, resolutions) =>
|
||||
analyzeCoverage(requirements as Requirement[], resolutions as Resolution<EdgeVerification>[]),
|
||||
{ usage: 'edge-probe.cjs <requirements.json> [resolutions.json]' },
|
||||
);
|
||||
}
|
||||
336
src/probe-core.cts
Normal file
336
src/probe-core.cts
Normal file
@@ -0,0 +1,336 @@
|
||||
/**
|
||||
* probe-core — generic spec-phase probe resolution model (ADR-550 Decision 7).
|
||||
*
|
||||
* Extracted from the edge-probe (the first adapter) once the prohibition probe (#644)
|
||||
* proved it the *second* adapter of the same model: one adapter is a hypothetical seam,
|
||||
* two is a real one. This module owns everything generic — the resolution lifecycle,
|
||||
* the status×verification re-cut, `validateResolution`/`validateRequirement`, the
|
||||
* `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject engine,
|
||||
* the `byVerification` rollup, and the `runProbeCli` I/O scaffold. Each probe is a thin
|
||||
* adapter: it supplies the proposal logic (deterministic for edge, LLM-recall for
|
||||
* prohibition) and its closed vocabularies via injected validators.
|
||||
*
|
||||
* Authored as strict TypeScript (`src/probe-core.cts`) and compiled by
|
||||
* `tsc -p tsconfig.build.json` to the gitignored runtime artifact
|
||||
* `gsd-core/bin/lib/probe-core.cjs`. Do NOT hand-write the `.cjs`; it is emitted.
|
||||
*
|
||||
* Two orthogonal axes (the re-cut):
|
||||
* - status: resolved | dismissed | unresolved — the resolution LIFECYCLE (shared)
|
||||
* - verification: <probe-defined> | null — HOW a resolved item is verified
|
||||
* The edge adapter declares `verification: explicit | backstop`; the prohibition adapter
|
||||
* (#644) will declare `test | judgment`. Splitting the axes keeps the lifecycle enum free
|
||||
* of a verification fact and lets a sibling probe add its own tiers without a parallel enum.
|
||||
*
|
||||
* Typing is hybrid (ADR-550 #5): generic type params for adapter DX, but enforcement runs
|
||||
* through injected runtime validators, because the CLI executes over JSON where TS types are
|
||||
* erased. The contract test pins the validators, not the types.
|
||||
*/
|
||||
|
||||
import fs from 'node:fs';
|
||||
|
||||
/** Resolution lifecycle — shared across every probe adapter. */
|
||||
export type Status = 'resolved' | 'dismissed' | 'unresolved';
|
||||
|
||||
/** The LOCKED set of valid lifecycle statuses (the re-cut: no covered/backstop). */
|
||||
export const VALID_STATUS: Status[] = ['resolved', 'dismissed', 'unresolved'];
|
||||
|
||||
/**
|
||||
* A proposed or resolved item for a requirement/category pair. Generic over the probe's
|
||||
* verification-tier vocabulary `V` (edge: `'explicit' | 'backstop'`). A freshly proposed
|
||||
* item is `{ status: 'unresolved', verification: null, resolution: null, reason: null }`.
|
||||
*/
|
||||
export interface Item<V extends string = string> {
|
||||
requirement_id: string;
|
||||
category: string;
|
||||
status: Status;
|
||||
verification: V | null;
|
||||
resolution: string | null;
|
||||
reason: string | null;
|
||||
probe: string;
|
||||
}
|
||||
|
||||
/** An author resolution merged onto a proposed item. */
|
||||
export interface Resolution<V extends string = string> {
|
||||
requirement_id: string;
|
||||
category: string;
|
||||
status: Status;
|
||||
verification?: V | null;
|
||||
resolution?: string | null;
|
||||
reason?: string | null;
|
||||
}
|
||||
|
||||
/** A coverage report: the merged items plus rollup counts (incl. the per-tier breakdown). */
|
||||
export interface CoverageReport<V extends string = string> {
|
||||
items: Item<V>[];
|
||||
coverage: {
|
||||
applicable: number;
|
||||
resolved: number;
|
||||
unresolved: number;
|
||||
byVerification: Record<string, number>;
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The injected runtime enforcement contract (ADR-550 #5). The probe declares its closed
|
||||
* vocabularies so `analyzeCoverage`/`validateResolution` can enforce them at the JSON
|
||||
* boundary where TS types no longer exist.
|
||||
* - `categories` — the probe's valid category ids (a proposed item outside this set is an
|
||||
* adapter bug, caught rather than silently rolled up).
|
||||
* - `verification` — the valid verification tiers for a `resolved` item.
|
||||
* - `requiredFieldsByVerification` — for each tier, which resolution fields MUST be a
|
||||
* non-empty string (edge: every tier needs `resolution` text for plan-phase to lift).
|
||||
*/
|
||||
export interface Validators {
|
||||
categories: string[];
|
||||
verification: string[];
|
||||
requiredFieldsByVerification: Record<string, Array<'resolution' | 'reason'>>;
|
||||
}
|
||||
|
||||
/** A generic requirement — every probe ingests at least `{ id, text }`. */
|
||||
export interface Requirement {
|
||||
id: string;
|
||||
text?: string;
|
||||
}
|
||||
|
||||
function errMessage(e: unknown): string {
|
||||
return e instanceof Error ? e.message : String(e);
|
||||
}
|
||||
|
||||
/**
|
||||
* Structural guard for the report an adapter's `analyze` returns. The scaffold types `analyze`
|
||||
* loosely (it runs over JSON-parsed input the adapter `as`-casts), so a future adapter (#644)
|
||||
* that forgets to validate inside its closure could hand back a malformed object. Rather than
|
||||
* stringify garbage as green output, `runProbeCli` checks the report shape and fails closed.
|
||||
*/
|
||||
function isValidReport(report: unknown): report is CoverageReport {
|
||||
if (report == null || typeof report !== 'object') return false;
|
||||
const r = report as { items?: unknown; coverage?: unknown };
|
||||
if (!Array.isArray(r.items)) return false;
|
||||
const c = r.coverage as
|
||||
| { applicable?: unknown; resolved?: unknown; unresolved?: unknown; byVerification?: unknown }
|
||||
| undefined;
|
||||
if (c == null || typeof c !== 'object') return false;
|
||||
if (typeof c.applicable !== 'number' || typeof c.resolved !== 'number' || typeof c.unresolved !== 'number') {
|
||||
return false;
|
||||
}
|
||||
if (c.byVerification == null || typeof c.byVerification !== 'object') return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate a requirement's generic structural fields — fail closed on malformed input rather
|
||||
* than coercing it. Probe-specific fields (e.g. the edge adapter's `shapes`) are validated by
|
||||
* the adapter. Typed loosely because the CLI casts arbitrary parsed JSON to `Requirement`.
|
||||
*/
|
||||
export function validateRequirement(requirement: Requirement): void {
|
||||
const r = requirement as unknown as { id?: unknown; text?: unknown };
|
||||
if (typeof r.id !== 'string' || !r.id.trim()) {
|
||||
throw new Error(`requirement id must be a non-empty string (got ${JSON.stringify(r.id)})`);
|
||||
}
|
||||
if (r.text != null && typeof r.text !== 'string') {
|
||||
throw new Error(`requirement ${r.id} text must be a string when present`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate a resolution against the probe's injected validators. Rejects an unknown status,
|
||||
* a dismissal without a non-empty reason, a `resolved` item with a missing/unknown
|
||||
* verification tier, and a `resolved` item missing any field its tier requires (per
|
||||
* `requiredFieldsByVerification`). Returns true on success.
|
||||
*/
|
||||
export function validateResolution<V extends string>(r: Resolution<V>, validators: Validators): true {
|
||||
const key = `${r.requirement_id}::${r.category}`;
|
||||
if (!VALID_STATUS.includes(r.status)) {
|
||||
throw new Error(`invalid status "${r.status}" for ${key}`);
|
||||
}
|
||||
// Invariant (this module's header): `verification` is null unless `status` is `resolved`.
|
||||
// Enforce it for EVERY status — a dismissed/unresolved resolution carrying a verification
|
||||
// tier would otherwise merge verbatim (`analyzeCoverage` below) and silently break the
|
||||
// model for the second adapter (#644) that inherits this seam. Fail closed across the full
|
||||
// status×verification space, not just `resolved`.
|
||||
if (r.status !== 'resolved' && r.verification != null) {
|
||||
throw new Error(`verification must be null unless status is "resolved" (got "${r.verification}") for ${key}`);
|
||||
}
|
||||
// An `unresolved` resolution is an UNACTED item: it must carry no resolution/reason payload.
|
||||
// A populated payload is an authoring mistake (the author meant resolved/dismissed) that
|
||||
// would otherwise be silently dropped into the unresolved count with no error pointing at
|
||||
// it. Reject it so the mistake surfaces.
|
||||
if (r.status === 'unresolved') {
|
||||
if (r.resolution != null && String(r.resolution).trim()) {
|
||||
throw new Error(`unresolved must not carry a resolution (${key})`);
|
||||
}
|
||||
if (r.reason != null && String(r.reason).trim()) {
|
||||
throw new Error(`unresolved must not carry a reason (${key})`);
|
||||
}
|
||||
}
|
||||
if (r.status === 'dismissed' && !(r.reason && String(r.reason).trim())) {
|
||||
throw new Error(`dismissed requires a reason (${key})`);
|
||||
}
|
||||
if (r.status === 'resolved') {
|
||||
const tier = r.verification;
|
||||
if (tier == null) {
|
||||
throw new Error(`resolved requires a verification tier (one of: ${validators.verification.join(', ')}) for ${key}`);
|
||||
}
|
||||
if (!validators.verification.includes(tier)) {
|
||||
throw new Error(`invalid verification "${tier}" for ${key} — must be one of: ${validators.verification.join(', ')}`);
|
||||
}
|
||||
const required = validators.requiredFieldsByVerification[tier] ?? [];
|
||||
for (const field of required) {
|
||||
// field is 'resolution' | 'reason'; both are `string | null | undefined` on Resolution,
|
||||
// so the indexed access is string-typed (no unknown-to-string coercion).
|
||||
const value = r[field];
|
||||
if (!(value != null && String(value).trim())) {
|
||||
throw new Error(`${tier} requires a ${field} (${key})`);
|
||||
}
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Merge author resolutions onto ALREADY-PROPOSED items and roll up coverage counts.
|
||||
*
|
||||
* Core operates on `items[]`, never a `proposeFn`: probes have different deterministic
|
||||
* surfaces (edge = deterministic propose + LLM resolve; prohibition = LLM propose + deterministic
|
||||
* validate/merge), so proposal stays in each adapter and core must not assume it is deterministic.
|
||||
*
|
||||
* `coverage.resolved` is the COUNT of CLOSED items (`resolved` + `dismissed` status) =
|
||||
* `applicable - unresolved` — the pre-re-cut "covered + dismissed + backstop" set,
|
||||
* count-preserved. `byVerification` breaks the `resolved`-status items down by tier (each tier
|
||||
* initialized to 0). Throws on any invalid resolution, a duplicate, an orphan (a resolution
|
||||
* matching no proposed item), or a proposed item whose category is outside `validators.categories`.
|
||||
*/
|
||||
export function analyzeCoverage<V extends string>(
|
||||
items: Item<V>[],
|
||||
resolutions: Resolution<V>[] = [],
|
||||
validators: Validators,
|
||||
): CoverageReport<V> {
|
||||
if (!Array.isArray(items)) {
|
||||
throw new Error('items must be an array');
|
||||
}
|
||||
const key = (r: { requirement_id: string; category: string }): string => `${r.requirement_id}::${r.category}`;
|
||||
const resMap = new Map<string, Resolution<V>>();
|
||||
for (const r of resolutions) {
|
||||
validateResolution(r, validators);
|
||||
if (resMap.has(key(r))) {
|
||||
throw new Error(`duplicate resolution for ${key(r)}`);
|
||||
}
|
||||
resMap.set(key(r), r);
|
||||
}
|
||||
const validCategories = new Set(validators.categories);
|
||||
const merged: Item<V>[] = [];
|
||||
const itemKeys = new Set<string>();
|
||||
for (const item of items) {
|
||||
if (!validCategories.has(item.category)) {
|
||||
throw new Error(`item ${key(item)} has unknown category "${item.category}" — not one of: ${validators.categories.join(', ')}`);
|
||||
}
|
||||
itemKeys.add(key(item));
|
||||
const o = resMap.get(key(item));
|
||||
if (o) {
|
||||
merged.push({ ...item, status: o.status, verification: o.verification ?? null, resolution: o.resolution ?? null, reason: o.reason ?? null });
|
||||
} else {
|
||||
// No author resolution: the item is rolled up VERBATIM, so its own status/fields must be
|
||||
// valid too. The edge adapter only proposes `unresolved` items, but the prohibition adapter
|
||||
// (#644) proposes LLM-generated items that arrive already populated — one carrying an
|
||||
// out-of-enum status (e.g. the dropped "covered") or `dismissed` with no reason would
|
||||
// otherwise be counted closed without validation. An Item is structurally a superset of a
|
||||
// Resolution, so the same fail-closed check guards both. (ADR-550 Decision 5 hardens this
|
||||
// shared seam for the second adapter; m1.)
|
||||
validateResolution(item as unknown as Resolution<V>, validators);
|
||||
merged.push(item);
|
||||
}
|
||||
}
|
||||
// Reject orphan resolutions — a resolution whose (requirement_id, category) matches no
|
||||
// proposed item (typo'd category or a non-applicable one) would otherwise be silently
|
||||
// dropped, leaving the author believing an item is resolved while the report shows it
|
||||
// unresolved (adversarial-review HIGH; preserved from the edge-probe's original engine).
|
||||
for (const k of resMap.keys()) {
|
||||
if (!itemKeys.has(k)) {
|
||||
throw new Error(`unknown resolution for ${k} — no matching proposed item (typo'd category or non-applicable shape?)`);
|
||||
}
|
||||
}
|
||||
const unresolved = merged.filter((i) => i.status === 'unresolved').length;
|
||||
const applicable = merged.length;
|
||||
const resolved = applicable - unresolved; // closed set: resolved-status + dismissed
|
||||
const byVerification: Record<string, number> = {};
|
||||
for (const tier of validators.verification) byVerification[tier] = 0;
|
||||
for (const i of merged) {
|
||||
if (i.status === 'resolved' && i.verification != null) {
|
||||
byVerification[i.verification] = (byVerification[i.verification] ?? 0) + 1;
|
||||
}
|
||||
}
|
||||
return { items: merged, coverage: { applicable, resolved, unresolved, byVerification } };
|
||||
}
|
||||
|
||||
/*
|
||||
* CLI scaffold (the EP-06 invokable surface, generalized). Each probe ships one bin that
|
||||
* calls `runProbeCli` with its own `analyze` (closing over the adapter's propose + validators)
|
||||
* and usage string; a single dispatcher CLI is a deferred follow-on. The I/O dependencies are
|
||||
* injectable so the generic plumbing is unit-testable without spawning a process.
|
||||
*
|
||||
* `tsconfig.build.json` sets `"types": ["node"]`, so `process` and `node:fs` are typed.
|
||||
*/
|
||||
|
||||
/** Injectable I/O for `runProbeCli` (defaults wire to the real process). */
|
||||
export interface ProbeCliOptions {
|
||||
usage: string;
|
||||
argv?: string[];
|
||||
readFile?: (path: string) => string;
|
||||
write?: (s: string) => void;
|
||||
writeErr?: (s: string) => void;
|
||||
exit?: (code: number) => void;
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the requirements file (and optional resolutions file), run the adapter's `analyze`,
|
||||
* and write the report as pretty JSON + newline. With no requirements path, writes the usage
|
||||
* line to stderr and exits 2. A JSON-parse failure or any `analyze` throw is a handled error:
|
||||
* stderr + exit 2, never an uncaught stack trace — so the engine's fail-closed validation
|
||||
* surfaces at the workflow boundary rather than failing open.
|
||||
*/
|
||||
export function runProbeCli(
|
||||
analyze: (requirements: unknown, resolutions: unknown) => CoverageReport,
|
||||
options: ProbeCliOptions,
|
||||
): void {
|
||||
const argv = options.argv ?? process.argv;
|
||||
const readFile = options.readFile ?? ((p: string) => fs.readFileSync(p, 'utf8'));
|
||||
const write = options.write ?? ((s: string) => { process.stdout.write(s); });
|
||||
const writeErr = options.writeErr ?? ((s: string) => { process.stderr.write(s); });
|
||||
const exit = options.exit ?? ((code: number) => { process.exit(code); });
|
||||
|
||||
const reqPath: string | undefined = argv[2];
|
||||
const resPath: string | undefined = argv[3];
|
||||
if (!reqPath) {
|
||||
writeErr(`usage: ${options.usage}\n`);
|
||||
exit(2);
|
||||
return;
|
||||
}
|
||||
let requirements: unknown;
|
||||
try {
|
||||
requirements = JSON.parse(readFile(reqPath));
|
||||
} catch (e: unknown) {
|
||||
writeErr(`error: cannot parse JSON from ${reqPath}: ${errMessage(e)}\n`);
|
||||
exit(2);
|
||||
return;
|
||||
}
|
||||
let resolutions: unknown = [];
|
||||
if (resPath) {
|
||||
try {
|
||||
resolutions = JSON.parse(readFile(resPath));
|
||||
} catch (e: unknown) {
|
||||
writeErr(`error: cannot parse JSON from ${resPath}: ${errMessage(e)}\n`);
|
||||
exit(2);
|
||||
return;
|
||||
}
|
||||
}
|
||||
try {
|
||||
const report = analyze(requirements, resolutions);
|
||||
if (!isValidReport(report)) {
|
||||
throw new Error('adapter returned a structurally-invalid coverage report (expected { items[], coverage{ applicable, resolved, unresolved, byVerification } })');
|
||||
}
|
||||
write(`${JSON.stringify(report, null, 2)}\n`);
|
||||
} catch (e: unknown) {
|
||||
writeErr(`error: ${errMessage(e)}\n`);
|
||||
exit(2);
|
||||
}
|
||||
}
|
||||
115
tests/edge-probe-docs-fixtures.test.cjs
Normal file
115
tests/edge-probe-docs-fixtures.test.cjs
Normal file
@@ -0,0 +1,115 @@
|
||||
// allow-test-rule: runtime-contract-is-the-product — the rendered reference/SPEC/ADR vocab surfaces are the runtime contract; this pins their bijection to the code (docs-parity)
|
||||
// Asserts the portable reference doc (gsd-core/references/edge-probe.md) keeps its
|
||||
// worked-example JSON blocks in sync with the source-of-truth fixture files under
|
||||
// gsd-core/references/edge-probe-fixtures/. The fixtures are the canonical data; the
|
||||
// doc embeds copies. Per the CONTRIBUTING.md exception matrix this is `docs-parity`: a
|
||||
// reference doc must mirror source-defined data and there is no runtime enumeration API.
|
||||
// The comparison is PARSED JSON (deepEqual of JSON.parse on both sides), never a raw-text
|
||||
// substring match — so a reformat that preserves the data does not fail, and any semantic
|
||||
// drift between doc and fixture does.
|
||||
'use strict';
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const docPath = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe.md');
|
||||
const fixturesRoot = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures');
|
||||
const adrPath = path.join(__dirname, '..', 'docs', 'adr', '550-spec-phase-probe-contract.md');
|
||||
const specTemplatePath = path.join(__dirname, '..', 'gsd-core', 'templates', 'spec.md');
|
||||
|
||||
// Extract fenced blocks tagged ```json edge-probe:<dir>/<file> from the doc, keyed by ref.
|
||||
// The \n? before the closing fence allows blocks whose closing fence has no preceding newline
|
||||
// (fixes the silent-skip bug where a trailing-fence-with-no-newline was not matched).
|
||||
function taggedJsonBlocks(md) {
|
||||
const re = /```json edge-probe:([^\n]+)\n([\s\S]*?)\n?```/g;
|
||||
const out = {};
|
||||
let m;
|
||||
while ((m = re.exec(md))) out[m[1].trim()] = m[2];
|
||||
return out;
|
||||
}
|
||||
|
||||
describe('edge-probe doc/fixture sync', () => {
|
||||
test('reference doc exists', () => {
|
||||
assert.ok(fs.existsSync(docPath), `${docPath} must exist`);
|
||||
});
|
||||
|
||||
test('doc embeds tagged fixture blocks for every expected-coverage.json fixture (count-equality)', () => {
|
||||
const md = fs.readFileSync(docPath, 'utf8');
|
||||
const blocks = taggedJsonBlocks(md);
|
||||
// Count the expected-coverage.json files under the fixtures root (one per fixture dir).
|
||||
const expectedCount = fs.readdirSync(fixturesRoot, { withFileTypes: true })
|
||||
.filter(e => e.isDirectory())
|
||||
.filter(dir => fs.existsSync(path.join(fixturesRoot, dir.name, 'expected-coverage.json')))
|
||||
.length;
|
||||
assert.strictEqual(
|
||||
Object.keys(blocks).length,
|
||||
expectedCount,
|
||||
`edge-probe.md must embed exactly ${expectedCount} tagged blocks (one per fixture expected-coverage.json)`
|
||||
);
|
||||
});
|
||||
|
||||
test('every tagged doc block parses and deepEquals its fixture file', () => {
|
||||
const md = fs.readFileSync(docPath, 'utf8');
|
||||
const blocks = taggedJsonBlocks(md);
|
||||
for (const [ref, body] of Object.entries(blocks)) {
|
||||
const fixtureFile = path.join(fixturesRoot, ref);
|
||||
const onDisk = fs.readFileSync(fixtureFile, 'utf8');
|
||||
assert.deepEqual(JSON.parse(body), JSON.parse(onDisk),
|
||||
`doc block edge-probe:${ref} must deepEqual ${fixtureFile}`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
// m2: lock the machine↔SPEC vocabulary mapping so the two layers cannot silently drift. The
|
||||
// machine contract uses orthogonal `status` × `verification`; the SPEC table and planner prose
|
||||
// render a flat `covered/dismissed/backstop/unresolved`. The migration is documented in ADR-550
|
||||
// Decision 7a and the reference's "Generic mapping" table, but the SPEC table is rendered by the
|
||||
// LLM workflow (no JS renderer to round-trip against). So this pins the canonical map as code AND
|
||||
// grounds it in every doc surface that renders the vocabulary — a renderer/parser drift fails here.
|
||||
describe('edge-probe machine↔SPEC vocabulary mapping (m2 drift lock)', () => {
|
||||
// The canonical migration map (ADR-550 Decision 7a): machine state → SPEC display label.
|
||||
// `verification` is null for the lifecycle-only states (dismissed/unresolved).
|
||||
const MACHINE_TO_DISPLAY = [
|
||||
{ status: 'resolved', verification: 'explicit', display: 'covered' },
|
||||
{ status: 'resolved', verification: 'backstop', display: 'backstop' },
|
||||
{ status: 'dismissed', verification: null, display: 'dismissed' },
|
||||
{ status: 'unresolved', verification: null, display: 'unresolved' },
|
||||
];
|
||||
const machineKey = (s, v) => `${s}|${v ?? '∅'}`;
|
||||
|
||||
test('the mapping is a bijection — machine→display→machine is identity, no shared labels', () => {
|
||||
const toDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [machineKey(m.status, m.verification), m.display]));
|
||||
const fromDisplay = new Map(MACHINE_TO_DISPLAY.map((m) => [m.display, machineKey(m.status, m.verification)]));
|
||||
// No two machine states collapse onto the same display label (the silent-drift failure mode).
|
||||
assert.equal(toDisplay.size, MACHINE_TO_DISPLAY.length, 'each machine state must have a distinct key');
|
||||
assert.equal(fromDisplay.size, MACHINE_TO_DISPLAY.length, 'each display label must map back to exactly one machine state');
|
||||
// Round-trip identity.
|
||||
for (const m of MACHINE_TO_DISPLAY) {
|
||||
const key = machineKey(m.status, m.verification);
|
||||
assert.equal(fromDisplay.get(toDisplay.get(key)), key, `${key} must round-trip through its display label`);
|
||||
}
|
||||
});
|
||||
|
||||
test('ADR-550 Decision 7a documents the resolved/explicit↔covered and resolved/backstop↔backstop migration', () => {
|
||||
const adr = fs.readFileSync(adrPath, 'utf8');
|
||||
assert.match(adr, /covered\b[^.]*resolved[^.]*explicit/i, 'ADR must document covered → {resolved, explicit}');
|
||||
assert.match(adr, /backstop\b[^.]*resolved[^.]*backstop/i, 'ADR must document backstop → {resolved, backstop}');
|
||||
assert.match(adr, /count-for-count|count-preserved/i, 'ADR must state coverage.resolved is count-preserved across the migration');
|
||||
});
|
||||
|
||||
test('the SPEC template legend renders exactly the four canonical display labels', () => {
|
||||
const spec = fs.readFileSync(specTemplatePath, 'utf8');
|
||||
for (const { display } of MACHINE_TO_DISPLAY) {
|
||||
assert.match(spec, new RegExp(`\\b${display}\\b`, 'i'), `spec.md Edge Coverage legend must render the "${display}" label`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the reference Generic-mapping table distinguishes the resolved/explicit and resolved/backstop tiers', () => {
|
||||
const md = fs.readFileSync(docPath, 'utf8');
|
||||
assert.match(md, /resolved`?\/`?explicit/i, 'reference must name the resolved/explicit tier in the mapping table');
|
||||
assert.match(md, /resolved`?\/`?backstop/i, 'reference must name the resolved/backstop tier in the mapping table');
|
||||
});
|
||||
});
|
||||
167
tests/edge-probe-planner-contract.test.cjs
Normal file
167
tests/edge-probe-planner-contract.test.cjs
Normal file
@@ -0,0 +1,167 @@
|
||||
// allow-test-rule: runtime-contract-is-the-product — plan-phase.md's planner prompt is the deployed runtime contract under assertion
|
||||
// plan-phase.md is the deployed planning workflow contract; these checks lock
|
||||
// the SPEC path wiring and quality-gate that the edge-probe review (RR-01/02/03)
|
||||
// requires — assertions scope to extracted sub-blocks to avoid false positives.
|
||||
|
||||
'use strict';
|
||||
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const PLAN_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-phase.md');
|
||||
const PROMPT_PATH = path.join(__dirname, '..', 'gsd-core', 'templates', 'planner-subagent-prompt.md');
|
||||
|
||||
function readPlanPhase() {
|
||||
return fs.readFileSync(PLAN_PHASE_PATH, 'utf8');
|
||||
}
|
||||
|
||||
// Extract the planner <downstream_consumer> block from plan-phase.md. The runtime planner is
|
||||
// spawned from plan-phase.md's own inline <planning_context>/<downstream_consumer>, so the
|
||||
// load-bearing lift instruction must live here — NOT in templates/planner-subagent-prompt.md,
|
||||
// which nothing loads at runtime (no @-import in agents/gsd-planner.md, no read in plan-phase.md).
|
||||
function extractDownstreamConsumerBlock(content) {
|
||||
const start = content.indexOf('<downstream_consumer>');
|
||||
if (start === -1) return '';
|
||||
const end = content.indexOf('</downstream_consumer>', start);
|
||||
if (end === -1) return '';
|
||||
return content.slice(start, end + '</downstream_consumer>'.length);
|
||||
}
|
||||
|
||||
// Extract the planner <files_to_read> block that contains {UI_SPEC_PATH}
|
||||
// There are multiple <files_to_read> blocks in plan-phase.md; we need the one
|
||||
// at ~line 890-912 inside the planning_context markdown block.
|
||||
function extractPlannerFilesBlock(content) {
|
||||
let pos = 0;
|
||||
while (true) {
|
||||
const start = content.indexOf('<files_to_read>', pos);
|
||||
if (start === -1) return '';
|
||||
const end = content.indexOf('</files_to_read>', start);
|
||||
if (end === -1) return '';
|
||||
const block = content.slice(start, end + '</files_to_read>'.length);
|
||||
if (block.includes('{UI_SPEC_PATH}')) {
|
||||
return block;
|
||||
}
|
||||
pos = end + 1;
|
||||
}
|
||||
}
|
||||
|
||||
// Extract the planner <quality_gate> block (the last one in plan-phase.md,
|
||||
// inside the planner prompt template)
|
||||
function extractQualityGateBlock(content) {
|
||||
const start = content.lastIndexOf('<quality_gate>');
|
||||
if (start === -1) return '';
|
||||
const end = content.indexOf('</quality_gate>', start);
|
||||
if (end === -1) return '';
|
||||
return content.slice(start, end + '</quality_gate>'.length);
|
||||
}
|
||||
|
||||
// Test A (RR-01): plan-phase.md resolves a phase *-SPEC.md into SPEC_FILE/SPEC_PATH
|
||||
// Uses new RegExp to correctly match literal $ and ( characters in bash snippets.
|
||||
// This MUST FAIL before the RR-01 fix (no SPEC_PATH resolution exists today)
|
||||
test('RR-01: plan-phase.md resolves phase *-SPEC.md (excluding AI/UI variants) into SPEC_FILE/SPEC_PATH', () => {
|
||||
const content = readPlanPhase();
|
||||
|
||||
// Assert the canonical SPEC_FILE resolution form is present.
|
||||
// new RegExp used so that \$ and \( are treated as literal dollar-sign and open-paren
|
||||
// (JS regex literals interpret \$ as end-anchor and \( as group open).
|
||||
assert.match(
|
||||
content,
|
||||
new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md'),
|
||||
'plan-phase.md must resolve SPEC_FILE using ls "${...PHASE_DIR...}"/*-SPEC.md pattern'
|
||||
);
|
||||
|
||||
// Assert {SPEC_PATH} token appears in the planner files_to_read block
|
||||
const filesBlock = extractPlannerFilesBlock(content);
|
||||
assert.match(
|
||||
filesBlock,
|
||||
/[{]SPEC_PATH[}]/,
|
||||
'The planner <files_to_read> block (containing {UI_SPEC_PATH}) must also contain {SPEC_PATH}'
|
||||
);
|
||||
});
|
||||
|
||||
// Test B (RR-01): The {SPEC_PATH} entry is labelled as carrying the ## Edge Coverage section
|
||||
// This MUST FAIL before the RR-01 fix (no {SPEC_PATH} entry exists today)
|
||||
test('RR-01: {SPEC_PATH} entry in files_to_read is labelled with Edge Coverage', () => {
|
||||
const content = readPlanPhase();
|
||||
const filesBlock = extractPlannerFilesBlock(content);
|
||||
|
||||
assert.match(
|
||||
filesBlock,
|
||||
/[{]SPEC_PATH[}][^\n]*Edge Coverage/,
|
||||
'{SPEC_PATH} entry in planner files_to_read must be labelled as carrying the ## Edge Coverage section'
|
||||
);
|
||||
});
|
||||
|
||||
// Extract a "## " section from its heading until the next "## " heading (or EOF).
|
||||
function extractSection(content, heading) {
|
||||
const start = content.indexOf(heading);
|
||||
if (start === -1) return '';
|
||||
const next = content.indexOf('\n## ', start + heading.length);
|
||||
return content.slice(start, next === -1 ? content.length : next);
|
||||
}
|
||||
|
||||
// Test E (RR-01 REACHABILITY — the assertion that catches the original no-op):
|
||||
// Token presence is not enough. The SPEC resolution must live on an UN-GATED path. §4.5
|
||||
// "Check AI-SPEC" is skipped on every non-AI phase (ai_integration_phase_enabled false /
|
||||
// --skip-ai-spec), so a resolution placed there leaves SPEC_PATH unbound and the planner
|
||||
// never receives the SPEC — exactly the #550 silent no-op. Assert it is NOT in §4.5.
|
||||
test('RR-01 reachability: SPEC_FILE resolution is NOT gated inside the skippable "## 4.5 Check AI-SPEC" section', () => {
|
||||
const content = readPlanPhase();
|
||||
const aiSpecSection = extractSection(content, '## 4.5. Check AI-SPEC');
|
||||
const specResolution = new RegExp('SPEC_FILE=\\$\\(ls "\\$\\{[A-Z_]*PHASE_DIR[A-Z_]*\\}"[/][*]-SPEC\\.md');
|
||||
assert.ok(aiSpecSection.length > 0, 'sanity: the §4.5 Check AI-SPEC section must exist to scope this test');
|
||||
assert.doesNotMatch(
|
||||
aiSpecSection,
|
||||
specResolution,
|
||||
'SPEC_FILE resolution must NOT live inside the skippable §4.5 "Check AI-SPEC" block — gating it there silently starves the planner of the SPEC on non-AI phases (the original #550 no-op)'
|
||||
);
|
||||
assert.match(content, specResolution, 'SPEC_FILE resolution must still exist on an un-gated path elsewhere in plan-phase.md');
|
||||
});
|
||||
|
||||
// Test C (RR-02 consumer end): the lift instruction lives in the RUNTIME planner surface.
|
||||
// plan-phase.md spawns the planner from its own inline <planning_context>; the
|
||||
// templates/planner-subagent-prompt.md file is orphaned (loaded by nothing), so asserting the
|
||||
// contract there is false assurance — the test stays green even if the runtime never consumes it.
|
||||
// Pin the contract to the <downstream_consumer> block plan-phase.md actually sends the planner.
|
||||
test('RR-02 consumer: plan-phase.md downstream_consumer instructs lifting covered/backstop edges into must_haves.truths', () => {
|
||||
const block = extractDownstreamConsumerBlock(readPlanPhase());
|
||||
|
||||
assert.ok(block.length > 0, 'sanity: plan-phase.md must contain a <downstream_consumer> block to scope this test');
|
||||
assert.match(
|
||||
block,
|
||||
/##\s*Edge Coverage/,
|
||||
'plan-phase.md <downstream_consumer> must reference ## Edge Coverage'
|
||||
);
|
||||
assert.match(
|
||||
block,
|
||||
/must_haves\.truths/,
|
||||
'plan-phase.md <downstream_consumer> must reference must_haves.truths as the lift target'
|
||||
);
|
||||
|
||||
// Regression guard against re-orphaning: the lift instruction must NOT be relocated back into
|
||||
// templates/planner-subagent-prompt.md (which nothing loads) and presented as the consumer
|
||||
// contract — that is exactly the false-assurance this retarget fixes.
|
||||
const prompt = fs.readFileSync(PROMPT_PATH, 'utf8');
|
||||
assert.doesNotMatch(
|
||||
prompt,
|
||||
/Edge Coverage/,
|
||||
'planner-subagent-prompt.md is not loaded at runtime — the Edge Coverage lift instruction must not live there'
|
||||
);
|
||||
});
|
||||
|
||||
// Test D (RR-03): plan-phase.md <quality_gate> contains a covered/backstop ↔ must_haves item
|
||||
// This MUST FAIL before the RR-03 fix (no such quality_gate item exists today)
|
||||
test('RR-03: planner quality_gate requires covered/backstop edges represented in must_haves', () => {
|
||||
const content = readPlanPhase();
|
||||
const qgBlock = extractQualityGateBlock(content);
|
||||
|
||||
assert.match(
|
||||
qgBlock,
|
||||
/covered.*backstop.*must_haves|backstop.*covered.*must_haves/i,
|
||||
'planner quality_gate must contain a checklist item tying covered/backstop edges to must_haves'
|
||||
);
|
||||
});
|
||||
193
tests/edge-probe-spec-phase-contract.test.cjs
Normal file
193
tests/edge-probe-spec-phase-contract.test.cjs
Normal file
@@ -0,0 +1,193 @@
|
||||
// allow-test-rule: runtime-contract-is-the-product — spec-phase.md Step 5.5 is the deployed workflow runtime contract under assertion
|
||||
// spec-phase.md is the deployed spec workflow contract; these checks lock
|
||||
// the Step 5.5 wiring so the edge-probe.cjs runtime invocation cannot
|
||||
// silently rot the way the original plan-phase no-op did (reviewer finding RR-11).
|
||||
// Assertions scope to the extracted Step 5.5 block to avoid false positives
|
||||
// from incidental mentions elsewhere in the file.
|
||||
|
||||
'use strict';
|
||||
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const SPEC_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'spec-phase.md');
|
||||
|
||||
function readSpecPhase() {
|
||||
return fs.readFileSync(SPEC_PHASE_PATH, 'utf8');
|
||||
}
|
||||
|
||||
// Slice the Step 5.5 block: from the "Step 5.5" heading to the next "## " or "Step " heading.
|
||||
// This scopes assertions to Step 5.5 only, preventing false positives from mentions elsewhere.
|
||||
function extractStep55Block(content) {
|
||||
const startIdx = content.indexOf('## Step 5.5');
|
||||
if (startIdx === -1) {
|
||||
// Also try without the ## prefix
|
||||
const altIdx = content.indexOf('Step 5.5');
|
||||
if (altIdx === -1) return '';
|
||||
// Find end: next heading starting with ## or Step N (not Step 5.5)
|
||||
const rest = content.slice(altIdx + 'Step 5.5'.length);
|
||||
const nextHeading = rest.search(/\n## |\nStep \d/);
|
||||
if (nextHeading === -1) return content.slice(altIdx);
|
||||
return content.slice(altIdx, altIdx + 'Step 5.5'.length + nextHeading);
|
||||
}
|
||||
const rest = content.slice(startIdx + '## Step 5.5'.length);
|
||||
const nextHeading = rest.search(/\n## /);
|
||||
if (nextHeading === -1) return content.slice(startIdx);
|
||||
return content.slice(startIdx, startIdx + '## Step 5.5'.length + nextHeading);
|
||||
}
|
||||
|
||||
// Test A (RR-11): Step 5.5 resolves and invokes edge-probe.cjs via node.
|
||||
// MUST FAIL before the RR-04 wire (Step 5.5 is prose-only today — no CLI invocation).
|
||||
test('RR-11: spec-phase Step 5.5 resolves edge-probe.cjs via path-fallback loop', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
|
||||
// Assert the path-fallback resolution loop for edge-probe.cjs is present in Step 5.5.
|
||||
// The token "edge-probe.cjs" must appear inside the block (the artifact being resolved).
|
||||
assert.match(
|
||||
block,
|
||||
/edge-probe\.cjs/,
|
||||
'Step 5.5 must reference edge-probe.cjs as the artifact being resolved'
|
||||
);
|
||||
|
||||
// Assert node invocation of edge-probe.cjs in Step 5.5.
|
||||
// Matches: node "$EDGE_PROBE_JS" or node ... edge-probe.cjs
|
||||
assert.match(
|
||||
block,
|
||||
/node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/,
|
||||
'Step 5.5 must invoke edge-probe.cjs via node (e.g. node "$EDGE_PROBE_JS" ...)'
|
||||
);
|
||||
});
|
||||
|
||||
// Test C (RR-11 FUNCTION — the assertion that catches "decorative bash"):
|
||||
// Token presence is not enough. The invocation is a no-op unless $REQS_JSON is actually
|
||||
// POPULATED before the engine runs (the original block only mktemp'd it + left a comment,
|
||||
// so the CLI parsed an empty file). Assert the block (a) writes $REQS_JSON via a redirect,
|
||||
// (b) does so BEFORE the node invocation, and (c) guards against an empty/invalid file.
|
||||
test('RR-11 function: Step 5.5 writes $REQS_JSON before invoking, and guards against empty input', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
|
||||
// (a) A redirect that writes the requirements into $REQS_JSON (e.g. `cat > "$REQS_JSON"`).
|
||||
const writeIdx = block.search(/>\s*"\$REQS_JSON"/);
|
||||
assert.ok(
|
||||
writeIdx !== -1,
|
||||
'Step 5.5 must WRITE requirements into $REQS_JSON (a redirect like `cat > "$REQS_JSON"`), not just mktemp it — an empty file makes the probe a silent no-op'
|
||||
);
|
||||
|
||||
// (b) The write must precede the node invocation of the engine.
|
||||
const invokeIdx = block.search(/node\s+["$].*[Ee][Dd][Gg][Ee][-_][Pp][Rr][Oo][Bb][Ee]/);
|
||||
assert.ok(invokeIdx !== -1, 'Step 5.5 must invoke the engine via node');
|
||||
assert.ok(
|
||||
writeIdx < invokeIdx,
|
||||
'Step 5.5 must populate $REQS_JSON BEFORE invoking edge-probe.cjs (write precedes the node call)'
|
||||
);
|
||||
|
||||
// (c) A guard that refuses to run on an empty/invalid requirements array.
|
||||
assert.match(
|
||||
block,
|
||||
/Array\.isArray|empty\/invalid|empty or invalid|REQS_JSON[^\n]*empty/i,
|
||||
'Step 5.5 must guard against an empty/invalid $REQS_JSON before invoking (fail loud, not silent no-op)'
|
||||
);
|
||||
});
|
||||
|
||||
// Test B (RR-11): Step 5.5 has an explicit not-found branch — build:lib or error token.
|
||||
// MUST FAIL before the RR-04 wire (no not-found handling today).
|
||||
test('RR-11: spec-phase Step 5.5 has an explicit not-found branch (build:lib or blocking error)', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
|
||||
// Assert either a build:lib invocation or an explicit "not found" / error message exists.
|
||||
// This prevents the wire from being added as a silent-skip with no fallback.
|
||||
assert.match(
|
||||
block,
|
||||
/build:lib|not found|ERROR.*edge-probe|edge-probe.*not found/i,
|
||||
'Step 5.5 must have an explicit not-found branch (build:lib attempt or clear blocking error)'
|
||||
);
|
||||
});
|
||||
|
||||
// Test D (review High): the build fallback must NEVER run the CONSUMING project's package
|
||||
// scripts. Every executable `build:lib` invocation must be pinned to the GSD dir with
|
||||
// `npm --prefix`, and the build must be gated behind a verified GSD source checkout. A bare
|
||||
// `npm run build:lib` (no --prefix) uses cwd — which, under the git-toplevel fallback, is the
|
||||
// consumer repo — and would execute its codegen/migrations during a spec workflow.
|
||||
test('review High: Step 5.5 build:lib is --prefix-pinned to the GSD dir and gated on a source checkout', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
|
||||
// Collect lines that actually INVOKE npm (trimmed start === "npm"), excluding echo/comment
|
||||
// mentions (e.g. the error message that quotes `npm run build:lib` for the user).
|
||||
const buildInvocations = block
|
||||
.split('\n')
|
||||
.filter((l) => l.trim().startsWith('npm') && l.includes('build:lib'));
|
||||
|
||||
assert.ok(
|
||||
buildInvocations.length > 0,
|
||||
'Step 5.5 must contain at least one npm build:lib invocation (the dev-checkout fallback)'
|
||||
);
|
||||
for (const line of buildInvocations) {
|
||||
assert.match(
|
||||
line,
|
||||
/npm\s+--prefix\s+"?\$?\{?_?GSD_RT/,
|
||||
`build:lib must be pinned with \`npm --prefix "$_GSD_RT"\` so it never runs the consuming project's scripts — offending line: ${line.trim()}`
|
||||
);
|
||||
}
|
||||
|
||||
// The build must be gated behind a verified GSD source checkout (tsconfig.build.json present),
|
||||
// so it cannot fire inside a plain consumer repo where the artifact merely happens to be absent.
|
||||
assert.match(
|
||||
block,
|
||||
/tsconfig\.build\.json/,
|
||||
'Step 5.5 must gate the build behind a GSD source checkout (e.g. test -f "$_GSD_RT/tsconfig.build.json")'
|
||||
);
|
||||
});
|
||||
|
||||
// Test E (review #4 High): the engine's fail-closed exit(2) must NOT be swallowed by command
|
||||
// substitution. A bare `COVERAGE=$(node "$EDGE_PROBE_JS" ...)` discards the exit status, leaves
|
||||
// $COVERAGE empty on an invalid-shapes failure, and lets the workflow proceed into prose
|
||||
// re-derivation — fail-OPEN at the very boundary the engine validation exists to protect.
|
||||
test('review #4 High: Step 5.5 exit-checks the engine capture and validates the report (no fail-open)', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
|
||||
// The engine capture must be FATAL — guarded by `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)`
|
||||
// (or an explicit exit-status check) that exits non-zero on failure.
|
||||
assert.match(
|
||||
block,
|
||||
/if\s+!\s+COVERAGE=\$\(node\s+"\$EDGE_PROBE_JS"/,
|
||||
'Step 5.5 must exit-check the engine invocation (e.g. `if ! COVERAGE=$(node "$EDGE_PROBE_JS" …)`) — a bare command substitution swallows the engine exit code and fails open'
|
||||
);
|
||||
|
||||
// And the captured report must be validated as JSON before the resolution loop consumes it
|
||||
// (guards against an exit-0-but-garbage capture).
|
||||
assert.match(
|
||||
block,
|
||||
/(COVERAGE[\s\S]{0,500}JSON\.parse)|(JSON\.parse[\s\S]{0,500}\$COVERAGE)/,
|
||||
'Step 5.5 must validate $COVERAGE parses as JSON before use (guard against exit-0-but-malformed output)'
|
||||
);
|
||||
});
|
||||
|
||||
// Test F (adversarial review): a report with ZERO applicable edges across all requirements is
|
||||
// the likely-classification-miss fail-open (same shape as an invalid shape yielding applicable:0).
|
||||
// Step 5.5 must surface it, not silently emit a green empty ## Edge Coverage section.
|
||||
test('adversarial review: Step 5.5 guards a zero-applicable coverage report', () => {
|
||||
const content = readSpecPhase();
|
||||
const block = extractStep55Block(content);
|
||||
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
|
||||
assert.match(
|
||||
block,
|
||||
/coverage\.applicable/,
|
||||
'Step 5.5 must read coverage.applicable and guard the zero-applicable case (warn/confirm, not silently proceed)'
|
||||
);
|
||||
});
|
||||
441
tests/edge-probe.test.cjs
Normal file
441
tests/edge-probe.test.cjs
Normal file
@@ -0,0 +1,441 @@
|
||||
/**
|
||||
* Edge-probe reference core unit tests.
|
||||
*
|
||||
* Asserts the LOCKED export surface of the spec-completeness edge-probe against
|
||||
* the BUILT artifact (`gsd-core/bin/lib/edge-probe.cjs`), which
|
||||
* `npm run build:lib` (run by pretest) emits from `src/edge-probe.cts`.
|
||||
*
|
||||
* Post ADR-550 Decision 7: the generic resolution model lives in `probe-core`;
|
||||
* edge-probe is its first adapter (shapes/TAXONOMY/proposeEdges + the
|
||||
* `{explicit, backstop}` verification validators). The resolution model is the
|
||||
* status×verification re-cut: `status: resolved | dismissed | unresolved` ×
|
||||
* `verification: explicit | backstop`. `covered`/`backstop` are no longer status
|
||||
* values — `covered → {resolved, explicit}`, `backstop → {resolved, backstop}`.
|
||||
*/
|
||||
'use strict';
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const path = require('node:path');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const { execFileSync, spawnSync } = require('node:child_process');
|
||||
const { cleanup } = require('./helpers.cjs');
|
||||
|
||||
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs');
|
||||
const ep = require(BUILT_SCRIPT);
|
||||
|
||||
describe('edge-probe: classifyShape', () => {
|
||||
test('detects numeric-range from rounding/threshold cues', () => {
|
||||
assert.deepEqual(ep.classifyShape('Round a number to N decimal places').sort(),
|
||||
['numeric-range']);
|
||||
});
|
||||
test('detects collection from interval/merge cues', () => {
|
||||
const shapes = ep.classifyShape('Merge a list of overlapping intervals');
|
||||
assert.ok(shapes.includes('collection'));
|
||||
});
|
||||
test('detects text from truncate/string cues', () => {
|
||||
const shapes = ep.classifyShape('Truncate a string to a maximum length');
|
||||
assert.ok(shapes.includes('text'));
|
||||
});
|
||||
test('returns [] when no cue matches', () => {
|
||||
assert.deepEqual(ep.classifyShape('Display the company logo'), []);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: TAXONOMY + applicableCategories', () => {
|
||||
test('TAXONOMY has the 8 documented categories in order', () => {
|
||||
assert.deepEqual(ep.TAXONOMY.map((c) => c.id),
|
||||
['boundary', 'adjacency', 'empty', 'encoding', 'ordering', 'precision', 'idempotency', 'concurrency']);
|
||||
});
|
||||
test('every category has name, shapes[], probe', () => {
|
||||
for (const c of ep.TAXONOMY) {
|
||||
assert.equal(typeof c.name, 'string');
|
||||
assert.ok(Array.isArray(c.shapes) && c.shapes.length >= 1);
|
||||
assert.equal(typeof c.probe, 'string');
|
||||
}
|
||||
});
|
||||
test('numeric-range raises boundary + precision only', () => {
|
||||
assert.deepEqual(ep.applicableCategories(['numeric-range']).sort(),
|
||||
['boundary', 'precision']);
|
||||
});
|
||||
test('collection raises adjacency, empty, ordering', () => {
|
||||
assert.deepEqual(ep.applicableCategories(['collection']).sort(),
|
||||
['adjacency', 'empty', 'ordering']);
|
||||
});
|
||||
test('text raises empty + encoding', () => {
|
||||
assert.deepEqual(ep.applicableCategories(['text']).sort(),
|
||||
['empty', 'encoding']);
|
||||
});
|
||||
test('no shapes raises nothing', () => {
|
||||
assert.deepEqual(ep.applicableCategories([]), []);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: proposeEdges', () => {
|
||||
test('rounding requirement proposes boundary + precision, all unresolved (verification null)', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places' });
|
||||
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
|
||||
for (const e of edges) {
|
||||
assert.equal(e.requirement_id, 'R1');
|
||||
assert.equal(e.status, 'unresolved');
|
||||
assert.equal(e.verification, null);
|
||||
assert.equal(e.resolution, null);
|
||||
assert.equal(e.reason, null);
|
||||
assert.equal(typeof e.probe, 'string');
|
||||
}
|
||||
});
|
||||
test('authored shapes override prose classification', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque label', shapes: ['collection'] });
|
||||
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: validateResolution', () => {
|
||||
test('rejects an unknown status', () => {
|
||||
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'maybe' }),
|
||||
/invalid status/i);
|
||||
});
|
||||
test('rejects a former covered status (re-cut: covered is no longer a status)', () => {
|
||||
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'covered', resolution: 'x' }),
|
||||
/invalid status/i);
|
||||
});
|
||||
test('rejects dismissed without a reason', () => {
|
||||
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: '' }),
|
||||
/dismissed requires a reason/i);
|
||||
});
|
||||
test('accepts dismissed with a reason', () => {
|
||||
assert.equal(ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'dismissed', reason: 'bounded enum' }), true);
|
||||
});
|
||||
test('rejects resolved with a missing verification tier', () => {
|
||||
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', resolution: 'AC' }),
|
||||
/verification/i);
|
||||
});
|
||||
test('rejects resolved with a verification tier outside {explicit, backstop}', () => {
|
||||
assert.throws(() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'judgment', resolution: 'AC' }),
|
||||
/invalid verification/i);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: analyzeCoverage', () => {
|
||||
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
|
||||
test('with no resolutions, every applicable edge is unresolved (byVerification zeroed)', () => {
|
||||
const rep = ep.analyzeCoverage(reqs, []);
|
||||
assert.deepEqual(rep.coverage, { applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 } });
|
||||
});
|
||||
test('merges a resolved/explicit resolution and counts it resolved', () => {
|
||||
const rep = ep.analyzeCoverage(reqs, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
|
||||
]);
|
||||
const adj = rep.items.find((i) => i.category === 'adjacency');
|
||||
assert.equal(adj.status, 'resolved');
|
||||
assert.equal(adj.verification, 'explicit');
|
||||
assert.equal(adj.resolution, 'AC#6: touching intervals merge');
|
||||
assert.equal(rep.coverage.resolved, 1);
|
||||
assert.equal(rep.coverage.unresolved, 2);
|
||||
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 });
|
||||
});
|
||||
test('throws if a resolution is invalid (dismissed w/o reason)', () => {
|
||||
assert.throws(() => ep.analyzeCoverage(reqs, [
|
||||
{ requirement_id: 'R1', category: 'empty', status: 'dismissed' },
|
||||
]), /dismissed requires a reason/i);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: CLI (built artifact)', () => {
|
||||
test('reads a requirements file and prints a coverage report as JSON', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-'));
|
||||
const reqPath = path.join(dir, 'requirements.json');
|
||||
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
|
||||
const out = execFileSync('node', [BUILT_SCRIPT, reqPath], { encoding: 'utf8' });
|
||||
const rep = JSON.parse(out);
|
||||
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
|
||||
});
|
||||
test('with no args exits with status 2 (assert on exit code, not stderr prose)', () => {
|
||||
let status;
|
||||
try {
|
||||
execFileSync('node', [BUILT_SCRIPT], { stdio: 'pipe' });
|
||||
status = 0;
|
||||
} catch (error) {
|
||||
status = error.status;
|
||||
}
|
||||
assert.equal(status, 2);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: CLI JSON.parse error handling (RR-10)', () => {
|
||||
test('invalid requirements JSON exits with status 2 (handled error, not uncaught throw)', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
|
||||
const badJson = path.join(dir, 'bad-req.json');
|
||||
fs.writeFileSync(badJson, 'not valid json {{{');
|
||||
try {
|
||||
const r = spawnSync(process.execPath, [BUILT_SCRIPT, badJson], { stdio: 'pipe' });
|
||||
assert.equal(r.status, 2);
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
test('invalid resolutions JSON exits with status 2 (handled error, not uncaught throw)', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
|
||||
const goodReq = path.join(dir, 'req.json');
|
||||
const badRes = path.join(dir, 'bad-res.json');
|
||||
fs.writeFileSync(goodReq, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
|
||||
fs.writeFileSync(badRes, 'not valid json {{{');
|
||||
try {
|
||||
const r = spawnSync(process.execPath, [BUILT_SCRIPT, goodReq, badRes], { stdio: 'pipe' });
|
||||
assert.equal(r.status, 2);
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
test('valid requirements file exits 0 and stdout is parseable JSON', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-rr10-'));
|
||||
const reqPath = path.join(dir, 'req.json');
|
||||
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number to N decimal places' }]));
|
||||
try {
|
||||
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe', encoding: 'utf8' });
|
||||
assert.equal(r.status, 0);
|
||||
const rep = JSON.parse(r.stdout);
|
||||
assert.deepEqual(rep.coverage, { applicable: 2, resolved: 0, unresolved: 2, byVerification: { explicit: 0, backstop: 0 } });
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: proposeEdges — empty-shapes override (RR-06)', () => {
|
||||
test('shapes: [] returns zero edges (explicit empty-shapes override)', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] });
|
||||
assert.deepEqual(edges, []);
|
||||
});
|
||||
test('absent shapes key classifies from prose (no override)', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R1', text: 'merge intervals' });
|
||||
assert.ok(edges.length > 0, 'should classify collection edges from prose');
|
||||
});
|
||||
test('shapes: [collection] overrides prose and proposes collection categories', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R9', text: 'opaque text with no cues', shapes: ['collection'] });
|
||||
assert.deepEqual(edges.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: proposeEdges — invalid authored shapes fail closed (re-review #3 High)', () => {
|
||||
// A non-empty but INVALID shapes array must NOT silently suppress every probe.
|
||||
// shapes:['numeric'] (typo for the locked 'numeric-range') previously passed
|
||||
// Array.isArray, matched no category, and returned applicable:0 — failing OPEN.
|
||||
test('rejects an unknown shape value (typo for a locked shape)', () => {
|
||||
assert.throws(
|
||||
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric'] }),
|
||||
/invalid shape/i,
|
||||
);
|
||||
});
|
||||
test('rejects a mixed array where one entry is invalid', () => {
|
||||
assert.throws(
|
||||
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: ['numeric-range', 'bogus'] }),
|
||||
/invalid shape/i,
|
||||
);
|
||||
});
|
||||
test('rejects a non-string shape entry', () => {
|
||||
assert.throws(
|
||||
() => ep.proposeEdges({ id: 'R1', text: 'Round a number', shapes: [42] }),
|
||||
/invalid shape/i,
|
||||
);
|
||||
});
|
||||
test('analyzeCoverage propagates the invalid-shape throw', () => {
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]),
|
||||
/invalid shape/i,
|
||||
);
|
||||
});
|
||||
test('CLI exits 2 (handled) on an invalid authored shape, not an uncaught trace', () => {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'edge-probe-shape-'));
|
||||
const reqPath = path.join(dir, 'req.json');
|
||||
fs.writeFileSync(reqPath, JSON.stringify([{ id: 'R1', text: 'Round a number', shapes: ['numeric'] }]));
|
||||
try {
|
||||
const r = spawnSync(process.execPath, [BUILT_SCRIPT, reqPath], { stdio: 'pipe' });
|
||||
assert.equal(r.status, 2);
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
});
|
||||
test('a valid locked shape still proposes its categories (no false rejection)', () => {
|
||||
const edges = ep.proposeEdges({ id: 'R1', text: 'opaque', shapes: ['numeric-range'] });
|
||||
assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']);
|
||||
});
|
||||
test('shapes: [] remains a valid zero-edge override (RR-06 intact)', () => {
|
||||
assert.deepEqual(ep.proposeEdges({ id: 'R1', text: 'merge intervals', shapes: [] }), []);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: input validation & orphan-resolution rejection (adversarial review)', () => {
|
||||
// HIGH: a resolution whose (requirement_id, category) matches no proposed edge — a typo'd
|
||||
// category or a non-applicable one — was silently DROPPED, so an author who typos `precison`
|
||||
// sees the precision edge as still-unresolved with no error (a confirmed money-rounding exploit).
|
||||
test('rejects an orphan resolution (typo category — no matching proposed edge)', () => {
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage(
|
||||
[{ id: 'R1', text: 'Round a number to N decimal places' }],
|
||||
[{ requirement_id: 'R1', category: 'precison', status: 'resolved', verification: 'explicit', resolution: 'AC: precision handled' }],
|
||||
),
|
||||
/unknown resolution|no matching proposed edge/i,
|
||||
);
|
||||
});
|
||||
test('rejects a resolution for a valid-but-non-applicable category', () => {
|
||||
// 'encoding' is a real taxonomy id but applies to text, not the numeric-range requirement.
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage(
|
||||
[{ id: 'R1', text: 'Round a number to N decimal places' }],
|
||||
[{ requirement_id: 'R1', category: 'encoding', status: 'resolved', verification: 'explicit', resolution: 'AC' }],
|
||||
),
|
||||
/unknown resolution|no matching proposed edge/i,
|
||||
);
|
||||
});
|
||||
test('a matching resolution still resolves (no false orphan rejection)', () => {
|
||||
const rep = ep.analyzeCoverage(
|
||||
[{ id: 'R1', text: 'Round a number to N decimal places' }],
|
||||
[{ requirement_id: 'R1', category: 'precision', status: 'resolved', verification: 'explicit', resolution: 'AC: precision tested' }],
|
||||
);
|
||||
assert.equal(rep.coverage.resolved, 1);
|
||||
});
|
||||
test('rejects requirements that is not an array', () => {
|
||||
assert.throws(() => ep.analyzeCoverage('nope'), /requirements must be an array/i);
|
||||
});
|
||||
test('rejects a duplicate requirement id', () => {
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage([{ id: 'R1', text: 'a' }, { id: 'R1', text: 'b' }]),
|
||||
/duplicate requirement/i,
|
||||
);
|
||||
});
|
||||
test('rejects a truthy non-array shapes (string instead of array)', () => {
|
||||
// A bare string `shapes: "numeric-range"` previously fell through to prose classification,
|
||||
// silently ignoring the authored override instead of honoring or rejecting it.
|
||||
assert.throws(
|
||||
() => ep.proposeEdges({ id: 'R1', text: 'x', shapes: 'numeric-range' }),
|
||||
/shapes must be an array/i,
|
||||
);
|
||||
});
|
||||
test('rejects a missing requirement id', () => {
|
||||
assert.throws(() => ep.proposeEdges({ text: 'x' }), /requirement id must be a non-empty string/i);
|
||||
});
|
||||
test('rejects an empty requirement id', () => {
|
||||
assert.throws(() => ep.proposeEdges({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i);
|
||||
});
|
||||
test('rejects a non-string requirement text', () => {
|
||||
assert.throws(() => ep.proposeEdges({ id: 'R1', text: 42 }), /text must be a string/i);
|
||||
});
|
||||
test('rejects a missing requirement text when no shapes override (M2 fail-open)', () => {
|
||||
// Without text or an authored shape, prose classification yields zero shapes → zero edges →
|
||||
// the requirement is silently DROPPED from coverage with no signal — the exact fail-open this
|
||||
// feature exists to eliminate. The edge adapter's `text` is required, so reject it.
|
||||
assert.throws(
|
||||
() => ep.proposeEdges({ id: 'R1' }),
|
||||
/text must be a non-empty string when no shapes override/i,
|
||||
);
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage([{ id: 'R1' }]),
|
||||
/text must be a non-empty string when no shapes override/i,
|
||||
);
|
||||
});
|
||||
test('rejects an empty/whitespace requirement text when no shapes override (M2)', () => {
|
||||
assert.throws(() => ep.proposeEdges({ id: 'R1', text: '' }), /text must be a non-empty string when no shapes override/i);
|
||||
assert.throws(() => ep.proposeEdges({ id: 'R1', text: ' ' }), /text must be a non-empty string when no shapes override/i);
|
||||
});
|
||||
test('allows missing/empty text WHEN an explicit shapes override is provided (M2 legitimate path)', () => {
|
||||
// An authored `shapes` array (including `[]` for "no applicable categories") opts out of prose
|
||||
// classification, so `text` is not required — this must remain valid.
|
||||
assert.deepEqual(ep.proposeEdges({ id: 'R1', shapes: [] }), []);
|
||||
const edges = ep.proposeEdges({ id: 'R1', shapes: ['numeric-range'] });
|
||||
assert.ok(edges.length > 0, 'an explicit shape override must still propose edges without text');
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: validateResolution — explicit-needs-resolution (RR-07, re-cut)', () => {
|
||||
test('rejects resolved/explicit with empty resolution string', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: '' }),
|
||||
/explicit requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('rejects resolved/explicit with whitespace-only resolution', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: ' ' }),
|
||||
/explicit requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('rejects resolved/explicit with missing resolution', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit' }),
|
||||
/explicit requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('accepts resolved/explicit with a non-empty resolution', () => {
|
||||
assert.equal(
|
||||
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC#3: boundary tested in suite' }),
|
||||
true,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: validateResolution — backstop-needs-resolution (RR-07 follow-up, re-cut)', () => {
|
||||
test('rejects resolved/backstop with empty resolution string', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: '' }),
|
||||
/backstop requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('rejects resolved/backstop with whitespace-only resolution', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: ' ' }),
|
||||
/backstop requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('rejects resolved/backstop with missing resolution', () => {
|
||||
assert.throws(
|
||||
() => ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop' }),
|
||||
/backstop requires a resolution/i,
|
||||
);
|
||||
});
|
||||
test('accepts resolved/backstop with a non-empty resolution note', () => {
|
||||
assert.equal(
|
||||
ep.validateResolution({ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'backstop', resolution: 'held-out: covered by integration fuzz suite' }),
|
||||
true,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: analyzeCoverage — duplicate rejection (RR-09)', () => {
|
||||
const reqs = [{ id: 'R1', text: 'Merge a list of overlapping intervals' }];
|
||||
test('rejects duplicate (requirement_id, category) resolution', () => {
|
||||
assert.throws(
|
||||
() => ep.analyzeCoverage(reqs, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' },
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' },
|
||||
]),
|
||||
/duplicate resolution/i,
|
||||
);
|
||||
});
|
||||
test('distinct pairs still analyze without throwing', () => {
|
||||
// mirrors fixture 06-resolved-mixed
|
||||
assert.doesNotThrow(() => ep.analyzeCoverage(reqs, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
|
||||
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', resolution: null, reason: 'output is canonically sorted; no tie possible' },
|
||||
]));
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge-probe: golden fixtures', () => {
|
||||
const root = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures');
|
||||
const fixtures = fs.readdirSync(root).filter((d) =>
|
||||
fs.statSync(path.join(root, d)).isDirectory());
|
||||
assert.ok(fixtures.length >= 6, 'expected at least 6 fixtures');
|
||||
for (const name of fixtures) {
|
||||
test(`fixture ${name} matches its golden coverage`, () => {
|
||||
const dir = path.join(root, name);
|
||||
const reqs = JSON.parse(fs.readFileSync(path.join(dir, 'requirements.json'), 'utf8'));
|
||||
const resPath = path.join(dir, 'resolutions.json');
|
||||
const res = fs.existsSync(resPath) ? JSON.parse(fs.readFileSync(resPath, 'utf8')) : [];
|
||||
const expected = JSON.parse(fs.readFileSync(path.join(dir, 'expected-coverage.json'), 'utf8'));
|
||||
assert.deepEqual(ep.analyzeCoverage(reqs, res), expected);
|
||||
});
|
||||
}
|
||||
});
|
||||
167
tests/probe-core.property.test.cjs
Normal file
167
tests/probe-core.property.test.cjs
Normal file
@@ -0,0 +1,167 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Property-based tests for probe-core.cjs (ADR-550 Decision 7).
|
||||
*
|
||||
* Module: gsd-core/bin/lib/probe-core.cjs (generated from src/probe-core.cts)
|
||||
* Exercised: analyzeCoverage(items, resolutions?, validators) — the generic
|
||||
* merge/rollup/orphan-reject engine shared by the edge probe and the #644
|
||||
* prohibition probe.
|
||||
*
|
||||
* trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
|
||||
* transformation/rollup module (items × resolutions → CoverageReport) — exactly the
|
||||
* class the predicate covers. The example-based suite (tests/probe-core.test.cjs)
|
||||
* pins specific scenarios; these properties pin the algebraic invariants that must
|
||||
* hold for EVERY valid scenario.
|
||||
*
|
||||
* Properties tested:
|
||||
* (a) closed-set identity: applicable === resolved + unresolved (and === items.length)
|
||||
* (b) byVerification sums ≤ resolved (dismissed is closed but unverified)
|
||||
* (c) per-tier byVerification ≤ resolved, and only `resolved`-status items are counted
|
||||
* (d) determinism: same input → identical CoverageReport (stable rollup)
|
||||
* (e) orphan rejection is stable: a resolution matching no proposed item always throws
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const path = require('node:path');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
|
||||
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
|
||||
const pc = require(BUILT_SCRIPT);
|
||||
|
||||
// The same representative validators bundle the edge adapter injects (see
|
||||
// tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe.
|
||||
const VALIDATORS = {
|
||||
categories: ['adjacency', 'empty', 'ordering'],
|
||||
verification: ['explicit', 'backstop'],
|
||||
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
|
||||
};
|
||||
|
||||
function bareItem(requirement_id, category) {
|
||||
return {
|
||||
requirement_id,
|
||||
category,
|
||||
status: 'unresolved',
|
||||
verification: null,
|
||||
resolution: null,
|
||||
reason: null,
|
||||
probe: `probe-for-${category}`,
|
||||
};
|
||||
}
|
||||
|
||||
const catArb = fc.constantFrom(...VALIDATORS.categories);
|
||||
const idArb = fc.constantFrom('R1', 'R2', 'R3', 'R4', 'R5');
|
||||
const keyArb = fc.record({ requirement_id: idArb, category: catArb });
|
||||
// Unique (requirement_id, category) keys — the merge keys analyzeCoverage maps on.
|
||||
const keyOf = (k) => `${k.requirement_id}::${k.category}`;
|
||||
const uniqueKeysArb = fc.uniqueArray(keyArb, { selector: keyOf, minLength: 0, maxLength: 12 });
|
||||
|
||||
// Each unique item key gets one resolution disposition. Resolution text/reason use fixed
|
||||
// non-empty literals — the counting invariants are independent of their content, and this
|
||||
// keeps the generator off the validateResolution rejection paths (which the example suite
|
||||
// already covers exhaustively).
|
||||
const DISPOSITIONS = ['none', 'resolved-explicit', 'resolved-backstop', 'dismissed', 'unresolved'];
|
||||
|
||||
function resolutionFor(k, disposition) {
|
||||
const base = { requirement_id: k.requirement_id, category: k.category };
|
||||
switch (disposition) {
|
||||
case 'resolved-explicit':
|
||||
return { ...base, status: 'resolved', verification: 'explicit', resolution: 'AC#1' };
|
||||
case 'resolved-backstop':
|
||||
return { ...base, status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' };
|
||||
case 'dismissed':
|
||||
return { ...base, status: 'dismissed', reason: 'bounded enum — not applicable' };
|
||||
case 'unresolved':
|
||||
return { ...base, status: 'unresolved' };
|
||||
default: // 'none' — author left no resolution; item rolls up verbatim (bare unresolved)
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
// A fully valid scenario: unique items (all bare-unresolved) + a per-item resolution choice.
|
||||
const scenarioArb = uniqueKeysArb.chain((keys) =>
|
||||
fc.tuple(...keys.map(() => fc.constantFrom(...DISPOSITIONS))).map((choices) => {
|
||||
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
|
||||
const resolutions = [];
|
||||
keys.forEach((k, i) => {
|
||||
const r = resolutionFor(k, choices[i]);
|
||||
if (r) resolutions.push(r);
|
||||
});
|
||||
return { items, resolutions };
|
||||
}),
|
||||
);
|
||||
|
||||
describe('probe-core property: analyzeCoverage algebraic invariants', () => {
|
||||
test('(a) closed-set identity: applicable === resolved + unresolved === items.length', () => {
|
||||
fc.assert(
|
||||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||||
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||||
return (
|
||||
coverage.applicable === coverage.resolved + coverage.unresolved &&
|
||||
coverage.applicable === items.length
|
||||
);
|
||||
}),
|
||||
);
|
||||
});
|
||||
|
||||
test('(b) sum(byVerification) ≤ resolved — dismissed counts closed but unverified', () => {
|
||||
fc.assert(
|
||||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||||
const { coverage } = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||||
const verifiedTotal = Object.values(coverage.byVerification).reduce((a, b) => a + b, 0);
|
||||
return verifiedTotal <= coverage.resolved && verifiedTotal >= 0;
|
||||
}),
|
||||
);
|
||||
});
|
||||
|
||||
test('(c) byVerification only counts resolved-status items, and matches a direct recount', () => {
|
||||
fc.assert(
|
||||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||||
const rep = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||||
for (const tier of VALIDATORS.verification) {
|
||||
const recount = rep.items.filter((i) => i.status === 'resolved' && i.verification === tier).length;
|
||||
if (rep.coverage.byVerification[tier] !== recount) return false;
|
||||
if (rep.coverage.byVerification[tier] > rep.coverage.resolved) return false;
|
||||
}
|
||||
return true;
|
||||
}),
|
||||
);
|
||||
});
|
||||
|
||||
test('(d) determinism: identical inputs produce an identical CoverageReport', () => {
|
||||
fc.assert(
|
||||
fc.property(scenarioArb, ({ items, resolutions }) => {
|
||||
const a = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||||
const b = pc.analyzeCoverage(items, resolutions, VALIDATORS);
|
||||
return JSON.stringify(a) === JSON.stringify(b);
|
||||
}),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('probe-core property: orphan rejection is stable', () => {
|
||||
// An orphan resolution carries an id ('Z9') that no generated item ever uses, so its
|
||||
// (requirement_id, category) key never matches a proposed item. The resolution is itself
|
||||
// structurally VALID (a bare unresolved), so it clears validateResolution and reaches the
|
||||
// orphan-reject guard — isolating that guard from input-validation throws.
|
||||
const orphanScenarioArb = fc.record({
|
||||
keys: uniqueKeysArb,
|
||||
orphanCategory: catArb,
|
||||
});
|
||||
|
||||
test('(e) a resolution matching no proposed item always throws', () => {
|
||||
fc.assert(
|
||||
fc.property(orphanScenarioArb, ({ keys, orphanCategory }) => {
|
||||
const items = keys.map((k) => bareItem(k.requirement_id, k.category));
|
||||
const orphan = { requirement_id: 'Z9', category: orphanCategory, status: 'unresolved' };
|
||||
let threwForOrphan = false;
|
||||
try {
|
||||
pc.analyzeCoverage(items, [orphan], VALIDATORS);
|
||||
} catch (e) {
|
||||
threwForOrphan = /unknown resolution|no matching proposed item/i.test(e.message);
|
||||
}
|
||||
return threwForOrphan;
|
||||
}),
|
||||
);
|
||||
});
|
||||
});
|
||||
335
tests/probe-core.test.cjs
Normal file
335
tests/probe-core.test.cjs
Normal file
@@ -0,0 +1,335 @@
|
||||
/**
|
||||
* probe-core reference-model unit tests (ADR-550 Decision 7).
|
||||
*
|
||||
* probe-core is the GENERIC spec-phase probe resolution model extracted from the
|
||||
* edge-probe (the first adapter): the resolution lifecycle, the two-axis
|
||||
* status×verification re-cut, `validateResolution`/`validateRequirement`, the
|
||||
* `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject
|
||||
* engine, the `byVerification` rollup, and the `runProbeCli` I/O scaffold.
|
||||
*
|
||||
* Asserts the LOCKED export surface against the BUILT artifact
|
||||
* (`gsd-core/bin/lib/probe-core.cjs`), which `npm run build:lib` (run by pretest /
|
||||
* the run-tests sentinel) emits from `src/probe-core.cts`.
|
||||
*
|
||||
* The injected runtime validators are the enforcement contract (ADR-550 #5): the
|
||||
* CLI runs over JSON where TS types are erased, so `analyzeCoverage` is told its
|
||||
* probe's closed vocabularies — `{ categories, verification, requiredFieldsByVerification }`
|
||||
* — rather than relying on the type system. These tests pin the validators' behavior.
|
||||
*/
|
||||
'use strict';
|
||||
process.env.GSD_TEST_MODE = '1';
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const path = require('node:path');
|
||||
|
||||
const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs');
|
||||
const pc = require(BUILT_SCRIPT);
|
||||
|
||||
// A representative validators bundle — the shape the edge adapter injects, used here
|
||||
// to exercise the generic engine independent of any one probe.
|
||||
const VALIDATORS = {
|
||||
categories: ['adjacency', 'empty', 'ordering'],
|
||||
verification: ['explicit', 'backstop'],
|
||||
requiredFieldsByVerification: { explicit: ['resolution'], backstop: ['resolution'] },
|
||||
};
|
||||
|
||||
function item(category, overrides = {}) {
|
||||
return {
|
||||
requirement_id: 'R1',
|
||||
category,
|
||||
status: 'unresolved',
|
||||
verification: null,
|
||||
resolution: null,
|
||||
reason: null,
|
||||
probe: `probe-for-${category}`,
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
const UNRESOLVED_ITEMS = [item('adjacency'), item('empty'), item('ordering')];
|
||||
|
||||
describe('probe-core: VALID_STATUS is the re-cut lifecycle enum', () => {
|
||||
test('exposes exactly resolved | dismissed | unresolved (no covered/backstop)', () => {
|
||||
assert.deepEqual([...ep_sorted(pc.VALID_STATUS)], ['dismissed', 'resolved', 'unresolved']);
|
||||
assert.ok(!pc.VALID_STATUS.includes('covered'), 'covered must not survive the re-cut as a status');
|
||||
assert.ok(!pc.VALID_STATUS.includes('backstop'), 'backstop must not survive the re-cut as a status');
|
||||
});
|
||||
});
|
||||
|
||||
function ep_sorted(arr) {
|
||||
return [...arr].sort();
|
||||
}
|
||||
|
||||
describe('probe-core: validateResolution (status×verification)', () => {
|
||||
const v = (r) => pc.validateResolution(r, VALIDATORS);
|
||||
test('rejects an unknown status', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'maybe' }), /invalid status/i);
|
||||
});
|
||||
test('rejects a former covered status (re-cut: no longer a valid status)', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'covered', resolution: 'x' }), /invalid status/i);
|
||||
});
|
||||
test('rejects dismissed without a reason', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: '' }), /dismissed requires a reason/i);
|
||||
});
|
||||
test('accepts dismissed with a reason', () => {
|
||||
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' }), true);
|
||||
});
|
||||
test('rejects resolved with a missing verification tier', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', resolution: 'AC' }), /verification/i);
|
||||
});
|
||||
test('rejects resolved with an unknown verification tier', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'judgment', resolution: 'AC' }), /invalid verification/i);
|
||||
});
|
||||
test('rejects resolved/explicit with empty resolution text', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: ' ' }), /explicit requires a resolution/i);
|
||||
});
|
||||
test('rejects resolved/backstop with missing resolution note', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop' }), /backstop requires a resolution/i);
|
||||
});
|
||||
test('accepts resolved/explicit with a resolution', () => {
|
||||
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' }), true);
|
||||
});
|
||||
test('accepts resolved/backstop with a resolution note', () => {
|
||||
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'backstop', resolution: 'held-out PBT suite' }), true);
|
||||
});
|
||||
// Re-review #5 Medium — fail closed across the FULL status×verification model, not just
|
||||
// `resolved`. The invariant (probe-core header): verification is null unless status is
|
||||
// resolved. A dismissed/unresolved resolution carrying a tier otherwise merges verbatim
|
||||
// (analyzeCoverage line ~189), breaking the model for the second adapter (#644) that
|
||||
// inherits this seam.
|
||||
test('rejects a dismissed resolution carrying a verification tier (null unless resolved)', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'n/a', verification: 'explicit' }), /verification must be null/i);
|
||||
});
|
||||
test('rejects an unresolved resolution carrying a verification tier', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', verification: 'backstop' }), /verification must be null/i);
|
||||
});
|
||||
// An unresolved resolution is an UNACTED item — a populated resolution/reason payload is an
|
||||
// authoring mistake (the author meant resolved/dismissed) that today is silently dropped into
|
||||
// the unresolved count with no error pointing at it. Reject the payload.
|
||||
test('rejects an unresolved resolution carrying a resolution payload', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', resolution: 'AC#1' }), /unresolved must not carry/i);
|
||||
});
|
||||
test('rejects an unresolved resolution carrying a reason payload', () => {
|
||||
assert.throws(() => v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved', reason: 'because' }), /unresolved must not carry/i);
|
||||
});
|
||||
test('accepts a bare unresolved resolution (no payload — a harmless no-op merge)', () => {
|
||||
assert.equal(v({ requirement_id: 'R1', category: 'adjacency', status: 'unresolved' }), true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('probe-core: validateRequirement (generic id/text only)', () => {
|
||||
test('rejects a missing id', () => {
|
||||
assert.throws(() => pc.validateRequirement({ text: 'x' }), /requirement id must be a non-empty string/i);
|
||||
});
|
||||
test('rejects an empty id', () => {
|
||||
assert.throws(() => pc.validateRequirement({ id: ' ', text: 'x' }), /requirement id must be a non-empty string/i);
|
||||
});
|
||||
test('rejects a non-string text', () => {
|
||||
assert.throws(() => pc.validateRequirement({ id: 'R1', text: 42 }), /text must be a string/i);
|
||||
});
|
||||
test('accepts a valid requirement', () => {
|
||||
assert.doesNotThrow(() => pc.validateRequirement({ id: 'R1', text: 'a testable statement' }));
|
||||
});
|
||||
});
|
||||
|
||||
describe('probe-core: analyzeCoverage (merge · rollup · byVerification)', () => {
|
||||
test('no resolutions → every item unresolved; resolved 0; byVerification zeroed per tier', () => {
|
||||
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [], VALIDATORS);
|
||||
assert.deepEqual(rep.coverage, {
|
||||
applicable: 3, resolved: 0, unresolved: 3, byVerification: { explicit: 0, backstop: 0 },
|
||||
});
|
||||
});
|
||||
test('merges a resolved/explicit resolution and counts byVerification.explicit', () => {
|
||||
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6: touching intervals merge' },
|
||||
], VALIDATORS);
|
||||
const adj = rep.items.find((i) => i.category === 'adjacency');
|
||||
assert.equal(adj.status, 'resolved');
|
||||
assert.equal(adj.verification, 'explicit');
|
||||
assert.equal(adj.resolution, 'AC#6: touching intervals merge');
|
||||
assert.equal(rep.coverage.resolved, 1);
|
||||
assert.equal(rep.coverage.unresolved, 2);
|
||||
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 0 });
|
||||
});
|
||||
test('a dismissed item counts toward coverage.resolved (closed set) but NOT byVerification', () => {
|
||||
// coverage.resolved preserves the pre-re-cut "closed" semantic = applicable - unresolved
|
||||
// (covered + dismissed + backstop), per edge-probe.md and the migration contract.
|
||||
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted; no tie' },
|
||||
], VALIDATORS);
|
||||
assert.equal(rep.coverage.resolved, 1);
|
||||
assert.equal(rep.coverage.unresolved, 2);
|
||||
assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 });
|
||||
const ord = rep.items.find((i) => i.category === 'ordering');
|
||||
assert.equal(ord.status, 'dismissed');
|
||||
assert.equal(ord.verification, null);
|
||||
assert.equal(ord.reason, 'canonically sorted; no tie');
|
||||
});
|
||||
test('mixed resolved/explicit + backstop + dismissed: resolved = closed = applicable - unresolved', () => {
|
||||
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' },
|
||||
{ requirement_id: 'R1', category: 'empty', status: 'resolved', verification: 'backstop', resolution: 'held-out empty-input PBT' },
|
||||
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' },
|
||||
], VALIDATORS);
|
||||
assert.equal(rep.coverage.applicable, 3);
|
||||
assert.equal(rep.coverage.unresolved, 0);
|
||||
assert.equal(rep.coverage.resolved, 3); // 2 resolved-status + 1 dismissed
|
||||
assert.deepEqual(rep.coverage.byVerification, { explicit: 1, backstop: 1 });
|
||||
});
|
||||
test('an all-dismissed run is CLOSED but NOT affirmatively covered (byVerification is the honest gate)', () => {
|
||||
// Re-review #5 Medium — coverage.resolved is the closed set (resolved + dismissed), kept
|
||||
// count-preserved per the blessed migration contract. So an all-dismissed spec has
|
||||
// resolved === applicable while NOTHING was affirmatively resolved/backstopped. A CI gate
|
||||
// keying on `resolved === applicable` would read green here; the honest signal is
|
||||
// byVerification (all tiers zero). This test locks that distinction.
|
||||
const rep = pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'dismissed', reason: 'bounded enum' },
|
||||
{ requirement_id: 'R1', category: 'empty', status: 'dismissed', reason: 'guaranteed non-empty' },
|
||||
{ requirement_id: 'R1', category: 'ordering', status: 'dismissed', reason: 'canonically sorted' },
|
||||
], VALIDATORS);
|
||||
assert.equal(rep.coverage.unresolved, 0);
|
||||
assert.equal(rep.coverage.resolved, 3); // closed set counts the dismissals
|
||||
assert.deepEqual(rep.coverage.byVerification, { explicit: 0, backstop: 0 }); // nothing affirmatively verified
|
||||
});
|
||||
test('rejects a duplicate (requirement_id, category) resolution', () => {
|
||||
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#1' },
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#2' },
|
||||
], VALIDATORS), /duplicate resolution/i);
|
||||
});
|
||||
test('rejects an orphan resolution (no matching proposed item)', () => {
|
||||
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'boundary', status: 'resolved', verification: 'explicit', resolution: 'AC' },
|
||||
], VALIDATORS), /unknown resolution|no matching proposed/i);
|
||||
});
|
||||
test('propagates an invalid-resolution throw from validateResolution', () => {
|
||||
assert.throws(() => pc.analyzeCoverage(UNRESOLVED_ITEMS, [
|
||||
{ requirement_id: 'R1', category: 'empty', status: 'dismissed' },
|
||||
], VALIDATORS), /dismissed requires a reason/i);
|
||||
});
|
||||
test('rejects items that is not an array', () => {
|
||||
assert.throws(() => pc.analyzeCoverage('nope', [], VALIDATORS), /items must be an array/i);
|
||||
});
|
||||
test('rejects a proposed item whose category is not in validators.categories', () => {
|
||||
// Adapter self-consistency: an item carrying a category outside the probe's closed
|
||||
// vocabulary is an adapter bug, caught here rather than silently rolled up.
|
||||
assert.throws(() => pc.analyzeCoverage([item('bogus-category')], [], VALIDATORS), /unknown category/i);
|
||||
});
|
||||
|
||||
// m1: a verbatim item (no matching author resolution) is rolled up as-is, so its OWN
|
||||
// status/fields must also be validated. The edge adapter only proposes `unresolved` items, but
|
||||
// the prohibition adapter (#644) proposes LLM-generated items that arrive already populated —
|
||||
// an out-of-enum status or a `dismissed` with no reason must fail closed, not count as closed.
|
||||
test('m1: rejects a verbatim item carrying an out-of-enum status (e.g. the dropped "covered")', () => {
|
||||
assert.throws(
|
||||
() => pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [], VALIDATORS),
|
||||
/invalid status/i,
|
||||
);
|
||||
});
|
||||
test('m1: rejects a verbatim item dismissed without a reason', () => {
|
||||
assert.throws(
|
||||
() => pc.analyzeCoverage([item('adjacency', { status: 'dismissed' })], [], VALIDATORS),
|
||||
/dismissed requires a reason/i,
|
||||
);
|
||||
});
|
||||
test('m1: rejects a verbatim resolved item missing its verification tier', () => {
|
||||
assert.throws(
|
||||
() => pc.analyzeCoverage([item('adjacency', { status: 'resolved', resolution: 'AC' })], [], VALIDATORS),
|
||||
/requires a verification tier/i,
|
||||
);
|
||||
});
|
||||
test('m1: a matching resolution still governs — item validation targets VERBATIM items only', () => {
|
||||
// When a resolution matches, the item is rebuilt from the (already-validated) resolution, so a
|
||||
// pre-populated item status is irrelevant and must NOT cause a throw. Guards against the m1
|
||||
// fix over-reaching into the merge path.
|
||||
const rep = pc.analyzeCoverage([item('adjacency', { status: 'covered' })], [
|
||||
{ requirement_id: 'R1', category: 'adjacency', status: 'resolved', verification: 'explicit', resolution: 'AC#6' },
|
||||
], VALIDATORS);
|
||||
const adj = rep.items.find((i) => i.category === 'adjacency');
|
||||
assert.equal(adj.status, 'resolved');
|
||||
assert.equal(adj.verification, 'explicit');
|
||||
assert.equal(rep.coverage.resolved, 1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('probe-core: runProbeCli (generic I/O scaffold, injected io)', () => {
|
||||
const report = { items: [], coverage: { applicable: 0, resolved: 0, unresolved: 0, byVerification: {} } };
|
||||
test('no requirements path → writes usage to stderr and exits 2', () => {
|
||||
let code; let err = '';
|
||||
pc.runProbeCli(() => report, {
|
||||
usage: 'demo-probe.cjs <requirements.json> [resolutions.json]',
|
||||
argv: ['node', 'demo'], writeErr: (s) => { err += s; }, exit: (c) => { code = c; },
|
||||
});
|
||||
assert.equal(code, 2);
|
||||
assert.match(err, /usage: demo-probe\.cjs/);
|
||||
});
|
||||
test('valid requirements path → calls analyze and prints the report as JSON', () => {
|
||||
let out = '';
|
||||
pc.runProbeCli((reqs, res) => {
|
||||
assert.deepEqual(reqs, [{ id: 'R1', text: 'x' }]);
|
||||
assert.deepEqual(res, []);
|
||||
return report;
|
||||
}, {
|
||||
usage: 'demo', argv: ['node', 'demo', '/fake/req.json'],
|
||||
readFile: () => '[{"id":"R1","text":"x"}]', write: (s) => { out += s; }, exit: () => {},
|
||||
});
|
||||
assert.deepEqual(JSON.parse(out), report);
|
||||
});
|
||||
test('reads the optional resolutions file when a second path is given', () => {
|
||||
let seenRes;
|
||||
pc.runProbeCli((reqs, res) => { seenRes = res; return report; }, {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json', '/res.json'],
|
||||
readFile: (p) => (p === '/res.json' ? '[{"requirement_id":"R1"}]' : '[{"id":"R1"}]'),
|
||||
write: () => {}, exit: () => {},
|
||||
});
|
||||
assert.deepEqual(seenRes, [{ requirement_id: 'R1' }]);
|
||||
});
|
||||
test('invalid requirements JSON → exits 2 (handled, not an uncaught throw)', () => {
|
||||
let code;
|
||||
pc.runProbeCli(() => report, {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json'],
|
||||
readFile: () => 'not json {{{', writeErr: () => {}, exit: (c) => { code = c; },
|
||||
});
|
||||
assert.equal(code, 2);
|
||||
});
|
||||
test('an analyze throw → exits 2 (the engine fail-closed surfaces, never silently passes)', () => {
|
||||
let code;
|
||||
pc.runProbeCli(() => { throw new Error('boom'); }, {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json'],
|
||||
readFile: () => '[]', writeErr: () => {}, exit: (c) => { code = c; },
|
||||
});
|
||||
assert.equal(code, 2);
|
||||
});
|
||||
// Re-review #5 Low — the scaffold trusts each adapter's `as` cast for the returned report.
|
||||
// A future adapter (#644) that forgets to validate inside its closure would otherwise have a
|
||||
// structurally-broken report written as green output. Guard the report shape structurally.
|
||||
test('a structurally-invalid report from analyze → exits 2 and writes nothing (no silent malformed output)', () => {
|
||||
let code; let out = '';
|
||||
pc.runProbeCli(() => ({ nope: true }), {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json'],
|
||||
readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; },
|
||||
});
|
||||
assert.equal(code, 2);
|
||||
assert.equal(out, '');
|
||||
});
|
||||
test('a report whose coverage object carries non-numeric counts → exits 2 (partial-guard case)', () => {
|
||||
// items[] is well-formed and `coverage` IS an object, so the cheaper checks pass — this
|
||||
// exercises the numeric-count branch of the structural guard specifically.
|
||||
let code; let out = '';
|
||||
pc.runProbeCli(() => ({ items: [], coverage: { applicable: 'x', resolved: null, unresolved: undefined, byVerification: {} } }), {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json'],
|
||||
readFile: () => '[]', write: (s) => { out += s; }, writeErr: () => {}, exit: (c) => { code = c; },
|
||||
});
|
||||
assert.equal(code, 2);
|
||||
assert.equal(out, '');
|
||||
});
|
||||
test('a well-formed report still writes and does not trip the structural guard', () => {
|
||||
let out = '';
|
||||
pc.runProbeCli(() => report, {
|
||||
usage: 'demo', argv: ['node', 'demo', '/req.json'],
|
||||
readFile: () => '[]', write: (s) => { out += s; }, exit: () => {},
|
||||
});
|
||||
assert.deepEqual(JSON.parse(out), report);
|
||||
});
|
||||
});
|
||||
@@ -51,7 +51,7 @@
|
||||
"note.md": 6563,
|
||||
"pause-work.md": 13654,
|
||||
"plan-milestone-gaps.md": 11765,
|
||||
"plan-phase.md": 93135,
|
||||
"plan-phase.md": 94253,
|
||||
"plan-review-convergence.md": 22949,
|
||||
"plant-seed.md": 11741,
|
||||
"pr-branch.md": 4994,
|
||||
@@ -72,7 +72,7 @@
|
||||
"ship.md": 20896,
|
||||
"sketch-wrap-up.md": 14223,
|
||||
"sketch.md": 19960,
|
||||
"spec-phase.md": 15131,
|
||||
"spec-phase.md": 23094,
|
||||
"spike-wrap-up.md": 15092,
|
||||
"spike.md": 24517,
|
||||
"stats.md": 6718,
|
||||
|
||||
Reference in New Issue
Block a user