Files
msd-core/docs/how-to/resolve-edge-coverage-findings.md
Tom Boucher 36375513b9 feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments

docs/FEATURES.md was hand-maintained, and every feature PR wrote into two
shared mutable cells: the '### N.' heading whose integer was hand-allocated at
authoring time, and the hand-maintained table of contents. Concurrent PRs all
picked the same next integer, and two PRs adding differently numbered features
still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across
successive rebases, each collision also costing a full matrix verification run
because the sha-keyed pass marker dies with the rebase.

Mechanism: one fragment per feature at docs/features/<slug>.md carrying
id/title/group (and an optional order) in frontmatter, consolidated by
scripts/gen-features.cjs --write|--check into a marker-delimited region of
docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings
and their order are derived too - a group sorts by its lowest-ordered member -
so there is no shared registry to edit either; optional per-group prose lives in
docs/features/_groups/<slug>.md. A contributor adds exactly one new file.
Wired into regen:derived and lint:generated-sync alongside the eight existing
generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting.

Migration froze all 168 existing numbers verbatim: identical section set,
identical order, identical bodies. Two defects found in the tree are fixed
inline rather than carried forward - the '## Related' block had been spliced
into the middle of the document, orphaning §142's Reference line, and four
inbound anchors were already broken on next (FEATURES.md#runtime-identity in
two files, and #143-spec-phase-edge-completeness-probe off by one). Since the
repo has no link checker, --check now validates every inbound
FEATURES.md#anchor by resolved target, so that class cannot ship silently
again; locale FEATURES.md files resolve elsewhere and stay out of scope.

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3840): carry upstream §69 delta into its fragment and harden the generator

Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus
origin/next. Root cause was a stale base, not extraction loss: those lines
landed in 394bf384b (#3844) AFTER this branch forked at 63abcface, and
'git diff 63abcface origin/next -- docs/FEATURES.md' is exactly that hunk.
Merging origin/next auto-applied the hunk into the GENERATED region, which
--check immediately reported as stale; the delta is now carried in
docs/features/statemd-consistency-gates.md and regenerated from there.

--write is now fail-closed. It previously rendered the region even with
violations outstanding, warning only on stderr and exiting 0, so a
'--write && git commit' chain could commit a FEATURES.md carrying two
colliding sections. It now refuses and exits 1; --force is the explicit
override and says so in the report. The test that pinned the old behavior now
pins the refusal, plus the --force override and its scoping.

Marker forgery is rejected at two layers. A fragment body containing
'<!-- FEATURES:START' or '<!-- FEATURES:END' is a typed
body_forges_region_marker violation (fragments and group notes alike), and
spliceIntoFeatures anchors the end boundary with lastIndexOf instead of
indexOf, so a marker that reaches the document by any other route can only
make the generated region grow, never shrink. Matching is on marker PREFIXES,
so a decorated variant comment cannot slip past.

Symlinked corpus entries are refused with a typed dirent_not_regular_file
rather than read. A fork PR could otherwise commit docs/features/evil.md as a
symlink to any readable path and have the generator inline those bytes into
the committed docs/FEATURES.md on the next regen.

Equivalence re-verified with a method that cannot cancel out. The first
check extracted both operands with the same body-normalising helper, so
anything that helper dropped was dropped on both sides. The replacement runs
two independent passes: a global content-line multiset diff with no
per-section logic at all (0 gained, 19 lost, all 19 the stale hand-written
mini-TOC links this change deliberately deletes), and a per-section
byte-exact body diff carrying a coverage assertion that fails loudly per file
when the extractor accounts for fewer lines than the file contains. That
assertion caught two blind spots in the checker itself. 168/168 sections
present, order identical, one intended body difference (§142 regains the
Reference line orphaned by the misplaced '## Related' block).

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3840): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 22:49:00 -04:00

6.0 KiB

How to resolve edge-coverage findings while writing a spec

Goal: Turn each domain-boundary edge the spec phase surfaces into an explicit, verifiable spec decision — so omitted boundaries (rounding ties, touching ranges, grapheme truncation) become checkable requirements before any code exists, instead of silent blind spots the verifier is confidently wrong about.

Prerequisites: A phase whose /gsd-spec-phase run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically and presents its findings — you do not invoke it separately.

For the category taxonomy and the reasoning behind front-of-pipeline edge analysis, see Spec-Phase Edge-Completeness Probe. This guide covers only how to act on the findings.


Read a finding

Each finding is one applicable edge for one requirement — a boundary the probe's relevance filter decided is in scope for that requirement's shape, and that you have not yet addressed. A finding is phrased as a probe question, for example:

R3 · precision — Where can precision loss, overflow, or rounding/tie-breaking occur — and what is the exact contract (e.g. half-up vs half-to-even, ceil/floor/truncate)?

You must resolve each finding into exactly one of four states. Claude presents them as a numbered choice (or an AskUserQuestion menu).

You may also see an unclassified — review manually finding. The relevance filter is a heuristic over prose cues, so a requirement whose wording is edge-relevant but matched no cue surfaces this single soft candidate instead of being silently dropped. Resolve it like any other finding — specify a criterion, or dismiss it with a reason if the requirement is genuinely edge-free (e.g. "static asset, no input"). It is a manual-review nudge, never a hard block.


Specify it — write an acceptance criterion

Choose this when you can state the correct behaviour as a pass/fail check. This is the strongest resolution: it produces a concrete assertion the planner and verifier can enforce.

Claude writes a new line into the spec's Acceptance Criteria and marks the edge covered. For the precision finding above:

  • Monetary amounts round half-to-even to 2 decimal places; 2.005 → 2.00

Prefer this state whenever a defensible criterion can be written. A covered edge is lifted into the plan's must_haves.truths, so the verifier checks it.


Dismiss it — record why it does not apply

Choose this when the edge genuinely cannot occur — and say why. A dismissal requires a non-empty reason; silence is rejected. The reason is the audit trail.

⛔ dismissed — input is a bounded enum; no boundary value exists

A wrong dismissal is the exact silent failure this probe exists to prevent, so dismiss only when the reason is solid.


Backstop it — defer the contract to a held-out test

Choose this when you know the edge matters but cannot fully articulate the correct behaviour in prose yet. Claude marks the edge backstop and notes that a held-out / property-based test stands in for the missing assertion.

🧪 backstop — held-out property test: dedupe output order is stable under input permutation

A backstop edge is also carried into the plan's must_haves.truths as a non-inferable check — the planner records the intent, and the test body is authored later during execution.


Defer it — leave it unresolved and flagged

Choose this only when you are not ready to decide. The edge stays unresolved and is flagged. Unlike a dismissal, deferring makes no claim that the edge is safe — it is an explicit, visible assumption the planner must surface, not silently drop.


Clear the soft gate

After you have worked through the findings, the probe runs a soft gate:

  • All applicable edges resolved → the spec proceeds to the next step.
  • One or more still unresolved → Claude asks what to do:
    • Resolve now — loop back and resolve the remaining edges.
    • Write the spec anyway — the spec is written with those rows marked ⚠ Edge unresolved — planner must treat as assumption. Use this deliberately; you are choosing to ship a known gap.

The gate is soft: it never blocks you, but every unresolved edge remains visible in the spec's ## Edge Coverage section.


Let Claude resolve them for you

If the edges are low-stakes or already implied by earlier phases, run the spec phase in auto mode:

/gsd-spec-phase 3 --auto

In --auto, Claude marks an edge covered where it can write a defensible acceptance criterion, and backstop otherwise. It never auto-dismisses — dismissing an edge requires a human reason, because a wrong auto-dismissal is precisely the silent failure being eliminated. Claude logs the tally, for example:

[auto] edge coverage: 4 covered, 2 backstop, 1 unresolved

Review the logged choices afterwards; auto mode is a fast first pass, not a substitute for judgement on edges that carry risk.


What happens to resolved findings downstream

When you next run /gsd-plan-phase, the planner reads the spec's ## Edge Coverage section and:

  • lifts every covered edge's acceptance criterion into must_haves.truths,
  • carries every backstop edge into must_haves.truths as a non-inferable check (needing a held-out/property test),
  • surfaces every unresolved edge as an explicit assumption.

This is the payoff: a resolved edge becomes a unit the goal-backward verifier actually checks, extending its reach to boundaries the requirement prose never stated.