* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local>
6.5 KiB
id, title, group
| id | title | group |
|---|---|---|
| 144 | Spec-Phase Edge-Completeness Probe | v1.42.1 Features |
Command: /gsd-spec-phase
Purpose: Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as Step 5.5 of spec-phase, after the ambiguity gate.
Behavior: For each SPEC requirement the probe classifies its data/behavior shape, then raises only the applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:
| State | Meaning | Downstream effect |
|---|---|---|
covered |
An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into plan-phase must_haves.truths |
dismissed |
The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
backstop |
Intent recorded, needs a held-out/property-based test | Lifted into must_haves.truths as a non-inferable check |
unresolved |
Deferred | Soft-gates the spec; row stamped ⚠ Edge unresolved — planner must treat as assumption |
When a requirement's prose matches no shape cue, the probe does not silently drop it (#1110): it emits a single unclassified — review manually candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block.
Non-English projects: the probe reads English via text_en, the SPEC does not have to (#2773, durable fix #3717). The shape cues are English word-boundary patterns, so a project running with response_language set would otherwise have every requirement match nothing, classify to zero shapes, and land in unclassified — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. spec-phase Step 5.5 therefore populates an optional text_en field alongside each requirement's text with a faithful English translation: text_en is engine input, never user-facing output, so it is translated while text keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to response_language. Translation makes the classifier applicable; it does not make it omniscient. A requirement carrying no shape cue in any language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit shapes array on the requirement instead of relying on prose classification.
The resolved edges populate a ## Edge Coverage section in SPEC.md. Unresolved applicable edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under --auto, the probe never auto-dismisses — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs [auto] edge coverage: C covered, B backstop, U unresolved. The one exception is an unclassified candidate: --auto leaves it unresolved (surfaced as a flagged assumption), never auto-backstop — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim.
The load-bearing wire is the plan-phase lift: covered and backstop edges become must_haves.truths the verifier can check, so the section is not merely documentation. A backstop edge is lifted as a structured non-inferable marker ({ statement, verification: backstop }, a flat scalar — not a prose note), which the honest verifier then consumes (see below) — closing the loop the edge-probe opened.
Honest verifier — abstention on non-inferable checks (#1154). A non-inferable (backstop) truth is one whose correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would confidently false-pass it (~100% of the time). At verify time, a backstop truth the verifier cannot confirm with explicit evidence (a passing wired held-out/property-based test, or a directly-observed behavior) abstains → human_needed with reason insufficient_spec (reported as unverified — held-out test recommended), never a silent passed. This is the verify-time, truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): exogenous (driven by the backstop tag, never a self-judged "abstain if unsure"), routing-not-diagnosis (the held-out test carries the omitted rule), and capable-tier dependent (reliable on sonnet+; the budget haiku tier degrades toward current behavior). An inferable truth is never abstained (the over-abstention guard). Reference: Honest Verifier.
Requirements:
- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a
## Edge CoverageSPEC section. - REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
- REQ-EDGE-03: A
dismissedresolution MUST require a non-empty reason. - REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
- REQ-EDGE-05:
--autoMUST never auto-dismiss — auto-cover or auto-backstop only. - REQ-EDGE-06:
plan-phaseMUST liftcoveredcriteria andbackstopnotes intomust_haves.truths. - REQ-EDGE-07: A requirement whose prose matches no shape cue MUST surface an
unclassified — review manuallycandidate (never silently dropped);--autoMUST leave itunresolved, never auto-backstop. - REQ-EDGE-08:
plan-phaseMUST lift abackstopedge intomust_haves.truthsas a structured flat-scalar marker ({ statement, verification: backstop }), never a prose parenthetical. - REQ-HONEST-01: At verify time a
backstoptruth that cannot be confirmed with explicit evidence MUST abstain →human_needed(reasoninsufficient_spec), neverpassed; an inferable truth MUST never be abstained (over-abstention guard); abstention MUST be exogenous (driven by thebackstoptag, not self-judgment).
Reference: Edge Probe