Files
msd-core/docs/how-to/probe-edges-in-a-non-english-project.md
Tom Boucher bf4485ada2 enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification

Adds unit tests for the not-yet-implemented text_en field on Requirement
(fallback selection, empty/whitespace/non-string rejection, shapes-override
precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX),
and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents
populating text_en for response_language projects. All new tests are RED
until src/edge-probe.cts and the workflow docs are updated.

* feat(#3717): make edge-probe shape classification read an optional text_en field

Requirement gains an optional text_en; classifyShape's own signature stays
untouched (a locked, directly-tested export), and the text_en ?? text
selection is pushed to proposeEdges' single call site instead. text_en is
validated fail-closed: an empty or whitespace-only value throws rather than
silently winning the ?? fallback and degrading classification to zero shapes.

This makes the #2773 doc-only translation convention an explicit,
validatable field instead of an invisible instruction, per the approved
Form-1 scope on #3717.

* docs(#3717): document the text_en field across spec-phase, reference and how-to docs

Updates Step 5.5's response_language instructions, the edge-probe reference
Inputs contract, the FEATURES.md fragment, and the non-English how-to guide
to describe the new text_en field: text keeps the requirement's own wording
in all cases, text_en (when populated) is the engine-only English rendering
the classifier prefers.

* docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550

Updates the Edge Probe Module glossary entry to describe the text_en field
and its fail-closed validation, and appends an ADR-550 amendment recording
why this is additive and does not re-open the #652 LLM-classifier rejection
(text_en is a plain field read by the existing deterministic regex
classifier, not a new model-dependent surface).

* docs(#3717): add changeset fragment and regenerate FEATURES.md

pr:0 placeholder — backfilled with the real PR number after the PR opens.

* docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests

Code-review (Spec axis) finding: the workflow-prose contract tests and the
ADR-550 amendment overclaimed themselves as "the machine check the #2773
doc-only stopgap lacked." That check is actually engine-level
(validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) —
the prose tests are the same style of assertion #2773 already used. Reworded
both to attribute the claim correctly.

* fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line

The #3717 rewrite of Step 5.5's response_language paragraph moved a line
break so "requirement `id`s" ended one physical line and "are never
translated" started the next. The pre-existing #2773 regression test
(tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never
translated" on the SAME line (no \n in between, matching git's own
line-oriented prose), so the reflow silently broke it. Rewrapped so the
sentence lands on one line again, verified against every #2773/#3717
regex assertion in that test file.

Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift.

* chore(#3717): backfill changeset PR number

pr:0 -> pr:4156 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:53 -04:00

5.1 KiB

How to probe edges in a non-English project

Goal: Get real edge-completeness coverage on a spec written in a language other than English, instead of every requirement landing in unclassified and the whole taxonomy quietly contributing nothing.

Prerequisites: A project with response_language set (see Configuration), and a phase whose /gsd-spec-phase run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically — you do not invoke it separately.

For the taxonomy and the reasoning behind front-of-pipeline edge analysis, see Spec-Phase Edge-Completeness Probe. For how to act on findings once they appear, see Resolve edge-coverage findings. This guide covers only what is different when your spec is not in English.


What happens, and why

The probe classifies each requirement's data/behavior shape by matching English word-boundary cues against the requirement text. A requirement written in another language matches nothing, classifies to zero shapes, raises zero categories, and surfaces as a single unclassified — review manually row.

That is not a rejection you can act on — it looks identical to a genuinely edge-free requirement. When it happens to every requirement, the probe has contributed nothing to the spec.

So Step 5.5 populates an optional text_en field alongside each requirement — a faithful English translation — which the engine reads in preference to text for classification (text_en ?? text). Your spec's own text stays in your language throughout. Concretely, for the same requirement:

Requirement text Requirement text_en Shapes Edges raised
O sistema mescla intervalos sobrepostos em uma lista ordenada (absent) none none — one unclassified row
O sistema mescla intervalos sobrepostos em uma lista ordenada The system merges overlapping intervals in a sorted list collection adjacency, empty, ordering

What you do

Nothing extra. The translation happens inside Step 5.5 as part of the run, populating the engine-only text_en field.

What you should see is the split: your spec stays in response_language — its requirements, its acceptance criteria, its ## Edge Coverage section — while the probe's findings are reasoned about from text_en, an English rendering of the requirement text that lives alongside (never in place of) your requirement's own text. Requirement ids (R1, R2, …) are never translated or renumbered, so a finding always names the same requirement you wrote.

If your spec comes back anglicized, that is a bug worth reporting — only the transient text_en field is translated, never the document.

Tell "no edges here" apart from "the probe could not read it"

This is the distinction that matters, because both look like an unclassified row.

What you see What it means What to do
A few unclassified rows among normally-classified ones Those requirements carry no shape cue in any language. This is the classifier's known recall gap, not a translation problem. Resolve each like any other finding — or author an explicit shapes array on the requirement (below).
Every requirement unclassified, and a WARNING: edge-probe proposed ZERO applicable edges The probe could not read your requirements at all. Confirm the run really is translating the probe input. Do not accept an empty ## Edge Coverage section.
Some requirements classified, the rest unclassified, no warning The silent case. The zero-applicable warning fires only when all requirements are unclassified, so a partly-classified spec raises nothing. Check the unclassified ones individually against the row above.

Translation makes the classifier applicable; it does not make it omniscient. A requirement carrying no shape cue in English either — for example "the command exits with code 1 on invalid input" — still classifies to zero. That is expected, and the fix is the same one an English-language project uses.

Force the shape when the prose carries no cue

When a requirement is genuinely edge-relevant but no cue fires, do not fight the wording. Author the shape explicitly — an authored shapes array bypasses prose classification entirely, in any language:

{ "id": "R4", "text": "The command exits with code 1 on invalid input", "shapes": ["stateful"] }

shapes accepts any of numeric-range, collection, text, stateful, io. The example above raises idempotency and concurrency.

An explicit empty array — "shapes": [] — is the opposite signal: your deliberate "this requirement has no edge surface", which stays silent rather than surfacing an unclassified row.