* test(#2773): failing-first contract and premise tests for translated edge-probe input Locks the Step 5.5 contract that a response_language project must feed the edge probe an English translation of each requirement's text, and binds that advice to measured engine behavior: the same requirement classifies to zero shapes in Portuguese and to collection/adjacency/empty/ordering in English. Also pins the honest limit — the issue's own repro sentence classifies to [] in English too, so translation is necessary but not sufficient and the authored shapes override is the documented fallback. Red before the doc change; the assertions are all false today. Refs #2773 * fix(#2773): feed the spec-phase edge probe English-translated requirement text The shape cues in src/edge-probe.cts are English word-boundary regexes, so a project running with response_language set wrote its SPEC requirements into the Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy contributed nothing and --auto left it all unresolved — the probe was a silent no-op for exactly the spec type it exists to harden. Step 5.5 now states that the $REQS_JSON payload is engine input rather than user-facing output, so the response_language rule does not govern it: each requirement's text carries a faithful English translation, the SPEC keeps its original language, and requirement ids are never translated or renumbered. The instruction sits before the heredoc on purpose — the downstream APPLICABLE=0 warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal at all. Measured against the compiled engine: the same requirement returns [] in Portuguese and collection -> adjacency/empty/ordering in English. Also measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient — the instruction therefore points at the authored shapes override for prose carrying no cue in any language rather than promising that translation restores classification. Doc scope only, per the triage disposition on the issue. The compiled engine is untouched; the lang-hint / per-language cue-set fix is a separate follow-up. Closes #2773 * fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path Surfaced by the isolated security review of this branch. Between the mktemp and the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before exiting, while the empty/placeholder guard directly above it exited without one. A spec run that tripped the placeholder check therefore stranded a temp file holding the SPEC's requirement text in TMPDIR, once per failed run. The added contract test walks the region between the mktemp and the unconditional cleanup and asserts no exit path leaves the file behind, so the two guards can no longer drift apart. Proven to bind: run against the pre-fix file the walker reports the leaking exit; against the fixed file it reports none. Refs #2773 * docs(#2773): record the edge probe's English-cue input constraint in the predicate store The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real gaps rather than incidental coupling. CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for classifyShape but did not record that SHAPE_CUES are English word-boundary patterns — so the predicate store implied text was language-agnostic, which is what a future agent reads before touching this seam. docs/CONFIGURATION.md's response_language row is what a non-English project reads when it turns the setting on; it now names the one deliberate exception and links to the FEATURES.md explanation, so the interaction is discoverable from the config key rather than only from the workflow. CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack fragment is updated for the final byte range and now also records the placeholder-guard cleanup fix folded into the same block. Refs #2773 * fix(#2773): append the growth rationale to the existing spec-phase.md ack entry The remote runner caught this: emitted-attribution.test.cjs pins the 0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration regression test asserts the exact '31987 -> 31997' delta text survives), so removing it to avoid a duplicate-key collision with a new fragment broke that test instead of satisfying the ratchet. The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102 were each appended to the same reason string by later PRs, which is how a shared growth key coexists with the rule that two ack sources may never name the same path. This appends the #2773 rationale the same way and drops the separate fragment, whose spec-phase.md key was the collision. Verified locally by reproducing both affected tests against the real fragment before re-dispatching: the pinned delta survives, grown[0].acked is true, staleAcks is empty, and all 35 entries still read as spent. Refs #2773 * docs(#2773): add a how-to for probing edges in a non-English project The phase gate's enablementSequence check caught a wrong call of mine. I had recorded that no how-to was owed because the user takes zero extra steps — the workflow translates the probe input itself. Written out, though, the sequence from off to value is two steps and step 1 depends on response_language, a setting owned by a different capability than the edge probe, which is exactly the condition the how-to test names. There is also real task content a reference table cannot carry: the three-way split between a few unclassified rows (the classifier's recall gap), every row unclassified (the probe could not read the spec at all), and the silent partly-classified case where the APPLICABLE=0 warning never fires. That last one is what a user would otherwise misread as a clean bill of health. Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard siblings and indexed from docs/README.md next to its closest relative. Refs #2773 * chore(#2773): backfill the changeset PR number pr:0 placeholder replaced with the real PR number now that #3713 exists. Refs #2773 --------- Co-authored-by: sim <sim@local>
4.8 KiB
How to probe edges in a non-English project
Goal: Get real edge-completeness coverage on a spec written in a language other than English, instead of every requirement landing in unclassified and the whole taxonomy quietly contributing nothing.
Prerequisites: A project with response_language set (see Configuration), and a phase whose /gsd-spec-phase run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically — you do not invoke it separately.
For the taxonomy and the reasoning behind front-of-pipeline edge analysis, see Spec-Phase Edge-Completeness Probe. For how to act on findings once they appear, see Resolve edge-coverage findings. This guide covers only what is different when your spec is not in English.
What happens, and why
The probe classifies each requirement's data/behavior shape by matching English word-boundary cues against the requirement text. A requirement written in another language matches nothing, classifies to zero shapes, raises zero categories, and surfaces as a single unclassified — review manually row.
That is not a rejection you can act on — it looks identical to a genuinely edge-free requirement. When it happens to every requirement, the probe has contributed nothing to the spec.
So Step 5.5 sends the probe an English rendering of each requirement, while your spec stays in your language. Concretely, for the same requirement:
| Requirement text handed to the probe | Shapes | Edges raised |
|---|---|---|
O sistema mescla intervalos sobrepostos em uma lista ordenada |
none | none — one unclassified row |
The system merges overlapping intervals in a sorted list |
collection |
adjacency, empty, ordering |
What you do
Nothing extra. The translation happens inside Step 5.5 as part of the run.
What you should see is the split: your spec stays in response_language — its requirements, its acceptance criteria, its ## Edge Coverage section — while the probe's findings are reasoned about from an English rendering of the requirement text. Requirement ids (R1, R2, …) are never translated or renumbered, so a finding always names the same requirement you wrote.
If your spec comes back anglicized, that is a bug worth reporting — only the probe's transient input is translated, never the document.
Tell "no edges here" apart from "the probe could not read it"
This is the distinction that matters, because both look like an unclassified row.
| What you see | What it means | What to do |
|---|---|---|
A few unclassified rows among normally-classified ones |
Those requirements carry no shape cue in any language. This is the classifier's known recall gap, not a translation problem. | Resolve each like any other finding — or author an explicit shapes array on the requirement (below). |
Every requirement unclassified, and a WARNING: edge-probe proposed ZERO applicable edges |
The probe could not read your requirements at all. | Confirm the run really is translating the probe input. Do not accept an empty ## Edge Coverage section. |
Some requirements classified, the rest unclassified, no warning |
The silent case. The zero-applicable warning fires only when all requirements are unclassified, so a partly-classified spec raises nothing. | Check the unclassified ones individually against the row above. |
Translation makes the classifier applicable; it does not make it omniscient. A requirement carrying no shape cue in English either — for example "the command exits with code 1 on invalid input" — still classifies to zero. That is expected, and the fix is the same one an English-language project uses.
Force the shape when the prose carries no cue
When a requirement is genuinely edge-relevant but no cue fires, do not fight the wording. Author the shape explicitly — an authored shapes array bypasses prose classification entirely, in any language:
{ "id": "R4", "text": "The command exits with code 1 on invalid input", "shapes": ["stateful"] }
shapes accepts any of numeric-range, collection, text, stateful, io. The example above raises idempotency and concurrency.
An explicit empty array — "shapes": [] — is the opposite signal: your deliberate "this requirement has no edge surface", which stays silent rather than surfacing an unclassified row.
Related
- Resolve edge-coverage findings — what to do with each finding once it is raised
- Spec-Phase Edge-Completeness Probe — the taxonomy and the rationale
- Configuration — the
response_languagesetting