diff --git a/.changeset/tidy-badgers-fly.md b/.changeset/tidy-badgers-fly.md new file mode 100644 index 000000000..9047ef0c6 --- /dev/null +++ b/.changeset/tidy-badgers-fly.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 4156 +--- +**Edge-completeness probe requirements accept an optional `text_en` field** — spec-phase Step 5.5 can now populate an explicit English translation for non-English SPEC requirements, which the shape classifier reads in preference to `text` (`text_en ?? text`). This replaces the #2773 doc-only convention where `text` silently carried the translation; `text` now always keeps the requirement's own wording. (#3717) diff --git a/CONTEXT.md b/CONTEXT.md index fade9b4ea..71d7ffc43 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -475,7 +475,7 @@ Deterministic eval-scoring projection (#10 / #1579) that moves the `gsd-eval-aud Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the Prohibition Probe Module (#644) next. Exports (generic surface): `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` — the prohibition adapter exports that also ship from this module (`projectProhibitions`, `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, `dispositionForProhibition`) are documented under the Prohibition Probe Module's own locked-surface line. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. Under ADR-857 (phase-6 boundary, settled 2026-06-12) this seam is classified **core verification substrate** on the *contract* side: its deterministic validators are the verifier↔predicate contract's CI-testable surface (ADR-550 Decision 5) — core and non-toggleable, never an off-by-default Feature Capability. (The recall-gapped *generator* is the probe adapters that propose predicates, not this resolution engine — see Edge Probe Module and Verification substrate (predicate boundary).) ### Edge Probe Module -First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. `SHAPE_CUES` are **English** word-boundary patterns, so `text` is an English-language input regardless of the project's `response_language`: spec-phase Step 5.5 passes a faithful English translation of each requirement while the SPEC itself keeps its original language and the `id` is never translated (#2773 — without it every requirement in a non-English project matched no cue and landed in `unclassified`, silently disabling the taxonomy). Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability. +First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved; an optional `text_en` field, when present, must also be a non-empty string (#3717) or `validateRequirement` throws — an empty `text_en` would otherwise win `text_en ?? text` unnoticed (`??` does not catch `''`) and silently degrade classification. `SHAPE_CUES` are **English** word-boundary patterns, so classification needs an English-language input regardless of the project's `response_language`: `classifyShape` reads `text_en ?? text` (#3717), so spec-phase Step 5.5 populates `text_en` with a faithful English translation while `text` keeps the requirement's own wording and the SPEC keeps its original language; `id` is never translated (#2773 — without a translated field every requirement in a non-English project matched no cue and landed in `unclassified`, silently disabling the taxonomy; #3717 made the translation an explicit, validated field rather than an invisible convention on `text` itself). Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). The `Requirement` shape is `{ id, text, text_en?, shapes? }` (#3717 added the optional `text_en`; per-language cue sets were considered and declined — `en` remains the only `SHAPE_CUES` vocabulary). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability. ### Verification substrate (predicate boundary) The ADR-857 classification (settled 2026-06-12, prompted by @davesienkowski's boundary analysis on #857) that predicate-generation — the must-NOT-have / edge predicates that set the verifier's reach — is **core**, not an off-by-default Feature Capability. Load-bearing premise: *verifier reach = spec reach* (the verifier can only catch what the spec concretely names). Decomposes into: the **verifier↔predicate contract** (the verifier always expects predicates and grades **exogenously** against them — core, non-toggleable, a stability contract alongside the Loop Extension Point names), and the **generator** (the probe *adapters* that propose predicates — edge-probe's `classifyShape`/`proposeEdges`, the prohibition probe's adversarial LLM-propose — core-default but independently versionable, kept their own module because of a measured recall gap; `probe-core`'s deterministic validators sit on the contract side, not the generator side). Altitude rule distinguishing it from `gate` hooks: a gate runs *against* the spec (hook); predicate-generation defines the spec's *reach* (core). The decision-#6 `produces`/`consumes` artifact flow is the internal rail from generation to the core verifier. See ADR-857 *Verification substrate vs. plug-in tier (the predicate boundary)* and the ADR-550 cross-reference. diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 1d080e165..88eadc2a9 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3212,7 +3212,7 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev When a requirement's prose matches **no** shape cue, the probe does not silently drop it (#1110): it emits a single `unclassified — review manually` candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block. -**Non-English projects: the probe reads English, the SPEC does not have to (#2773).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore feeds the probe a faithful **English translation** of each requirement's `text`: that payload is engine input, never user-facing output, so it is translated while the SPEC itself stays in the original language, requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification. +**Non-English projects: the probe reads English via `text_en`, the SPEC does not have to (#2773, durable fix #3717).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore populates an optional `text_en` field alongside each requirement's `text` with a faithful **English translation**: `text_en` is engine input, never user-facing output, so it is translated while `text` keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification. The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim. diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md index 6fe1861ff..97fe19629 100644 --- a/docs/adr/550-spec-phase-probe-contract.md +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -197,3 +197,42 @@ prohibition-enforcement verify-time seam.) this addendum adds to 550's decision record; the other two are cross-references to existing decisions, collected so the probe family's full "alternatives considered" set is readable in one place. + +## Amendment (2026-09-01, #3717): `text_en` field — durable fix for the #2773 language stopgap + +**What changed.** `Requirement` (the Edge Probe Module's locked input shape) gains an optional +`text_en: string` field. `classifyShape`'s own signature is unchanged (still `(text: string) => +Shape[]`, a directly-tested export); the `text_en ?? text` selection happens once, at +`proposeEdges`' single call site. `validateRequirement` fail-closes on a present-but-empty or +whitespace-only `text_en` (an unvalidated `''` would otherwise win `??` silently, since nullish +coalescing does not treat `''` as nullish). + +**Why this is a locked-surface change, not a silent one.** `text` previously carried a dual, +undocumented meaning under the #2773 stopgap: the requirement's own text for English projects, +but silently the *translation* for `response_language` projects (spec-phase Step 5.5 wrote the +English rendering into `text`, never the SPEC's own words). `text_en` removes that overload — +`text` always means the requirement's own text, in whatever language the SPEC uses; `text_en`, +when present, is the engine-only English rendering the classifier prefers. + +**Not a re-open of the #652 rejection (see the Addendum above, `docs/adr/550-spec-phase-probe-contract.md:172-177`).** That addendum rejects a *standalone, model-dependent* LLM +classifier surface. `text_en` is a plain optional string field read by the existing +deterministic regex classifier — no new model-dependent surface, no change to the taxonomy or +`SHAPE_CUES`'s cue-matching mechanism. The maintainer's approval on #3717 confirmed this reading +explicitly and scoped the change to Form 1 only (the `text_en` field); per-language `SHAPE_CUES` +tables (a form that WAS considered — a `lang` hint + a cue-table-per-language) were evaluated and +declined as an unwarranted maintenance burden for a solo-maintainer project. `en` remains the +only `SHAPE_CUES` vocabulary. + +**Acceptance bar carried into `tests/edge-probe.test.cjs` / `tests/edge-probe-spec-phase-contract.test.cjs`:** a `SHAPE_CUES`/`VALID_SHAPES` parity assertion (`RULESET.GENERATIVE-FIX`), a fixture +pair proving a non-English requirement with `text_en` classifies identically to its English +equivalent, boundary coverage for `text_en`'s absence/presence/empty-string cases, and a +workflow-prose contract test — in the same style as the existing #2773 Step 5.5 assertions — +confirming `spec-phase.md` documents populating `text_en` for `response_language` projects. +The actual machine check the #2773 doc-only stopgap lacked is engine-level: `text_en` is now +a real, validated `Requirement` field (`validateRequirement` fail-closes on an empty value) +that `classifyShape`/`proposeEdges` demonstrably prefer over `text` — a property #2773's +prose-only convention had no way to enforce. + +**Explicitly out of scope (unchanged from #2773).** The ADR-857 §98 / Decision 7b recall gap — a +requirement carrying no shape cue in *any* language, including English, still classifies to +zero shapes — is untouched. The authored `shapes` override remains the remedy there. diff --git a/docs/features/spec-phase-edge-completeness-probe.md b/docs/features/spec-phase-edge-completeness-probe.md index b83ebaeea..de7b83257 100644 --- a/docs/features/spec-phase-edge-completeness-probe.md +++ b/docs/features/spec-phase-edge-completeness-probe.md @@ -19,7 +19,7 @@ group: v1.42.1 Features When a requirement's prose matches **no** shape cue, the probe does not silently drop it (#1110): it emits a single `unclassified — review manually` candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block. -**Non-English projects: the probe reads English, the SPEC does not have to (#2773).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore feeds the probe a faithful **English translation** of each requirement's `text`: that payload is engine input, never user-facing output, so it is translated while the SPEC itself stays in the original language, requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification. +**Non-English projects: the probe reads English via `text_en`, the SPEC does not have to (#2773, durable fix #3717).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore populates an optional `text_en` field alongside each requirement's `text` with a faithful **English translation**: `text_en` is engine input, never user-facing output, so it is translated while `text` keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification. The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim. diff --git a/docs/how-to/probe-edges-in-a-non-english-project.md b/docs/how-to/probe-edges-in-a-non-english-project.md index e72164faf..4d7818b2f 100644 --- a/docs/how-to/probe-edges-in-a-non-english-project.md +++ b/docs/how-to/probe-edges-in-a-non-english-project.md @@ -14,20 +14,20 @@ The probe classifies each requirement's data/behavior shape by matching **Englis That is not a rejection you can act on — it looks identical to a genuinely edge-free requirement. When it happens to *every* requirement, the probe has contributed nothing to the spec. -So Step 5.5 sends the probe an English rendering of each requirement, while your spec stays in your language. Concretely, for the same requirement: +So Step 5.5 populates an optional `text_en` field alongside each requirement — a faithful English translation — which the engine reads in preference to `text` for classification (`text_en ?? text`). Your spec's own `text` stays in your language throughout. Concretely, for the same requirement: -| Requirement text handed to the probe | Shapes | Edges raised | -|---|---|---| -| `O sistema mescla intervalos sobrepostos em uma lista ordenada` | none | none — one `unclassified` row | -| `The system merges overlapping intervals in a sorted list` | `collection` | `adjacency`, `empty`, `ordering` | +| Requirement `text` | Requirement `text_en` | Shapes | Edges raised | +|---|---|---|---| +| `O sistema mescla intervalos sobrepostos em uma lista ordenada` | *(absent)* | none | none — one `unclassified` row | +| `O sistema mescla intervalos sobrepostos em uma lista ordenada` | `The system merges overlapping intervals in a sorted list` | `collection` | `adjacency`, `empty`, `ordering` | ## What you do -Nothing extra. The translation happens inside Step 5.5 as part of the run. +Nothing extra. The translation happens inside Step 5.5 as part of the run, populating the engine-only `text_en` field. -What you should *see* is the split: your **spec stays in `response_language`** — its requirements, its acceptance criteria, its `## Edge Coverage` section — while the probe's findings are reasoned about from an English rendering of the requirement text. Requirement ids (`R1`, `R2`, …) are never translated or renumbered, so a finding always names the same requirement you wrote. +What you should *see* is the split: your **spec stays in `response_language`** — its requirements, its acceptance criteria, its `## Edge Coverage` section — while the probe's findings are reasoned about from `text_en`, an English rendering of the requirement text that lives alongside (never in place of) your requirement's own `text`. Requirement ids (`R1`, `R2`, …) are never translated or renumbered, so a finding always names the same requirement you wrote. -If your spec comes back anglicized, that is a bug worth reporting — only the probe's transient input is translated, never the document. +If your spec comes back anglicized, that is a bug worth reporting — only the transient `text_en` field is translated, never the document. ## Tell "no edges here" apart from "the probe could not read it" diff --git a/gsd-core/references/edge-probe.md b/gsd-core/references/edge-probe.md index 75acfbcdb..2844b22a9 100644 --- a/gsd-core/references/edge-probe.md +++ b/gsd-core/references/edge-probe.md @@ -41,19 +41,23 @@ core finding that the spec layer is the measured weak point: ## Inputs -A list of requirements, each a `{ id, text, shapes? }` record where `text` is a testable -statement and `shapes` is an optional author-supplied override of the data/behavior shape. -The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. When -`shapes` is absent, a heuristic classifier proposes them from the requirement prose -(propose-then-confirm) — the author may correct the shape. +A list of requirements, each a `{ id, text, text_en?, shapes? }` record where `text` is a +testable statement, `text_en` is an optional English translation of `text`, and `shapes` is an +optional author-supplied override of the data/behavior shape. The five shapes are: +`numeric-range`, `collection`, `text`, `stateful`, `io`. When `shapes` is absent, a heuristic +classifier proposes them from the requirement prose — reading `text_en` in preference to +`text` when present (`text_en ?? text`) — (propose-then-confirm); the author may correct the +shape. -**`text` is English, whatever language the SPEC is in.** The heuristic cues are English -word-boundary patterns, so a project running with `response_language` set must pass `text` as a -faithful English translation of the requirement; the SPEC itself keeps the original language and -the `id` is never translated. Prose in another language matches no cue, classifies to zero -shapes, and surfaces as `unclassified` (#1110) — the probe contributes nothing. Where a -requirement carries no cue even in English, author `shapes` explicitly rather than leaning on -the classifier. +**The classifier reads English; `text` does not have to be.** The heuristic cues (`SHAPE_CUES`) +are English word-boundary patterns, so a requirement whose `text` is not English classifies to +zero shapes unless `text_en` supplies a faithful English translation. A project running with +`response_language` set should populate `text_en` for every requirement; `text` keeps its own +meaning — the requirement's own text, in whatever language the SPEC uses — and is never +translated or overwritten. `text_en` is engine input, not part of the SPEC. `id` is never +translated. Prose that matches no cue in either field surfaces as `unclassified` (#1110) — the +probe contributes nothing for that requirement. Where a requirement carries no cue even in +English, author `shapes` explicitly rather than leaning on the classifier. ## Taxonomy (8 categories) diff --git a/gsd-core/workflows/spec-phase.md b/gsd-core/workflows/spec-phase.md index 0d886c2fa..423f9d075 100644 --- a/gsd-core/workflows/spec-phase.md +++ b/gsd-core/workflows/spec-phase.md @@ -192,21 +192,25 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met): Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague ones). Reference: @~/.claude/gsd-core/references/edge-probe.md. -**Non-English projects — the probe input is translated, the SPEC is not.** The shape cues the -classifier matches are **English** word-boundary patterns, so requirement prose written in -another language matches nothing, classifies to zero shapes, and lands every row in -`unclassified` (#1110) — the taxonomy contributes nothing and `--auto` leaves it all -`unresolved`. The `$REQS_JSON` payload below is **engine input, never user-facing output**, so -the `response_language` rule at the top of this workflow does not govern it: write each -requirement's `text` as a faithful **English** translation of the SPEC requirement. -The SPEC keeps the original language — only the probe input is translated, and requirement +**Non-English projects — `text_en` carries the classifier-facing translation; the SPEC is +not.** The shape cues the classifier matches are **English** word-boundary patterns, so +requirement prose written in another language matches nothing, classifies to zero shapes, and +lands every row in `unclassified` (#1110) — the taxonomy contributes nothing and `--auto` +leaves it all `unresolved`. When this project has `response_language` set, add an optional +`text_en` key to each `$REQS_JSON` entry: a faithful **English** translation of that +requirement's `text`. `text_en` is **engine input, never user-facing output**, so the +`response_language` rule at the top of this workflow does not govern it — but `text` itself is +NOT translated: write it as the requirement's own text, exactly as it appears in the SPEC. +The SPEC keeps the original language — only `text_en` is translated, and requirement `id`s are never translated or renumbered (coverage rows join back on `id`, and any Acceptance -Criteria you write from the resolved edges go into the SPEC in `response_language`). -Translate **every** requirement, not only the ones that look edge-relevant: the +Criteria you write from the resolved edges go into the SPEC in `response_language`). Populate +`text_en` for **every** requirement, not only the ones that look edge-relevant: the `$APPLICABLE = 0` warning below fires only when *all* requirements are unclassified, so a -partly-classified spec slips through with no signal at all. -If a requirement still classifies to zero shapes after translation, it carries no cue in any -language (the recorded recall gap — ADR-857 §98 / ADR-550 D7b, not a translation failure); +partly-classified spec slips through with no signal at all. When `response_language` is unset +(an English-language project), omit `text_en` — `text` is already English and the engine +falls back to it automatically (`text_en ?? text`). +If a requirement still classifies to zero shapes with `text_en` populated, it carries no cue in +any language (the recorded recall gap — ADR-857 §98 / ADR-550 D7b, not a translation failure); author an explicit `shapes` array on that requirement instead of relying on the prose classifier. @@ -251,11 +255,12 @@ fi # Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the # canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object -# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes -# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04). -# When `response_language` is set, `text` carries a faithful ENGLISH translation (see above) — -# the shape cues are English-only, so original-language prose classifies to zero shapes and the -# probe becomes a silent no-op. Keep every `id` exactly as it appears in the SPEC. +# per requirement: {"id","text","text_en"?,"shapes"?}. This is the load-bearing step: an empty +# file makes the probe a no-op, so the guard below fails loud rather than silently skipping +# (RR-04). When `response_language` is set, ALSO add `text_en` — a faithful ENGLISH +# translation of `text` (see above) — the shape cues are English-only, so original-language +# `text` alone classifies to zero shapes; `text` itself stays the SPEC's own requirement text +# and is never translated. Keep every `id` exactly as it appears in the SPEC. # BSD/macOS mktemp only randomizes XXXXXX when it is the final path component, so make a # suffixless temp then append the extension — portable across BSD + GNU (#1520). REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX") && mv "$REQS_JSON" "${REQS_JSON}.json" && REQS_JSON="${REQS_JSON}.json" || exit 1 diff --git a/src/edge-probe.cts b/src/edge-probe.cts index f7fd20045..66f30e65d 100644 --- a/src/edge-probe.cts +++ b/src/edge-probe.cts @@ -46,10 +46,21 @@ export interface TaxonomyEntry { probe: string; } -/** A SPEC requirement; `shapes` is an optional authored override of classification. */ +/** + * A SPEC requirement; `shapes` is an optional authored override of classification. + * + * `text_en` (#3717) is an optional English translation of `text`, read by shape + * classification in preference to `text` when present (`text_en ?? text`). `SHAPE_CUES` + * are English-only word-boundary patterns, so a non-English `text` (e.g. a project running + * with `response_language` set) classifies to zero shapes unless `text_en` supplies an + * English rendering. `text` itself is unaffected and keeps its own meaning (the + * requirement's own text, in whatever language the SPEC uses) — only classification reads + * `text_en` preferentially. + */ export interface Requirement { id: string; text: string; + text_en?: string; shapes?: Shape[]; } @@ -135,7 +146,7 @@ export const EDGE_VALIDATORS: Validators = { */ export function validateRequirement(requirement: Requirement): void { coreValidateRequirement(requirement); - const r = requirement as unknown as { shapes?: unknown; text?: unknown }; + const r = requirement as unknown as { shapes?: unknown; text?: unknown; text_en?: unknown }; if (r.shapes != null && !Array.isArray(r.shapes)) { throw new Error(`requirement ${requirement.id} shapes must be an array when present`); } @@ -144,6 +155,15 @@ export function validateRequirement(requirement: Requirement): void { `requirement ${requirement.id} text must be a non-empty string when no shapes override is provided`, ); } + // text_en (#3717) is optional, but when present it must be a non-empty string. An empty + // string is NOT caught by `??` (only null/undefined are), so an unvalidated `text_en: ''` + // would silently win `text_en ?? text` and classify against '' — the same fail-open shape + // #1110/#2773 already exist to eliminate, just moved one field over. Validated + // unconditionally (not gated on whether `shapes` will make it unused) so bad data fails + // closed even when it happens to be dead for this particular call. + if (r.text_en != null && !(typeof r.text_en === 'string' && r.text_en.trim())) { + throw new Error(`requirement ${requirement.id} text_en must be a non-empty string when present`); + } } /** Validate an edge resolution against the edge verification vocabulary. */ @@ -172,7 +192,11 @@ export function proposeEdges(requirement: Requirement): Edge[] { } shapes = requirement.shapes; } else { - shapes = classifyShape(requirement.text); + // #3717: prefer the English translation when present — SHAPE_CUES are English-only + // word-boundary patterns, so a non-English `text` (e.g. response_language projects) + // would otherwise classify to zero shapes. validateRequirement (called above) has + // already guaranteed text_en, if present, is a non-empty string. + shapes = classifyShape(requirement.text_en ?? requirement.text); if (shapes.length === 0) { // Prose present but no shape cue matched. Do NOT silently drop it (#1110): an // edge-relevant requirement whose phrasing missed every cue would otherwise vanish from diff --git a/tests/edge-probe-spec-phase-contract.test.cjs b/tests/edge-probe-spec-phase-contract.test.cjs index 4f1956b29..5294d11a3 100644 --- a/tests/edge-probe-spec-phase-contract.test.cjs +++ b/tests/edge-probe-spec-phase-contract.test.cjs @@ -504,3 +504,59 @@ test('#2773: no Step 5.5 exit path leaks the $REQS_JSON temp file', () => { `every exit between the mktemp and the unconditional cleanup must rm -f "$REQS_JSON" first; leaking exits: ${JSON.stringify(leaks)}`, ); }); + +// ─── #3717: durable text_en field (follow-up to the #2773 doc-only stopgap) ─── +// +// #2773 shipped a documented convention: for a response_language project, Step 5.5's `text` +// field silently carried the English TRANSLATION rather than the SPEC's own requirement +// prose. #3717 (approved scope: Form 1 only, `text_en` field) makes this an explicit, +// validatable field: `text` goes back to meaning "the requirement's own text" (in +// response_language, when set), and `text_en` carries the translation the classifier reads. +// +// These three tests only check the WORKFLOW PROSE instructs populating text_en correctly — +// the same prose-assertion pattern the #2773 tests above already use for Step 5.5, not a new +// testing technique. The actual machine check the #2773 doc-only stopgap lacked is +// ENGINE-level, not workflow-level: text_en is now a real, validated Requirement field — +// validateRequirement fail-closes on an empty value, and classifyShape/proposeEdges +// demonstrably prefer it over text — a property the #2773 prose-only convention had no way +// to enforce. See tests/edge-probe.test.cjs for that engine-level coverage. + +test('#3717: Step 5.5 documents populating text_en for response_language projects', () => { + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + assert.match( + block, + /text_en/, + 'Step 5.5 must name the text_en field — the durable #3717 fix for non-English requirement classification', + ); + assert.match( + block, + /response_language/, + 'Step 5.5 must still name response_language as the setting that triggers populating text_en', + ); +}); + +test('#3717: Step 5.5 documents that text stays the original SPEC-language requirement text', () => { + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // The #2773-era stopgap had `text` secretly carry the English translation. The #3717 + // durable fix keeps `text` meaning the requirement's own text (original language) and + // moves the translation into the new text_en field — this must be documented, not just + // implemented, or the next reader re-introduces the #2773 overload by habit. + assert.match( + block, + /`?text`?[^\n]*(original language|SPEC's own|requirement's own|unchanged)|(original language)[^\n]*`?text`?/i, + 'Step 5.5 must state that text carries the SPEC requirement in its own (original) language — not the translation', + ); +}); + +test('#3717: the edge-probe reference documents the text_en field', () => { + const ref = fs.readFileSync(EDGE_PROBE_REF_PATH, 'utf8'); + assert.match( + ref, + /text_en/, + 'the edge-probe reference `## Inputs` contract must document the optional text_en field and its text_en ?? text fallback', + ); +}); diff --git a/tests/edge-probe.test.cjs b/tests/edge-probe.test.cjs index 90f1a4e30..2a982c3fb 100644 --- a/tests/edge-probe.test.cjs +++ b/tests/edge-probe.test.cjs @@ -471,6 +471,115 @@ describe('edge-probe: analyzeCoverage — duplicate rejection (RR-09)', () => { }); }); +describe('edge-probe: text_en language-aware classification (#3717)', () => { + // The prose the #2773 stopgap test file already uses as its canonical Portuguese/English + // pair, kept identical here so the two test files agree on the same fixture (no drift). + const pt = 'O sistema mescla intervalos sobrepostos em uma lista ordenada'; + const en = 'The system merges overlapping intervals in a sorted list'; + + test('proposeEdges: text_en absent falls back to text (back-compat)', () => { + const withoutTextEn = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places' }); + assert.deepEqual(withoutTextEn.map((e) => e.category).sort(), ['boundary', 'precision']); + }); + + test('proposeEdges: text_en present is used for classification instead of text', () => { + // text alone (non-English) classifies to zero shapes -> the unclassified sentinel. + const nonEnglishOnly = ep.proposeEdges({ id: 'R1', text: pt }); + assert.deepEqual(nonEnglishOnly.map((e) => e.category), ['unclassified']); + + // text_en present -> classification runs against the English translation. + const withTextEn = ep.proposeEdges({ id: 'R1', text: pt, text_en: en }); + assert.deepEqual(withTextEn.map((e) => e.category).sort(), ['adjacency', 'empty', 'ordering']); + }); + + test('#3717: a non-English requirement with text_en classifies identically to its English equivalent', () => { + const englishOnly = ep.proposeEdges({ id: 'R1', text: en }); + const nonEnglishWithTranslation = ep.proposeEdges({ id: 'R1', text: pt, text_en: en }); + assert.deepEqual( + nonEnglishWithTranslation.map((e) => e.category).sort(), + englishOnly.map((e) => e.category).sort(), + 'a translated non-English requirement must raise the same categories as the English original', + ); + }); + + test('validateRequirement: text_en: null is treated as absent (no throw)', () => { + assert.doesNotThrow(() => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: null })); + }); + + test('proposeEdges: text_en: null falls back to text', () => { + const edges = ep.proposeEdges({ id: 'R1', text: 'Round a number to N decimal places', text_en: null }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']); + }); + + test('validateRequirement: rejects empty-string text_en (?? does not catch \'\')', () => { + // Nullish coalescing only falls back on null/undefined — an empty string would + // otherwise win `text_en ?? text` and silently classify against '', degrading to + // zero shapes with no signal (the exact fail-open #1110/#2773 exist to eliminate). + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: '' }), + /text_en must be a non-empty string when present/i, + ); + }); + + test('validateRequirement: rejects whitespace-only text_en', () => { + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: ' ' }), + /text_en must be a non-empty string when present/i, + ); + }); + + test('validateRequirement: rejects non-string text_en (number/array/object)', () => { + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: 42 }), + /text_en must be a non-empty string when present/i, + ); + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: ['x'] }), + /text_en must be a non-empty string when present/i, + ); + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'Round a number', text_en: {} }), + /text_en must be a non-empty string when present/i, + ); + }); + + test('proposeEdges: authored shapes override still wins when text_en is also present', () => { + const edges = ep.proposeEdges({ id: 'R9', text: pt, text_en: en, shapes: ['numeric-range'] }); + assert.deepEqual(edges.map((e) => e.category).sort(), ['boundary', 'precision']); + }); + + test('validateRequirement: rejects empty text_en even when shapes override makes it unused', () => { + // Validation is unconditional — it does not skip the text_en check just because the + // classify branch would never run. Bad data fails closed regardless of whether it + // happens to be dead for this particular call. + assert.throws( + () => ep.validateRequirement({ id: 'R1', text: 'x', text_en: '', shapes: ['collection'] }), + /text_en must be a non-empty string when present/i, + ); + }); +}); + +describe('edge-probe: SHAPE_CUES/VALID_SHAPES/Shape vocabulary stay single-sourced (DEFECT.GENERATIVE-FIX)', () => { + // RULESET.GENERATIVE-FIX (CONTEXT.md): parallel implementations diverge silently when no + // parity test enforces equality at the test layer. VALID_SHAPES is derived from + // Object.keys(SHAPE_CUES) in source, but that construction alone is not a regression + // guard — this test fails if a future edit ever hardcodes one of them independently or + // adds/removes a shape from only one side. + const LOCKED_SHAPES = ['numeric-range', 'collection', 'text', 'stateful', 'io']; + + test('SHAPE_CUES keys match the locked 5-shape vocabulary exactly', () => { + assert.deepEqual(Object.keys(ep.SHAPE_CUES).sort(), [...LOCKED_SHAPES].sort()); + }); + + test('VALID_SHAPES matches SHAPE_CUES keys exactly (single source of truth)', () => { + assert.deepEqual([...ep.VALID_SHAPES].sort(), Object.keys(ep.SHAPE_CUES).sort()); + }); + + test('VALID_SHAPES matches the locked 5-shape vocabulary exactly', () => { + assert.deepEqual([...ep.VALID_SHAPES].sort(), [...LOCKED_SHAPES].sort()); + }); +}); + describe('edge-probe: golden fixtures', () => { const root = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe-fixtures'); const fixtures = fs.readdirSync(root).filter((d) =>