fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)

* test(#2773): failing-first contract and premise tests for translated edge-probe input

Locks the Step 5.5 contract that a response_language project must feed the
edge probe an English translation of each requirement's text, and binds that
advice to measured engine behavior: the same requirement classifies to zero
shapes in Portuguese and to collection/adjacency/empty/ordering in English.

Also pins the honest limit — the issue's own repro sentence classifies to []
in English too, so translation is necessary but not sufficient and the
authored shapes override is the documented fallback.

Red before the doc change; the assertions are all false today.

Refs #2773

* fix(#2773): feed the spec-phase edge probe English-translated requirement text

The shape cues in src/edge-probe.cts are English word-boundary regexes, so a
project running with response_language set wrote its SPEC requirements into the
Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero
shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy
contributed nothing and --auto left it all unresolved — the probe was a silent
no-op for exactly the spec type it exists to harden.

Step 5.5 now states that the $REQS_JSON payload is engine input rather than
user-facing output, so the response_language rule does not govern it: each
requirement's text carries a faithful English translation, the SPEC keeps its
original language, and requirement ids are never translated or renumbered. The
instruction sits before the heredoc on purpose — the downstream APPLICABLE=0
warning fires only when every requirement is unclassified, so a partly-classified
non-English spec would otherwise slip through with no signal at all.

Measured against the compiled engine: the same requirement returns [] in
Portuguese and collection -> adjacency/empty/ordering in English. Also measured:
the issue's own repro sentence returns [] in English too, so translation is
necessary but not sufficient — the instruction therefore points at the authored
shapes override for prose carrying no cue in any language rather than promising
that translation restores classification.

Doc scope only, per the triage disposition on the issue. The compiled engine is
untouched; the lang-hint / per-language cue-set fix is a separate follow-up.

Closes #2773

* fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path

Surfaced by the isolated security review of this branch. Between the mktemp and
the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about
their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before
exiting, while the empty/placeholder guard directly above it exited without one.
A spec run that tripped the placeholder check therefore stranded a temp file
holding the SPEC's requirement text in TMPDIR, once per failed run.

The added contract test walks the region between the mktemp and the
unconditional cleanup and asserts no exit path leaves the file behind, so the
two guards can no longer drift apart. Proven to bind: run against the pre-fix
file the walker reports the leaking exit; against the fixed file it reports none.

Refs #2773

* docs(#2773): record the edge probe's English-cue input constraint in the predicate store

The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and
docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real
gaps rather than incidental coupling.

CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for
classifyShape but did not record that SHAPE_CUES are English word-boundary
patterns — so the predicate store implied text was language-agnostic, which is
what a future agent reads before touching this seam.

docs/CONFIGURATION.md's response_language row is what a non-English project
reads when it turns the setting on; it now names the one deliberate exception
and links to the FEATURES.md explanation, so the interaction is discoverable
from the config key rather than only from the workflow.

CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack
fragment is updated for the final byte range and now also records the
placeholder-guard cleanup fix folded into the same block.

Refs #2773

* fix(#2773): append the growth rationale to the existing spec-phase.md ack entry

The remote runner caught this: emitted-attribution.test.cjs pins the
0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration
regression test asserts the exact '31987 -> 31997' delta text survives), so
removing it to avoid a duplicate-key collision with a new fragment broke that
test instead of satisfying the ratchet.

The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102
were each appended to the same reason string by later PRs, which is how a shared
growth key coexists with the rule that two ack sources may never name the same
path. This appends the #2773 rationale the same way and drops the separate
fragment, whose spec-phase.md key was the collision.

Verified locally by reproducing both affected tests against the real fragment
before re-dispatching: the pinned delta survives, grown[0].acked is true,
staleAcks is empty, and all 35 entries still read as spent.

Refs #2773

* docs(#2773): add a how-to for probing edges in a non-English project

The phase gate's enablementSequence check caught a wrong call of mine. I had
recorded that no how-to was owed because the user takes zero extra steps — the
workflow translates the probe input itself. Written out, though, the sequence
from off to value is two steps and step 1 depends on response_language, a
setting owned by a different capability than the edge probe, which is exactly
the condition the how-to test names.

There is also real task content a reference table cannot carry: the three-way
split between a few unclassified rows (the classifier's recall gap), every row
unclassified (the probe could not read the spec at all), and the silent
partly-classified case where the APPLICABLE=0 warning never fires. That last
one is what a user would otherwise misread as a clean bill of health.

Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard
siblings and indexed from docs/README.md next to its closest relative.

Refs #2773

* chore(#2773): backfill the changeset PR number

pr:0 placeholder replaced with the real PR number now that #3713 exists.

Refs #2773

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-08-20 13:36:00 -04:00
committed by GitHub
parent adb46cdd85
commit 77fa08f1e8
10 changed files with 306 additions and 3 deletions

View File

@@ -0,0 +1,5 @@
---
type: Fixed
pr: 3713
---
**The spec-phase edge probe now classifies requirements in non-English projects** — a project running with `response_language` set had every requirement fall through the English-only shape cues into `unclassified`, silently disabling the whole edge taxonomy; Step 5.5 now feeds the probe an English translation of each requirement while the SPEC keeps its original language. (#2773)

View File

@@ -432,7 +432,7 @@ Deterministic eval-scoring projection (#10 / #1579) that moves the `gsd-eval-aud
Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the Prohibition Probe Module (#644) next. Exports (generic surface): `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` — the prohibition adapter exports that also ship from this module (`projectProhibitions`, `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, `dispositionForProhibition`) are documented under the Prohibition Probe Module's own locked-surface line. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. Under ADR-857 (phase-6 boundary, settled 2026-06-12) this seam is classified **core verification substrate** on the *contract* side: its deterministic validators are the verifier↔predicate contract's CI-testable surface (ADR-550 Decision 5) — core and non-toggleable, never an off-by-default Feature Capability. (The recall-gapped *generator* is the probe adapters that propose predicates, not this resolution engine — see Edge Probe Module and Verification substrate (predicate boundary).)
### Edge Probe Module
First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability.
First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. `SHAPE_CUES` are **English** word-boundary patterns, so `text` is an English-language input regardless of the project's `response_language`: spec-phase Step 5.5 passes a faithful English translation of each requirement while the SPEC itself keeps its original language and the `id` is never translated (#2773 — without it every requirement in a non-English project matched no cue and landed in `unclassified`, silently disabling the taxonomy). Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability.
### Verification substrate (predicate boundary)
The ADR-857 classification (settled 2026-06-12, prompted by @davesienkowski's boundary analysis on #857) that predicate-generation — the must-NOT-have / edge predicates that set the verifier's reach — is **core**, not an off-by-default Feature Capability. Load-bearing premise: *verifier reach = spec reach* (the verifier can only catch what the spec concretely names). Decomposes into: the **verifier↔predicate contract** (the verifier always expects predicates and grades **exogenously** against them — core, non-toggleable, a stability contract alongside the Loop Extension Point names), and the **generator** (the probe *adapters* that propose predicates — edge-probe's `classifyShape`/`proposeEdges`, the prohibition probe's adversarial LLM-propose — core-default but independently versionable, kept their own module because of a measured recall gap; `probe-core`'s deterministic validators sit on the contract side, not the generator side). Altitude rule distinguishing it from `gate` hooks: a gate runs *against* the spec (hook); predicate-generation defines the spec's *reach* (core). The decision-#6 `produces`/`consumes` artifact flow is the internal rail from generation to the core verifier. See ADR-857 *Verification substrate vs. plug-in tier (the predicate boundary)* and the ADR-550 cross-reference.

View File

@@ -187,7 +187,7 @@ project one is reported, since that is the file you are most likely able to fix.
| `dynamic_routing.provider_escalation` | string[] | ordered model IDs | (none) | Opt-in fallback providers tried when a run dies on a quota / rate limit — see [provider escalation](#provider-escalation-on-quota-exceeded--added-in-v143). Added in v1.43 ([#2296](https://github.com/open-gsd/gsd-core/issues/2296)) |
| `project_code` | string | any short string | (none) | Prefix for phase directory names (e.g., `"ABC"` produces `ABC-01-setup/`). Added in v1.31 |
| `phase_id_convention` | enum | `"milestone-prefixed"`, `null` | `null` | Phase ID naming convention. `null` = legacy numeric IDs (`Phase 1`, `Phase 2`). `"milestone-prefixed"` = globally unique IDs that encode the enclosing milestone (`Phase 1-01`, `Phase 1-02`). Run `gsd-tools roadmap upgrade --convention milestone-prefixed` to migrate an existing ROADMAP.md. |
| `response_language` | string | language code | (none) | Language for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`). Propagates to all spawned agents for cross-phase language consistency. Added in v1.32. UAT checkpoint frames (`/gsd-verify-work`) render a localized banner/instruction for English, Spanish, French, German, Portuguese, Japanese, Chinese, Korean, Italian, Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian (endonyms and ISO codes also accepted); any other value falls back to the English frame. |
| `response_language` | string | language code | (none) | Language for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`). Propagates to all spawned agents for cross-phase language consistency. Added in v1.32. UAT checkpoint frames (`/gsd-verify-work`) render a localized banner/instruction for English, Spanish, French, German, Portuguese, Japanese, Chinese, Korean, Italian, Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian (endonyms and ISO codes also accepted); any other value falls back to the English frame. One deliberate exception: the `spec-phase` edge-completeness probe is fed an English translation of each requirement's text, because its shape cues are English-only — the SPEC itself stays in this language. See [Spec-Phase Edge-Completeness Probe](FEATURES.md#144-spec-phase-edge-completeness-probe). |
| `context_window` | number | any integer | `200000` | Context window size in tokens. Set `1000000` for 1M-context models (e.g., `claude-fable-5`). Values `>= 500000` enable adaptive context enrichment (full-body reads of prior SUMMARY.md, deeper anti-pattern reads). Configured via `/gsd-config --advanced`. |
| `context_profile` | string | `dev`, `research`, `review` | (none) | Execution context preset that applies a pre-configured bundle of mode, model, and workflow settings for the current type of work. Added in v1.34 |
| `claude_md_path` | string | any file path | `./.claude/CLAUDE.md` | Custom output path for the generated CLAUDE.md file. Useful for monorepos or projects that need CLAUDE.md in a non-root location. Defaults to `./.claude/CLAUDE.md` — a valid project-scoped memory location that keeps GSD-generated content from polluting a hand-crafted repo-root `CLAUDE.md` ([#1098](https://github.com/open-gsd/gsd-core/issues/1098)). An existing file without GSD markers is never overwritten unless `--force` is passed. Default changed from `./CLAUDE.md` in v1.5. Added in v1.36 |

View File

@@ -3179,6 +3179,8 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev
When a requirement's prose matches **no** shape cue, the probe does not silently drop it (#1110): it emits a single `unclassified — review manually` candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block.
**Non-English projects: the probe reads English, the SPEC does not have to (#2773).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore feeds the probe a faithful **English translation** of each requirement's `text`: that payload is engine input, never user-facing output, so it is translated while the SPEC itself stays in the original language, requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification.
The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim.
The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. A `backstop` edge is lifted as a **structured non-inferable marker** (`{ statement, verification: backstop }`, a flat scalar — not a prose note), which the **honest verifier** then consumes (see below) — closing the loop the edge-probe opened.

View File

@@ -22,6 +22,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [Attach a plugin-provided skill to a GSD agent](how-to/attach-a-plugin-skill-to-a-gsd-agent.md) — use the `global:plugin:skill` entry form to load Claude Code plugin skills into agent prompts
- [Discuss a phase](how-to/discuss-a-phase.md) — capture implementation decisions before planning begins
- [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- [Probe edges in a non-English project](how-to/probe-edges-in-a-non-english-project.md) — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it"
- [Resolve prohibition findings](how-to/resolve-prohibition-findings.md) — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- [Resolve an unreachable-workflow finding](how-to/resolve-unreachable-workflow-findings.md) — wire or fully sweep a shipped workflow that no command, agent, or skill references
- [Resolve verify-command path findings](how-to/resolve-verify-command-path-findings.md) — fix an `<automated>` verify command whose target directory does not resolve from the executor's cwd

View File

@@ -0,0 +1,60 @@
# How to probe edges in a non-English project
**Goal:** Get real edge-completeness coverage on a spec written in a language other than English, instead of every requirement landing in `unclassified` and the whole taxonomy quietly contributing nothing.
**Prerequisites:** A project with `response_language` set (see [Configuration](../CONFIGURATION.md)), and a phase whose `/gsd-spec-phase` run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically — you do not invoke it separately.
For the taxonomy and the reasoning behind front-of-pipeline edge analysis, see [Spec-Phase Edge-Completeness Probe](../FEATURES.md#144-spec-phase-edge-completeness-probe). For how to act on findings once they appear, see [Resolve edge-coverage findings](resolve-edge-coverage-findings.md). This guide covers only what is different when your spec is not in English.
---
## What happens, and why
The probe classifies each requirement's data/behavior shape by matching **English** word-boundary cues against the requirement text. A requirement written in another language matches nothing, classifies to zero shapes, raises zero categories, and surfaces as a single `unclassified — review manually` row.
That is not a rejection you can act on — it looks identical to a genuinely edge-free requirement. When it happens to *every* requirement, the probe has contributed nothing to the spec.
So Step 5.5 sends the probe an English rendering of each requirement, while your spec stays in your language. Concretely, for the same requirement:
| Requirement text handed to the probe | Shapes | Edges raised |
|---|---|---|
| `O sistema mescla intervalos sobrepostos em uma lista ordenada` | none | none — one `unclassified` row |
| `The system merges overlapping intervals in a sorted list` | `collection` | `adjacency`, `empty`, `ordering` |
## What you do
Nothing extra. The translation happens inside Step 5.5 as part of the run.
What you should *see* is the split: your **spec stays in `response_language`** — its requirements, its acceptance criteria, its `## Edge Coverage` section — while the probe's findings are reasoned about from an English rendering of the requirement text. Requirement ids (`R1`, `R2`, …) are never translated or renumbered, so a finding always names the same requirement you wrote.
If your spec comes back anglicized, that is a bug worth reporting — only the probe's transient input is translated, never the document.
## Tell "no edges here" apart from "the probe could not read it"
This is the distinction that matters, because both look like an `unclassified` row.
| What you see | What it means | What to do |
|---|---|---|
| A few `unclassified` rows among normally-classified ones | Those requirements carry no shape cue **in any language**. This is the classifier's known recall gap, not a translation problem. | Resolve each like any other finding — or author an explicit `shapes` array on the requirement (below). |
| **Every** requirement `unclassified`, and a `WARNING: edge-probe proposed ZERO applicable edges` | The probe could not read your requirements at all. | Confirm the run really is translating the probe input. Do not accept an empty `## Edge Coverage` section. |
| Some requirements classified, the rest `unclassified`, no warning | **The silent case.** The zero-applicable warning fires only when *all* requirements are unclassified, so a partly-classified spec raises nothing. | Check the `unclassified` ones individually against the row above. |
Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in English either — for example "the command exits with code 1 on invalid input" — still classifies to zero. That is expected, and the fix is the same one an English-language project uses.
## Force the shape when the prose carries no cue
When a requirement is genuinely edge-relevant but no cue fires, do not fight the wording. Author the shape explicitly — an authored `shapes` array bypasses prose classification entirely, in any language:
```json
{ "id": "R4", "text": "The command exits with code 1 on invalid input", "shapes": ["stateful"] }
```
`shapes` accepts any of `numeric-range`, `collection`, `text`, `stateful`, `io`. The example above raises `idempotency` and `concurrency`.
An explicit empty array — `"shapes": []` — is the opposite signal: your deliberate "this requirement has no edge surface", which stays silent rather than surfacing an `unclassified` row.
## Related
- [Resolve edge-coverage findings](resolve-edge-coverage-findings.md) — what to do with each finding once it is raised
- [Spec-Phase Edge-Completeness Probe](../FEATURES.md#144-spec-phase-edge-completeness-probe) — the taxonomy and the rationale
- [Configuration](../CONFIGURATION.md) — the `response_language` setting

View File

@@ -47,6 +47,14 @@ The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. Wh
`shapes` is absent, a heuristic classifier proposes them from the requirement prose
(propose-then-confirm) — the author may correct the shape.
**`text` is English, whatever language the SPEC is in.** The heuristic cues are English
word-boundary patterns, so a project running with `response_language` set must pass `text` as a
faithful English translation of the requirement; the SPEC itself keeps the original language and
the `id` is never translated. Prose in another language matches no cue, classifies to zero
shapes, and surfaces as `unclassified` (#1110) — the probe contributes nothing. Where a
requirement carries no cue even in English, author `shapes` explicitly rather than leaning on
the classifier.
## Taxonomy (8 categories)
Closed and small by design: a fixed eight the author must explicitly clear beats thirty

View File

@@ -192,6 +192,24 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met):
Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague
ones). Reference: @~/.claude/gsd-core/references/edge-probe.md.
**Non-English projects — the probe input is translated, the SPEC is not.** The shape cues the
classifier matches are **English** word-boundary patterns, so requirement prose written in
another language matches nothing, classifies to zero shapes, and lands every row in
`unclassified` (#1110) — the taxonomy contributes nothing and `--auto` leaves it all
`unresolved`. The `$REQS_JSON` payload below is **engine input, never user-facing output**, so
the `response_language` rule at the top of this workflow does not govern it: write each
requirement's `text` as a faithful **English** translation of the SPEC requirement.
The SPEC keeps the original language — only the probe input is translated, and requirement
`id`s are never translated or renumbered (coverage rows join back on `id`, and any Acceptance
Criteria you write from the resolved edges go into the SPEC in `response_language`).
Translate **every** requirement, not only the ones that look edge-relevant: the
`$APPLICABLE = 0` warning below fires only when *all* requirements are unclassified, so a
partly-classified spec slips through with no signal at all.
If a requirement still classifies to zero shapes after translation, it carries no cue in any
language (the recorded recall gap — ADR-857 §98 / ADR-550 D7b, not a translation failure);
author an explicit `shapes` array on that requirement instead of relying on the prose
classifier.
**Runtime coverage compute — resolve and invoke edge-probe.cjs:**
```bash
@@ -235,6 +253,9 @@ fi
# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object
# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes
# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04).
# When `response_language` is set, `text` carries a faithful ENGLISH translation (see above) —
# the shape cues are English-only, so original-language prose classifies to zero shapes and the
# probe becomes a silent no-op. Keep every `id` exactly as it appears in the SPEC.
# BSD/macOS mktemp only randomizes XXXXXX when it is the final path component, so make a
# suffixless temp then append the extension — portable across BSD + GNU (#1520).
REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX") && mv "$REQS_JSON" "${REQS_JSON}.json" && REQS_JSON="${REQS_JSON}.json" || exit 1
@@ -247,6 +268,7 @@ JSON
# `<replace: …>` placeholder (a forgotten substitution would otherwise yield a
# meaningful-looking but bogus coverage report). Fail loud, not silent no-op.
if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("<replace:")))process.exit(1)' "$REQS_JSON" 2>/dev/null; then
rm -f "$REQS_JSON"
echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the <replace: …> placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2
exit 1
fi

View File

@@ -14,8 +14,13 @@ const { test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const fc = require('fast-check');
const SPEC_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'spec-phase.md');
const EDGE_PROBE_REF_PATH = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe.md');
const { classifyShape, applicableCategories } = require(
path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs'),
);
function readSpecPhase() {
return fs.readFileSync(SPEC_PHASE_PATH, 'utf8');
@@ -299,3 +304,203 @@ test('#3132: specless-probe-fallback.md uses resolved+verification — not cover
assert.match(content, /verification: explicit/,
'specless-probe-fallback.md must reference "verification: explicit"');
});
// ─── #2773: non-English requirements and the English-only shape cues ──────────
//
// `SHAPE_CUES` (src/edge-probe.cts) are English word-boundary regexes. A project that sets
// `response_language` writes its SPEC Requirements in that language, so transcribing them
// verbatim into the Step 5.5 `$REQS_JSON` heredoc classifies every requirement to zero shapes
// -> every row becomes the `unclassified` sentinel (#1110) and the 8-category taxonomy
// contributes nothing. Approved scope for #2773 is doc-only: Step 5.5 must instruct that the
// probe's `text` field carries a faithful English translation (engine input, never
// user-facing), while the SPEC itself stays in the original language.
test('#2773: Step 5.5 documents the English-translation step for response_language projects', () => {
const block = extractStep55Block(readSpecPhase());
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
assert.match(
block,
/response_language/,
'Step 5.5 must name `response_language` — it is the setting that makes the requirement prose non-English',
);
assert.match(
block,
/translat/i,
'Step 5.5 must instruct that the probe input carries a translation, not the original-language prose',
);
assert.match(
block,
/English/,
'Step 5.5 must say the translation target is English (the cue set the classifier actually speaks)',
);
// The SPEC must NOT be anglicized — only the transient probe payload is translated.
assert.match(
block,
/SPEC[^\n]*(original language|stays in|keeps)|(original language)[^\n]*SPEC/i,
'Step 5.5 must state the SPEC keeps the original language — only the probe input is translated',
);
// Requirement ids are the join key for coverage rows; translating them breaks the mapping.
assert.match(
block,
/`?id`?s?[^\n]*(unchanged|not translated|never translated|stable)/i,
'Step 5.5 must state requirement ids are NOT translated or renumbered (coverage rows join on id)',
);
});
test('#2773: Step 5.5 names the authored shapes override as the zero-cue fallback', () => {
const block = extractStep55Block(readSpecPhase());
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
// Translation is necessary but NOT sufficient: prose carrying no cue in ANY language still
// classifies to []. The engine already accepts an authored `shapes` override for exactly
// that case, so the instruction must point at it rather than over-promise.
assert.match(
block,
/`shapes`/,
'Step 5.5 must name the authored `shapes` override as the fallback for prose that still classifies to zero',
);
});
test('#2773: the translation instruction precedes the $REQS_JSON write', () => {
const block = extractStep55Block(readSpecPhase());
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
const langIdx = block.search(/response_language/);
const writeIdx = block.search(/>\s*"\$REQS_JSON"/);
assert.ok(langIdx !== -1, 'Step 5.5 must mention response_language');
assert.ok(writeIdx !== -1, 'Step 5.5 must write $REQS_JSON');
assert.ok(
langIdx < writeIdx,
'the translation instruction must come BEFORE the $REQS_JSON write — the downstream `$APPLICABLE = 0` warning only fires when EVERY requirement is unclassified, so a partially-classified non-English spec would slip through silently',
);
});
test('#2773: translating a non-English requirement is what makes the cue matcher apply', () => {
// Same requirement, two languages. This is the premise the Step 5.5 instruction rests on;
// if a future SHAPE_CUES edit breaks it, the documented advice becomes false and this fails.
const pt = 'O sistema mescla intervalos sobrepostos em uma lista ordenada';
const en = 'The system merges overlapping intervals in a sorted list';
assert.deepEqual(classifyShape(pt), [], 'non-English prose matches no English cue — the silent no-op #2773 reports');
assert.deepEqual(applicableCategories(classifyShape(pt)), [], 'zero shapes raise zero categories');
const enShapes = classifyShape(en);
assert.ok(enShapes.includes('collection'), 'the English rendering must classify as a collection');
const enCategories = applicableCategories(enShapes);
for (const expected of ['adjacency', 'empty', 'ordering']) {
assert.ok(enCategories.includes(expected), `translated requirement must raise \`${expected}\``);
}
// A second shape, so the row is not a single-cue coincidence.
const ptText = 'O nome do usuario e truncado em 50 caracteres';
const enText = 'The user name is truncated at 50 characters';
assert.deepEqual(classifyShape(ptText), [], 'non-English text-shape prose also matches nothing');
assert.ok(classifyShape(enText).includes('text'), 'the English rendering must classify as text');
// Multi-cue: a sentence carrying cues for two shapes yields the UNION, not one of them.
const multi = 'The API request uploads a sorted list of items';
const multiShapes = classifyShape(multi);
assert.ok(multiShapes.includes('io'), 'multi-cue prose must include io');
assert.ok(multiShapes.includes('collection'), 'multi-cue prose must include collection');
});
test('#2773: translation alone does not rescue genuinely zero-cue prose', () => {
// The issue's own repro sentence classifies to [] in ENGLISH too — it carries no shape cue
// in any language. That is the recorded recall gap (ADR-857 §98 / ADR-550 D7b), not a
// language failure, which is why Step 5.5 must point at the `shapes` override rather than
// promise that translation restores classification.
const enZeroCue = 'The command exits with code 1 and prints to stderr on invalid input';
assert.deepEqual(
classifyShape(enZeroCue),
[],
'a genuinely zero-cue requirement stays unclassified even in English — the doc must not over-promise',
);
// The authored `shapes` override is the deterministic escape hatch for exactly this case.
assert.deepEqual(
applicableCategories(['stateful']),
['idempotency', 'concurrency'],
'an authored `shapes` override raises categories with no dependence on prose cues',
);
});
test('#2773 property: a cue word classifies regardless of surrounding text', () => {
// A translated sentence carries its cue word amid arbitrary other words. The Step 5.5
// advice is only sound if classification is robust to that surrounding context rather than
// anchored to a fixed sentence shape.
const cueForShape = {
'numeric-range': 'threshold',
collection: 'items',
text: 'unicode',
stateful: 'persist',
io: 'endpoints',
};
const filler = fc.string({ minLength: 0, maxLength: 24 }).filter((s) => !/[A-Za-z]/.test(s));
fc.assert(
fc.property(
fc.constantFrom(...Object.keys(cueForShape)),
filler,
filler,
(shape, before, after) => {
const sentence = `${before} ${cueForShape[shape]} ${after}`;
assert.ok(
classifyShape(sentence).includes(shape),
`cue "${cueForShape[shape]}" must classify as ${shape} inside ${JSON.stringify(sentence)}`,
);
},
),
{ numRuns: 100, seed: 2773 },
);
});
test('#2773: the edge-probe reference documents the English-cue assumption', () => {
const ref = fs.readFileSync(EDGE_PROBE_REF_PATH, 'utf8');
assert.match(
ref,
/English/,
'the edge-probe reference `## Inputs` contract must state that the heuristic classifier is English-cue based',
);
assert.match(
ref,
/response_language/,
'the edge-probe reference must point non-English projects at the translated-input requirement',
);
});
test('#2773: no Step 5.5 exit path leaks the $REQS_JSON temp file', () => {
// The temp file holds the SPEC's requirement text. Every guard between its creation and
// the unconditional cleanup must `rm -f` it before `exit 1`, or a failed spec run strands
// requirement content in TMPDIR. The engine-failure guard always did; the empty/placeholder
// guard directly above it did not, so the two siblings disagreed about their own invariant.
const block = extractStep55Block(readSpecPhase());
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
const lines = block.split('\n');
const createIdx = lines.findIndex((l) => /REQS_JSON=\$\(mktemp/.test(l));
assert.ok(createIdx !== -1, 'Step 5.5 must create $REQS_JSON via mktemp');
// The region ends at the first UNCONDITIONAL cleanup (a bare `rm -f "$REQS_JSON"` at column
// zero); past that the file is already gone and later exits cannot leak it.
const afterCreate = lines.slice(createIdx + 1);
const endOffset = afterCreate.findIndex((l) => /^rm -f "\$REQS_JSON"/.test(l));
assert.ok(endOffset !== -1, 'Step 5.5 must unconditionally rm -f "$REQS_JSON" after the engine run');
const region = afterCreate.slice(0, endOffset);
// Walk the region tracking whether the current guard branch has cleaned up. `then`/`else`
// opens a fresh branch; a cleanup inside it arms the branch; an `exit` must find it armed.
let cleanedInBranch = false;
const leaks = [];
for (const line of region) {
if (/\bthen\b|^\s*else\b|^\s*elif\b/.test(line)) cleanedInBranch = false;
if (/rm -f "\$REQS_JSON"/.test(line)) cleanedInBranch = true;
if (/^\s*exit\s+\d+/.test(line) && !cleanedInBranch) leaks.push(line.trim());
}
assert.deepEqual(
leaks,
[],
`every exit between the mktemp and the unconditional cleanup must rm -f "$REQS_JSON" first; leaking exits: ${JSON.stringify(leaks)}`,
);
});

View File

@@ -36,7 +36,7 @@
"agents/gsd-verifier.toml": "#2834: Codex agent TOML now carries model-routing fields on first install.",
"gsd-code-fixer.md": "#2647: the three worktree-path sites (setup_worktree bash, concrete-steps prose, critical_rules) replaced the hardcoded /tmp/sv- mktemp path with a repo-relative .claude/worktrees/rf-<phase>-<pid>-<epoch> path, and added a defense-in-depth padded_phase validation at the sink (the agent prompt is a literal bash contract any caller can spawn; the orchestrator validates upstream but the sink now self-defends against path-traversal/branch-name injection). On Windows/Git Bash the /tmp path landed outside the project tree (outside the session permission allowlist, prompting on every read) and mktemp's MAX_PATH substitute was un-removable; .claude/worktrees/ is the same dir the harness-managed executor worktrees use (gitignored via .claude/, inside the permission scope). Growth is the path-resolution bash (main_repo via `git worktree list --porcelain | awk`) + the $$-PID/epoch uniqueness replacing mktemp's XXXXXX + the padded_phase guard + the #2647 rationale comments at each site. Supersedes the prior #2825 attribution, whose gated-bash + guardrail growth is already in next.",
"spec-phase.md": {
"reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification. #3102: Step 5.5 now RENDERS the edge-probe coverage report into the model's context (a raw printf of $COVERAGE after the well-formedness guard) and binds those rows in the resolution loop and --auto as a deterministic FLOOR, plus the corrected block comment; load-bearing workflow instruction that makes ADR-550 D7b (deterministic propose + LLM resolve) real at runtime and honors ADR-857 sec98's recall gap (floor, not ceiling). 32238 -> 34056 bytes (+1818), DEFAULT tier, cap 40960."
"reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification. #3102: Step 5.5 now RENDERS the edge-probe coverage report into the model's context (a raw printf of $COVERAGE after the well-formedness guard) and binds those rows in the resolution loop and --auto as a deterministic FLOOR, plus the corrected block comment; load-bearing workflow instruction that makes ADR-550 D7b (deterministic propose + LLM resolve) real at runtime and honors ADR-857 sec98's recall gap (floor, not ceiling). 32238 -> 34056 bytes (+1818), DEFAULT tier, cap 40960. #2773: Step 5.5 now tells a `response_language` project that the edge-probe `$REQS_JSON` payload is engine input rather than user-facing output, so each requirement's `text` carries a faithful English translation while the SPEC keeps its original language and requirement ids stay unchanged. The shape cues in src/edge-probe.cts are English word-boundary regexes, so prose in another language matched nothing, classified to zero shapes, and put every row in the `unclassified` sentinel (#1110) — the taxonomy contributed nothing. Growth is that instruction plus the pointer to the authored `shapes` override for prose that classifies to zero even in English (measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient), and a one-line fix an isolated security review surfaced in the same block — the empty/placeholder guard exited without `rm -f \"$REQS_JSON\"` while its engine-failure sibling below it did, stranding the SPEC requirement text in TMPDIR on every failed run. The instruction sits BEFORE the heredoc deliberately: the `$APPLICABLE = 0` warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal. Load-bearing runtime instruction; the compiled engine is deliberately untouched (the `lang`-hint / per-language cue-set fix is out of scope per the #2773 triage). 34020 -> 35730 bytes (+1710), DEFAULT tier, cap 40960."
}
}
}