diff --git a/.changeset/steady-zebras-leap.md b/.changeset/steady-zebras-leap.md new file mode 100644 index 000000000..dbdd59d6a --- /dev/null +++ b/.changeset/steady-zebras-leap.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 3713 +--- +**The spec-phase edge probe now classifies requirements in non-English projects** — a project running with `response_language` set had every requirement fall through the English-only shape cues into `unclassified`, silently disabling the whole edge taxonomy; Step 5.5 now feeds the probe an English translation of each requirement while the SPEC keeps its original language. (#2773) diff --git a/CONTEXT.md b/CONTEXT.md index 78b7b4bf4..4f92aea2a 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -432,7 +432,7 @@ Deterministic eval-scoring projection (#10 / #1579) that moves the `gsd-eval-aud Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the Prohibition Probe Module (#644) next. Exports (generic surface): `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` — the prohibition adapter exports that also ship from this module (`projectProhibitions`, `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, `dispositionForProhibition`) are documented under the Prohibition Probe Module's own locked-surface line. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. Under ADR-857 (phase-6 boundary, settled 2026-06-12) this seam is classified **core verification substrate** on the *contract* side: its deterministic validators are the verifier↔predicate contract's CI-testable surface (ADR-550 Decision 5) — core and non-toggleable, never an off-by-default Feature Capability. (The recall-gapped *generator* is the probe adapters that propose predicates, not this resolution engine — see Edge Probe Module and Verification substrate (predicate boundary).) ### Edge Probe Module -First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability. +First adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase edge-completeness probe wired into spec-phase Step 5.5. Owns shape classification (`classifyShape`), the applicable-category relevance filter (`applicableCategories` over the 8-category edge `TAXONOMY`), edge proposal (`proposeEdges`), and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`. Fail-closed input contract: an edge requirement with missing/empty `text` and no `shapes` override is rejected (a `{ id }`-only requirement no longer classifies to zero edges and silently drops), while the legitimate `shapes: []` opt-out is preserved. `SHAPE_CUES` are **English** word-boundary patterns, so `text` is an English-language input regardless of the project's `response_language`: spec-phase Step 5.5 passes a faithful English translation of each requirement while the SPEC itself keeps its original language and the `id` is never translated (#2773 — without it every requirement in a non-English project matched no cue and landed in `unclassified`, silently disabling the taxonomy). Downstream, the `plan-phase` planner lifts every `covered`/`backstop` edge from the SPEC `## Edge Coverage` section into `must_haves.truths`. Exports (locked surface): `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `validateRequirement`, plus the constants `TAXONOMY` (the **closed** 8 edge categories), `UNCLASSIFIED_CATEGORY` (the `unclassified` review-manually sentinel for zero-cue prose, #1110 — deliberately **not** a 9th taxonomy entry: it stays out of `TAXONOMY` but is present in `EDGE_VALIDATORS.categories`), `VALID_SHAPES`, `SHAPE_CUES`, and `EDGE_VALIDATORS` (the `{explicit, backstop}` validators bundle injected into `probe-core`'s generic engine; its `categories` = the `TAXONOMY` ids plus `UNCLASSIFIED_CATEGORY`). Source of truth: `gsd-core/bin/lib/edge-probe.cjs` (generated from `src/edge-probe.cts`, gitignored per ADR-457). Tests: `tests/edge-probe.test.cjs`, `tests/edge-probe-spec-phase-contract.test.cjs`, `tests/edge-probe-planner-contract.test.cjs`. See ADR-550 and Probe Core Module. Per ADR-857's phase-6 boundary (2026-06-12) the predicates this module generates are **core verification substrate** (they set the verifier's reach), so it is wired onto the core predicate rail as a core-default module rather than migrated to an off-by-default `capabilities/edge-probe/` Feature Capability. ### Verification substrate (predicate boundary) The ADR-857 classification (settled 2026-06-12, prompted by @davesienkowski's boundary analysis on #857) that predicate-generation — the must-NOT-have / edge predicates that set the verifier's reach — is **core**, not an off-by-default Feature Capability. Load-bearing premise: *verifier reach = spec reach* (the verifier can only catch what the spec concretely names). Decomposes into: the **verifier↔predicate contract** (the verifier always expects predicates and grades **exogenously** against them — core, non-toggleable, a stability contract alongside the Loop Extension Point names), and the **generator** (the probe *adapters* that propose predicates — edge-probe's `classifyShape`/`proposeEdges`, the prohibition probe's adversarial LLM-propose — core-default but independently versionable, kept their own module because of a measured recall gap; `probe-core`'s deterministic validators sit on the contract side, not the generator side). Altitude rule distinguishing it from `gate` hooks: a gate runs *against* the spec (hook); predicate-generation defines the spec's *reach* (core). The decision-#6 `produces`/`consumes` artifact flow is the internal rail from generation to the core verifier. See ADR-857 *Verification substrate vs. plug-in tier (the predicate boundary)* and the ADR-550 cross-reference. diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index 614153037..9de353f68 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -187,7 +187,7 @@ project one is reported, since that is the file you are most likely able to fix. | `dynamic_routing.provider_escalation` | string[] | ordered model IDs | (none) | Opt-in fallback providers tried when a run dies on a quota / rate limit — see [provider escalation](#provider-escalation-on-quota-exceeded--added-in-v143). Added in v1.43 ([#2296](https://github.com/open-gsd/gsd-core/issues/2296)) | | `project_code` | string | any short string | (none) | Prefix for phase directory names (e.g., `"ABC"` produces `ABC-01-setup/`). Added in v1.31 | | `phase_id_convention` | enum | `"milestone-prefixed"`, `null` | `null` | Phase ID naming convention. `null` = legacy numeric IDs (`Phase 1`, `Phase 2`). `"milestone-prefixed"` = globally unique IDs that encode the enclosing milestone (`Phase 1-01`, `Phase 1-02`). Run `gsd-tools roadmap upgrade --convention milestone-prefixed` to migrate an existing ROADMAP.md. | -| `response_language` | string | language code | (none) | Language for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`). Propagates to all spawned agents for cross-phase language consistency. Added in v1.32. UAT checkpoint frames (`/gsd-verify-work`) render a localized banner/instruction for English, Spanish, French, German, Portuguese, Japanese, Chinese, Korean, Italian, Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian (endonyms and ISO codes also accepted); any other value falls back to the English frame. | +| `response_language` | string | language code | (none) | Language for agent responses (e.g., `"pt"`, `"ko"`, `"ja"`). Propagates to all spawned agents for cross-phase language consistency. Added in v1.32. UAT checkpoint frames (`/gsd-verify-work`) render a localized banner/instruction for English, Spanish, French, German, Portuguese, Japanese, Chinese, Korean, Italian, Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian (endonyms and ISO codes also accepted); any other value falls back to the English frame. One deliberate exception: the `spec-phase` edge-completeness probe is fed an English translation of each requirement's text, because its shape cues are English-only — the SPEC itself stays in this language. See [Spec-Phase Edge-Completeness Probe](FEATURES.md#144-spec-phase-edge-completeness-probe). | | `context_window` | number | any integer | `200000` | Context window size in tokens. Set `1000000` for 1M-context models (e.g., `claude-fable-5`). Values `>= 500000` enable adaptive context enrichment (full-body reads of prior SUMMARY.md, deeper anti-pattern reads). Configured via `/gsd-config --advanced`. | | `context_profile` | string | `dev`, `research`, `review` | (none) | Execution context preset that applies a pre-configured bundle of mode, model, and workflow settings for the current type of work. Added in v1.34 | | `claude_md_path` | string | any file path | `./.claude/CLAUDE.md` | Custom output path for the generated CLAUDE.md file. Useful for monorepos or projects that need CLAUDE.md in a non-root location. Defaults to `./.claude/CLAUDE.md` — a valid project-scoped memory location that keeps GSD-generated content from polluting a hand-crafted repo-root `CLAUDE.md` ([#1098](https://github.com/open-gsd/gsd-core/issues/1098)). An existing file without GSD markers is never overwritten unless `--force` is passed. Default changed from `./CLAUDE.md` in v1.5. Added in v1.36 | diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 55f86439a..0e66ce31a 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3179,6 +3179,8 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev When a requirement's prose matches **no** shape cue, the probe does not silently drop it (#1110): it emits a single `unclassified — review manually` candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block. +**Non-English projects: the probe reads English, the SPEC does not have to (#2773).** The shape cues are English word-boundary patterns, so a project running with [`response_language`](CONFIGURATION.md) set would otherwise have *every* requirement match nothing, classify to zero shapes, and land in `unclassified` — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. `spec-phase` Step 5.5 therefore feeds the probe a faithful **English translation** of each requirement's `text`: that payload is engine input, never user-facing output, so it is translated while the SPEC itself stays in the original language, requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to `response_language`. Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in **any** language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit `shapes` array on the requirement instead of relying on prose classification. + The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim. The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. A `backstop` edge is lifted as a **structured non-inferable marker** (`{ statement, verification: backstop }`, a flat scalar — not a prose note), which the **honest verifier** then consumes (see below) — closing the loop the edge-probe opened. diff --git a/docs/README.md b/docs/README.md index 922a859da..3dba62bf1 100644 --- a/docs/README.md +++ b/docs/README.md @@ -22,6 +22,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md) - [Attach a plugin-provided skill to a GSD agent](how-to/attach-a-plugin-skill-to-a-gsd-agent.md) — use the `global:plugin:skill` entry form to load Claude Code plugin skills into agent prompts - [Discuss a phase](how-to/discuss-a-phase.md) — capture implementation decisions before planning begins - [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions +- [Probe edges in a non-English project](how-to/probe-edges-in-a-non-english-project.md) — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it" - [Resolve prohibition findings](how-to/resolve-prohibition-findings.md) — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions - [Resolve an unreachable-workflow finding](how-to/resolve-unreachable-workflow-findings.md) — wire or fully sweep a shipped workflow that no command, agent, or skill references - [Resolve verify-command path findings](how-to/resolve-verify-command-path-findings.md) — fix an `` verify command whose target directory does not resolve from the executor's cwd diff --git a/docs/how-to/probe-edges-in-a-non-english-project.md b/docs/how-to/probe-edges-in-a-non-english-project.md new file mode 100644 index 000000000..e72164faf --- /dev/null +++ b/docs/how-to/probe-edges-in-a-non-english-project.md @@ -0,0 +1,60 @@ +# How to probe edges in a non-English project + +**Goal:** Get real edge-completeness coverage on a spec written in a language other than English, instead of every requirement landing in `unclassified` and the whole taxonomy quietly contributing nothing. + +**Prerequisites:** A project with `response_language` set (see [Configuration](../CONFIGURATION.md)), and a phase whose `/gsd-spec-phase` run has passed the ambiguity gate. The edge-completeness probe (Step 5.5) then runs automatically — you do not invoke it separately. + +For the taxonomy and the reasoning behind front-of-pipeline edge analysis, see [Spec-Phase Edge-Completeness Probe](../FEATURES.md#144-spec-phase-edge-completeness-probe). For how to act on findings once they appear, see [Resolve edge-coverage findings](resolve-edge-coverage-findings.md). This guide covers only what is different when your spec is not in English. + +--- + +## What happens, and why + +The probe classifies each requirement's data/behavior shape by matching **English** word-boundary cues against the requirement text. A requirement written in another language matches nothing, classifies to zero shapes, raises zero categories, and surfaces as a single `unclassified — review manually` row. + +That is not a rejection you can act on — it looks identical to a genuinely edge-free requirement. When it happens to *every* requirement, the probe has contributed nothing to the spec. + +So Step 5.5 sends the probe an English rendering of each requirement, while your spec stays in your language. Concretely, for the same requirement: + +| Requirement text handed to the probe | Shapes | Edges raised | +|---|---|---| +| `O sistema mescla intervalos sobrepostos em uma lista ordenada` | none | none — one `unclassified` row | +| `The system merges overlapping intervals in a sorted list` | `collection` | `adjacency`, `empty`, `ordering` | + +## What you do + +Nothing extra. The translation happens inside Step 5.5 as part of the run. + +What you should *see* is the split: your **spec stays in `response_language`** — its requirements, its acceptance criteria, its `## Edge Coverage` section — while the probe's findings are reasoned about from an English rendering of the requirement text. Requirement ids (`R1`, `R2`, …) are never translated or renumbered, so a finding always names the same requirement you wrote. + +If your spec comes back anglicized, that is a bug worth reporting — only the probe's transient input is translated, never the document. + +## Tell "no edges here" apart from "the probe could not read it" + +This is the distinction that matters, because both look like an `unclassified` row. + +| What you see | What it means | What to do | +|---|---|---| +| A few `unclassified` rows among normally-classified ones | Those requirements carry no shape cue **in any language**. This is the classifier's known recall gap, not a translation problem. | Resolve each like any other finding — or author an explicit `shapes` array on the requirement (below). | +| **Every** requirement `unclassified`, and a `WARNING: edge-probe proposed ZERO applicable edges` | The probe could not read your requirements at all. | Confirm the run really is translating the probe input. Do not accept an empty `## Edge Coverage` section. | +| Some requirements classified, the rest `unclassified`, no warning | **The silent case.** The zero-applicable warning fires only when *all* requirements are unclassified, so a partly-classified spec raises nothing. | Check the `unclassified` ones individually against the row above. | + +Translation makes the classifier *applicable*; it does not make it omniscient. A requirement carrying no shape cue in English either — for example "the command exits with code 1 on invalid input" — still classifies to zero. That is expected, and the fix is the same one an English-language project uses. + +## Force the shape when the prose carries no cue + +When a requirement is genuinely edge-relevant but no cue fires, do not fight the wording. Author the shape explicitly — an authored `shapes` array bypasses prose classification entirely, in any language: + +```json +{ "id": "R4", "text": "The command exits with code 1 on invalid input", "shapes": ["stateful"] } +``` + +`shapes` accepts any of `numeric-range`, `collection`, `text`, `stateful`, `io`. The example above raises `idempotency` and `concurrency`. + +An explicit empty array — `"shapes": []` — is the opposite signal: your deliberate "this requirement has no edge surface", which stays silent rather than surfacing an `unclassified` row. + +## Related + +- [Resolve edge-coverage findings](resolve-edge-coverage-findings.md) — what to do with each finding once it is raised +- [Spec-Phase Edge-Completeness Probe](../FEATURES.md#144-spec-phase-edge-completeness-probe) — the taxonomy and the rationale +- [Configuration](../CONFIGURATION.md) — the `response_language` setting diff --git a/gsd-core/references/edge-probe.md b/gsd-core/references/edge-probe.md index ae6171591..75acfbcdb 100644 --- a/gsd-core/references/edge-probe.md +++ b/gsd-core/references/edge-probe.md @@ -47,6 +47,14 @@ The five shapes are: `numeric-range`, `collection`, `text`, `stateful`, `io`. Wh `shapes` is absent, a heuristic classifier proposes them from the requirement prose (propose-then-confirm) — the author may correct the shape. +**`text` is English, whatever language the SPEC is in.** The heuristic cues are English +word-boundary patterns, so a project running with `response_language` set must pass `text` as a +faithful English translation of the requirement; the SPEC itself keeps the original language and +the `id` is never translated. Prose in another language matches no cue, classifies to zero +shapes, and surfaces as `unclassified` (#1110) — the probe contributes nothing. Where a +requirement carries no cue even in English, author `shapes` explicitly rather than leaning on +the classifier. + ## Taxonomy (8 categories) Closed and small by design: a fixed eight the author must explicitly clear beats thirty diff --git a/gsd-core/workflows/spec-phase.md b/gsd-core/workflows/spec-phase.md index fb95e6e69..821e35ca7 100644 --- a/gsd-core/workflows/spec-phase.md +++ b/gsd-core/workflows/spec-phase.md @@ -192,6 +192,24 @@ If gate passes (ambiguity ≤ 0.20 AND all minimums met): Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague ones). Reference: @~/.claude/gsd-core/references/edge-probe.md. +**Non-English projects — the probe input is translated, the SPEC is not.** The shape cues the +classifier matches are **English** word-boundary patterns, so requirement prose written in +another language matches nothing, classifies to zero shapes, and lands every row in +`unclassified` (#1110) — the taxonomy contributes nothing and `--auto` leaves it all +`unresolved`. The `$REQS_JSON` payload below is **engine input, never user-facing output**, so +the `response_language` rule at the top of this workflow does not govern it: write each +requirement's `text` as a faithful **English** translation of the SPEC requirement. +The SPEC keeps the original language — only the probe input is translated, and requirement +`id`s are never translated or renumbered (coverage rows join back on `id`, and any Acceptance +Criteria you write from the resolved edges go into the SPEC in `response_language`). +Translate **every** requirement, not only the ones that look edge-relevant: the +`$APPLICABLE = 0` warning below fires only when *all* requirements are unclassified, so a +partly-classified spec slips through with no signal at all. +If a requirement still classifies to zero shapes after translation, it carries no cue in any +language (the recorded recall gap — ADR-857 §98 / ADR-550 D7b, not a translation failure); +author an explicit `shapes` array on that requirement instead of relying on the prose +classifier. + **Runtime coverage compute — resolve and invoke edge-probe.cjs:** ```bash @@ -235,6 +253,9 @@ fi # canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object # per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes # the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04). +# When `response_language` is set, `text` carries a faithful ENGLISH translation (see above) — +# the shape cues are English-only, so original-language prose classifies to zero shapes and the +# probe becomes a silent no-op. Keep every `id` exactly as it appears in the SPEC. # BSD/macOS mktemp only randomizes XXXXXX when it is the final path component, so make a # suffixless temp then append the extension — portable across BSD + GNU (#1520). REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX") && mv "$REQS_JSON" "${REQS_JSON}.json" && REQS_JSON="${REQS_JSON}.json" || exit 1 @@ -247,6 +268,7 @@ JSON # `` placeholder (a forgotten substitution would otherwise yield a # meaningful-looking but bogus coverage report). Fail loud, not silent no-op. if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("/dev/null; then + rm -f "$REQS_JSON" echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2 exit 1 fi diff --git a/tests/edge-probe-spec-phase-contract.test.cjs b/tests/edge-probe-spec-phase-contract.test.cjs index 33d9eecc4..4f1956b29 100644 --- a/tests/edge-probe-spec-phase-contract.test.cjs +++ b/tests/edge-probe-spec-phase-contract.test.cjs @@ -14,8 +14,13 @@ const { test } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('node:fs'); const path = require('node:path'); +const fc = require('fast-check'); const SPEC_PHASE_PATH = path.join(__dirname, '..', 'gsd-core', 'workflows', 'spec-phase.md'); +const EDGE_PROBE_REF_PATH = path.join(__dirname, '..', 'gsd-core', 'references', 'edge-probe.md'); +const { classifyShape, applicableCategories } = require( + path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'edge-probe.cjs'), +); function readSpecPhase() { return fs.readFileSync(SPEC_PHASE_PATH, 'utf8'); @@ -299,3 +304,203 @@ test('#3132: specless-probe-fallback.md uses resolved+verification — not cover assert.match(content, /verification: explicit/, 'specless-probe-fallback.md must reference "verification: explicit"'); }); + +// ─── #2773: non-English requirements and the English-only shape cues ────────── +// +// `SHAPE_CUES` (src/edge-probe.cts) are English word-boundary regexes. A project that sets +// `response_language` writes its SPEC Requirements in that language, so transcribing them +// verbatim into the Step 5.5 `$REQS_JSON` heredoc classifies every requirement to zero shapes +// -> every row becomes the `unclassified` sentinel (#1110) and the 8-category taxonomy +// contributes nothing. Approved scope for #2773 is doc-only: Step 5.5 must instruct that the +// probe's `text` field carries a faithful English translation (engine input, never +// user-facing), while the SPEC itself stays in the original language. + +test('#2773: Step 5.5 documents the English-translation step for response_language projects', () => { + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + assert.match( + block, + /response_language/, + 'Step 5.5 must name `response_language` — it is the setting that makes the requirement prose non-English', + ); + assert.match( + block, + /translat/i, + 'Step 5.5 must instruct that the probe input carries a translation, not the original-language prose', + ); + assert.match( + block, + /English/, + 'Step 5.5 must say the translation target is English (the cue set the classifier actually speaks)', + ); + // The SPEC must NOT be anglicized — only the transient probe payload is translated. + assert.match( + block, + /SPEC[^\n]*(original language|stays in|keeps)|(original language)[^\n]*SPEC/i, + 'Step 5.5 must state the SPEC keeps the original language — only the probe input is translated', + ); + // Requirement ids are the join key for coverage rows; translating them breaks the mapping. + assert.match( + block, + /`?id`?s?[^\n]*(unchanged|not translated|never translated|stable)/i, + 'Step 5.5 must state requirement ids are NOT translated or renumbered (coverage rows join on id)', + ); +}); + +test('#2773: Step 5.5 names the authored shapes override as the zero-cue fallback', () => { + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + // Translation is necessary but NOT sufficient: prose carrying no cue in ANY language still + // classifies to []. The engine already accepts an authored `shapes` override for exactly + // that case, so the instruction must point at it rather than over-promise. + assert.match( + block, + /`shapes`/, + 'Step 5.5 must name the authored `shapes` override as the fallback for prose that still classifies to zero', + ); +}); + +test('#2773: the translation instruction precedes the $REQS_JSON write', () => { + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + const langIdx = block.search(/response_language/); + const writeIdx = block.search(/>\s*"\$REQS_JSON"/); + assert.ok(langIdx !== -1, 'Step 5.5 must mention response_language'); + assert.ok(writeIdx !== -1, 'Step 5.5 must write $REQS_JSON'); + assert.ok( + langIdx < writeIdx, + 'the translation instruction must come BEFORE the $REQS_JSON write — the downstream `$APPLICABLE = 0` warning only fires when EVERY requirement is unclassified, so a partially-classified non-English spec would slip through silently', + ); +}); + +test('#2773: translating a non-English requirement is what makes the cue matcher apply', () => { + // Same requirement, two languages. This is the premise the Step 5.5 instruction rests on; + // if a future SHAPE_CUES edit breaks it, the documented advice becomes false and this fails. + const pt = 'O sistema mescla intervalos sobrepostos em uma lista ordenada'; + const en = 'The system merges overlapping intervals in a sorted list'; + + assert.deepEqual(classifyShape(pt), [], 'non-English prose matches no English cue — the silent no-op #2773 reports'); + assert.deepEqual(applicableCategories(classifyShape(pt)), [], 'zero shapes raise zero categories'); + + const enShapes = classifyShape(en); + assert.ok(enShapes.includes('collection'), 'the English rendering must classify as a collection'); + const enCategories = applicableCategories(enShapes); + for (const expected of ['adjacency', 'empty', 'ordering']) { + assert.ok(enCategories.includes(expected), `translated requirement must raise \`${expected}\``); + } + + // A second shape, so the row is not a single-cue coincidence. + const ptText = 'O nome do usuario e truncado em 50 caracteres'; + const enText = 'The user name is truncated at 50 characters'; + assert.deepEqual(classifyShape(ptText), [], 'non-English text-shape prose also matches nothing'); + assert.ok(classifyShape(enText).includes('text'), 'the English rendering must classify as text'); + + // Multi-cue: a sentence carrying cues for two shapes yields the UNION, not one of them. + const multi = 'The API request uploads a sorted list of items'; + const multiShapes = classifyShape(multi); + assert.ok(multiShapes.includes('io'), 'multi-cue prose must include io'); + assert.ok(multiShapes.includes('collection'), 'multi-cue prose must include collection'); +}); + +test('#2773: translation alone does not rescue genuinely zero-cue prose', () => { + // The issue's own repro sentence classifies to [] in ENGLISH too — it carries no shape cue + // in any language. That is the recorded recall gap (ADR-857 §98 / ADR-550 D7b), not a + // language failure, which is why Step 5.5 must point at the `shapes` override rather than + // promise that translation restores classification. + const enZeroCue = 'The command exits with code 1 and prints to stderr on invalid input'; + assert.deepEqual( + classifyShape(enZeroCue), + [], + 'a genuinely zero-cue requirement stays unclassified even in English — the doc must not over-promise', + ); + + // The authored `shapes` override is the deterministic escape hatch for exactly this case. + assert.deepEqual( + applicableCategories(['stateful']), + ['idempotency', 'concurrency'], + 'an authored `shapes` override raises categories with no dependence on prose cues', + ); +}); + +test('#2773 property: a cue word classifies regardless of surrounding text', () => { + // A translated sentence carries its cue word amid arbitrary other words. The Step 5.5 + // advice is only sound if classification is robust to that surrounding context rather than + // anchored to a fixed sentence shape. + const cueForShape = { + 'numeric-range': 'threshold', + collection: 'items', + text: 'unicode', + stateful: 'persist', + io: 'endpoints', + }; + const filler = fc.string({ minLength: 0, maxLength: 24 }).filter((s) => !/[A-Za-z]/.test(s)); + + fc.assert( + fc.property( + fc.constantFrom(...Object.keys(cueForShape)), + filler, + filler, + (shape, before, after) => { + const sentence = `${before} ${cueForShape[shape]} ${after}`; + assert.ok( + classifyShape(sentence).includes(shape), + `cue "${cueForShape[shape]}" must classify as ${shape} inside ${JSON.stringify(sentence)}`, + ); + }, + ), + { numRuns: 100, seed: 2773 }, + ); +}); + +test('#2773: the edge-probe reference documents the English-cue assumption', () => { + const ref = fs.readFileSync(EDGE_PROBE_REF_PATH, 'utf8'); + assert.match( + ref, + /English/, + 'the edge-probe reference `## Inputs` contract must state that the heuristic classifier is English-cue based', + ); + assert.match( + ref, + /response_language/, + 'the edge-probe reference must point non-English projects at the translated-input requirement', + ); +}); + +test('#2773: no Step 5.5 exit path leaks the $REQS_JSON temp file', () => { + // The temp file holds the SPEC's requirement text. Every guard between its creation and + // the unconditional cleanup must `rm -f` it before `exit 1`, or a failed spec run strands + // requirement content in TMPDIR. The engine-failure guard always did; the empty/placeholder + // guard directly above it did not, so the two siblings disagreed about their own invariant. + const block = extractStep55Block(readSpecPhase()); + assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md'); + + const lines = block.split('\n'); + const createIdx = lines.findIndex((l) => /REQS_JSON=\$\(mktemp/.test(l)); + assert.ok(createIdx !== -1, 'Step 5.5 must create $REQS_JSON via mktemp'); + + // The region ends at the first UNCONDITIONAL cleanup (a bare `rm -f "$REQS_JSON"` at column + // zero); past that the file is already gone and later exits cannot leak it. + const afterCreate = lines.slice(createIdx + 1); + const endOffset = afterCreate.findIndex((l) => /^rm -f "\$REQS_JSON"/.test(l)); + assert.ok(endOffset !== -1, 'Step 5.5 must unconditionally rm -f "$REQS_JSON" after the engine run'); + const region = afterCreate.slice(0, endOffset); + + // Walk the region tracking whether the current guard branch has cleaned up. `then`/`else` + // opens a fresh branch; a cleanup inside it arms the branch; an `exit` must find it armed. + let cleanedInBranch = false; + const leaks = []; + for (const line of region) { + if (/\bthen\b|^\s*else\b|^\s*elif\b/.test(line)) cleanedInBranch = false; + if (/rm -f "\$REQS_JSON"/.test(line)) cleanedInBranch = true; + if (/^\s*exit\s+\d+/.test(line) && !cleanedInBranch) leaks.push(line.trim()); + } + + assert.deepEqual( + leaks, + [], + `every exit between the mktemp and the unconditional cleanup must rm -f "$REQS_JSON" first; leaking exits: ${JSON.stringify(leaks)}`, + ); +}); diff --git a/tests/emitted-drift-acks/0000-legacy-migration.json b/tests/emitted-drift-acks/0000-legacy-migration.json index 73b90aa66..041730733 100644 --- a/tests/emitted-drift-acks/0000-legacy-migration.json +++ b/tests/emitted-drift-acks/0000-legacy-migration.json @@ -36,7 +36,7 @@ "agents/gsd-verifier.toml": "#2834: Codex agent TOML now carries model-routing fields on first install.", "gsd-code-fixer.md": "#2647: the three worktree-path sites (setup_worktree bash, concrete-steps prose, critical_rules) replaced the hardcoded /tmp/sv- mktemp path with a repo-relative .claude/worktrees/rf--- path, and added a defense-in-depth padded_phase validation at the sink (the agent prompt is a literal bash contract any caller can spawn; the orchestrator validates upstream but the sink now self-defends against path-traversal/branch-name injection). On Windows/Git Bash the /tmp path landed outside the project tree (outside the session permission allowlist, prompting on every read) and mktemp's MAX_PATH substitute was un-removable; .claude/worktrees/ is the same dir the harness-managed executor worktrees use (gitignored via .claude/, inside the permission scope). Growth is the path-resolution bash (main_repo via `git worktree list --porcelain | awk`) + the $$-PID/epoch uniqueness replacing mktemp's XXXXXX + the padded_phase guard + the #2647 rationale comments at each site. Supersedes the prior #2825 attribution, whose gated-bash + guardrail growth is already in next.", "spec-phase.md": { - "reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification. #3102: Step 5.5 now RENDERS the edge-probe coverage report into the model's context (a raw printf of $COVERAGE after the well-formedness guard) and binds those rows in the resolution loop and --auto as a deterministic FLOOR, plus the corrected block comment; load-bearing workflow instruction that makes ADR-550 D7b (deterministic propose + LLM resolve) real at runtime and honors ADR-857 sec98's recall gap (floor, not ceiling). 32238 -> 34056 bytes (+1818), DEFAULT tier, cap 40960." + "reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification. #3102: Step 5.5 now RENDERS the edge-probe coverage report into the model's context (a raw printf of $COVERAGE after the well-formedness guard) and binds those rows in the resolution loop and --auto as a deterministic FLOOR, plus the corrected block comment; load-bearing workflow instruction that makes ADR-550 D7b (deterministic propose + LLM resolve) real at runtime and honors ADR-857 sec98's recall gap (floor, not ceiling). 32238 -> 34056 bytes (+1818), DEFAULT tier, cap 40960. #2773: Step 5.5 now tells a `response_language` project that the edge-probe `$REQS_JSON` payload is engine input rather than user-facing output, so each requirement's `text` carries a faithful English translation while the SPEC keeps its original language and requirement ids stay unchanged. The shape cues in src/edge-probe.cts are English word-boundary regexes, so prose in another language matched nothing, classified to zero shapes, and put every row in the `unclassified` sentinel (#1110) — the taxonomy contributed nothing. Growth is that instruction plus the pointer to the authored `shapes` override for prose that classifies to zero even in English (measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient), and a one-line fix an isolated security review surfaced in the same block — the empty/placeholder guard exited without `rm -f \"$REQS_JSON\"` while its engine-failure sibling below it did, stranding the SPEC requirement text in TMPDIR on every failed run. The instruction sits BEFORE the heredoc deliberately: the `$APPLICABLE = 0` warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal. Load-bearing runtime instruction; the compiled engine is deliberately untouched (the `lang`-hint / per-language cue-set fix is out of scope per the #2773 triage). 34020 -> 35730 bytes (+1710), DEFAULT tier, cap 40960." } } }