fix(#3102): render edge-probe coverage report so the resolution loop consumes it (#3391)

* fix(#3102): render edge-probe coverage report so the resolution loop consumes it

Step 5.5 captured the edge-probe report into $COVERAGE, shape-checked it, and
reduced it to coverage.applicable — the engine's per-requirement items[] never
reached the model, so the resolution loop re-derived edge categories from prose
(the data-flow twin of #2733's control-flow discard). The block's own comment
claimed the opposite.

Render $COVERAGE RAW into context after the well-formedness guard (schema-agnostic
so an ADR-550 D7a-style re-cut cannot desync a bespoke renderer), and bind the rows
in the resolution loop as a deterministic FLOOR the model unions with its own
classification — floor, never ceiling, since the classifier has a measured recall
gap (ADR-857 §98 / ADR-550 D7b). --auto consumes the same floor. Comment corrected
to match. Regression test asserts a bare render of $COVERAGE, not a count-only cross.

* chore(#3102): add changeset for the Step 5.5 edge-coverage render fix

* chore(#3102): re-arm spec-phase.md emitted-drift ack for the Step 5.5 render growth

The render + floor-binding prose grows spec-phase.md ~1818 bytes (32238 -> 34056,
under the 40960 cap). Re-arms the existing spent spec-phase.md ack rather than adding
a new fragment (a second key would collide with the base-relative duplicate check).
This commit is contained in:
Rezolv
2026-08-12 20:44:57 -04:00
committed by GitHub
parent bd31cf2bba
commit 3dff700aaa
4 changed files with 84 additions and 10 deletions

View File

@@ -0,0 +1,5 @@
---
type: Fixed
pr: 3391
---
**`spec-phase` Step 5.5 now surfaces the edge-probe's proposed edges to the resolution loop instead of discarding them** — the deterministic coverage report was computed, validated, then reduced to a single applicable-count, so the resolution loop re-derived edge categories from requirement prose and the engine's proposals never reached it. The report is now rendered into context and its rows are consumed as a *floor* the model unions with its own classification (still adding any category the classifier missed), so the written `## Edge Coverage` reflects the engine's deterministic taxonomy rather than model-invented categories; `--auto` gets the same floor. (#3102)

View File

@@ -251,9 +251,11 @@ if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===
exit 1
fi
# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per
# requirement. The resolved/dismissed/unresolved rows in $COVERAGE (resolved items carry
# verification: explicit|backstop) drive the
# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose).
# requirement. The report is RENDERED into context below (#3102); its resolved/dismissed/
# unresolved rows (resolved items carry verification: explicit|backstop) are the deterministic
# FLOOR the resolution loop consumes and unions with its own classification — the loop no longer
# re-derives the taxonomy from prose unaided. Floor, never ceiling: the classifier has a measured
# recall gap (ADR-857 §98 / ADR-550 D7b), so the model still ADDS any category the engine missed.
# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture
# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE
# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary
@@ -270,6 +272,16 @@ if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d
echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2
exit 1
fi
# Render the validated report into the model's visible context (#3102). Until here $COVERAGE was
# captured, shape-checked, and reduced to coverage.applicable — the engine's per-requirement
# items[] never reached the model, so the resolution loop below re-derived edge categories from
# requirement PROSE (the data-flow twin of #2733's control-flow discard). These rows are the
# deterministic FLOOR the resolution loop consumes. Printed RAW (not a bespoke table) so this
# step holds NO knowledge of the item schema: an ADR-550 D7a-style re-cut of the item/coverage
# shape cannot silently desync a hand-rolled renderer here — the engine stays the single source.
echo "### Edge-probe coverage report (deterministic proposals — the FLOOR for the resolution loop below):"
printf '%s\n' "$COVERAGE"
echo "### (end edge-probe coverage report)"
# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY
# requirement is far more likely a shape-classification miss (or malformed requirements) than
# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding
@@ -287,9 +299,14 @@ edge-free spec, or should we revisit the requirement wording / authored shapes?"
an empty `## Edge Coverage` section after explicit confirmation.
For each Requirement gathered so far:
1. Classify its shape and raise only applicable edge categories (relevance filter — see
the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already
surfaced as pre-resolved.
1. Start from the edge-probe rows RENDERED above — the deterministic `items[]` are the FLOOR:
every proposed `(requirement_id, category)` MUST be resolved below (Specify / Dismiss-with-
reason / Backstop / Defer), none silently dropped. Then raise any applicable category the
engine MISSED — the rows are a floor, never a ceiling: the classifier has a measured recall
gap on terse prose (ADR-857 §98 / ADR-550 D7b), e.g. a CSV-export requirement whose
`encoding` edge the shape cue under-fires. Union the engine's rows with your own
classification (relevance filter — see the taxonomy in the reference); do not narrow to them.
Reuse any edges the Round-4 Failure Analyst already surfaced as pre-resolved.
2. For each raised category, propose a CONCRETE candidate edge (not "consider
boundaries" — e.g. "R2 merges intervals; what about `[[1,2],[2,3]]` that only touch?").
3. Resolve each with the user (AskUserQuestion; text mode → numbered list):
@@ -313,9 +330,11 @@ For each Requirement gathered so far:
- On "anyway": write SPEC.md with those rows marked `⚠ Edge unresolved — planner must
treat as assumption`.
**`--auto` mode:** auto-`resolved` (verification: explicit) where a defensible acceptance
criterion can be written; otherwise auto-`resolved` (verification: backstop) (never
auto-dismiss — a wrong dismissal is the exact silent failure being eliminated). Log:
**`--auto` mode:** resolve over the **same rendered floor** (#3102) — every engine-proposed row
from Step 5.5's report (step 1) plus any category the classifier missed, never a narrower set.
For each: auto-`resolved` (verification: explicit) where a defensible acceptance criterion can be
written; otherwise auto-`resolved` (verification: backstop) (never auto-dismiss — a wrong
dismissal is the exact silent failure being eliminated). Log:
`[auto] edge coverage: E explicit, B backstop, U unresolved`.
**`unclassified` exception (#1110):** `--auto` leaves an `unclassified` candidate

View File

@@ -193,6 +193,56 @@ test('adversarial review: Step 5.5 guards a zero-applicable coverage report', ()
);
});
// #3102 (data-flow reachability): the validated $COVERAGE report must be RENDERED into the
// model's visible context, not merely captured and reduced to `coverage.applicable`. Before
// #3102, every emission of $COVERAGE on the success path piped it into a `node -e` consumer
// (the shape guard, the count extract) whose output the model never sees — so the engine's
// items[] were computed, validated, then discarded, and the resolution loop re-derived edge
// categories from requirement prose (the sibling of #2733's control-flow discard, one layer
// down). This asserts a BARE render: $COVERAGE emitted to stdout as the leading command, not
// captured into a variable (`=$(`) and not piped into a consumer (`| node`).
test('#3102: Step 5.5 renders the $COVERAGE report to stdout (not only the applicable count)', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
const rendersReport = block.split('\n').some((raw) => {
const line = raw.trim();
if (!/\$COVERAGE\b/.test(line)) return false; // must reference the captured report
if (line.startsWith('#')) return false; // not a comment mention
if (!/^(printf|echo|cat)\b/.test(line)) return false; // emitted as the leading command
if (/=\s*\$\(/.test(line)) return false; // a capture reaches a variable, not the model
if (/\|\s*node\b/.test(line)) return false; // piped into a consumer (shape guard / count extract)
return true;
});
assert.ok(
rendersReport,
'Step 5.5 must RENDER $COVERAGE to the model context — a bare `printf`/`echo`/`cat` of $COVERAGE that is not captured (`=$(`) and not piped into `node` — so the engine items[] reach the resolution loop (#3102 data-flow discard)'
);
});
// #3102: rendering without binding leaves the rows decorative. The resolution loop must
// consume the rendered engine rows as a deterministic FLOOR — every proposed (requirement_id,
// category) is resolved, and the model ADDS any category the classifier missed (floor, never
// ceiling — the classifier has a measured recall gap: ADR-857 §98 / ADR-550 D7b). This guards
// a future edit that renders the report but leaves the loop re-deriving categories from prose.
test('#3102: Step 5.5 resolution loop binds the engine rows as a floor, not a ceiling', () => {
const content = readSpecPhase();
const block = extractStep55Block(content);
assert.ok(block.length > 0, 'Step 5.5 block must be extractable from spec-phase.md');
assert.match(
block,
/floor/i,
'Step 5.5 resolution loop must describe the rendered engine rows as a FLOOR the model unions with its own classification (ADR-550 D7b) — surfacing the report without binding it leaves it decorative'
);
assert.match(
block,
/recall gap|never a ceiling|not a ceiling|add(?:ing)? (?:any|the missed|categor)/i,
'the floor must NOT be a ceiling — the loop must instruct the model to add categories the classifier missed (ADR-857 §98 recall gap), not narrow to the engine rows'
);
});
// #3132: the retired covered/backstop-as-status vocabulary must not appear in
// the workflow prose. probe-core.cts locks Status to resolved|dismissed|unresolved;
// backstop survives only as a verification tier on a resolved item.

View File

@@ -36,7 +36,7 @@
"agents/gsd-verifier.toml": "#2834: Codex agent TOML now carries model-routing fields on first install.",
"gsd-code-fixer.md": "#2647: the three worktree-path sites (setup_worktree bash, concrete-steps prose, critical_rules) replaced the hardcoded /tmp/sv- mktemp path with a repo-relative .claude/worktrees/rf-<phase>-<pid>-<epoch> path, and added a defense-in-depth padded_phase validation at the sink (the agent prompt is a literal bash contract any caller can spawn; the orchestrator validates upstream but the sink now self-defends against path-traversal/branch-name injection). On Windows/Git Bash the /tmp path landed outside the project tree (outside the session permission allowlist, prompting on every read) and mktemp's MAX_PATH substitute was un-removable; .claude/worktrees/ is the same dir the harness-managed executor worktrees use (gitignored via .claude/, inside the permission scope). Growth is the path-resolution bash (main_repo via `git worktree list --porcelain | awk`) + the $$-PID/epoch uniqueness replacing mktemp's XXXXXX + the padded_phase guard + the #2647 rationale comments at each site. Supersedes the prior #2825 attribution, whose gated-bash + guardrail growth is already in next.",
"spec-phase.md": {
"reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification."
"reason": "#2733: five transitions in gsd-core/workflows/spec-phase.md were re-pointed so control reaches the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes, which no path could reach before. Four upstream gate-passed jumps went from 'Jump to Step 6' to 'Jump to Step 5.5', and Step 5.5's own terminal soft gate at :305 went from 'proceed to Step 6' to 'proceed to Step 5.6' so the common all-edges-resolved path stops skipping the prohibition probe. The +10 bytes is exactly those five targets growing by 2 bytes each ('Step 6' -> 'Step 5.5' / 'Step 5.6'); it is the literal fix, not incidental prose growth, and cannot be avoided without leaving a probe unreachable. Verified: 31987 -> 31997 bytes, DEFAULT tier, cap 40960. #3132: realigned retired covered/backstop-as-status vocab to resolved+verification. #3102: Step 5.5 now RENDERS the edge-probe coverage report into the model's context (a raw printf of $COVERAGE after the well-formedness guard) and binds those rows in the resolution loop and --auto as a deterministic FLOOR, plus the corrected block comment; load-bearing workflow instruction that makes ADR-550 D7b (deterministic propose + LLM resolve) real at runtime and honors ADR-857 sec98's recall gap (floor, not ceiling). 32238 -> 34056 bytes (+1818), DEFAULT tier, cap 40960."
}
}
}