* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
7.2 KiB
Honest Verifier — Abstention on Non-Inferable Checks
Shared reference for the verify phase. The verify-time companion to the spec-time
@~/.claude/gsd-core/references/edge-probe.md (which classifies non-inferable checks) and
@~/.claude/gsd-core/references/prohibition-probe.md (whose judgment-tier disposition this mirrors).
This doc is written in generic spec → predicate → verifier terms with no tool-specific vocabulary,
so it is portable: copy it into any verification process.
The problem it solves
A verifier is trustworthy on inferable checks — defects determined by the stated spec. On a
non-inferable check the correct answer is not derivable from the spec alone (e.g. "does [1,2]
touching [2,3] merge?", "is a 'character' a grapheme or a code unit?"). On these the verifier does
not know that it does not know: measured behavior is a confident PASS on the blind-spot check ~100%
of the time (mean confidence ~0.93), because a model cannot self-detect a gap it does not perceive.
The edge-probe already detects these at spec time and tags them verification: backstop (ADR-550
D7a). The honest verifier consumes that tag so the verifier abstains instead of confidently
false-passing — converting a silent false-pass (the worst failure: you don't know to look) into an
explicit, actionable "write a held-out test." Measured: the confident-false-pass rate on the blind
spot drops 100% → 17% (N17).
The two properties that define the design
- Exogenous, not endogenous. The trigger is the external tag (
backstop), never the verifier's self-judgment. Asking the verifier to "abstain if unsure" barely moves the number (100% → 67%) and only on ambiguity it already notices; on a true blind spot it stays confidently wrong. A confidence gate cannot reach a blind spot the model does not feel — so there is no "are you sure?" prompt; routing is on the pre-existing tag only. - Routing, not diagnosis. The verifier need not name the omitted rule (if it could, it wouldn't be a blind spot). In testing, verifiers abstained correctly while citing the wrong edge. The honest verdict requires only "I was told this is under-specified and I cannot rule it out." The omitted rule is carried by a human-authored held-out test, not by the verifier.
The disposition (the protocol)
For each must_haves.truths item:
| Item | Confirmable with explicit evidence? | Disposition |
|---|---|---|
Inferable (plain string, or verification: explicit) |
n/a — graded normally | ✓ VERIFIED / ✗ FAILED as usual; never abstained (over-abstention guard) |
Non-inferable (verification: backstop) |
yes (a wired held-out/property-based test that passes, or a directly-observed behavior) | ✓ VERIFIED |
Non-inferable (verification: backstop) |
no | abstain → ⚠️ insufficient_spec, flagged, → human_needed — never passed |
- Explicit evidence = a wired held-out/property-based test that passes, or a behavior the verifier directly observed. Symbol presence + wiring is not explicit evidence for a non-inferable truth.
- Never silent, never a hard halt. Interactive: the abstained item routes to the end-of-phase
human checkpoint. Autonomous (AFK): it produces a prominent
unverified — held-out test recommendedflag and the completion line reads "complete with N unverified non-inferable checks"; the run neither silently passes the blind spot nor hard-halts. - Distinguishable reason. The abstain disposition carries
reason: insufficient_specso thehuman_neededoutcome is never conflated with an ordinary manual-UAThuman_needed.
This is the verify-time half of ADR-550 Decision 4 (the never-silent-pass disposition), applied to the
edge backstop truth tier instead of the prohibition judgment tier — the same machinery, opposite
polarity (must-HAVE under-specified vs must-NOT irreducible).
Deterministic engine surface
The CI-testable surface is the deterministic disposition + projection, never the LLM's judgment
(ADR-550 D5 — a test asserting the model's verdict is vacuous and rejected). In probe-core:
truthStatement(t)/truthVerification(t)— normalizers; read a truth's statement and tier from either the plain-string or object form (a truth-reader MUST normalize, never assume a string).projectTruths(items)— conservative serializer: abackstoptruth → flat-scalar object{ statement, verification: backstop }; every inferable truth → a bare string.dispositionForUnverifiableTruth(truth, { evidence })→{ status, flagged, tier, reason }:backstop+ no evidence →unverified/flagged/insufficient_spec;backstop+ evidence →green; non-backstop→green(over-abstention guard).
Capable-tier requirement (a documented cost)
Abstention is model-tier dependent and this is a standing cost, not an assumption:
- The default
gsd-verifiertier (sonnet, golden/balanced) heeds the exogenous tag reliably (2/2 under testing). - The budget tier (
haiku) is the least flag-responsive (1/2, inconsistent) and degrades toward current behavior (confident false-pass). Run honest-verifier on a capable tier; treat the budget tier as best-effort. Re-validate when thegsd-verifiermodel tier changes or a new budget model is adopted (captured as a test so a tier regression is caught, not discovered in production).
Evidence and scope (stated honestly)
- Evidence strength. N17 is n=27 verdicts (3 models × 3 conditions × 3 tasks), 1 rep — direction-finding, not powered. The blind-spot effect is large and monotone (100% → 67% → 17%); the two costs are clean single events (a false tag made the strongest model over-abstain on a real spec-determined bug; the weakest tier was flag-deaf) and they name exactly the failure modes the over-abstention guard and the capable-tier requirement defend against.
- Tag-precision coupling. Quality is bounded by the edge-probe's
backstoprecall/precision — a false non-inferable flag causes over-abstention. Positive coupling: improving the probe (#1110) improves this for free. It adds no independent burden. - Explicit non-goals. Does NOT identify the omitted rule; does NOT recalibrate decisive verdicts; does NOT defend against malicious compliance (a self-graded review rationalizing away its own findings). It raises the floor on honest uncertainty about non-inferable checks — that is the whole claim.
Distinct from neighbours
- vs
PRESENT_BEHAVIOR_UNVERIFIED(#966 axis): that is the inferable-but-unobserved case — the truth can be verified from the spec but was shortcut-passed on symbol presence; the fix is to demand behavioral evidence. Honest-verifier is the non-inferable case — the truth cannot be verified from the spec at all; the fix is to abstain and route to a held-out test. Orthogonal axes (insufficient evidence vs insufficient spec); both feed the samehuman_neededsink. - vs prohibition judgment-tier (#644): that disposes must-NOT constraints; honest-verifier disposes non-inferable positive truths. Opposite polarity, same never-silent disposition.