diff --git a/.changeset/merry-mice-travel.md b/.changeset/merry-mice-travel.md new file mode 100644 index 000000000..544e94c8e --- /dev/null +++ b/.changeset/merry-mice-travel.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1738 +--- +**Honest verifier — verify-phase now abstains on non-inferable `backstop` truths instead of confidently false-passing them (#1154).** When the spec's edge-probe marks a truth non-inferable (`verification: backstop`) and the verifier cannot confirm it with explicit evidence (a passing wired held-out/property test, or a directly-observed behavior), it now reports `human_needed` with reason `insufficient_spec` ("unverified — held-out test recommended") rather than a silent `passed`. Autonomous runs complete with "N unverified non-inferable checks"; interactive runs route to the end-of-phase human checkpoint. Inferable truths are never abstained (over-abstention guard); abstention is exogenous (driven by the tag, not self-judgment). Truth-axis mirror of the prohibition judgment-tier (ADR-550 D4). diff --git a/agents/gsd-verifier.md b/agents/gsd-verifier.md index 55a4c427c..b30969917 100644 --- a/agents/gsd-verifier.md +++ b/agents/gsd-verifier.md @@ -197,6 +197,7 @@ For each truth: - A pre-existing test exercises the transition/invariant and passes (confirm via Step 7b's single-named-test path) → ✓ VERIFIED. - No such test exists, or it can't run without a server/state mutation → ⚠️ PRESENT_BEHAVIOR_UNVERIFIED. Emit a human-verification item (Step 8) and do not count it toward the verified score (Step 9). - An accepted override (Step 3b) carries the truth as PASSED (override), exactly as it does for a FAILED truth. +5b. **Non-inferable (`backstop`) truths:** a `verification: backstop` truth (via `truthVerification()`) abstains unless confirmed by explicit evidence — mark `insufficient_spec` -> a human-verification item -> `human_needed`. See `references/honest-verifier.md`. 6. Determine truth status ## Step 3b: Check Verification Overrides diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index 8bfd35fae..caa4eeb58 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -315,6 +315,8 @@ For browser-backed UAT, use a configured browser MCP server. The current Open GS **Coverage-aware UAT routing (#1602).** When a SUMMARY.md carries a `coverage:` frontmatter block, `verify-work` classifies each deliverable deterministically instead of prompting for every prose bullet: deliverables proven by passing tests are auto-passed (recorded with `source: automated`, no prompt) and only judgment-dependent deliverables are presented for human sign-off. SUMMARYs without a `coverage:` block fall back to the previous prose-based extraction unchanged. See the [`coverage:` block reference](#summary-coverage-block) below. +**Honest verifier — `insufficient_spec` abstention (#1154).** A `must_haves.truths` item carrying the `verification: backstop` marker (a *non-inferable* check the edge-probe surfaced at spec time) is graded specially: if the verifier cannot confirm it with **explicit evidence** (a passing wired held-out/property-based test, or a directly-observed behavior), it **abstains** — the item is reported `unverified — held-out test recommended` and the phase verdict becomes `human_needed` (with reason `insufficient_spec`, distinct from ordinary manual-UAT `human_needed`), **never a silent `passed`**. Autonomous runs complete with "N unverified non-inferable checks" rather than hard-halting; interactive runs route the item to the end-of-phase human checkpoint. Abstention is exogenous (driven by the `backstop` tag, never a self-judged "abstain if unsure") and an inferable truth is never abstained. Reliable on capable verifier tiers (`sonnet`+); the budget `haiku` tier degrades toward current behavior. See [Honest Verifier](../gsd-core/references/honest-verifier.md). + #### SUMMARY `coverage:` block A SUMMARY.md may carry an optional `coverage:` frontmatter block — a list of per-deliverable entries that joins requirements → tests → verification status: diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 1409b652c..9ed6de2ed 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3115,7 +3115,9 @@ When a requirement's prose matches **no** shape cue, the probe does not silently The resolved edges populate a `## Edge Coverage` section in `SPEC.md`. Unresolved *applicable* edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under `--auto`, the probe **never auto-dismisses** — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs `[auto] edge coverage: C covered, B backstop, U unresolved`. The one exception is an `unclassified` candidate: `--auto` leaves it **`unresolved`** (surfaced as a flagged assumption), never auto-`backstop` — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim. -The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. +The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges become `must_haves.truths` the verifier can check, so the section is not merely documentation. A `backstop` edge is lifted as a **structured non-inferable marker** (`{ statement, verification: backstop }`, a flat scalar — not a prose note), which the **honest verifier** then consumes (see below) — closing the loop the edge-probe opened. + +**Honest verifier — abstention on non-inferable checks (#1154).** A non-inferable (`backstop`) truth is one whose correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would confidently false-pass it (~100% of the time). At verify time, a `backstop` truth the verifier cannot confirm with **explicit evidence** (a passing wired held-out/property-based test, or a directly-observed behavior) **abstains** → `human_needed` with reason `insufficient_spec` (reported as `unverified — held-out test recommended`), **never a silent `passed`**. This is the verify-time, truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): exogenous (driven by the `backstop` tag, never a self-judged "abstain if unsure"), routing-not-diagnosis (the held-out test carries the omitted rule), and capable-tier dependent (reliable on `sonnet`+; the budget `haiku` tier degrades toward current behavior). An inferable truth is never abstained (the over-abstention guard). Reference: [Honest Verifier](../gsd-core/references/honest-verifier.md). **Requirements:** - REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a `## Edge Coverage` SPEC section. @@ -3125,6 +3127,8 @@ The load-bearing wire is the `plan-phase` lift: `covered` and `backstop` edges b - REQ-EDGE-05: `--auto` MUST never auto-dismiss — auto-cover or auto-backstop only. - REQ-EDGE-06: `plan-phase` MUST lift `covered` criteria and `backstop` notes into `must_haves.truths`. - REQ-EDGE-07: A requirement whose prose matches no shape cue MUST surface an `unclassified — review manually` candidate (never silently dropped); `--auto` MUST leave it `unresolved`, never auto-`backstop`. +- REQ-EDGE-08: `plan-phase` MUST lift a `backstop` edge into `must_haves.truths` as a structured flat-scalar marker (`{ statement, verification: backstop }`), never a prose parenthetical. +- REQ-HONEST-01: At verify time a `backstop` truth that cannot be confirmed with explicit evidence MUST abstain → `human_needed` (reason `insufficient_spec`), never `passed`; an inferable truth MUST never be abstained (over-abstention guard); abstention MUST be exogenous (driven by the `backstop` tag, not self-judgment). **Reference:** [Edge Probe](../gsd-core/references/edge-probe.md) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index b835db63b..0395aa953 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -222,6 +222,7 @@ "gates.md", "git-integration.md", "git-planning-commit.md", + "honest-verifier.md", "ios-scaffold.md", "loop-hook-dispatch.md", "mandatory-initial-read.md", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index 71cdcd970..8db0f0946 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -308,6 +308,7 @@ Full roster at `gsd-core/references/*.md`. References are shared knowledge docum | `domain-probes.md` | Domain-specific probing questions for discuss-phase. | | `edge-probe.md` | Spec-phase edge-completeness probe — 8-category edge taxonomy, shape classification, and the `requirements → checks → verifier` resolution model (Step 5.5). | | `prohibition-probe.md` | Spec-phase prohibition-completeness probe — the two-stage adversarial-recall → precision protocol that surfaces the unwritten *must-NOT* constraints (values/safety/ethics), with status×verification (`test`/`judgment`) tiering and canon-referral breadcrumbs (Step 5.6); second adapter of the `probe-core` resolution model. | +| `honest-verifier.md` | Verify-time abstention on non-inferable (`backstop`) truths — the truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): a `backstop` truth the verifier can't confirm with explicit evidence abstains → `human_needed` (reason `insufficient_spec`), never a silent pass (#1154). | | `gate-prompts.md` | Gate/checkpoint prompt templates. | | `loop-hook-dispatch.md` | Generic dispatch contract for consuming `gsd_run loop render-hooks --raw` output in any host-loop workflow — envelope shape, per-kind dispatch rules (contribution/step/gate), and liveness banner. | | `scout-codebase.md` | Phase-type→codebase-map selection table for discuss-phase scout step (extracted via the discuss-phase/modes progressive-disclosure split, #717). | diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md index 8389b1245..02e49143d 100644 --- a/docs/adr/550-spec-phase-probe-contract.md +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -143,6 +143,22 @@ This ratifies the **deterministic SOURCE** for the test-tier `CheckDescriptor` t Net effect on D3: the prohibition-item shape is extended with three optional, backward-compatible flat-scalar keys that give the test-tier locate a deterministic spec-phase source; the contract's CI-testable surface (D5) gains the projection round-trip parity (CHK-03), the fail-closed guard (CHK-06), and the byte-stable backward-compat fixture (CHK-07). The decision also lives in `src/probe-core.cts` / `src/prohibition-enforcement.cts` comments, the `verify-phase.md` / `spec-phase.md` prose, and the #1278 changeset. +## Addendum (2026-06-25, #1154) — honest verifier: the truth-axis disposition mirror of D4 + +This records the **truth-axis half of Decision 4** that the original ADR deliberately scoped out. D3 left `truths` untouched (no `polarity` field — that is a prohibition concern), and the Lineage note parked the N17 abstention experiment as the verify-time half of the **prohibition** judgment-tier only. D7a, however, already gives the **edge** axis an orthogonal `verification` tier (`explicit | backstop`), and `plan-phase` already lifts a `backstop` edge into `must_haves.truths`. Until now that tier was flattened to a prose parenthetical at the projection, so the verifier had nothing structured to branch on and graded a `backstop` truth `passed` like any inferable one — confidently false-passing a non-inferable check ~100% of the time (the exact "verifier reach = spec reach" failure ADR-857 names). This addendum closes that gap by giving the `backstop` **truth** tier the same abstain-and-flag disposition D4 gave the prohibition `judgment` tier — **opposite polarity (must-HAVE under-specified vs must-NOT irreducible), same never-silent-pass machinery.** It cross-references **ADR-857** (the verifier↔predicate contract this rides is core, non-toggleable substrate, graded exogenously — `:65`); this completes the already-endorsed edge branch of that rail and adds no parallel mechanism. + +1. **The D3 truth-item shape gains an OPTIONAL flat-scalar `verification` marker.** A `must_haves.truths` item is normally a plain string (an inferable truth — today's shape, unchanged). A **non-inferable** truth MAY instead be an object item carrying `statement` + a flat scalar `verification: backstop`. The marker is additive and default-absent: a string truth, or an object with no marker, behaves byte-identically to today (Hyrum's Law backward-compat). This extends the **Decision 3 truth shape** the same way #1278 extended the prohibition shape — it does **not** add a `polarity` field (D3's "truths untouched" holds); `verification` is D7a's pre-existing orthogonal axis, now carried through to the truth projection. + +2. **Flat scalars — NOT a nested object (load-bearing, #1278 precedent).** The marker is a flat `verification:` continuation key on the truth item, never a nested object — the shared flat `parseMustHavesBlock` round-trips it with **no parser change** (the round-trip parity test is the proof; the untouched frontmatter suite confirms `truths`/`artifacts`/`key_links`/`prohibitions` readers stay regression-free). + +3. **Deterministic disposition + projection.** `projectTruths` (`src/probe-core.cts`) emits the flat-scalar marker ONLY for a `backstop` truth and collapses every inferable truth to a bare string (conservative serializer). `dispositionForUnverifiableTruth(truth, { evidence })` is the pure, fail-closed verdict: a `backstop` truth with no **explicit evidence** (a passing wired held-out/property-based test, or a directly-observed behavior) → `{ status: 'unverified', flagged: true, reason: 'insufficient_spec' }`, **never green**; with evidence → green; any non-`backstop` truth → green (the over-abstention guard). No LLM judgment is tested (D5) — the helper owns routing once evidence-existence is known; the LLM verifier's only job is to decide whether explicit evidence exists. + +4. **`insufficient_spec` feeds the EXISTING `human_needed` outcome — no new verifier status (maintainer Decision 1).** Abstention reuses the locked 3-value `VERIFIER_STATUSES` (`['passed','gaps_found','human_needed']`, `src/verification.cts`) with **zero change** and no new downstream routing in `ship`/`execute-phase`. The abstain cause rides as a **distinguishable report reason** (`human_needed` + `reason: insufficient_spec`) so it is never conflated with an ordinary manual-UAT `human_needed` — asserted in a test (review condition-1 caveat). *Interactive:* the item routes to the end-of-phase human checkpoint. *Autonomous (AFK):* a prominent `unverified — held-out test recommended` flag; completion reads "complete with N unverified non-inferable checks" — never a silent pass, never a hard halt (the D4 guarantee, now on the truth axis). + +5. **Two measured properties define the design (maintainer Decision 2; caveats to record).** *Exogenous, not endogenous:* abstention is triggered by the external `backstop` tag, never a self-judged "abstain if unsure" — endogenous abstention was measured near-useless on true blind spots (100% → 67% vs exogenous 100% → 17%; N17). *Routing, not diagnosis:* the verdict does not name the omitted rule (the held-out test carries it). **Evidence honesty:** N17 is n=27, 1 rep — **direction-finding, not powered**; the effect is large and monotone but real-world precision depends on the edge-probe's *true* `backstop` recall/precision (the experiment modeled a perfect tagger), which is why the over-abstention guard and the capable-tier requirement are load-bearing acceptance criteria. **Model-tier coupling:** abstention is reliable on the default `gsd-verifier` tier (`sonnet`+); the budget tier (`haiku`) heeds the tag only inconsistently and degrades toward current behavior — captured as a documented cost (and a test) so a tier regression is caught, not discovered in production. + +Net effect: the truth-axis `backstop` tier gains the verify-time disposition D4 gave the prohibition judgment tier; the contract's CI-testable surface (D5) gains the truth-axis projection round-trip parity and the abstain-on-unconfirmed-backstop regression. The decision also lives in `src/probe-core.cts` comments, `gsd-core/references/honest-verifier.md`, the `plan-phase.md` / `verify-phase.md` / `agents/gsd-verifier.md` prose, and the #1154 changeset. + ## Addendum (2026-06-22) — Alternatives considered (recall / representation / packaging side) This consolidates the spec-phase-side rejected and deferred alternatives for the probe family, diff --git a/gsd-core/references/honest-verifier.md b/gsd-core/references/honest-verifier.md new file mode 100644 index 000000000..85c932f56 --- /dev/null +++ b/gsd-core/references/honest-verifier.md @@ -0,0 +1,105 @@ +# Honest Verifier — Abstention on Non-Inferable Checks + +Shared reference for the **verify** phase. The verify-time companion to the spec-time +`@~/.claude/gsd-core/references/edge-probe.md` (which *classifies* non-inferable checks) and +`@~/.claude/gsd-core/references/prohibition-probe.md` (whose judgment-tier disposition this mirrors). +This doc is written in generic `spec → predicate → verifier` terms with no tool-specific vocabulary, +so it is portable: copy it into any verification process. + +## The problem it solves + +A verifier is trustworthy on **inferable** checks — defects determined by the stated spec. On a +**non-inferable** check the correct answer is *not derivable from the spec alone* (e.g. "does `[1,2]` +touching `[2,3]` merge?", "is a 'character' a grapheme or a code unit?"). On these the verifier *does +not know that it does not know*: measured behavior is a **confident PASS on the blind-spot check ~100% +of the time** (mean confidence ~0.93), because a model cannot self-detect a gap it does not perceive. + +The edge-probe already detects these at spec time and tags them `verification: backstop` (ADR-550 +D7a). The honest verifier consumes that tag so the verifier **abstains** instead of confidently +false-passing — converting a silent false-pass (the worst failure: you don't know to look) into an +explicit, actionable "write a held-out test." Measured: the confident-false-pass rate on the blind +spot drops **100% → 17%** (N17). + +## The two properties that define the design + +1. **Exogenous, not endogenous.** The trigger is the *external tag* (`backstop`), never the verifier's + self-judgment. Asking the verifier to "abstain if unsure" barely moves the number (100% → 67%) and + only on ambiguity it already notices; on a true blind spot it stays confidently wrong. A confidence + gate cannot reach a blind spot the model does not feel — so there is **no "are you sure?" prompt**; + routing is on the pre-existing tag only. +2. **Routing, not diagnosis.** The verifier need not name the omitted rule (if it could, it wouldn't + be a blind spot). In testing, verifiers abstained correctly while citing the *wrong* edge. The + honest verdict requires only "I was told this is under-specified and I cannot rule it out." The + omitted rule is carried by a human-authored held-out test, not by the verifier. + +## The disposition (the protocol) + +For each `must_haves.truths` item: + +| Item | Confirmable with explicit evidence? | Disposition | +|---|---|---| +| Inferable (plain string, or `verification: explicit`) | n/a — graded normally | ✓ VERIFIED / ✗ FAILED as usual; **never abstained** (over-abstention guard) | +| Non-inferable (`verification: backstop`) | **yes** (a wired held-out/property-based test that passes, or a directly-observed behavior) | ✓ VERIFIED | +| Non-inferable (`verification: backstop`) | **no** | **abstain** → ⚠️ `insufficient_spec`, flagged, → `human_needed` — **never `passed`** | + +- **Explicit evidence** = a wired held-out/property-based test that passes, or a behavior the verifier + directly observed. Symbol presence + wiring is **not** explicit evidence for a non-inferable truth. +- **Never silent, never a hard halt.** *Interactive:* the abstained item routes to the end-of-phase + human checkpoint. *Autonomous (AFK):* it produces a prominent `unverified — held-out test + recommended` flag and the completion line reads "complete with N unverified non-inferable checks"; + the run neither silently passes the blind spot nor hard-halts. +- **Distinguishable reason.** The abstain disposition carries `reason: insufficient_spec` so the + `human_needed` outcome is never conflated with an ordinary manual-UAT `human_needed`. + +This is the verify-time half of ADR-550 Decision 4 (the never-silent-pass disposition), applied to the +edge `backstop` truth tier instead of the prohibition judgment tier — the same machinery, opposite +polarity (must-HAVE under-specified vs must-NOT irreducible). + +## Deterministic engine surface + +The CI-testable surface is the **deterministic disposition + projection**, never the LLM's judgment +(ADR-550 D5 — a test asserting the model's verdict is vacuous and rejected). In `probe-core`: + +- `truthStatement(t)` / `truthVerification(t)` — normalizers; read a truth's statement and tier from + either the plain-string or object form (a truth-reader MUST normalize, never assume a string). +- `projectTruths(items)` — conservative serializer: a `backstop` truth → flat-scalar object + `{ statement, verification: backstop }`; every inferable truth → a bare string. +- `dispositionForUnverifiableTruth(truth, { evidence })` → `{ status, flagged, tier, reason }`: + `backstop` + no evidence → `unverified`/`flagged`/`insufficient_spec`; `backstop` + evidence → + `green`; non-`backstop` → `green` (over-abstention guard). + +## Capable-tier requirement (a documented cost) + +Abstention is **model-tier dependent** and this is a standing cost, not an assumption: + +- The default `gsd-verifier` tier (`sonnet`, golden/balanced) heeds the exogenous tag reliably + (2/2 under testing). +- The **budget tier (`haiku`)** is the least flag-responsive (1/2, inconsistent) and **degrades toward + current behavior** (confident false-pass). Run honest-verifier on a capable tier; treat the budget + tier as best-effort. Re-validate when the `gsd-verifier` model tier changes or a new budget model is + adopted (captured as a test so a tier regression is caught, not discovered in production). + +## Evidence and scope (stated honestly) + +- **Evidence strength.** N17 is n=27 verdicts (3 models × 3 conditions × 3 tasks), 1 rep — + **direction-finding, not powered.** The blind-spot effect is large and monotone + (100% → 67% → 17%); the two costs are clean single events (a *false* tag made the strongest model + over-abstain on a real spec-determined bug; the weakest tier was flag-deaf) and they name exactly + the failure modes the over-abstention guard and the capable-tier requirement defend against. +- **Tag-precision coupling.** Quality is bounded by the edge-probe's `backstop` recall/precision — a + false non-inferable flag causes over-abstention. Positive coupling: improving the probe (#1110) + improves this for free. It adds no independent burden. +- **Explicit non-goals.** Does NOT identify the omitted rule; does NOT recalibrate decisive verdicts; + does NOT defend against *malicious compliance* (a self-graded review rationalizing away its own + findings). It raises the floor on *honest* uncertainty about non-inferable checks — that is the + whole claim. + +## Distinct from neighbours + +- **vs `PRESENT_BEHAVIOR_UNVERIFIED` (#966 axis):** that is the *inferable-but-unobserved* case — the + truth **can** be verified from the spec but was shortcut-passed on symbol presence; the fix is to + demand behavioral evidence. Honest-verifier is the *non-inferable* case — the truth **cannot** be + verified from the spec at all; the fix is to abstain and route to a held-out test. Orthogonal axes + (insufficient *evidence* vs insufficient *spec*); both feed the same `human_needed` sink. +- **vs prohibition judgment-tier (#644):** that disposes **must-NOT** constraints; honest-verifier + disposes **non-inferable positive truths**. Opposite polarity, same never-silent disposition. diff --git a/gsd-core/workflows/plan-phase.md b/gsd-core/workflows/plan-phase.md index 6432bfd22..c9dbacb84 100644 --- a/gsd-core/workflows/plan-phase.md +++ b/gsd-core/workflows/plan-phase.md @@ -912,7 +912,7 @@ Output consumed by /gsd:execute-phase. Plans need: - Tasks in XML format with read_first and acceptance_criteria fields (MANDATORY on every task) - Verification criteria - must_haves for goal-backward verification -- If the SPEC has an `## Edge Coverage` section, lift every `covered` edge's acceptance criterion into `must_haves.truths`, and every `backstop` edge into `must_haves.truths` as a non-inferable check (note it needs a held-out/property-based test). `unresolved` edges are explicit assumptions — surface them in the plan, do not silently drop them. +- If the SPEC has an `## Edge Coverage` section, lift every `covered` edge's acceptance criterion into `must_haves.truths` as a plain string, and every `backstop` edge **as a structured flat-scalar marker** — an object item `{ statement: , verification: backstop }`, NOT a prose note (the verifier branches deterministically on the `verification: backstop` field; a parenthetical is unparseable — the #1110 fragility). Use a flat scalar `verification:` continuation key, never a nested object (ADR-550 #1278). At verify time a `backstop` truth the verifier cannot confirm with explicit evidence abstains → `human_needed` (reason `insufficient_spec`), never a silent pass (#1154; see `references/honest-verifier.md`). `unresolved` edges are explicit assumptions — surface them in the plan, do not silently drop them. - If the SPEC has a `## Prohibitions` section, lift every resolved prohibition into the `must_haves.prohibitions:` sibling block (NOT `truths` — ADR-550 D3) carrying `statement` + `status` + `verification`; unresolved prohibitions are explicit assumptions — surface them in the plan, do not silently drop them. A prohibition is a must-NOT (negative) check that belongs in its own `must_haves.prohibitions` block. Never place a must-NOT under `must_haves.truths` — that block keeps positive-observable semantics only. - **"Artifacts this phase produces" section (MANDATORY)** — list every symbol this phase creates: decorators, classes, functions, CLI flags, struct/dataclass fields, new file paths. The plan-review-convergence source-grounding pass reads this section to exclude newly-created symbols from drift verification; omitting it causes new symbols to be flagged for acknowledgement. diff --git a/gsd-core/workflows/verify-phase.md b/gsd-core/workflows/verify-phase.md index b028c6dce..09e9a1d45 100644 --- a/gsd-core/workflows/verify-phase.md +++ b/gsd-core/workflows/verify-phase.md @@ -117,6 +117,8 @@ For each truth: identify supporting artifacts → check artifact status → chec **Behavior-dependent truths:** when a truth asserts a state transition or a cancellation/cleanup/ordering invariant, symbol presence + wiring is necessary but not sufficient — the code can be present and wired yet still leak state on the path the invariant covers. Mark such a truth ✓ VERIFIED only when a pre-existing test exercises the transition/invariant and passes (one named test, never the full suite); otherwise mark it ⚠️ PRESENT_BEHAVIOR_UNVERIFIED, emit a human-verification item, and exclude it from the verified score. +**Non-inferable (`backstop`) truths (#1154):** a `must_haves.truths` item in object form `{ statement, verification: backstop }` is non-inferable — the correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would false-pass it confidently. Branch on the `verification: backstop` field (read via `truthVerification()`, never prose): if confirmable with **explicit evidence** (a passing wired held-out/property test, or a directly-observed behavior) → ✓ VERIFIED; otherwise **abstain** — mark ⚠️ `insufficient_spec`, emit an `unverified — held-out test recommended` human-verification item, exclude from the verified score (routes to `human_needed`). Exogenous only (never a self-judged "abstain if unsure"); an inferable truth is never abstained. See `references/honest-verifier.md`. + **Example:** Truth "User can see existing messages" depends on Chat.tsx (renders), /api/chat GET (provides), Message model (schema). If Chat.tsx is a stub or API returns hardcoded [] → FAILED. If all exist, are substantive, and connected → VERIFIED. @@ -488,17 +490,22 @@ Classify status using this decision tree IN ORDER (most restrictive first): - **judgment-tier, autonomous run** (non-authoritative LLM-judge verdict): emit the `unverified-prohibition — human review recommended` flag and classify → **human_needed** (autonomous completion reads "complete with N flagged prohibitions"; never a silent pass, never a hard halt). - **judgment-tier, interactive run**: route to the end-of-phase human checkpoint → **human_needed**. -3. IF the previous step produced ANY human verification items — this includes every ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth: +2b. IF any `must_haves.truths` item carries the `verification: backstop` marker (#1154 — the verify-time truth-axis mirror of ADR-550 D4) AND the verifier cannot confirm it with **explicit evidence** (a wired held-out/property-based test that PASSES, or a directly-observed behavior — i.e. `dispositionForUnverifiableTruth()` returns `status: 'unverified'`, `flagged: true`, `reason: 'insufficient_spec'`): + - **abstain → human_needed**, NEVER `passed` and never silently graded green. Emit a prominent `unverified — held-out test recommended` flag carrying the distinguishable `reason: insufficient_spec` (so it is not conflated with ordinary manual-UAT `human_needed`). + - *Autonomous run:* record it and continue — completion reads "complete with N unverified non-inferable checks"; never a hard halt of an AFK run. *Interactive run:* route to the end-of-phase human checkpoint. + - **Exogenous only:** abstention fires SOLELY on the `backstop` tag, never a self-judged "abstain if unsure" (N17). An **inferable** truth is NEVER abstained (over-abstention guard); a `backstop` truth WITH a passing wired held-out test reaches **passed**. Reliable on capable tiers (`sonnet`+); the budget `haiku` tier degrades — see `references/honest-verifier.md`. + +3. IF the previous step produced ANY human verification items — this includes every ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth and every abstained `insufficient_spec` backstop truth: → **human_needed** (even if all other truths VERIFIED) -4. IF all checks pass AND no human verification items AND no flagged prohibitions: +4. IF all checks pass AND no human verification items AND no flagged prohibitions AND no abstained (`insufficient_spec`) truths: → **passed** -**passed is ONLY valid when no human verification items AND no flagged prohibitions exist.** A prohibition (must-NOT) can never be silently absorbed into a `passed` verdict — that is the core failure mode ADR-550 D4 forbids. +**passed is ONLY valid when no human verification items, no flagged prohibitions, AND no abstained `insufficient_spec` truths exist.** Neither a prohibition (must-NOT) nor an unconfirmable non-inferable truth can ever be silently absorbed into a `passed` verdict — that is the core failure mode ADR-550 D4 forbids (now closed on both the prohibition and truth axes). A ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth is never FAILED and never VERIFIED: it does not trigger gaps_found (the code is present and wired) and is not counted as verified (its runtime behavior was not exercised). It routes through the existing human_needed sink — no new overall status. -**Score:** `verified_truths / total_truths` — `verified_truths` counts ✓ VERIFIED truths plus PASSED (override) truths; ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths are the only ones excluded, reported separately as the `behavior_unverified` count. A headline N/N therefore certifies behavioral evidence for every behavior-dependent truth, not merely symbol presence. +**Score:** `verified_truths / total_truths` — `verified_truths` counts ✓ VERIFIED truths plus PASSED (override) truths; excluded are ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths (the `behavior_unverified` count) and abstained ⚠️ `insufficient_spec` backstop truths (#1154) — both are not ✓ VERIFIED and both route to `human_needed`. A headline N/N therefore certifies behavioral evidence for every behavior-dependent truth and explicit evidence for every non-inferable one, not merely symbol presence. diff --git a/src/probe-core.cts b/src/probe-core.cts index d51271e21..0297f9eaa 100644 --- a/src/probe-core.cts +++ b/src/probe-core.cts @@ -489,6 +489,140 @@ export function dispositionForProhibition( }; } +/* ───────────────────────────────────────────────────────────────────────────── + * Honest verifier (#1154) — the truth-axis abstention disposition. + * + * The verify-time MIRROR of the prohibition judgment-tier (ADR-550 D4), applied to the edge + * `backstop` truth tier (D7a). The edge probe already CLASSIFIES a non-inferable check as + * `verification: 'backstop'` and plan-phase lifts it into `must_haves.truths`. This gives that tier + * the same abstain-and-flag disposition D4 gave prohibitions: a `backstop` truth the verifier cannot + * confirm with explicit evidence disposes UNVERIFIED+flagged → `human_needed` (reason + * `insufficient_spec`), NEVER a silent green. An inferable (explicit/plain) truth never abstains (the + * over-abstention guard, AC#3). Exogenous, not endogenous: the trigger is the external `backstop` + * tag, not the verifier's self-judgment (ADR-550 Lineage / N17 — confidence-gating cannot reach a + * blind spot the model does not perceive). + * ───────────────────────────────────────────────────────────────────────────── */ + +/** The truth verification tier — the SAME orthogonal axis the edge adapter uses (ADR-550 D7a). */ +export type TruthVerification = 'explicit' | 'backstop'; + +/** + * A `must_haves.truths` item is EITHER a plain string (an inferable truth — today's shape and the + * overwhelmingly common case) OR a flat-scalar object carrying the non-inferable marker. Object form + * is additive and default-absent (Hyrum's Law): a reader that only ever sees strings behaves + * byte-identically. New readers MUST normalize via `truthStatement`/`truthVerification`. + */ +export type TruthItem = string | { statement: string; verification?: TruthVerification | null }; + +/** Extract a truth's statement text from either the string or the object form (the Hyrum normalizer). */ +export function truthStatement(truth: unknown): string { + if (typeof truth === 'string') return truth; + if (truth != null && typeof truth === 'object') { + const s = (truth as { statement?: unknown }).statement; + if (typeof s === 'string') return s; + } + return ''; +} + +/** + * Extract a truth's verification tier, or `null` when it carries none (a plain string, or an object + * with no/garbled marker). Failing toward `null` is the Postel-safe direction: an unrecognized marker + * grades NORMALLY (never a spurious abstention — the over-abstention guard, AC#3), and the marker is + * machine-emitted from validated edge data so garbling is not a live input path. + */ +export function truthVerification(truth: unknown): TruthVerification | null { + if (truth == null || typeof truth !== 'object') return null; + const v = (truth as { verification?: unknown }).verification; + return v === 'explicit' || v === 'backstop' ? v : null; +} + +/** + * Conservative serializer (Postel: "send well-formed, minimal data") for projecting truths into a + * `must_haves.truths` block — the truth-axis analogue of `projectProhibitions`. A `backstop` truth is + * emitted as a flat-scalar object `{ statement, verification: 'backstop' }` (ADR-550 #1278: flat + * scalars round-trip the existing `parseMustHavesBlock`; a nested object would mangle it). Every other + * truth collapses to a bare statement string — only the non-inferable tier needs a structured marker, + * so an `explicit`/inferable truth never carries one (no spurious markers). Empty statements are dropped. + */ +export function projectTruths( + items: unknown, +): Array { + if (!Array.isArray(items)) return []; + const out: Array = []; + for (const item of items) { + const statement = truthStatement(item); + if (!statement) continue; + if (truthVerification(item) === 'backstop') { + out.push({ statement, verification: 'backstop' }); + } else { + out.push(statement); + } + } + return out; +} + +/** + * The structured verify-time disposition of a single truth. Same shape as `ProhibitionDisposition` + * (status the verifier reads + `flagged` for SUMMARY surfacing + `tier` echo + human-readable + * `reason`), but `reason` carries the STABLE token `insufficient_spec` on abstention so the + * `human_needed` outcome is distinguishable from an ordinary manual-UAT `human_needed` (review + * condition-1 caveat). + */ +export interface TruthDisposition { + status: 'green' | 'unverified'; + flagged: boolean; + tier: TruthVerification | null; + reason: string; +} + +/** Optional context: evidence that a `backstop` truth IS confirmable (a passing wired held-out/PBT test, or a directly-observed behavior). */ +export interface TruthDispositionContext { + evidence?: unknown[]; +} + +/** The stable, distinguishable verdict-reason token for an abstained non-inferable truth (review condition 1). */ +export const INSUFFICIENT_SPEC = 'insufficient_spec'; + +/** + * Deterministic verify-time disposition for a single truth (ADR-550 D4 truth-axis mirror, #1154). + * PURE — no LLM judgment (ADR-550 D5); the LLM verifier's only job is to decide whether `evidence` + * exists, this helper owns the routing once that is known. + * + * - A `backstop` (non-inferable) truth with NO explicit evidence → `{ unverified, flagged }`, + * reason `insufficient_spec`. NEVER green — the verify-time companion to D4's never-silent-pass. + * - A `backstop` truth WITH explicit evidence (a passing wired held-out/property test) → `green`. + * Abstention is for the *unconfirmable*, not for every non-inferable check. + * - Any non-`backstop` truth (explicit, or a plain inferable string) → `green`, never flagged. + * This is the over-abstention guard (AC#3): abstention fires ONLY on the exogenous backstop tag. + */ +export function dispositionForUnverifiableTruth( + truth: unknown, + context: TruthDispositionContext = {}, +): TruthDisposition { + const tier = truthVerification(truth); + // Over-abstention guard (AC#3): only a backstop (non-inferable) truth is ever a candidate to abstain. + if (tier !== 'backstop') { + return { + status: 'green', + flagged: false, + tier, + reason: 'inferable truth — verified normally (no abstention; ADR-550 D4 over-abstention guard)', + }; + } + const evidence = Array.isArray(context.evidence) ? context.evidence : []; + if (evidence.length === 0) { + // ABSTAIN: a non-inferable truth the verifier cannot confirm with explicit evidence. Routes to + // human_needed with the distinguishable insufficient_spec reason — never a silent pass (ADR-550 D4). + return { status: 'unverified', flagged: true, tier, reason: INSUFFICIENT_SPEC }; + } + return { + status: 'green', + flagged: false, + tier, + reason: 'backstop truth confirmed by explicit evidence (a passing wired held-out/property test or directly-observed behavior)', + }; +} + /* * CLI scaffold (the EP-06 invokable surface, generalized). Each probe ships one bin that * calls `runProbeCli` with its own `analyze` (closing over the adapter's propose + validators) diff --git a/src/roadmap.cts b/src/roadmap.cts index 510edf647..2fbe15df2 100644 --- a/src/roadmap.cts +++ b/src/roadmap.cts @@ -78,8 +78,11 @@ function coerceTruthToString(t: unknown): string { return String(t); } if (typeof t === 'object') { - // Prefer common title-bearing keys produced by parseMustHavesBlock - for (const k of ['title', 'text', 'name', 'rule', 'path', 'provides']) { + // Prefer common title-bearing keys produced by parseMustHavesBlock. `statement` is the canonical + // truth/prohibition payload field — and the carrier of #1154's object-form backstop truth + // `{ statement, verification: backstop }`, so it leads (a non-inferable truth must be coerced by + // its statement, never dropped — the Hyrum backward-compat guard for the new marker). + for (const k of ['statement', 'title', 'text', 'name', 'rule', 'path', 'provides']) { const v = (t as Record)[k]; if (typeof v === 'string' && v.trim()) return v; if (typeof v === 'number' || typeof v === 'boolean') return String(v); diff --git a/tests/agent-size-baseline.json b/tests/agent-size-baseline.json index 782aaa5dc..ca6d51408 100644 --- a/tests/agent-size-baseline.json +++ b/tests/agent-size-baseline.json @@ -32,5 +32,5 @@ "gsd-ui-checker.md": 11088, "gsd-ui-researcher.md": 19332, "gsd-user-profiler.md": 8516, - "gsd-verifier.md": 48859 + "gsd-verifier.md": 49124 } diff --git a/tests/enh-2447-roadmap-wave-deps.test.cjs b/tests/enh-2447-roadmap-wave-deps.test.cjs index 7aec5ddd5..e8e725b6a 100644 --- a/tests/enh-2447-roadmap-wave-deps.test.cjs +++ b/tests/enh-2447-roadmap-wave-deps.test.cjs @@ -133,6 +133,34 @@ Plans: assert.ok(roadmap.includes(sharedTruth), 'shared truth listed'); }); + test('#1154: surfaces a cross-cutting backstop (object-form) truth by its statement, not dropped', () => { + // An object-form backstop truth `{ statement, verification: backstop }` (the #1154 non-inferable + // marker on must_haves.truths) shared across 2 plans must be coerced by its `statement` — the + // Hyrum backward-compat guard: a truth-reader must tolerate the new object form, never drop it. + const backstopTruth = 'statement: Adjacent touching intervals merge\n verification: backstop'; + tmpDir = makePlanProject({ + '.planning/ROADMAP.md': `# Roadmap + +### Phase 1: Foundation +**Goal:** Set up project +**Plans:** 2 plans + +Plans: +- [ ] 01-01-PLAN.md — Set up DB +- [ ] 01-02-PLAN.md — Build API +`, + '.planning/phases/01-foundation/01-01-PLAN.md': PLAN_TEMPLATE(1, [backstopTruth, 'DB schema is correct']), + '.planning/phases/01-foundation/01-02-PLAN.md': PLAN_TEMPLATE(2, [backstopTruth, 'API returns 200']), + }); + + const result = runGsdTools('roadmap annotate-dependencies 1', tmpDir); + assert.ok(result.success, `Command failed: ${result.error}`); + const out = JSON.parse(result.output); + assert.strictEqual(out.cross_cutting_constraints, 1, 'the shared backstop truth is surfaced, not dropped'); + const roadmap = fs.readFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), 'utf-8'); + assert.ok(roadmap.includes('Adjacent touching intervals merge'), 'surfaced by its statement text, not [object Object]'); + }); + test('does not surface constraints that appear in only one plan', () => { tmpDir = makePlanProject({ '.planning/ROADMAP.md': `# Roadmap diff --git a/tests/fixtures/golden-install-parity/antigravity.json b/tests/fixtures/golden-install-parity/antigravity.json index 5eb4dc538..b8d76746d 100644 --- a/tests/fixtures/golden-install-parity/antigravity.json +++ b/tests/fixtures/golden-install-parity/antigravity.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "dbbe694a26265473", "agents/gsd-ui-researcher.md": "8c7e91c85e7099f5", "agents/gsd-user-profiler.md": "25d65f6458454764", - "agents/gsd-verifier.md": "2a64590bb09e21e6", + "agents/gsd-verifier.md": "5902e27c7091f4b1", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -86,6 +86,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "9e6076a137f9e156", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "a31a4bf42d82d5e9", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -263,7 +264,7 @@ "gsd-core/workflows/note.md": "3ce09c0aa0a20599", "gsd-core/workflows/pause-work.md": "3196d681d4dd8c71", "gsd-core/workflows/plan-milestone-gaps.md": "94b193dfc9ca3681", - "gsd-core/workflows/plan-phase.md": "d99fb8159a2db1a9", + "gsd-core/workflows/plan-phase.md": "fde84f67730f7131", "gsd-core/workflows/plan-review-convergence.md": "b1623082557cdca5", "gsd-core/workflows/plant-seed.md": "1fb45cd49f66f572", "gsd-core/workflows/pr-branch.md": "c8fd9fa250cb39fd", @@ -297,7 +298,7 @@ "gsd-core/workflows/undo.md": "6ab639d1fc7e0721", "gsd-core/workflows/update.md": "dc93f366e2156e37", "gsd-core/workflows/validate-phase.md": "5bac28c71d21c740", - "gsd-core/workflows/verify-phase.md": "1c6a2e1128966675", + "gsd-core/workflows/verify-phase.md": "e7059c04c816e62d", "gsd-core/workflows/verify-work.md": "dd7f78f947b86976", "hooks/gsd-check-update-worker.js": "668c24ea284ff623", "hooks/gsd-check-update.js": "7e42f76b2bcdd764", diff --git a/tests/fixtures/golden-install-parity/augment.json b/tests/fixtures/golden-install-parity/augment.json index 366d078da..004b272f2 100644 --- a/tests/fixtures/golden-install-parity/augment.json +++ b/tests/fixtures/golden-install-parity/augment.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "4cf947a98db6410e", "agents/gsd-ui-researcher.md": "3f8646572e9c3ec1", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "b1108277a4e858e3", + "agents/gsd-verifier.md": "4046b8d4ca23e342", "commands/gsd-add-tests.md": "3608d0cf4b515103", "commands/gsd-ai-integration-phase.md": "70843d4904743f7e", "commands/gsd-audit-fix.md": "1b805946362c4f19", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "77bf9dff38b2c9d4", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "5a99eb396c744619", "gsd-core/workflows/pause-work.md": "7bcbdf27ba957c8b", "gsd-core/workflows/plan-milestone-gaps.md": "02fee851c82e3b25", - "gsd-core/workflows/plan-phase.md": "fe6b786141eab878", + "gsd-core/workflows/plan-phase.md": "5e3c949ca65cd672", "gsd-core/workflows/plan-review-convergence.md": "10007f8382864bcd", "gsd-core/workflows/plant-seed.md": "7b795d7a1b4c9f06", "gsd-core/workflows/pr-branch.md": "f2a35833fe784a53", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "96d2775f008b3a85", "gsd-core/workflows/update.md": "231e7c305b40c417", "gsd-core/workflows/validate-phase.md": "50f37b705b6e44fb", - "gsd-core/workflows/verify-phase.md": "968d569ea4ef377f", + "gsd-core/workflows/verify-phase.md": "7579af6cd757f644", "gsd-core/workflows/verify-work.md": "6f9c666386cb6d7e", "hooks/gsd-check-update-worker.js": "e42be7414a05e99d", "hooks/gsd-check-update.js": "b3333951b2091807", diff --git a/tests/fixtures/golden-install-parity/claude.json b/tests/fixtures/golden-install-parity/claude.json index 22cf4df49..1c5edb297 100644 --- a/tests/fixtures/golden-install-parity/claude.json +++ b/tests/fixtures/golden-install-parity/claude.json @@ -33,7 +33,7 @@ "agents/gsd-ui-checker.md": "dd06843892f6b0c8", "agents/gsd-ui-researcher.md": "85d7d6cc36388435", "agents/gsd-user-profiler.md": "003276f85792cfda", - "agents/gsd-verifier.md": "0a0c618959bb00d2", + "agents/gsd-verifier.md": "76bcf41aa9fd9b53", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -85,6 +85,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "5c70ef3203b7c9ce", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -262,7 +263,7 @@ "gsd-core/workflows/note.md": "42b66686b2c102cb", "gsd-core/workflows/pause-work.md": "0be71264eafd16dc", "gsd-core/workflows/plan-milestone-gaps.md": "1976bf2001969719", - "gsd-core/workflows/plan-phase.md": "bce0904c3d3b7d59", + "gsd-core/workflows/plan-phase.md": "b0373c2198f4acf7", "gsd-core/workflows/plan-review-convergence.md": "414df04b7ddff73c", "gsd-core/workflows/plant-seed.md": "936a848f1d6c409e", "gsd-core/workflows/pr-branch.md": "c8827e8a15426bf5", @@ -296,7 +297,7 @@ "gsd-core/workflows/undo.md": "791e0bf96d9a057f", "gsd-core/workflows/update.md": "0b389258dcc09332", "gsd-core/workflows/validate-phase.md": "2aa540f2c2479501", - "gsd-core/workflows/verify-phase.md": "452968b6becb18a1", + "gsd-core/workflows/verify-phase.md": "15a999f82868ad29", "gsd-core/workflows/verify-work.md": "d2e8f5d5f2b8f050", "hooks/gsd-check-update-worker.js": "f0c2b5b7169642ba", "hooks/gsd-check-update.js": "a0e4882e66670e4d", diff --git a/tests/fixtures/golden-install-parity/cline.json b/tests/fixtures/golden-install-parity/cline.json index 65afcfb83..d5755c024 100644 --- a/tests/fixtures/golden-install-parity/cline.json +++ b/tests/fixtures/golden-install-parity/cline.json @@ -37,7 +37,7 @@ "agents/gsd-ui-checker.md": "35a6b14813aa03ff", "agents/gsd-ui-researcher.md": "878c7d0e82fa861a", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "64cc793b5f0110bc", + "agents/gsd-verifier.md": "0a437cf3ed90693f", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -89,6 +89,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "f5403e1470e46e69", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -266,7 +267,7 @@ "gsd-core/workflows/note.md": "5a99eb396c744619", "gsd-core/workflows/pause-work.md": "9c1acf8c30a244fd", "gsd-core/workflows/plan-milestone-gaps.md": "95ca791b0867fe2d", - "gsd-core/workflows/plan-phase.md": "2716d6636e37ce6a", + "gsd-core/workflows/plan-phase.md": "ab9ed0b5acfe7d8a", "gsd-core/workflows/plan-review-convergence.md": "2cdba6216184aac4", "gsd-core/workflows/plant-seed.md": "d0d63f83ae7c939b", "gsd-core/workflows/pr-branch.md": "15ccec6bd303ea46", @@ -300,7 +301,7 @@ "gsd-core/workflows/undo.md": "96d2775f008b3a85", "gsd-core/workflows/update.md": "9cad8a8f4baff922", "gsd-core/workflows/validate-phase.md": "54687a0f2a562fc4", - "gsd-core/workflows/verify-phase.md": "68a74f247fdce962", + "gsd-core/workflows/verify-phase.md": "5caca0023bc88a9a", "gsd-core/workflows/verify-work.md": "6f482d32e0c8d49a", "scripts/changeset/README.md": "86ff89331dfd94b2", "scripts/changeset/cli.cjs": "68f92a344b199271", diff --git a/tests/fixtures/golden-install-parity/codebuddy.json b/tests/fixtures/golden-install-parity/codebuddy.json index 84bc21f8c..253339c25 100644 --- a/tests/fixtures/golden-install-parity/codebuddy.json +++ b/tests/fixtures/golden-install-parity/codebuddy.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "68ee3eb209c9f0d6", "agents/gsd-ui-researcher.md": "507ed07fc9ba7d33", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "b62b7966b4aafd5b", + "agents/gsd-verifier.md": "b3983159ba46fcbf", "commands/gsd-add-tests.md": "aa65032141557fb7", "commands/gsd-ai-integration-phase.md": "373053d7cd887aa9", "commands/gsd-audit-fix.md": "c21ef925131fade7", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "77bf9dff38b2c9d4", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "5a99eb396c744619", "gsd-core/workflows/pause-work.md": "7bcbdf27ba957c8b", "gsd-core/workflows/plan-milestone-gaps.md": "02fee851c82e3b25", - "gsd-core/workflows/plan-phase.md": "3b30ccd918b5f703", + "gsd-core/workflows/plan-phase.md": "58d02dbdb533603a", "gsd-core/workflows/plan-review-convergence.md": "10007f8382864bcd", "gsd-core/workflows/plant-seed.md": "7b795d7a1b4c9f06", "gsd-core/workflows/pr-branch.md": "f2a35833fe784a53", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "96d2775f008b3a85", "gsd-core/workflows/update.md": "6a593863d1b0a287", "gsd-core/workflows/validate-phase.md": "50f37b705b6e44fb", - "gsd-core/workflows/verify-phase.md": "968d569ea4ef377f", + "gsd-core/workflows/verify-phase.md": "7579af6cd757f644", "gsd-core/workflows/verify-work.md": "6f9c666386cb6d7e", "hooks/gsd-check-update-worker.js": "084c3f7109bb4d29", "hooks/gsd-check-update.js": "23074675b7c31a47", diff --git a/tests/fixtures/golden-install-parity/codex.json b/tests/fixtures/golden-install-parity/codex.json index 2ebc6cae0..daddd245e 100644 --- a/tests/fixtures/golden-install-parity/codex.json +++ b/tests/fixtures/golden-install-parity/codex.json @@ -67,8 +67,8 @@ "agents/gsd-ui-researcher.toml": "ffe2ca0b232df3df", "agents/gsd-user-profiler.md": "1bf5033c929181c1", "agents/gsd-user-profiler.toml": "b9c244bb8fbf8140", - "agents/gsd-verifier.md": "994e2a4f22bc3c8b", - "agents/gsd-verifier.toml": "9743a551e04bc0a3", + "agents/gsd-verifier.md": "3471118de7e985c7", + "agents/gsd-verifier.toml": "68f4dcc6c0311e00", "config.toml": "a34316b9ac61a620", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", @@ -121,6 +121,7 @@ "gsd-core/references/gates.md": "11fd2bdf27df57a5", "gsd-core/references/git-integration.md": "94715b69448bb818", "gsd-core/references/git-planning-commit.md": "5aad099c51135a0f", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -298,7 +299,7 @@ "gsd-core/workflows/note.md": "664da466ab989f9d", "gsd-core/workflows/pause-work.md": "3c2ee96295959527", "gsd-core/workflows/plan-milestone-gaps.md": "8ee841bcc7adc836", - "gsd-core/workflows/plan-phase.md": "4bef5b532a009f9b", + "gsd-core/workflows/plan-phase.md": "f9a1846e30603d38", "gsd-core/workflows/plan-review-convergence.md": "f27818025e279aa4", "gsd-core/workflows/plant-seed.md": "bfa729ff2ad3441a", "gsd-core/workflows/pr-branch.md": "3b13915429c0e8d9", @@ -332,7 +333,7 @@ "gsd-core/workflows/undo.md": "5ff7d63b0a2f46d5", "gsd-core/workflows/update.md": "4166fd3e3ca72a7b", "gsd-core/workflows/validate-phase.md": "4e94708c9ca15d70", - "gsd-core/workflows/verify-phase.md": "b81bd1c061c9e6d3", + "gsd-core/workflows/verify-phase.md": "59bb0b12ebb47f5e", "gsd-core/workflows/verify-work.md": "92fb4e876d75ded5", "hooks/gsd-check-update.js": "79c846cd8dd54caa", "hooks/gsd-context-monitor.js": "caa8614524452835", diff --git a/tests/fixtures/golden-install-parity/copilot.json b/tests/fixtures/golden-install-parity/copilot.json index 7452f8766..caa6e7c47 100644 --- a/tests/fixtures/golden-install-parity/copilot.json +++ b/tests/fixtures/golden-install-parity/copilot.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.agent.md": "4e46f787d6420062", "agents/gsd-ui-researcher.agent.md": "9dd2acff7b24230e", "agents/gsd-user-profiler.agent.md": "ae16a248e18dd42b", - "agents/gsd-verifier.agent.md": "ebd2c921c24e2d9b", + "agents/gsd-verifier.agent.md": "6fe2c4f008ebfa22", "copilot-instructions.md": "1fb04111759f1645", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", @@ -87,6 +87,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "657a93c539c16cab", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "08b1e82f59c18078", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -264,7 +265,7 @@ "gsd-core/workflows/note.md": "4a5ee74cf2fc1f54", "gsd-core/workflows/pause-work.md": "324e04e675dc7f7f", "gsd-core/workflows/plan-milestone-gaps.md": "c0eb896eb42e22d4", - "gsd-core/workflows/plan-phase.md": "50042ca359933c51", + "gsd-core/workflows/plan-phase.md": "16ff6a162cbd4c01", "gsd-core/workflows/plan-review-convergence.md": "c238b5858ceb4e57", "gsd-core/workflows/plant-seed.md": "56451bdf104983b3", "gsd-core/workflows/pr-branch.md": "a080aed95785cf32", @@ -298,7 +299,7 @@ "gsd-core/workflows/undo.md": "ba1ef7aa80bef6bd", "gsd-core/workflows/update.md": "712ab18a9b7c14f5", "gsd-core/workflows/validate-phase.md": "f923eac9442b6053", - "gsd-core/workflows/verify-phase.md": "543845af6f701518", + "gsd-core/workflows/verify-phase.md": "eea9b642393b6b90", "gsd-core/workflows/verify-work.md": "e253fb8e33c68fa4", "hooks/gsd-session.json": "0a462834f2a28fee", "scripts/changeset/README.md": "86ff89331dfd94b2", diff --git a/tests/fixtures/golden-install-parity/cursor.json b/tests/fixtures/golden-install-parity/cursor.json index 7709df89b..9fff3e76f 100644 --- a/tests/fixtures/golden-install-parity/cursor.json +++ b/tests/fixtures/golden-install-parity/cursor.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "66b380ffcdfec7b5", "agents/gsd-ui-researcher.md": "fe237aa42ccd9ad2", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "5ec35372f0a58d48", + "agents/gsd-verifier.md": "7bacf6ef32186213", "commands/gsd-add-tests.md": "1f89b16ab2cca426", "commands/gsd-ai-integration-phase.md": "3ccac39673ffb92c", "commands/gsd-audit-fix.md": "d88502ffa2151cec", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "9840f521ad6612b5", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "1c1e466c764e3deb", "gsd-core/workflows/pause-work.md": "0be71264eafd16dc", "gsd-core/workflows/plan-milestone-gaps.md": "1976bf2001969719", - "gsd-core/workflows/plan-phase.md": "8975739befb097bc", + "gsd-core/workflows/plan-phase.md": "1f40b9c15ebef2ff", "gsd-core/workflows/plan-review-convergence.md": "783c5171101b26b9", "gsd-core/workflows/plant-seed.md": "b28224f37faa9b95", "gsd-core/workflows/pr-branch.md": "1f77535b6b156ad9", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "18dec684fb1076f9", "gsd-core/workflows/update.md": "a27dcfd2814bf2b1", "gsd-core/workflows/validate-phase.md": "f513c28c01a44cb7", - "gsd-core/workflows/verify-phase.md": "452968b6becb18a1", + "gsd-core/workflows/verify-phase.md": "15a999f82868ad29", "gsd-core/workflows/verify-work.md": "24d323b667d15d01", "hooks/gsd-cursor-post-tool.js": "019d503aee8b4a3f", "hooks/gsd-cursor-session-start.js": "c6e04ed597ea7020", diff --git a/tests/fixtures/golden-install-parity/gemini.json b/tests/fixtures/golden-install-parity/gemini.json index 71581553b..7a0cf1724 100644 --- a/tests/fixtures/golden-install-parity/gemini.json +++ b/tests/fixtures/golden-install-parity/gemini.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "cb122369806c5467", "agents/gsd-ui-researcher.md": "e1422e3b0f142f74", "agents/gsd-user-profiler.md": "d6cb5430d841cea6", - "agents/gsd-verifier.md": "6f3d8f7d56b72475", + "agents/gsd-verifier.md": "e52203ef59021e9c", "commands/gsd/add-tests.toml": "297ba4d4c285fd99", "commands/gsd/ai-integration-phase.toml": "411ba33b9a7d9e78", "commands/gsd/audit-fix.toml": "f4a198a455f668a9", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "77bf9dff38b2c9d4", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "95199087905ba3e6", "gsd-core/workflows/pause-work.md": "7bcbdf27ba957c8b", "gsd-core/workflows/plan-milestone-gaps.md": "02fee851c82e3b25", - "gsd-core/workflows/plan-phase.md": "b2904cffb7ff0127", + "gsd-core/workflows/plan-phase.md": "81c6ef4292e5f9bf", "gsd-core/workflows/plan-review-convergence.md": "57c284df29121220", "gsd-core/workflows/plant-seed.md": "339890021123e017", "gsd-core/workflows/pr-branch.md": "f2a35833fe784a53", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "993bb0a31edca7f8", "gsd-core/workflows/update.md": "51c3af33474bdc24", "gsd-core/workflows/validate-phase.md": "2d998cb78c918d08", - "gsd-core/workflows/verify-phase.md": "968d569ea4ef377f", + "gsd-core/workflows/verify-phase.md": "7579af6cd757f644", "gsd-core/workflows/verify-work.md": "1265e07f2c74a218", "hooks/gsd-check-update-worker.js": "7a3eba8c1c166dd2", "hooks/gsd-check-update.js": "c1c78326299f6eae", diff --git a/tests/fixtures/golden-install-parity/hermes.json b/tests/fixtures/golden-install-parity/hermes.json index 2c04105c0..8263bee3a 100644 --- a/tests/fixtures/golden-install-parity/hermes.json +++ b/tests/fixtures/golden-install-parity/hermes.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "32b2605c7254f959", "agents/gsd-ui-researcher.md": "87cf9eebc37c4a37", "agents/gsd-user-profiler.md": "ca3bf75581f211a0", - "agents/gsd-verifier.md": "4decc9b596090ca6", + "agents/gsd-verifier.md": "c1d1448bf31039ec", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -86,6 +86,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "12d23fcbaa7fcf06", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -263,7 +264,7 @@ "gsd-core/workflows/note.md": "42b66686b2c102cb", "gsd-core/workflows/pause-work.md": "2c3abcaa1fa6d2e2", "gsd-core/workflows/plan-milestone-gaps.md": "419744c1191354af", - "gsd-core/workflows/plan-phase.md": "7d6cbb6730d852d2", + "gsd-core/workflows/plan-phase.md": "ad2d30c92b495bcb", "gsd-core/workflows/plan-review-convergence.md": "8ebcd05fcd5a0d15", "gsd-core/workflows/plant-seed.md": "76379e2aeb18d53b", "gsd-core/workflows/pr-branch.md": "79f4d93e31af9c49", @@ -297,7 +298,7 @@ "gsd-core/workflows/undo.md": "791e0bf96d9a057f", "gsd-core/workflows/update.md": "472d4905fc8c5ce5", "gsd-core/workflows/validate-phase.md": "067acd7aeed52c5b", - "gsd-core/workflows/verify-phase.md": "fecef63dcecc4104", + "gsd-core/workflows/verify-phase.md": "4e1d99795e8e9413", "gsd-core/workflows/verify-work.md": "aa0b4d000b15dbf1", "hooks/gsd-check-update-worker.js": "80865e926e22623e", "hooks/gsd-check-update.js": "57433c9f9b224e3e", diff --git a/tests/fixtures/golden-install-parity/kilo.json b/tests/fixtures/golden-install-parity/kilo.json index c281079c3..38632837f 100644 --- a/tests/fixtures/golden-install-parity/kilo.json +++ b/tests/fixtures/golden-install-parity/kilo.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "2a6c5551e354414b", "agents/gsd-ui-researcher.md": "384af56f90784422", "agents/gsd-user-profiler.md": "b8cb09319c517701", - "agents/gsd-verifier.md": "b16febf0774e2db7", + "agents/gsd-verifier.md": "0cd3672d8ad81dd0", "command/gsd-add-tests.md": "c314f9a6be312bfa", "command/gsd-ai-integration-phase.md": "bbab86a540e00db5", "command/gsd-audit-fix.md": "4106e9dce9b71585", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "fbdf814a3af9c051", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "f8c2842a2217f776", "gsd-core/workflows/pause-work.md": "8b81699a46ca8e9b", "gsd-core/workflows/plan-milestone-gaps.md": "1976bf2001969719", - "gsd-core/workflows/plan-phase.md": "2d8eda8f283735d2", + "gsd-core/workflows/plan-phase.md": "38d57cf05c12ab50", "gsd-core/workflows/plan-review-convergence.md": "88e294a5cc616e90", "gsd-core/workflows/plant-seed.md": "249d3c6b4d474106", "gsd-core/workflows/pr-branch.md": "c8827e8a15426bf5", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "7cd2153f8b15e30d", "gsd-core/workflows/update.md": "5d0a2356dadf2017", "gsd-core/workflows/validate-phase.md": "7915e3a261f7c9c9", - "gsd-core/workflows/verify-phase.md": "452968b6becb18a1", + "gsd-core/workflows/verify-phase.md": "15a999f82868ad29", "gsd-core/workflows/verify-work.md": "7cb5144cc74986e7", "hooks/gsd-check-update-worker.js": "8c48db40d2d74193", "hooks/gsd-check-update.js": "1c863b30953b47c8", diff --git a/tests/fixtures/golden-install-parity/kimi.json b/tests/fixtures/golden-install-parity/kimi.json index 42372d0da..54faf0b4e 100644 --- a/tests/fixtures/golden-install-parity/kimi.json +++ b/tests/fixtures/golden-install-parity/kimi.json @@ -69,7 +69,7 @@ "agents/subagents/gsd-ui-researcher.yaml": "15e215874bc2c37d", "agents/subagents/gsd-user-profiler.md": "c16f94e09e433394", "agents/subagents/gsd-user-profiler.yaml": "645826a29d079159", - "agents/subagents/gsd-verifier.md": "3345b59f7a3b8de3", + "agents/subagents/gsd-verifier.md": "1536f3fa5a2d4419", "agents/subagents/gsd-verifier.yaml": "2d2bd6b37626f382", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", @@ -122,6 +122,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "77bf9dff38b2c9d4", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -299,7 +300,7 @@ "gsd-core/workflows/note.md": "5a99eb396c744619", "gsd-core/workflows/pause-work.md": "7bcbdf27ba957c8b", "gsd-core/workflows/plan-milestone-gaps.md": "02fee851c82e3b25", - "gsd-core/workflows/plan-phase.md": "f210c67e41f89f4a", + "gsd-core/workflows/plan-phase.md": "f40fb8f961a35925", "gsd-core/workflows/plan-review-convergence.md": "10007f8382864bcd", "gsd-core/workflows/plant-seed.md": "7b795d7a1b4c9f06", "gsd-core/workflows/pr-branch.md": "f2a35833fe784a53", @@ -333,7 +334,7 @@ "gsd-core/workflows/undo.md": "96d2775f008b3a85", "gsd-core/workflows/update.md": "6369691790864147", "gsd-core/workflows/validate-phase.md": "50f37b705b6e44fb", - "gsd-core/workflows/verify-phase.md": "968d569ea4ef377f", + "gsd-core/workflows/verify-phase.md": "7579af6cd757f644", "gsd-core/workflows/verify-work.md": "6f9c666386cb6d7e", "scripts/changeset/README.md": "86ff89331dfd94b2", "scripts/changeset/cli.cjs": "68f92a344b199271", diff --git a/tests/fixtures/golden-install-parity/opencode.json b/tests/fixtures/golden-install-parity/opencode.json index e238d774c..a2ef57ddf 100644 --- a/tests/fixtures/golden-install-parity/opencode.json +++ b/tests/fixtures/golden-install-parity/opencode.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "b98b00e586610657", "agents/gsd-ui-researcher.md": "7ff5c528bef6af09", "agents/gsd-user-profiler.md": "d825b4c0a6431f8b", - "agents/gsd-verifier.md": "1e0ff1d8d15a48ab", + "agents/gsd-verifier.md": "47816fa1ba8a9b8b", "command/gsd-add-tests.md": "b9cd93ed01945f75", "command/gsd-ai-integration-phase.md": "9d4bc4dcce7f1ee0", "command/gsd-audit-fix.md": "5615403843d5e491", @@ -155,6 +155,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "fbdf814a3af9c051", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "b67482409f896a10", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -332,7 +333,7 @@ "gsd-core/workflows/note.md": "0d1374f2a2257858", "gsd-core/workflows/pause-work.md": "c9f0b8826845dda7", "gsd-core/workflows/plan-milestone-gaps.md": "5ec459734bf7570c", - "gsd-core/workflows/plan-phase.md": "1aea2bd181a782cf", + "gsd-core/workflows/plan-phase.md": "6db3e766b7703874", "gsd-core/workflows/plan-review-convergence.md": "4f884d793d72d3fd", "gsd-core/workflows/plant-seed.md": "507e886d0f70d84c", "gsd-core/workflows/pr-branch.md": "3abaa18246facdc3", @@ -366,7 +367,7 @@ "gsd-core/workflows/undo.md": "0bba5e7f6196c894", "gsd-core/workflows/update.md": "af51041a3172c523", "gsd-core/workflows/validate-phase.md": "0fc5991f6d7c6c39", - "gsd-core/workflows/verify-phase.md": "d8999a0a1cd7b6de", + "gsd-core/workflows/verify-phase.md": "13a1a43395b5e47b", "gsd-core/workflows/verify-work.md": "52f6011634cff599", "hooks/gsd-check-update-worker.js": "5882dbd41918c863", "hooks/gsd-check-update.js": "5fb0027b7b76986c", diff --git a/tests/fixtures/golden-install-parity/qwen.json b/tests/fixtures/golden-install-parity/qwen.json index 6b9f86ec2..b26736970 100644 --- a/tests/fixtures/golden-install-parity/qwen.json +++ b/tests/fixtures/golden-install-parity/qwen.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "7a383f6a3fbba34b", "agents/gsd-ui-researcher.md": "67596ca5ee2f4547", "agents/gsd-user-profiler.md": "ca3bf75581f211a0", - "agents/gsd-verifier.md": "82b7967e24c2065b", + "agents/gsd-verifier.md": "4bc6c8e197ee56e0", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -86,6 +86,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "bd249d9024c39c0d", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -263,7 +264,7 @@ "gsd-core/workflows/note.md": "42b66686b2c102cb", "gsd-core/workflows/pause-work.md": "54f1c0a9e79e2c59", "gsd-core/workflows/plan-milestone-gaps.md": "1b820f4e70acce86", - "gsd-core/workflows/plan-phase.md": "fcc0bfeadf4aa5c4", + "gsd-core/workflows/plan-phase.md": "140b9f4360646c96", "gsd-core/workflows/plan-review-convergence.md": "b809588bff4f48b5", "gsd-core/workflows/plant-seed.md": "e3949fcbcf5d375f", "gsd-core/workflows/pr-branch.md": "95850381230787ed", @@ -297,7 +298,7 @@ "gsd-core/workflows/undo.md": "791e0bf96d9a057f", "gsd-core/workflows/update.md": "baae1a977fbeba49", "gsd-core/workflows/validate-phase.md": "0e153ccb3bdb9dda", - "gsd-core/workflows/verify-phase.md": "33a17f66c8291096", + "gsd-core/workflows/verify-phase.md": "0ef74d5b7f7646f6", "gsd-core/workflows/verify-work.md": "8a4157d179fce9c4", "hooks/gsd-check-update-worker.js": "74edb6f6b1010c85", "hooks/gsd-check-update.js": "81962511d619d038", diff --git a/tests/fixtures/golden-install-parity/trae.json b/tests/fixtures/golden-install-parity/trae.json index 3c0f90b05..1ec2a976b 100644 --- a/tests/fixtures/golden-install-parity/trae.json +++ b/tests/fixtures/golden-install-parity/trae.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "843c1deeaa1cb5e7", "agents/gsd-ui-researcher.md": "e09f0f7034eaa366", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "1e8a6492303366a0", + "agents/gsd-verifier.md": "3065cc4cf897078d", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -86,6 +86,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "1bd40c3f94d712b1", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -263,7 +264,7 @@ "gsd-core/workflows/note.md": "acc9130fb1e94f0b", "gsd-core/workflows/pause-work.md": "09a6b8980f7b771a", "gsd-core/workflows/plan-milestone-gaps.md": "022822b3b6971b75", - "gsd-core/workflows/plan-phase.md": "c3e494c63f5c04c5", + "gsd-core/workflows/plan-phase.md": "9781d27e30242ed8", "gsd-core/workflows/plan-review-convergence.md": "8fb889570c17db15", "gsd-core/workflows/plant-seed.md": "dc911406ada4d188", "gsd-core/workflows/pr-branch.md": "e939128047e02395", @@ -297,7 +298,7 @@ "gsd-core/workflows/undo.md": "59b8baa54efc4110", "gsd-core/workflows/update.md": "e259898f86ae7133", "gsd-core/workflows/validate-phase.md": "70d102b4d5113dd4", - "gsd-core/workflows/verify-phase.md": "ba797ce64e593a10", + "gsd-core/workflows/verify-phase.md": "67eefbd1f6417281", "gsd-core/workflows/verify-work.md": "1e2a10168f8fdc2b", "scripts/changeset/README.md": "86ff89331dfd94b2", "scripts/changeset/cli.cjs": "68f92a344b199271", diff --git a/tests/fixtures/golden-install-parity/windsurf.json b/tests/fixtures/golden-install-parity/windsurf.json index 6e3743cc8..4afdb5467 100644 --- a/tests/fixtures/golden-install-parity/windsurf.json +++ b/tests/fixtures/golden-install-parity/windsurf.json @@ -34,7 +34,7 @@ "agents/gsd-ui-checker.md": "bd8e9be4c75dcdcf", "agents/gsd-ui-researcher.md": "f04d24e6459bd23d", "agents/gsd-user-profiler.md": "622220df0654b6bf", - "agents/gsd-verifier.md": "88eac9841f52dd8c", + "agents/gsd-verifier.md": "e6353ed8a4baaa32", "gsd-core/CHANGELOG.md": "e141e3fb369ff712", "gsd-core/VERSION": "562368b20a64be95", "gsd-core/bin/check-latest-version.cjs": "e4a224058c8f4d74", @@ -86,6 +86,7 @@ "gsd-core/references/gates.md": "7dc9fd3a3d6217c6", "gsd-core/references/git-integration.md": "6ec36ea6b868655f", "gsd-core/references/git-planning-commit.md": "f897a15ebfc3f5a7", + "gsd-core/references/honest-verifier.md": "8815c9fc18c35719", "gsd-core/references/ios-scaffold.md": "5ef0cb7e0fac891f", "gsd-core/references/loop-hook-dispatch.md": "32e5dfb4dba76987", "gsd-core/references/mandatory-initial-read.md": "fe59abce693717cf", @@ -263,7 +264,7 @@ "gsd-core/workflows/note.md": "1c1e466c764e3deb", "gsd-core/workflows/pause-work.md": "5ce6a137bd9fc0f8", "gsd-core/workflows/plan-milestone-gaps.md": "19911e87fd4185f2", - "gsd-core/workflows/plan-phase.md": "98210e414d5d9e0b", + "gsd-core/workflows/plan-phase.md": "55ddbe990fcaa74f", "gsd-core/workflows/plan-review-convergence.md": "7e01c19b7b9a2aad", "gsd-core/workflows/plant-seed.md": "9d36ecd08093a494", "gsd-core/workflows/pr-branch.md": "cc28ab5cd9db16b9", @@ -297,7 +298,7 @@ "gsd-core/workflows/undo.md": "18dec684fb1076f9", "gsd-core/workflows/update.md": "cbf978622ad0f577", "gsd-core/workflows/validate-phase.md": "a36cf2c688c261ac", - "gsd-core/workflows/verify-phase.md": "6adfc47bee438b5d", + "gsd-core/workflows/verify-phase.md": "8a89a915dad5960b", "gsd-core/workflows/verify-work.md": "7586bdeeeb9ecba6", "scripts/changeset/README.md": "86ff89331dfd94b2", "scripts/changeset/cli.cjs": "68f92a344b199271", diff --git a/tests/probe-core.test.cjs b/tests/probe-core.test.cjs index ac1d00ce3..15e65a8dc 100644 --- a/tests/probe-core.test.cjs +++ b/tests/probe-core.test.cjs @@ -513,3 +513,111 @@ describe('probe-core: projectProhibitions descriptor projection (CHK-02)', () => assert.ok(!('check_rule' in projected[0]), 'descriptor-less item gains no check_rule'); }); }); + +// ─── Honest verifier (#1154): truth-axis abstention — the verify-time MIRROR of #644's +// prohibition judgment-tier (ADR-550 D4), applied to the edge `backstop` truth tier (D7a). +// Per ADR-550 D5 the deterministic, CI-testable surface is the disposition helper + the +// projection ROUND-TRIP — NEVER the LLM verdict (a test asserting the model's judgment is +// vacuous and rejected). The abstain-on-unconfirmed-backstop case is the REGRESSION +// (trek-e review condition 5) that must fail RED on `next` before the fix. +const FM_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'frontmatter.cjs'); +const fm = require(FM_SCRIPT); + +// Serialize projected truths into a plan-frontmatter `must_haves.truths` block exactly as +// plan-phase emits it — a backstop truth as a flat-scalar object (ADR-550 #1278: flat scalars, +// NEVER a nested object), a plain inferable truth as a bare string. +function renderTruthsBlock(projected) { + const lines = ['---', 'must_haves:', ' truths:']; + for (const t of projected) { + if (typeof t === 'string') { + lines.push(` - ${t}`); + } else { + lines.push(` - statement: ${t.statement}`); + if (t.verification) lines.push(` verification: ${t.verification}`); + } + } + lines.push('---'); + return lines.join('\n'); +} + +describe('probe-core: truthStatement / truthVerification normalizers (#1154, Hyrum backward-compat)', () => { + test('truthStatement extracts text from a plain-string truth and an object-form truth identically', () => { + assert.equal(pc.truthStatement('Overlapping intervals are merged'), 'Overlapping intervals are merged'); + assert.equal( + pc.truthStatement({ statement: 'Adjacent intervals merge', verification: 'backstop' }), + 'Adjacent intervals merge', + ); + }); + + test('truthVerification returns the tier for an object-form truth and null for a plain string (no spurious tier)', () => { + assert.equal(pc.truthVerification({ statement: 'Adjacent intervals merge', verification: 'backstop' }), 'backstop'); + assert.equal(pc.truthVerification('Overlapping intervals are merged'), null); + assert.equal(pc.truthVerification({ statement: 'x', verification: 'explicit' }), 'explicit'); + }); +}); + +describe('probe-core: dispositionForUnverifiableTruth (#1154, ADR-550 D4 truth-axis mirror)', () => { + test('abstain-on-unconfirmed-backstop (condition 5, REGRESSION): a backstop truth with no explicit evidence disposes insufficient_spec/unverified/flagged — never green', () => { + const d = pc.dispositionForUnverifiableTruth( + { statement: 'Adjacent/touching intervals [1,2],[2,3] merge', verification: 'backstop' }, + { evidence: [] }, + ); + assert.equal(d.status, 'unverified', 'a backstop truth with no explicit evidence is unverified'); + assert.equal(d.flagged, true, 'and flagged — never a silent pass (ADR-550 D4)'); + assert.equal(d.tier, 'backstop', 'the backstop tier is echoed unchanged'); + assert.equal( + d.reason, + 'insufficient_spec', + 'carries the distinguishable insufficient_spec reason code so human_needed is not conflated with ordinary manual-UAT', + ); + }); + + test('pass-on-wired-backstop: a backstop truth WITH explicit evidence (a wired held-out/property test) disposes green — abstention is for the unconfirmable only', () => { + const d = pc.dispositionForUnverifiableTruth( + { statement: 'Adjacent/touching intervals [1,2],[2,3] merge', verification: 'backstop' }, + { evidence: [{ test: 'tests/intervals.property.test.cjs', passed: true }] }, + ); + assert.equal(d.status, 'green'); + assert.equal(d.flagged, false); + assert.equal(d.tier, 'backstop'); + }); + + test('over-abstention guard (AC#3): a plain inferable truth is NEVER routed to abstention', () => { + const d = pc.dispositionForUnverifiableTruth('Overlapping intervals are merged', { evidence: [] }); + assert.equal(d.status, 'green', 'a plain inferable truth is graded normally, never abstained'); + assert.equal(d.flagged, false); + }); + + test('over-abstention guard (AC#3): an explicit-tier truth NEVER abstains even with no evidence (only backstop triggers it)', () => { + const d = pc.dispositionForUnverifiableTruth({ statement: 'Symbol X is wired', verification: 'explicit' }, { evidence: [] }); + assert.equal(d.status, 'green', 'an explicit (inferable) truth never abstains'); + assert.equal(d.flagged, false); + }); +}); + +describe('probe-core: projectTruths (#1154, conservative serializer — Postel) + round-trip parity (condition 4, ADR-550 D5b)', () => { + test('projectTruths emits the flat-scalar backstop marker and collapses inferable truths to plain strings', () => { + const projected = pc.projectTruths([ + { statement: 'Adjacent intervals merge', verification: 'backstop' }, + 'Overlapping intervals are merged', + { statement: 'Symbol X is wired', verification: 'explicit' }, + ]); + assert.deepEqual(projected, [ + { statement: 'Adjacent intervals merge', verification: 'backstop' }, + 'Overlapping intervals are merged', + 'Symbol X is wired', + ], 'only a backstop truth carries a structured marker; explicit/inferable truths stay bare strings (no spurious markers)'); + }); + + test('round-trip parity (SPEC backstop edge → must_haves.truths marker → read-back): the marker survives as a structured field, plain truths byte-identically', () => { + const projected = pc.projectTruths([ + { statement: 'Adjacent intervals merge', verification: 'backstop' }, + 'Overlapping intervals are merged', + ]); + const parsed = fm.parseMustHavesBlock(renderTruthsBlock(projected), 'truths'); + assert.equal(pc.truthVerification(parsed[0]), 'backstop', 'the backstop marker survives the round-trip as a structured field, not prose (#1110 fragility avoided)'); + assert.equal(pc.truthStatement(parsed[0]), 'Adjacent intervals merge'); + assert.equal(pc.truthStatement(parsed[1]), 'Overlapping intervals are merged'); + assert.equal(pc.truthVerification(parsed[1]), null, 'a plain truth round-trips with no marker (Hyrum byte-identity backward-compat)'); + }); +}); diff --git a/tests/workflow-size-baseline.json b/tests/workflow-size-baseline.json index f5e697b36..c0e2ef9dc 100644 --- a/tests/workflow-size-baseline.json +++ b/tests/workflow-size-baseline.json @@ -52,7 +52,7 @@ "note.md": 6563, "pause-work.md": 14397, "plan-milestone-gaps.md": 11765, - "plan-phase.md": 93973, + "plan-phase.md": 94459, "plan-review-convergence.md": 23468, "plant-seed.md": 11741, "pr-branch.md": 15919, @@ -86,6 +86,6 @@ "undo.md": 10431, "update.md": 21053, "validate-phase.md": 10745, - "verify-phase.md": 38228, + "verify-phase.md": 40728, "verify-work.md": 35212 }