Files
msd-core/gsd-core/references/verifier-phase-gates.md
0xdhx 285cd41be0 fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites (#3435)
* fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites

Step 3 item 5b abstained on non-inferable (backstop) truths "unless
confirmed by explicit evidence" with the term undefined — its definition
lived only in the non-included gsd-core/references/honest-verifier.md,
behind a stale bare `references/` cite that 404s. Undefined, the term
falls back to presence + wiring, the exact false-pass the #1154
abstention protocol refuses.

- 5b: inline the compressed definition (a passing wired
  held-out/property-based test or directly observed behavior; presence +
  wiring never qualifies) and fix the cite. +84 B on the rewritten line;
  file lands at 49,151 of the 49,152 LARGE cap.
- verifier-phase-gates.md (already <required_reading>): gains the
  backstop-abstention reporting contract — AFK completion line
  ("complete with N unverified non-inferable checks", never silent,
  never a halt) and reason-distinctness (insufficient_spec vs manual-UAT
  human_needed). New content, no relocation of measured prose.
- 5c (line 204) and MVP-mode (line 644) bare cites repaired to
  gsd-core/references/ (+9 B each).
- Drift acks per ADR-2719 §4; two entries merge-appended into existing
  fragments (two ack sources may never name the same path).
- Changeset fragment with the sanctioned pr: 0 placeholder (post-create
  backfill).

Sibling census at next@7976b1ca0: 7 bare-cite instances in 4 agent
files; the 3 in gsd-verifier.md are fixed here, gsd-executor.md:429,439
and gsd-doc-synthesizer.md:20,176 stay with the epic #1891 follow-up.

Refs #1891

* chore(#3206): set changeset fragment pr to 3435

* fix(#3206): drop stale emitted-drift-ack entries that trip the ADR-2719 ratchet

The round's ack bookkeeping explained ripples that were already
self-attributed, so `tests/emitted-attribution.test.cjs` failed
deterministically on the PR head with 5 stale acknowledgments.

`agents/gsd-verifier.md` and `gsd-core/references/verifier-phase-gates.md`
appear directly in `git diff --name-only`, so PROVENANCE_RULES attributes
their emitted deltas without an ack; `agents/gsd-verifier.agent.md`,
`agents/gsd-verifier.toml` and `agents/subagents/gsd-verifier.md` are
derived emissions of a changed source and are attributed the same way.
None of the five entries could ever be consumed, so all five were stale.

Removed: the whole `3206-verifier-explicit-evidence.json` fragment (all
four entries) and the `#3206 append` to `0000-legacy-migration.json`.

Deliberately KEPT: the `#3206 append` to
`1955-verifier-coincidental-reliance.json`. Its `gsd-verifier.md` entry is
consumed by the size-growth ratchet, not the hash pass — the agent grew
49049 -> 49151 bytes, and `diffEmitted` treats a base-identical ack as
spent and excludes it from `ackEntries`. Reverting that append as well
turns the stale-ack failure into `1 file(s) grew without an
acknowledgment` (verified both ways locally).

* fix(#3206): compress 5b and re-acknowledge growth after rebase onto next

The rebase onto next (159145435, PR #3558) invalidated two things at once:

- #3558 grew agents/gsd-verifier.md to 49,098 bytes, leaving 54 bytes of
  LARGE-cap headroom where this PR's +102 no longer fits. Compressed the
  5b rewrite to +34 net by dropping the trailing honest-verifier cite —
  superseded by the now-inline definition; honest-verifier.md stays
  cited at the adjacent 5c line. File lands at 49,150 (2 under the cap).
- #3558 deleted tests/emitted-drift-acks/1955-verifier-coincidental-
  reliance.json, which carried this PR's growth acknowledgment, and its
  own 3409 fragment now names gsd-verifier.md. Re-homed the #3206 growth
  ack as an append to that entry (two ack sources may never name the
  same path).

The changeset is updated to match: two bare references/ cites repaired
(5c honest-verifier.md, MVP-mode verify-mvp-mode.md), the third (5b's)
superseded by the inline definition rather than repaired.

Reversion controls: restoring the uncompressed 5b fails
agent-size-budget.test.cjs (LARGE hard cap); reverting the ack append
fails emitted-attribution.test.cjs (differential attribution over the
real tree, GSD_EMITTED_BASE=upstream/next). Both re-verified green at
this tree: 213/213 (size + attribution), 992/992 across the 14 suites
reading the touched files.

Refs #3206

* test(#3206): pin the 5b explicit-evidence definition and cite resolution

* fix(#3206): pin regression tests to the shipped contract text

---------

Co-authored-by: sim <sim@local>
2026-08-16 14:14:27 -04:00

11 KiB

Verifier Phase Gates

Loaded eagerly by agents/gsd-verifier.md (<required_reading>). Carries the three verification-time gates that lived in the retired verify-phase workflow (#1892 / epic #1891 F7): decision-coverage validation (#2492), the test-quality audit, and infrastructure-phase human-verification scoping (#2504) — plus the backstop-abstention reporting contract (#3206). Run each gate at its named agent step; gsd_run is the launcher shim defined in the agent's own Step 1 block.

verify_decisions — Decision Coverage Gate (run after Step 6, requirements coverage)

**Decision coverage validation gate (issue #2492).**

After requirements coverage, also check that each trackable CONTEXT.md <decisions> entry shows up somewhere in the shipped artifacts (plans, SUMMARY.md, files modified by the phase, or recent commit subjects on the phase branch).

This gate is non-blocking / warning only by deliberate asymmetry with the plan-phase translation gate. The plan-phase gate already blocked at translation time, so by the time verification runs every decision has either been translated or explicitly deferred. This gate's job is to surface decisions that were translated but vanished during execution — that's a soft signal because "honors a decision" is a fuzzy substring heuristic, and we don't want a paraphrase miss to fail an otherwise good phase.

Skip if workflow.context_coverage_gate is explicitly set to false (absent key = enabled). Also skip cleanly when CONTEXT.md is missing or has no <decisions> block.

GATE_CFG=$(gsd_run query config-get workflow.context_coverage_gate 2>/dev/null || echo "true")
if [ "$GATE_CFG" != "false" ]; then
  CONTEXT_PATH=$(ls "${PHASE_DIR}"/*-CONTEXT.md 2>/dev/null | head -1)  # #2962: not a for-glob (zsh aborts)
  DECISION_RESULT=$(gsd_run query check.decision-coverage-verify "${PHASE_DIR}" "${CONTEXT_PATH}")
fi

The handler returns JSON { skipped, blocking: false, total, honored, not_honored: [...], message }.

Reporting: Append the handler's message (a ### Decision Coverage section) to VERIFICATION.md regardless of outcome — even when all decisions are honored, recording the count helps reviewers spot drift over time. Set decision_coverage in the verification result to {honored, total, not_honored: [...]} so downstream tooling can read it.

Status impact: none. The decision gate does NOT influence the gaps_found / human_needed / passed decision tree in Step 9. Its findings are warnings the user reviews and may act on by re-opening the phase or by acknowledging the decision was abandoned intentionally.

audit_test_quality (run after Step 7b, alongside anti-patterns)

**Verify that tests PROVE what they claim to prove.**

This step catches test-level deceptions that pass all prior checks: files exist, are substantive, are wired, and tests pass — but the tests don't actually validate the requirement.

1. Identify requirement-linked test files

From PLAN and SUMMARY files, map each requirement to the test files that are supposed to prove it.

2. Disabled test scan

For ALL test files linked to requirements, search for disabled/skipped patterns:

grep -rn -E "it\.skip|describe\.skip|test\.skip|xit\(|xdescribe\(|xtest\(|@pytest\.mark\.skip|@unittest\.skip|#\[ignore\]|\.pending|it\.todo|test\.todo" "$TEST_FILE"

Rule: A disabled test linked to a requirement = requirement NOT tested.

  • 🛑 BLOCKER if the disabled test is the only test proving that requirement
  • ⚠️ WARNING if other active tests also cover the requirement

3. Circular test detection

Search for scripts/utilities that generate expected values by running the system under test:

grep -rn -E "writeFileSync|writeFile|fs\.write|open\(.*w\)" "$TEST_DIRS"

For each match, check if it also imports the system/service/module being tested. If a script both imports the system-under-test AND writes expected output values → CIRCULAR.

Circular test indicators:

  • Script imports a service AND writes to fixture files
  • Expected values have comments like "computed from engine", "captured from baseline"
  • Script filename contains "capture", "baseline", "generate", "snapshot" in test context
  • Expected values were added in the same commit as the test assertions

Rule: A test comparing system output against values generated by the same system is circular. It proves consistency, not correctness.

4. Expected value provenance (for comparison/parity/migration requirements)

When a requirement demands comparison with an external source ("identical to X", "matches Y", "same output as Z"):

  • Is the external source actually invoked or referenced in the test pipeline?
  • Do fixture files contain data sourced from the external system?
  • Or do all expected values come from the new system itself or from mathematical formulas?

Provenance classification:

  • VALID: Expected value from external/legacy system output, manual capture, or independent oracle
  • PARTIAL: Expected value from mathematical derivation (proves formula, not system match)
  • CIRCULAR: Expected value from the system being tested
  • UNKNOWN: No provenance information — treat as SUSPECT

5. Assertion strength

For each test linked to a requirement, classify the strongest assertion:

Level Examples Proves
Existence toBeDefined(), != null Something returned
Type typeof x === 'number' Correct shape
Status code === 200 No error
Value toEqual(expected), toBeCloseTo(x) Specific value
Behavioral Multi-step workflow assertions End-to-end correctness

If a requirement demands value-level or behavioral-level proof and the test only has existence/type/status assertions → INSUFFICIENT.

6. Coverage quantity

If a requirement specifies a quantity of test cases (e.g., "30 calculations"), check if the actual number of active (non-skipped) test cases meets the requirement.

Reporting — add to VERIFICATION.md:

### Test Quality Audit

| Test File | Linked Req | Active | Skipped | Circular | Assertion Level | Verdict |
|-----------|-----------|--------|---------|----------|-----------------|---------|

**Disabled tests on requirements:** {N} → {BLOCKER if any req has ONLY disabled tests}
**Circular patterns detected:** {N} → {BLOCKER if any}
**Insufficient assertions:** {N} → {WARNING}

Impact on status: Any BLOCKER from test quality audit → overall status = gaps_found (Step 9 rule 1), regardless of other checks passing.

identify_human_verification — infrastructure/foundation scoping (apply at Step 8)

First: determine if this is an infrastructure/foundation phase.

Infrastructure and foundation phases — code foundations, database schema, internal APIs, data models, build tooling, CI/CD, internal service integrations — have no user-facing elements by definition. For these phases:

  • Do NOT invent artificial manual steps (e.g., "manually run git commits", "manually invoke methods", "manually check database state").
  • Mark human verification as N/A with rationale: "Infrastructure/foundation phase — no user-facing elements to test manually."
  • Set human_verification: [] and do not produce a human_needed status solely due to lack of user-facing features.
  • Only add human verification items if the phase goal or success criteria explicitly describe something a user would interact with (UI, CLI command output visible to end users, external service UX).
  • Exception — behavior-unverified truths still count. A truth marked ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (a state transition or a cancellation/cleanup/ordering invariant with no test exercising it) is a behavioral-evidence gap, not an artificial user-facing step. Record it in behavior_unverified_items and emit a human-verification item for it even on an infrastructure/foundation phase — these invariants are exactly where infra phases hide runtime state leaks. Such a truth drives human_needed; the auto-pass-UAT shortcut applies only to the absence of user-facing UX, never to a behavior-unverified invariant. The same carve-out covers an abstained non-inferable truth (⚠️ insufficient_spec, § Backstop abstention below) — an insufficient-spec gap is an evidence gap, not a user-facing step, so it too still emits its human-verification item and drives human_needed on an infrastructure phase.

How to determine if a phase is infrastructure/foundation:

  • Phase goal or name contains: "foundation", "infrastructure", "schema", "database", "internal API", "data model", "scaffolding", "pipeline", "tooling", "CI", "migrations", "service layer", "backend", "core library"
  • Phase success criteria describe only technical artifacts (files exist, tests pass, schema is valid) with no user interaction required
  • There is no UI, CLI output visible to end users, or real-time behavior to observe

If the phase IS infrastructure/foundation: auto-pass UAT — skip the human verification items list entirely, except any ⚠️ PRESENT_BEHAVIOR_UNVERIFIED or abstained ⚠️ insufficient_spec truth (see exception above), which still emits a human-verification item and drives human_needed. Only when no such excepted truth exists, log:

## Human Verification

N/A — Infrastructure/foundation phase with no user-facing elements.
All acceptance criteria are verifiable programmatically.

If the phase IS user-facing: only flag items that genuinely require a human — per the Step 8 always/uncertain lists already in the agent. Do not invent steps.

Backstop abstention — reporting contract (#3206, companion to agent Step 3 item 5b)

When a non-inferable (verification: backstop) truth abstains for lack of explicit evidence:

  • Never silent, never a hard halt. Interactive: the abstained item routes to the end-of-phase human checkpoint. Autonomous (AFK): it produces a prominent unverified — held-out test recommended flag and the completion line reads "complete with N unverified non-inferable checks"; the run neither silently passes the blind spot nor hard-halts.
  • Distinguishable reason. The abstain disposition carries reason: insufficient_spec so its human_needed outcome is never conflated with an ordinary manual-UAT human_needed.
  • Infrastructure phases included. This rides the same carve-out as ⚠️ PRESENT_BEHAVIOR_UNVERIFIED in the infrastructure-phase gate above: an abstention is an evidence gap, not a user-facing step, so the infra auto-pass-UAT shortcut never absorbs it.

Full protocol and rationale: gsd-core/references/honest-verifier.md.

Lazy references

  • Per-stack verification patterns: before Step 4 (artifact verification) on an unfamiliar stack, Read ~/.claude/gsd-core/references/verification-patterns.md — the grep catalog for React/Next.js components, API routes, database schema, and the universal stub patterns. Read it lazily (only the sections for the stack under verification); it is too large to load wholesale on every run.
  • Canonical report shape: the emitted VERIFICATION.md follows @~/.claude/gsd-core/templates/verification-report.md — the template whose Guidelines and row shapes src/uat.cts treats as canonical when consuming verification output.