* fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites
Step 3 item 5b abstained on non-inferable (backstop) truths "unless
confirmed by explicit evidence" with the term undefined — its definition
lived only in the non-included gsd-core/references/honest-verifier.md,
behind a stale bare `references/` cite that 404s. Undefined, the term
falls back to presence + wiring, the exact false-pass the #1154
abstention protocol refuses.
- 5b: inline the compressed definition (a passing wired
held-out/property-based test or directly observed behavior; presence +
wiring never qualifies) and fix the cite. +84 B on the rewritten line;
file lands at 49,151 of the 49,152 LARGE cap.
- verifier-phase-gates.md (already <required_reading>): gains the
backstop-abstention reporting contract — AFK completion line
("complete with N unverified non-inferable checks", never silent,
never a halt) and reason-distinctness (insufficient_spec vs manual-UAT
human_needed). New content, no relocation of measured prose.
- 5c (line 204) and MVP-mode (line 644) bare cites repaired to
gsd-core/references/ (+9 B each).
- Drift acks per ADR-2719 §4; two entries merge-appended into existing
fragments (two ack sources may never name the same path).
- Changeset fragment with the sanctioned pr: 0 placeholder (post-create
backfill).
Sibling census at next@7976b1ca0: 7 bare-cite instances in 4 agent
files; the 3 in gsd-verifier.md are fixed here, gsd-executor.md:429,439
and gsd-doc-synthesizer.md:20,176 stay with the epic #1891 follow-up.
Refs #1891
* chore(#3206): set changeset fragment pr to 3435
* fix(#3206): drop stale emitted-drift-ack entries that trip the ADR-2719 ratchet
The round's ack bookkeeping explained ripples that were already
self-attributed, so `tests/emitted-attribution.test.cjs` failed
deterministically on the PR head with 5 stale acknowledgments.
`agents/gsd-verifier.md` and `gsd-core/references/verifier-phase-gates.md`
appear directly in `git diff --name-only`, so PROVENANCE_RULES attributes
their emitted deltas without an ack; `agents/gsd-verifier.agent.md`,
`agents/gsd-verifier.toml` and `agents/subagents/gsd-verifier.md` are
derived emissions of a changed source and are attributed the same way.
None of the five entries could ever be consumed, so all five were stale.
Removed: the whole `3206-verifier-explicit-evidence.json` fragment (all
four entries) and the `#3206 append` to `0000-legacy-migration.json`.
Deliberately KEPT: the `#3206 append` to
`1955-verifier-coincidental-reliance.json`. Its `gsd-verifier.md` entry is
consumed by the size-growth ratchet, not the hash pass — the agent grew
49049 -> 49151 bytes, and `diffEmitted` treats a base-identical ack as
spent and excludes it from `ackEntries`. Reverting that append as well
turns the stale-ack failure into `1 file(s) grew without an
acknowledgment` (verified both ways locally).
* fix(#3206): compress 5b and re-acknowledge growth after rebase onto next
The rebase onto next (159145435, PR #3558) invalidated two things at once:
- #3558 grew agents/gsd-verifier.md to 49,098 bytes, leaving 54 bytes of
LARGE-cap headroom where this PR's +102 no longer fits. Compressed the
5b rewrite to +34 net by dropping the trailing honest-verifier cite —
superseded by the now-inline definition; honest-verifier.md stays
cited at the adjacent 5c line. File lands at 49,150 (2 under the cap).
- #3558 deleted tests/emitted-drift-acks/1955-verifier-coincidental-
reliance.json, which carried this PR's growth acknowledgment, and its
own 3409 fragment now names gsd-verifier.md. Re-homed the #3206 growth
ack as an append to that entry (two ack sources may never name the
same path).
The changeset is updated to match: two bare references/ cites repaired
(5c honest-verifier.md, MVP-mode verify-mvp-mode.md), the third (5b's)
superseded by the inline definition rather than repaired.
Reversion controls: restoring the uncompressed 5b fails
agent-size-budget.test.cjs (LARGE hard cap); reverting the ack append
fails emitted-attribution.test.cjs (differential attribution over the
real tree, GSD_EMITTED_BASE=upstream/next). Both re-verified green at
this tree: 213/213 (size + attribution), 992/992 across the 14 suites
reading the touched files.
Refs #3206
* test(#3206): pin the 5b explicit-evidence definition and cite resolution
* fix(#3206): pin regression tests to the shipped contract text
---------
Co-authored-by: sim <sim@local>
11 KiB
Verifier Phase Gates
Loaded eagerly by
agents/gsd-verifier.md(<required_reading>). Carries the three verification-time gates that lived in the retiredverify-phaseworkflow (#1892 / epic #1891 F7): decision-coverage validation (#2492), the test-quality audit, and infrastructure-phase human-verification scoping (#2504) — plus the backstop-abstention reporting contract (#3206). Run each gate at its named agent step;gsd_runis the launcher shim defined in the agent's own Step 1 block.
verify_decisions — Decision Coverage Gate (run after Step 6, requirements coverage)
**Decision coverage validation gate (issue #2492).**After requirements coverage, also check that each trackable CONTEXT.md
<decisions> entry shows up somewhere in the shipped artifacts (plans,
SUMMARY.md, files modified by the phase, or recent commit subjects on the
phase branch).
This gate is non-blocking / warning only by deliberate asymmetry with the plan-phase translation gate. The plan-phase gate already blocked at translation time, so by the time verification runs every decision has either been translated or explicitly deferred. This gate's job is to surface decisions that were translated but vanished during execution — that's a soft signal because "honors a decision" is a fuzzy substring heuristic, and we don't want a paraphrase miss to fail an otherwise good phase.
Skip if workflow.context_coverage_gate is explicitly set to false
(absent key = enabled). Also skip cleanly when CONTEXT.md is missing or has
no <decisions> block.
GATE_CFG=$(gsd_run query config-get workflow.context_coverage_gate 2>/dev/null || echo "true")
if [ "$GATE_CFG" != "false" ]; then
CONTEXT_PATH=$(ls "${PHASE_DIR}"/*-CONTEXT.md 2>/dev/null | head -1) # #2962: not a for-glob (zsh aborts)
DECISION_RESULT=$(gsd_run query check.decision-coverage-verify "${PHASE_DIR}" "${CONTEXT_PATH}")
fi
The handler returns JSON { skipped, blocking: false, total, honored, not_honored: [...], message }.
Reporting: Append the handler's message (a ### Decision Coverage
section) to VERIFICATION.md regardless of outcome — even when all
decisions are honored, recording the count helps reviewers spot drift over
time. Set decision_coverage in the verification result to
{honored, total, not_honored: [...]} so downstream tooling can read it.
Status impact: none. The decision gate does NOT influence the
gaps_found / human_needed / passed decision tree in Step 9. Its
findings are warnings the user reviews and may act on by re-opening the
phase or by acknowledging the decision was abandoned intentionally.
audit_test_quality (run after Step 7b, alongside anti-patterns)
**Verify that tests PROVE what they claim to prove.**This step catches test-level deceptions that pass all prior checks: files exist, are substantive, are wired, and tests pass — but the tests don't actually validate the requirement.
1. Identify requirement-linked test files
From PLAN and SUMMARY files, map each requirement to the test files that are supposed to prove it.
2. Disabled test scan
For ALL test files linked to requirements, search for disabled/skipped patterns:
grep -rn -E "it\.skip|describe\.skip|test\.skip|xit\(|xdescribe\(|xtest\(|@pytest\.mark\.skip|@unittest\.skip|#\[ignore\]|\.pending|it\.todo|test\.todo" "$TEST_FILE"
Rule: A disabled test linked to a requirement = requirement NOT tested.
- 🛑 BLOCKER if the disabled test is the only test proving that requirement
- ⚠️ WARNING if other active tests also cover the requirement
3. Circular test detection
Search for scripts/utilities that generate expected values by running the system under test:
grep -rn -E "writeFileSync|writeFile|fs\.write|open\(.*w\)" "$TEST_DIRS"
For each match, check if it also imports the system/service/module being tested. If a script both imports the system-under-test AND writes expected output values → CIRCULAR.
Circular test indicators:
- Script imports a service AND writes to fixture files
- Expected values have comments like "computed from engine", "captured from baseline"
- Script filename contains "capture", "baseline", "generate", "snapshot" in test context
- Expected values were added in the same commit as the test assertions
Rule: A test comparing system output against values generated by the same system is circular. It proves consistency, not correctness.
4. Expected value provenance (for comparison/parity/migration requirements)
When a requirement demands comparison with an external source ("identical to X", "matches Y", "same output as Z"):
- Is the external source actually invoked or referenced in the test pipeline?
- Do fixture files contain data sourced from the external system?
- Or do all expected values come from the new system itself or from mathematical formulas?
Provenance classification:
- VALID: Expected value from external/legacy system output, manual capture, or independent oracle
- PARTIAL: Expected value from mathematical derivation (proves formula, not system match)
- CIRCULAR: Expected value from the system being tested
- UNKNOWN: No provenance information — treat as SUSPECT
5. Assertion strength
For each test linked to a requirement, classify the strongest assertion:
| Level | Examples | Proves |
|---|---|---|
| Existence | toBeDefined(), != null |
Something returned |
| Type | typeof x === 'number' |
Correct shape |
| Status | code === 200 |
No error |
| Value | toEqual(expected), toBeCloseTo(x) |
Specific value |
| Behavioral | Multi-step workflow assertions | End-to-end correctness |
If a requirement demands value-level or behavioral-level proof and the test only has existence/type/status assertions → INSUFFICIENT.
6. Coverage quantity
If a requirement specifies a quantity of test cases (e.g., "30 calculations"), check if the actual number of active (non-skipped) test cases meets the requirement.
Reporting — add to VERIFICATION.md:
### Test Quality Audit
| Test File | Linked Req | Active | Skipped | Circular | Assertion Level | Verdict |
|-----------|-----------|--------|---------|----------|-----------------|---------|
**Disabled tests on requirements:** {N} → {BLOCKER if any req has ONLY disabled tests}
**Circular patterns detected:** {N} → {BLOCKER if any}
**Insufficient assertions:** {N} → {WARNING}
Impact on status: Any BLOCKER from test quality audit → overall status = gaps_found (Step 9 rule 1), regardless of other checks passing.
identify_human_verification — infrastructure/foundation scoping (apply at Step 8)
First: determine if this is an infrastructure/foundation phase.
Infrastructure and foundation phases — code foundations, database schema, internal APIs, data models, build tooling, CI/CD, internal service integrations — have no user-facing elements by definition. For these phases:
- Do NOT invent artificial manual steps (e.g., "manually run git commits", "manually invoke methods", "manually check database state").
- Mark human verification as N/A with rationale: "Infrastructure/foundation phase — no user-facing elements to test manually."
- Set
human_verification: []and do not produce ahuman_neededstatus solely due to lack of user-facing features. - Only add human verification items if the phase goal or success criteria explicitly describe something a user would interact with (UI, CLI command output visible to end users, external service UX).
- Exception — behavior-unverified truths still count. A truth marked ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (a state transition or a cancellation/cleanup/ordering invariant with no test exercising it) is a behavioral-evidence gap, not an artificial user-facing step. Record it in
behavior_unverified_itemsand emit a human-verification item for it even on an infrastructure/foundation phase — these invariants are exactly where infra phases hide runtime state leaks. Such a truth driveshuman_needed; the auto-pass-UAT shortcut applies only to the absence of user-facing UX, never to a behavior-unverified invariant. The same carve-out covers an abstained non-inferable truth (⚠️insufficient_spec, § Backstop abstention below) — an insufficient-spec gap is an evidence gap, not a user-facing step, so it too still emits its human-verification item and driveshuman_neededon an infrastructure phase.
How to determine if a phase is infrastructure/foundation:
- Phase goal or name contains: "foundation", "infrastructure", "schema", "database", "internal API", "data model", "scaffolding", "pipeline", "tooling", "CI", "migrations", "service layer", "backend", "core library"
- Phase success criteria describe only technical artifacts (files exist, tests pass, schema is valid) with no user interaction required
- There is no UI, CLI output visible to end users, or real-time behavior to observe
If the phase IS infrastructure/foundation: auto-pass UAT — skip the human verification items list entirely, except any ⚠️ PRESENT_BEHAVIOR_UNVERIFIED or abstained ⚠️ insufficient_spec truth (see exception above), which still emits a human-verification item and drives human_needed. Only when no such excepted truth exists, log:
## Human Verification
N/A — Infrastructure/foundation phase with no user-facing elements.
All acceptance criteria are verifiable programmatically.
If the phase IS user-facing: only flag items that genuinely require a human — per the Step 8 always/uncertain lists already in the agent. Do not invent steps.
Backstop abstention — reporting contract (#3206, companion to agent Step 3 item 5b)
When a non-inferable (verification: backstop) truth abstains for lack of explicit evidence:
- Never silent, never a hard halt. Interactive: the abstained item routes to the end-of-phase
human checkpoint. Autonomous (AFK): it produces a prominent
unverified — held-out test recommendedflag and the completion line reads "complete with N unverified non-inferable checks"; the run neither silently passes the blind spot nor hard-halts. - Distinguishable reason. The abstain disposition carries
reason: insufficient_specso itshuman_neededoutcome is never conflated with an ordinary manual-UAThuman_needed. - Infrastructure phases included. This rides the same carve-out as ⚠️ PRESENT_BEHAVIOR_UNVERIFIED in the infrastructure-phase gate above: an abstention is an evidence gap, not a user-facing step, so the infra auto-pass-UAT shortcut never absorbs it.
Full protocol and rationale: gsd-core/references/honest-verifier.md.
Lazy references
- Per-stack verification patterns: before Step 4 (artifact verification) on an unfamiliar stack, Read
~/.claude/gsd-core/references/verification-patterns.md— the grep catalog for React/Next.js components, API routes, database schema, and the universal stub patterns. Read it lazily (only the sections for the stack under verification); it is too large to load wholesale on every run. - Canonical report shape: the emitted VERIFICATION.md follows
@~/.claude/gsd-core/templates/verification-report.md— the template whose Guidelines and row shapessrc/uat.ctstreats as canonical when consuming verification output.