diff --git a/.changeset/merry-badgers-tumble.md b/.changeset/merry-badgers-tumble.md new file mode 100644 index 000000000..ec87944c6 --- /dev/null +++ b/.changeset/merry-badgers-tumble.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 3250 +--- +**`gsd-verifier` now says *why* a verified truth holds, not just that it does** — a truth that reaches `✓ VERIFIED` is additionally classified against three incidental-reliance patterns (an undeclared precondition, an ordering or side effect nothing enforces, a truth that is only true under the test fixture) and, when one matches, is reported as `✓ VERIFIED (coincidental-reliance)` with an entry in the new `coincidental_reliance_items` frontmatter list naming what to harden. Purely advisory: the base `✓ VERIFIED` token is unchanged, the truth still counts toward the score, the overall `status` is unaffected, and no human-verification item is emitted — a passing phase still passes. Only a consumer matching the truth-row verdict cell for exact equality (rather than as a substring) needs to tolerate the suffix. Two limits stated up front: the check is endogenous, and so measurably weaker than the exogenous `backstop` tag `gsd-core/references/honest-verifier.md` routes on — advisory status is the consequence, and its precision is unmeasured; and `gsd-core/workflows/verify-phase.md` is not edited, receiving the rule through its eager import of the verification-report template rather than a second inline copy, because it sits 29 bytes under its size hard cap. (#1955) diff --git a/agents/gsd-verifier.md b/agents/gsd-verifier.md index 6894aba9a..4020dda21 100644 --- a/agents/gsd-verifier.md +++ b/agents/gsd-verifier.md @@ -200,6 +200,7 @@ For each truth: - No such test exists, or it can't run without a server/state mutation → ⚠️ PRESENT_BEHAVIOR_UNVERIFIED. Emit a human-verification item (Step 8) and do not count it toward the verified score (Step 9). - An accepted override (Step 3b) carries the truth as PASSED (override), exactly as it does for a FAILED truth. 5b. **Non-inferable (`backstop`) truths:** a `verification: backstop` truth (via `truthVerification()`) abstains unless confirmed by explicit evidence — mark `insufficient_spec` -> a human-verification item -> `human_needed`. See `references/honest-verifier.md`. +5c. **Reliance check (advisory, #1955).** Before finalizing a ✓ VERIFIED truth, ask *why* it holds. Classify the evidence already recorded, not your confidence in it. Endogenous, and so weaker than the exogenous `backstop` tag (`references/honest-verifier.md`) — advisory for exactly that reason. Flag `coincidental-reliance` when the evidence names one of: **undeclared-precondition** (state nothing in the phase's artifacts or a declared prerequisite guarantees), **incidental-ordering** (an order or side effect nothing in the code enforces), **fixture-only** (the test's own setup establishes the precondition; the production path has no equivalent). **Do NOT flag:** a precondition the code establishes or explicitly defaults; ordering the code enforces (await, explicit sequencing); a fixture merely supplying input the real caller also supplies; unease naming no specific state, ordering, or fixture. Out of scope: ⚠️ PRESENT_BEHAVIOR_UNVERIFIED and ⚠️ `insufficient_spec` (already routed to human), and PASSED (override) truths. Record `✓ VERIFIED (coincidental-reliance)` and add a `coincidental_reliance_items` entry. **Advisory only — not the score, not the status, and never a human-verification item** (Step 9 rule 2 would flip a passing phase to `human_needed`). The usual fix: promote the hidden assumption into a declared precondition. 6. Determine truth status ## Step 3b: Check Verification Overrides @@ -554,6 +555,7 @@ Classify status using this decision tree IN ORDER (most restrictive first): - `verified_truths` counts ✓ VERIFIED truths plus PASSED (override) truths (Step 3b). For a behavior-dependent truth, VERIFIED means a behavioral test passed, not just that symbols are present. - ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths are the *only* ones excluded from `verified_truths`; they are reported separately as `behavior_unverified`. +- `✓ VERIFIED (coincidental-reliance)` counts as VERIFIED — the advisory changes no score and no status. ```text score: verified_truths / total_truths # e.g. 6/7 @@ -701,6 +703,10 @@ behavior_unverified_items: # Only if behavior_unverified > 0 — emitted regardl test: "What to trigger" expected: "What state must hold afterward" why_human: "Why presence checks can't see it" +coincidental_reliance_items: # Only if a ✓ VERIFIED truth holds incidentally — emitted regardless of overall status (survives gaps_found) + - truth: "Observable truth that holds incidentally" + reason: undeclared-precondition | incidental-ordering | fixture-only + harden: "Precondition/ordering to declare or enforce" human_verification: # Only if status: human_needed - test: "What to do" expected: "What should happen" @@ -723,6 +729,7 @@ human_verification: # Only if status: human_needed | 1 | {truth} | ✓ VERIFIED | {evidence} | | 2 | {truth} | ✗ FAILED | {what's wrong} | | 3 | {truth} | ⚠️ PRESENT_BEHAVIOR_UNVERIFIED | {present + wired; no test exercises the transition/invariant — see Human Verification} | +| 4 | {truth} | ✓ VERIFIED (coincidental-reliance) | {holds, but incidentally — see coincidental_reliance_items} | **Score:** {N}/{M} truths verified ({P} present, behavior-unverified) diff --git a/docs/AGENTS.md b/docs/AGENTS.md index 62b30ee17..ff51f0855 100644 --- a/docs/AGENTS.md +++ b/docs/AGENTS.md @@ -319,6 +319,9 @@ Two further dimensions carry no number: **Verify Command Format Sanity** and - **Test quality audit** (v1.32): verifies that tests prove what they claim by checking for disabled/skipped tests on requirements, circular test patterns (system generating its own expected values), assertion strength (existence vs. value vs. behavioral), and expected value provenance. Blockers from test quality audit override an otherwise passing verification - Runs the full workspace test suite at most once per verification — proves a test *exists* by enumeration and that it *passes* via a single named test, never re-running the whole suite per must-have. - **Behavior-dependent calibration (#966):** a must-have that asserts a state transition or a cancellation/cleanup/ordering invariant is marked `⚠️ PRESENT_BEHAVIOR_UNVERIFIED` (not `VERIFIED`) when no test exercises it — excluded from the `verified_truths` score, counted in the `behavior_unverified` frontmatter field, and routed to human verification, so a clean `N/N` certifies behavioral evidence rather than mere symbol presence. +- **Coincidental-reliance advisory (#1955):** a truth that reaches `✓ VERIFIED` is additionally asked *why* it holds. When the recorded evidence shows the truth holding for an incidental reason — `undeclared-precondition`, `incidental-ordering`, or `fixture-only` — the verdict is qualified as `✓ VERIFIED (coincidental-reliance)` and the truth is listed in the `coincidental_reliance_items` frontmatter field with what to harden. This is **advisory**: the base `✓ VERIFIED` token is unchanged, the truth still counts toward `verified_truths`, the overall `status` is unaffected, and no human-verification item is emitted — a passing phase still passes. It classifies evidence the verifier already gathered rather than asking it to rate its own confidence — but it is honestly an **endogenous** check, and `gsd-core/references/honest-verifier.md` records that endogenous gates are measurably weaker than the exogenous `backstop` tag it routes on. Advisory status is the consequence, not a coincidence: a miss costs exactly today's behaviour (a plain `✓ VERIFIED`) and a false positive costs one line of prose, never a failed phase, so a weaker mechanism is affordable here in a way it would not be on a pass/fail axis. Its precision is unmeasured. It complements the two existing axes: `PRESENT_BEHAVIOR_UNVERIFIED` is *no* behavioral evidence, `insufficient_spec` is an under-specified truth, and this is evidence that exists and passes for the wrong reason. + + The advisory is carried by two surfaces. `agents/gsd-verifier.md` (Step 3, sub-step 5c) holds the detection rule for the subagent path. The non-subagent path, `gsd-core/workflows/verify-phase.md`, receives it through the eager `@`-import of `gsd-core/templates/verification-report.md`, whose `## Guidelines` carry the same instruction — `verify-phase.md` itself is deliberately **not** edited, because it sits 29 bytes under the DEFAULT tier hard cap in `tests/workflow-size-budget.test.cjs` and absorbing rubric prose there requires a lazy extraction first. --- diff --git a/docs/reference/planning-artifacts.md b/docs/reference/planning-artifacts.md index 3034f6c6d..fe97ee046 100644 --- a/docs/reference/planning-artifacts.md +++ b/docs/reference/planning-artifacts.md @@ -204,7 +204,7 @@ actuals: | | | |---|---| -| **Purpose** | Phase goal verification report. Checks `must_haves.truths`, `must_haves.artifacts`, and `must_haves.key_links` from all plans against the actual codebase after execution. Records `status: passed \| gaps_found \| human_needed`. A truth whose correctness depends on runtime behaviour — a state transition or a cancellation/cleanup/ordering invariant — is marked `⚠️ PRESENT_BEHAVIOR_UNVERIFIED` (not `VERIFIED`) when no test exercises it: it is excluded from `score`, counted in the `behavior_unverified` frontmatter field, and routed to `human_needed`, so a behaviour-dependent gap can no longer count toward a clean N/N. | +| **Purpose** | Phase goal verification report. Checks `must_haves.truths`, `must_haves.artifacts`, and `must_haves.key_links` from all plans against the actual codebase after execution. Records `status: passed \| gaps_found \| human_needed`. A truth whose correctness depends on runtime behaviour — a state transition or a cancellation/cleanup/ordering invariant — is marked `⚠️ PRESENT_BEHAVIOR_UNVERIFIED` (not `VERIFIED`) when no test exercises it: it is excluded from `score`, counted in the `behavior_unverified` frontmatter field, and routed to `human_needed`, so a behaviour-dependent gap can no longer count toward a clean N/N. A truth that *does* reach `✓ VERIFIED` but holds for an incidental reason (`undeclared-precondition`, `incidental-ordering`, `fixture-only`) is qualified `✓ VERIFIED (coincidental-reliance)` and listed in `coincidental_reliance_items` — **advisory only**: it still counts toward `score`, does not change `status`, and emits no human-verification item (#1955). | | **Produced by** | `/gsd-verify-work` (or the verify step within `/gsd-execute-phase`). | | **Consumed by** | `plan-phase` closed-phase gate (a `status: passed` VERIFICATION.md marks the phase `Complete` and blocks replanning without `--force`); `/gsd-progress`; human review. | diff --git a/gsd-core/templates/verification-report.md b/gsd-core/templates/verification-report.md index 14982d45c..ba4f725c7 100644 --- a/gsd-core/templates/verification-report.md +++ b/gsd-core/templates/verification-report.md @@ -18,6 +18,10 @@ behavior_unverified_items: # Only if behavior_unverified > 0 — the truths abov test: "What to trigger" expected: "What state must hold afterward" why_human: "Why presence checks can't see it" +coincidental_reliance_items: # Only if a ✓ VERIFIED truth holds incidentally — emitted regardless of overall status (survives gaps_found) + - truth: "Observable truth that holds incidentally" + reason: undeclared-precondition | incidental-ordering | fixture-only + harden: "Precondition/ordering to declare or enforce" --- # Phase {X}: {Name} Verification Report @@ -35,7 +39,8 @@ behavior_unverified_items: # Only if behavior_unverified > 0 — the truths abov | 1 | {truth from must_haves} | ✓ VERIFIED | {what confirmed it} | | 2 | {truth from must_haves} | ✗ FAILED | {what's wrong} | | 3 | {truth from must_haves} | ⚠️ PRESENT_BEHAVIOR_UNVERIFIED | {present + wired; transition/invariant not exercised by a test — see Human Verification} | -| 4 | {truth from must_haves} | ? UNCERTAIN | {why can't verify} | +| 4 | {truth from must_haves} | ✓ VERIFIED (coincidental-reliance) | {holds, but incidentally — see coincidental_reliance_items} | +| 5 | {truth from must_haves} | ? UNCERTAIN | {why can't verify} | **Score:** {N}/{M} truths verified ({P} present, behavior-unverified) @@ -176,6 +181,9 @@ None — all verifiable items checked programmatically. **Per-truth states (Observable Truths `Status` column):** - `✓ VERIFIED` — supporting artifacts pass all checks; for a behavior-dependent truth, a behavioral test exercised the asserted behavior - `⚠️ PRESENT_BEHAVIOR_UNVERIFIED` — present + wired, but a state transition or cancellation/cleanup/ordering invariant was not exercised by any test. Counts toward `behavior_unverified`, routes to human verification, and is *excluded* from the verified score. Per-truth only — on its own the overall `status:` becomes `human_needed` (unless a higher-precedence `gaps_found` also applies); the item is preserved in `behavior_unverified_items` regardless. +- `✓ VERIFIED (coincidental-reliance)` — an **advisory** qualifier on a truth that *is* verified but holds for an incidental reason rather than a guaranteed one (#1955): `undeclared-precondition` (state nothing in the phase's artifacts or a declared prerequisite guarantees), `incidental-ordering` (an order or side effect nothing in the code enforces), or `fixture-only` (the test's own setup establishes the precondition; the production path has no equivalent). The base `✓ VERIFIED` token is kept verbatim and leading, so it counts toward the verified score exactly as before — the advisory changes no score and no status, and never produces a human-verification item. Each flagged truth is listed in `coincidental_reliance_items` with the reason and what to harden. Not applied to a truth that never reached `✓ VERIFIED`, nor to a `PASSED (override)` truth. + + **Filling this column — apply the reliance check to every `✓ VERIFIED` truth before writing the row.** Ask why the truth holds and classify the evidence you already recorded, not your confidence in it. Flag it when the evidence names one of the three reasons above. Do NOT flag: a precondition the code establishes or explicitly defaults; ordering the code enforces (await, explicit sequencing); a fixture merely supplying input the real caller also supplies; unease naming no specific state, ordering, or fixture. The check is endogenous and so weaker than an exogenous tag (`references/honest-verifier.md`) — which is why it is advisory and never a gate. The usual fix is to promote the hidden assumption into a declared precondition. - `✗ FAILED` — artifact missing, stub, or unwired - `? UNCERTAIN` — can't verify programmatically diff --git a/tests/emitted-drift-acks/1955-verifier-coincidental-reliance.json b/tests/emitted-drift-acks/1955-verifier-coincidental-reliance.json new file mode 100644 index 000000000..d0ce65e29 --- /dev/null +++ b/tests/emitted-drift-acks/1955-verifier-coincidental-reliance.json @@ -0,0 +1,6 @@ +{ + "version": 1, + "paths": { + "gsd-verifier.md": "#1955: Step 3 gains sub-step 5c (the coincidental-reliance advisory), one Step 9 score bullet, one Observable Truths example row, and the coincidental_reliance_items frontmatter block. Growth is those four additions only — no existing text was rewritten. Deliberately dense rather than extracted: the issue's approved scope says 'No new files', and ADR-1610 Decision 4 / tests/workflow-size-budget.test.cjs:32-40 name eager @-import relocation as gaming the size proxy (it shrinks the measured file while total loaded context is unchanged or larger); this agent has no lazy-read seam. After the change the file sits at 48994 bytes against the LARGE tier hard cap of 49152 (tests/agent-size-budget.test.cjs), i.e. 158 bytes of headroom. That is deliberate and disclosed: the cap is not crossed and is not raised, but the next contributor who needs room in this agent must do a lazy extraction rather than add prose. Content justification: the verifier's goal-backward pass grades THAT a truth holds and never WHY, so a truth can read VERIFIED while resting on a coincidence — a precondition nothing guarantees, an ordering nothing enforces, or a fixture-only truth. 5c classifies the evidence already recorded rather than asking for a confidence rating, but it is honestly an endogenous check and so weaker than the exogenous backstop tag gsd-core/references/honest-verifier.md routes on — which is exactly why it is advisory only: it changes no score, no status, and emits no human-verification item, so a passing phase still passes. gsd-core/workflows/verify-phase.md is deliberately NOT edited (it has 29 bytes under its DEFAULT tier cap); it receives the rule through its existing eager @-import of gsd-core/templates/verification-report.md, whose Guidelines now carry the imperative check rather than only the row shape. tests/verifier-coincidental-reliance.test.cjs asserts both halves of that route and locks out a silent second copy appearing in the workflow." + } +} diff --git a/tests/verifier-coincidental-reliance.test.cjs b/tests/verifier-coincidental-reliance.test.cjs new file mode 100644 index 000000000..821ddbe1d --- /dev/null +++ b/tests/verifier-coincidental-reliance.test.cjs @@ -0,0 +1,288 @@ +// allow-test-rule: source-text-is-the-product (see #1955) +// allow-test-rule: docs-parity (see #1955) +// Two categories from CONTRIBUTING.md's exception matrix, deliberately both: +// source-text-is-the-product — agents/gsd-verifier.md and +// gsd-core/templates/verification-report.md are shipped .md whose text IS +// what the runtime loads, so asserting on the text asserts the deployed +// contract. Same category tests/verifier-behavior-unverified.test.cjs uses +// for the sibling #966 axis. +// docs-parity — the single docs/AGENTS.md assertion is not runtime-loaded +// text; it holds human documentation in sync with the shipped contract, +// which has no runtime enumeration API. Claiming the first category for it +// would overstate what that category covers. + +'use strict'; + +/** + * Issue #1955 — the verifier grades THAT a truth holds, never WHY. + * + * This suite locks the `coincidental-reliance` advisory: for a truth that + * already reached ✓ VERIFIED, the verifier additionally classifies whether it + * holds for a guaranteed reason or an incidental one, and names the incidental + * ones so they can be hardened instead of shipping invisibly. + * + * The load-bearing assertions are the INVARIANTS, not the feature: + * - the advisory changes neither the verified score nor the overall status; + * - it never emits a human-verification item (Step 9 rule 2 would escalate a + * passing phase to `human_needed`, contradicting "the phase can still pass"); + * - ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (#966) stays the ONLY truth state excluded + * from `verified_truths`; + * - the per-truth token never enters the overall-status vocabulary; + * - the qualifier SUFFIXES `✓ VERIFIED` rather than replacing it, so every + * substring matcher on the old verdict still hits (Hyrum's Law). + * + * Deliberately absent: any assertion on the model's verdict. ADR-550 Decision 5 + * rejects that as vacuous — the CI-testable surface is the deterministic + * contract, which for a prose rubric is the shipped text. + */ + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const ROOT = path.join(__dirname, '..'); + +// Canonical sources of truth (CONTRIBUTING.md "Source of truth for agents") — +// never .claude/agents/, which is a gitignored install-sync output and would +// false-pass against a stale local sync. +const read = (...p) => fs.readFileSync(path.join(ROOT, ...p), 'utf-8'); + +const verifier = read('agents', 'gsd-verifier.md'); +const template = read('gsd-core', 'templates', 'verification-report.md'); +const agentsDoc = read('docs', 'AGENTS.md'); +const verifyPhase = read('gsd-core', 'workflows', 'verify-phase.md'); + +const QUALIFIER = '✓ VERIFIED (coincidental-reliance)'; +const REASONS = ['undeclared-precondition', 'incidental-ordering', 'fixture-only']; + +/** + * Return a window of `source` centred on the first occurrence of `needle`, + * extending `radius` characters either side. Used to assert that a clause sits + * WITH the rule rather than anywhere in a 47 KB file — a bare + * `assert.match(verifier, /score/)` would pass on unrelated prose. + */ +function sectionAround(source, needle, radius) { + const at = source.indexOf(needle); + assert.notEqual(at, -1, `expected to find '${needle}' in the source`); + return source.slice(Math.max(0, at - radius), at + needle.length + radius); +} + +describe('#1955: coincidental-reliance advisory — the rule', () => { + test('defines the coincidental-reliance advisory', () => { + assert.match(verifier, /coincidental-reliance/); + }); + + test('names the three reliance triggers', () => { + for (const reason of REASONS) { + assert.ok( + verifier.includes(reason), + `verifier must name the '${reason}' trigger`, + ); + } + }); + + test('advisory applies only to truths that reached VERIFIED', () => { + // The advisory explains why a PASSING truth passes. A truth that never + // reached ✓ VERIFIED has nothing to explain. Asserted on the window around + // the rule rather than a forward-only regex: the scoping clause precedes + // the first `coincidental-reliance` token, so `token[\s\S]{0,N}?✓ VERIFIED` + // structurally cannot see it. + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /✓ VERIFIED truth/); + }); + + test('does not double-report truths already routed to human verification', () => { + // PRESENT_BEHAVIOR_UNVERIFIED (#966) and insufficient_spec (honest-verifier) + // already route to human_needed; re-flagging them is noise. + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /PRESENT_BEHAVIOR_UNVERIFIED/); + assert.match(window, /insufficient_spec/); + }); + + test('an accepted override is not re-flagged', () => { + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /override/i); + }); + + test('carries a do-not-flag list bounding the false-positive surface', () => { + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /(do not flag|don't flag|never flag)/i); + // The negative space must name at least the two cases that most look like + // the trigger and are not: code-established preconditions and enforced + // ordering. + assert.match(window, /(enforce|await|explicit sequenc)/i); + assert.match(window, /default/i); + }); + + test('rule routes on recorded evidence, not self-rated confidence', () => { + // The .out-of-scope/general-purpose-agent-prompt-skills.md reason-4 bar: + // self-rated confidence is measured weak (honest-verifier.md:25-29). The + // shipped prose must carry that constraint, not just the design doc. + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /honest-verifier/); + assert.match(window, /confiden/i); + }); +}); + +describe('#1955: coincidental-reliance advisory — the invariants', () => { + test('advisory does not change the verified score', () => { + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /score/i); + assert.match(window, /(not the score|never the score|does not (change|affect) the score|no score)/i); + }); + + test('advisory does not change the overall status', () => { + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /(not the status|never the status|does not (change|affect) the status|status is unchanged)/i); + }); + + test('advisory never emits a human-verification item', () => { + // Step 9 rule 2: ANY human-verification item forces `human_needed`. An + // advisory that emitted one would stop a passing phase from passing, which + // is precisely what issue #1955 says must NOT happen. + const window = sectionAround(verifier, 'coincidental-reliance', 1400); + assert.match(window, /never[^.]{0,80}human-verification item/i); + }); + + test('PRESENT_BEHAVIOR_UNVERIFIED remains the only score-excluded truth state', () => { + // #966's invariant, asserted here so #1955 cannot silently erode it. + assert.match( + verifier, + /PRESENT_BEHAVIOR_UNVERIFIED truths are the \*only\* ones excluded from `verified_truths`/, + ); + }); + + test('PARITY: the per-truth advisory never leaks into the status vocabulary', () => { + const unionLines = verifier.match(/^status:\s+[a-z_]+(?:\s*\|\s*[a-z_]+)+\s*$/gm) || []; + assert.ok(unionLines.length > 0, 'expected at least one status union line'); + for (const line of unionLines) { + assert.doesNotMatch(line, /coincidental/i); + for (const s of ['passed', 'gaps_found', 'human_needed']) { + assert.ok(line.includes(s), `status union must still contain ${s}: ${line}`); + } + } + }); + + test('overall-status enum in verification.cts is unchanged', () => { + const cts = read('src', 'verification.cts'); + const m = cts.match(/VERIFIER_STATUSES[^=]*=\s*\[([^\]]*)\]/); + assert.ok(m, 'VERIFIER_STATUSES array must be present'); + assert.doesNotMatch(m[1], /coincidental/i); + for (const s of ['passed', 'gaps_found', 'human_needed']) { + assert.match(m[1], new RegExp(`'${s}'`)); + } + }); + + test('qualifier suffixes VERIFIED, never replaces it (Hyrum)', () => { + // Every rendered verdict carrying the qualifier must keep the `✓ VERIFIED` + // token verbatim and leading, so a consumer matching the old verdict as a + // substring still hits. A bare `(coincidental-reliance)` verdict cell, or + // any form that puts the qualifier before the token, is a break. + for (const source of [verifier, template]) { + const verdictCells = source.match(/\|\s*[^|\n]*coincidental-reliance[^|\n]*\|/g) || []; + for (const cell of verdictCells) { + assert.ok( + cell.includes(QUALIFIER), + `verdict cell must render as "${QUALIFIER}": ${cell.trim()}`, + ); + } + } + assert.ok( + verifier.includes(QUALIFIER), + 'the agent must show the suffixed verdict form at least once', + ); + }); +}); + +describe('#1955: coincidental-reliance advisory — the report surface', () => { + test('report frontmatter carries coincidental_reliance_items', () => { + assert.match(verifier, /coincidental_reliance_items/); + }); + + test('items list is omitted when nothing is flagged', () => { + // 0-flag boundary: the key follows the `overrides:` / `deferred:` "only if" + // convention rather than emitting an empty array. + assert.match(verifier, /coincidental_reliance_items:[^\n]*(only if|Only if)/); + }); + + test('item shape names the truth, the reason, and the hardening', () => { + // 1-flag boundary: an advisory that names no fix is not actionable. + // `coincidental_reliance_items: #` anchors the FRONTMATTER block. The bare + // key also appears earlier in the rule prose, and a window around that + // occurrence contains none of the item fields. + const window = sectionAround(verifier, 'coincidental_reliance_items: #', 500); + assert.match(window, /truth:/); + assert.match(window, /reason:/); + assert.match(window, /(harden|fix|precondition)/i); + }); + + test('items survive a gaps_found phase', () => { + // many-flag / mixed boundary: same survival rule behavior_unverified_items + // carries, so an advisory is not lost when the phase also has gaps. + const window = sectionAround(verifier, 'coincidental_reliance_items: #', 500); + assert.match(window, /(regardless of (overall )?status|survive|never lost)/i); + }); +}); + +describe('#1955: cross-surface parity (agent, template, verify-phase workflow)', () => { + test('PARITY: agent and standalone template agree on the advisory vocabulary', () => { + // Generative-fix-divergence gate: two surfaces render the same report, so a + // token added to one and not the other is the defect this test exists for. + for (const token of ['coincidental-reliance', ...REASONS]) { + assert.ok(verifier.includes(token), `agent must carry '${token}'`); + assert.ok(template.includes(token), `standalone template must carry '${token}'`); + } + assert.ok( + template.includes('coincidental_reliance_items'), + 'standalone template must carry the frontmatter items list', + ); + }); + + test('template guidelines document the advisory per-truth state', () => { + const guidelines = template.slice(template.indexOf('**Per-truth states')); + assert.ok(guidelines.length > 0, 'template must keep its per-truth states guideline'); + assert.match(guidelines, /coincidental-reliance/); + }); + + test('verify-phase workflow reaches the rule through its eager template import', () => { + // The third surface. `gsd-core/workflows/verify-phase.md` is the + // non-subagent verification path and reimplements the truth rubric inline, + // but it sits 29 bytes under the DEFAULT tier hard cap in + // tests/workflow-size-budget.test.cjs, so the rule is NOT duplicated into + // it. It reaches the rule instead through the eager `@`-import of the + // template, whose Guidelines carry the instruction — not merely the output + // shape. Both halves of that claim are asserted here, because either one + // silently failing turns the workflow surface into an undetected + // divergence. + assert.match( + verifyPhase, + /@[^\n]*gsd-core\/templates\/verification-report\.md/, + 'verify-phase.md must eagerly import the verification-report template', + ); + const guidelines = template.slice(template.indexOf('**Per-truth states')); + assert.match( + guidelines, + /apply the reliance check to every `✓ VERIFIED` truth/i, + 'the template Guidelines must carry the imperative check, not just the row shape', + ); + }); + + test('the workflow surface carries no divergent copy of the rule', () => { + // Characterization, not aspiration: verify-phase.md deliberately holds NO + // copy of the detection prose today. If a future change adds one, this + // assertion fails and forces a decision — duplicate it deliberately and + // update this test, or keep the single template-carried source. Silent + // partial duplication across the two surfaces is the failure mode + // (generative fix divergence) this locks out. + assert.doesNotMatch( + verifyPhase, + /coincidental-reliance/, + 'verify-phase.md must not grow a second copy of the rule without a deliberate decision', + ); + }); + + test('docs/AGENTS.md documents the coincidental-reliance advisory', () => { + assert.match(agentsDoc, /coincidental-reliance/); + }); +}); diff --git a/tests/workflow-size-budget.test.cjs b/tests/workflow-size-budget.test.cjs index 1b68ebd4b..589ea6e94 100644 --- a/tests/workflow-size-budget.test.cjs +++ b/tests/workflow-size-budget.test.cjs @@ -89,11 +89,16 @@ const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows'); // only as the outer bound where the correct response is lazy extraction, never // a raise. Each sits above its tier's current high-water mark with real // headroom (vs the old GRACE=3000 hug): -// XL 96 KiB — high-water plan-phase.md 92,965 → ~5.2 KB headroom -// LARGE 60 KiB — high-water docs-update.md 54,600 → ~6.6 KB headroom -// DEFAULT 40 KiB — high-water settings-advanced.md 39,160 → ~1.8 KB headroom +// XL 96 KiB — high-water execute-phase.md 93,400 → ~4.8 KB headroom +// LARGE 60 KiB — high-water docs-update.md 55,468 → ~5.8 KB headroom +// DEFAULT 40 KiB — high-water verify-phase.md 40,931 → 29 BYTES headroom // (DEFAULT is deliberately the tightest: a single-purpose workflow approaching -// 40 KiB is the strongest extraction signal of the three.) +// 40 KiB is the strongest extraction signal of the three. verify-phase.md is +// effectively AT the red line — the next edit to it must be preceded by a lazy +// extraction, not absorbed. Measured 2026-08-09 via measureWorkflows(); the +// previous note here named settings-advanced.md at 39,160 with ~1.8 KB of +// headroom, which was stale on both the file and the number and invited an +// edit that would have crossed the cap.) const XL_CAP = 98304; // 96 KiB const LARGE_CAP = 61440; // 60 KiB const DEFAULT_CAP = 40960; // 40 KiB