enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)

* test(#2951): pin the absent-evidence provenance contract (failing first)

17 tests / 22 anchors on the deployed agent text. Measured against the parent
commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity
guards on the package-name and in-repo-value rules -- green before and after is
their intended signature.

Refs #2951

* enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata

A claim of the form "X does not support Y" drawn from MISSING metadata -- no
python_requires, no engines field, no per-version classifier, no changelog entry,
no matching support-matrix row -- no longer earns [VERIFIED] however
authoritative the source consulted. An absence is silence about every value, so
the same evidence would "prove" both the version being ruled out and the version
being standardized on. The only route from an absence to [VERIFIED] is a positive
falsification attempt with its failing output pasted; everything short of that is
[ASSUMED], which the file already routes to "needs user confirmation before
becoming a locked decision".

Third member of the family beside the package-name and in-repo-value provenance
rules, mirroring PR #2768's shape. A present declared constraint and an
affirmatively documented incompatibility are untouched.

Closes #2951

* fix(#2951): close the allow-list ambiguity and the mutation gap review found

Findings from the isolated adversarial pass and the two-axis review, all fixed:

MAJOR (x2, one root cause) -- the absence clause and the present-constraint
carve-out gave opposite verdicts on the same evidence for the commonest real
case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A
researcher could read the enumerated list as a "declared" positive constraint
and re-earn [VERIFIED], which is also the evasion vector. The rule now states
the decision procedure -- does the declaration bound EVERY value or only the
ones it names -- and closes the positive-reframing restatement explicitly. New
contract test pins all four clauses.

MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]'
asserted two independent substrings and never that the route lands on
[VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the
rule and survived all 17 tests. Now pinned as one joined sentence.

MINOR -- the attributable-failure test regex-matched illustrative examples
("a missing certificate, a wrong host"), so a copy-edit would break it for no
reason; relaxed to the substantive clause. The no-paraphrase guard counted only
the heading, missing the drift mode in its own name; it now also pins the core
proposition to one occurrence, and the test name matches what it checks. An
off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed.

MINOR -- docs/AGENTS.md listed four of the five governed absence forms while
the agent prose, docs/COMMANDS.md and the changeset listed five; three copies
disagreeing on list membership is the drift this repo treats as a defect.

SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its
docs/README.md index line. Both reviewers flagged them as a seventh and eighth
surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill
change" row requires only docs/AGENTS.md. The actionable four-case guidance is
retained in docs/COMMANDS.md, which is inside the approved scope.

Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550
bytes headroom under the LARGE cap of 49152.

Refs #2951

* docs(#2951): restore the how-to the phase gate requires

Reverses the removal in 6404b43d3. Both /code-review axes had flagged
docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md
index line as a seventh and eighth surface beyond the six the requester capped,
and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names
only docs/AGENTS.md, so they were dropped.

gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded
enablement sequence has more than one step and the how-to quadrant is empty.
The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED]
claim, probe or cite or accept it unlocked, then answer discuss-phase's
checkpoint -- and the last step lands on a different capability's surface, so a
reference table cannot carry it. Compressing the sequence to one step to unlock
howToSkipReason would be gaming the gate, which is the same Goodhart failure
this whole change exists to close.

A machine-enforced repo gate outranks two reviewers' scope preference and my own
reading, so the page is restored and the PR body discloses the two extra
surfaces instead of hiding them. Reverting is a one-file change if a maintainer
prefers the tighter scope.

Refs #2951

* chore(#2951): backfill the changeset PR number

pr: 0 -> 3718 now that the real PR exists. The placeholder fails
scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks
lint-docs-required from consuming the fragment.

Refs #2951

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-08-20 14:35:14 -04:00
committed by GitHub
parent 8da2dd3ad2
commit 8df5cb36c2
8 changed files with 347 additions and 1 deletions

View File

@@ -0,0 +1,5 @@
---
type: Changed
pr: 3718
---
**The phase researcher no longer treats missing metadata as a compatibility constraint** — a claim like "this library does not support that runtime version", drawn from an absent `python_requires`, `engines` field, per-version classifier, changelog entry, or support-matrix row, no longer earns `[VERIFIED]` however authoritative the registry or docs consulted. An absence says nothing about the version being ruled out and nothing about the version being standardized on, so the same evidence would "prove" both; the only route from an absence to `[VERIFIED]` is a positive falsification attempt with its failing output pasted, and everything short of that stays `[ASSUMED]`, which already routes through a confirmation checkpoint before it can lock a decision. A present declared constraint and an affirmatively documented incompatibility are untouched. Previously an honestly-tagged absence could lock a CONTEXT.md decision and produce a real version downgrade that no downstream stage re-derived. (#2951)

View File

@@ -34,6 +34,8 @@ Spawned by `/gsd:plan-phase` (integrated) or `/gsd:plan-phase --research-phase <
**In-repo value provenance rule:** A claim about an in-repo *discrete value* — an enum, a schema or type union, an error code, a status constant, or a filesystem path — may be tagged `[VERIFIED: …]` only if you opened the source-of-truth file with `Read` **this session**. A codebase `grep` is not sufficient on its own: it confirms a string occurs, not that you read the definition. Cite the path **and line range** (`[VERIFIED: src/types/order.ts:14-22]`), and quote the values **verbatim** in RESEARCH.md beside the claim — paraphrase is forbidden. The quote is what makes the tag checkable — a citation with no quote beside it does not earn `[VERIFIED]`, however precise the line range looks. Every value appearing in a code example or skeleton must also appear in that verbatim quote; a value that does not is `[ASSUMED]`. For a filesystem path, cite the line in the script that creates it, not the location you expect it to occupy. Training memory and a web search are not substitutes for reading the file — a discrete value that merely looks right fails at the executor's `parse()`/typecheck, the most expensive place to discover it. **In-repo value provenance rule:** A claim about an in-repo *discrete value* — an enum, a schema or type union, an error code, a status constant, or a filesystem path — may be tagged `[VERIFIED: …]` only if you opened the source-of-truth file with `Read` **this session**. A codebase `grep` is not sufficient on its own: it confirms a string occurs, not that you read the definition. Cite the path **and line range** (`[VERIFIED: src/types/order.ts:14-22]`), and quote the values **verbatim** in RESEARCH.md beside the claim — paraphrase is forbidden. The quote is what makes the tag checkable — a citation with no quote beside it does not earn `[VERIFIED]`, however precise the line range looks. Every value appearing in a code example or skeleton must also appear in that verbatim quote; a value that does not is `[ASSUMED]`. For a filesystem path, cite the line in the script that creates it, not the location you expect it to occupy. Training memory and a web search are not substitutes for reading the file — a discrete value that merely looks right fails at the executor's `parse()`/typecheck, the most expensive place to discover it.
**Absent-evidence provenance rule:** A compatibility claim resting on **missing** metadata — no `python_requires`, no `engines` field, no per-version classifier, no changelog entry, no matching row in a support matrix — does not earn `[VERIFIED: …]`, however authoritative the source you consulted. Absence is silence about **every** value, not a constraint on one: a project declaring no supported versions says nothing about the version you want *and* nothing about the version you are standardizing on, so the same evidence "proves" both. The rule keys on the **evidence, not the wording** — "does not support 3.14" rephrased as "supports only up to 3.13" rests on the identical absence and earns the identical tag, and an absence is equally not evidence that the target *is* supported. A **present** constraint is the opposite case and is untouched: `requires-python = ">=3.9,<3.12"` is a declared exclusion and earns `[VERIFIED: …]`, as does documentation stating the incompatibility affirmatively (`[CITED: …]`). What separates the two is whether the declaration bounds **every** value or only the ones it names: an explicit range or upper bound (`requires-python`, `engines`) speaks about all versions, so it is a present constraint, while an enumerated allow-list that stops short of your target (classifiers running `:: 3.9` through `:: 3.13` with no `:: 3.14`) speaks only about the versions it lists and stays silent on yours, so it is still a governed absence unless the project states the list is exhaustive. Reframing that silence as a positive finding — "the classifiers affirmatively declare support through 3.13" — is the same absence in different clothes and earns the same tag. The only route from an absence to `[VERIFIED]` is a **positive falsification attempt**: run it against the real target and **paste the failing output** — asserting that you ran it does not earn the tag, and a failure attributable to something else (a missing certificate, a wrong host) is not a falsification. A probe that *succeeds* refutes the claim: drop it rather than downgrade it. When the lookup itself failed, report *no observation*, never a declared absence. Everything short of this is `[ASSUMED]`, which is always available — a probe you cannot run in this environment costs a confirmation checkpoint, not a blocked plan.
Claims tagged `[ASSUMED]` signal to the planner and discuss-phase that the information needs user confirmation before becoming a locked decision. Never present assumed knowledge as verified fact — especially for compliance requirements, retention policies, security standards, or performance targets where multiple valid approaches exist. Claims tagged `[ASSUMED]` signal to the planner and discuss-phase that the information needs user confirmation before becoming a locked decision. Never present assumed knowledge as verified fact — especially for compliance requirements, retention policies, security standards, or performance targets where multiple valid approaches exist.
</role> </role>

View File

@@ -72,6 +72,7 @@ GSD uses a multi-agent architecture where thin orchestrators (workflow files) sp
- Investigates implementation patterns for the specific phase domain - Investigates implementation patterns for the specific phase domain
- Detects test infrastructure for Nyquist validation mapping - Detects test infrastructure for Nyquist validation mapping
- Tags in-repo discrete values (enums, schema unions, error codes, status constants, paths) `[VERIFIED]` only after reading the source-of-truth file that run, citing path and line range, and quoting the values verbatim - Tags in-repo discrete values (enums, schema unions, error codes, status constants, paths) `[VERIFIED]` only after reading the source-of-truth file that run, citing path and line range, and quoting the values verbatim
- Refuses `[VERIFIED]` for a compatibility claim resting on *missing* metadata (no `python_requires`, no `engines` field, no per-version classifier, no changelog entry, no matching support-matrix row) — an absence constrains no version, and an enumerated allow-list that stops short of the target is still an absence, so only a positive falsification attempt with its failing output pasted earns the tag; anything less stays `[ASSUMED]`
--- ---

View File

@@ -240,6 +240,9 @@ See [Package Legitimacy Gate in the User Guide](USER-GUIDE.md#package-legitimacy
**In-repo value citation:** **In-repo value citation:**
For any in-repo *discrete value* the researcher reports — an enum, a schema or type union, an error code, a status constant, or a filesystem path — a `[VERIFIED: …]` tag requires that it opened the source-of-truth file with `Read` during the run and cited the path **and line range** (`[VERIFIED: src/types/order.ts:14-22]`). The values are quoted verbatim in RESEARCH.md beside the claim, and any value used in a code example must also appear in that quote; anything else stays `[ASSUMED]`. A codebase `grep`, training memory, or a web search do not earn the tag on their own. This stops a plausible-but-drifted enum from reaching PLAN.md — where the planner lifts it into the plan's `<interfaces>` context block and the executor trusts it as ground truth — and surfacing only as a mid-execution deviation at typecheck. For any in-repo *discrete value* the researcher reports — an enum, a schema or type union, an error code, a status constant, or a filesystem path — a `[VERIFIED: …]` tag requires that it opened the source-of-truth file with `Read` during the run and cited the path **and line range** (`[VERIFIED: src/types/order.ts:14-22]`). The values are quoted verbatim in RESEARCH.md beside the claim, and any value used in a code example must also appear in that quote; anything else stays `[ASSUMED]`. A codebase `grep`, training memory, or a web search do not earn the tag on their own. This stops a plausible-but-drifted enum from reaching PLAN.md — where the planner lifts it into the plan's `<interfaces>` context block and the executor trusts it as ground truth — and surfacing only as a mid-execution deviation at typecheck.
**Absent-evidence citation:**
A compatibility claim the researcher reports — "this library does not support that runtime version" — earns a `[VERIFIED: …]` tag only from *positive* evidence. Metadata that is simply **missing** (no `python_requires`, no `engines` field, no per-version classifier, no changelog entry, no matching row in a support matrix) does not qualify, however authoritative the registry or documentation consulted: an absence says nothing about the version you are ruling out *and* nothing about the version you are standardizing on, so the same evidence would "prove" both. The rule keys on the evidence rather than the wording, so rephrasing the claim positively ("supports only up to 3.13") changes nothing, and an absence is equally not evidence that the version *is* supported. A **present** constraint is the opposite case and still earns the tag — `requires-python = ">=3.9,<3.12"` is a declared exclusion — as does documentation stating the incompatibility affirmatively, which is `[CITED: …]`. What separates the two is whether the declaration bounds every value or only the ones it names: an explicit range or upper bound speaks about all versions, while an enumerated allow-list that stops short of the target (classifiers running `:: 3.9` through `:: 3.13` with no `:: 3.14`) stays silent about the target and remains a governed absence unless the project says the list is exhaustive. See [How-to: verify a dependency-compatibility claim](how-to/verify-a-dependency-compatibility-claim.md). The one route from an absence to `[VERIFIED]` is a positive falsification attempt: run it against the real target and paste the failing output. Everything short of that stays `[ASSUMED]`, which routes the claim through the usual confirmation checkpoint before it can lock a decision in CONTEXT.md — so a probe you cannot run in this environment costs a checkpoint, not a blocked plan.
```bash ```bash
/gsd-plan-phase 1 # Research + plan + verify phase 1 /gsd-plan-phase 1 # Research + plan + verify phase 1
/gsd-plan-phase 3 --skip-research # Plan without research (familiar domain) /gsd-plan-phase 3 --skip-research # Plan without research (familiar domain)

View File

@@ -33,6 +33,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [Consume the planning snapshot](how-to/consume-the-planning-snapshot.md) — read `planning inspect` from a dashboard or harness, and tell "nothing to report" apart from "could not look" - [Consume the planning snapshot](how-to/consume-the-planning-snapshot.md) — read `planning inspect` from a dashboard or harness, and tell "nothing to report" apart from "could not look"
- [Keep planning docs out of a shared repo](how-to/keep-planning-docs-private.md) — make `.planning/` local-only, including untracking files git already tracks (the step `.gitignore` alone cannot do) - [Keep planning docs out of a shared repo](how-to/keep-planning-docs-private.md) — make `.planning/` local-only, including untracking files git already tracks (the step `.gitignore` alone cannot do)
- [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality - [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality
- [Verify a dependency-compatibility claim](how-to/verify-a-dependency-compatibility-claim.md) — act on a compatibility claim the researcher left `[ASSUMED]`, and tell "nothing declared" apart from "a constraint is declared" and "the lookup failed"
- [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents - [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents
- [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR - [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR
- [Catch complexity before it compounds](how-to/act-on-a-refactor-proposal.md) — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it - [Catch complexity before it compounds](how-to/act-on-a-refactor-proposal.md) — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it

View File

@@ -0,0 +1,98 @@
# How to verify a dependency-compatibility claim the researcher would not verify
**Goal:** Turn a compatibility claim that came back `[ASSUMED]` — "this library does not support that runtime version" — into either a `[VERIFIED]` claim backed by a real probe, a `[CITED]` claim backed by an affirmative sentence in the vendor's own docs, or a decision you deliberately take without locking. The point is that an *absence* of metadata never becomes a constraint by default, so the version bound you end up pinning is the one the world actually imposes rather than the one a missing field seemed to imply.
**Prerequisites:** A phase whose `/gsd-plan-phase` research pass has produced a RESEARCH.md. The absent-evidence provenance rule runs inside `gsd-phase-researcher` automatically — there is no flag and nothing to enable. You reach this guide because a claim you expected to be settled is tagged `[ASSUMED]` and `/gsd-discuss-phase` is asking you to confirm it before it can lock a decision.
For the rule itself and the other two provenance rules beside it, see [`/gsd-plan-phase` in COMMANDS.md](../COMMANDS.md#gsd-plan-phase). This guide covers only how to *act* on the claim.
---
## Read the claim
A governed claim looks like this in RESEARCH.md:
> `[ASSUMED]` ldap3 publishes no `python_requires` and no per-minor classifier for 3.14 — support for 3.14 is unconfirmed.
Two things are true of it at once, and holding both is the whole skill:
- The **lookup was real.** The registry was consulted, the field genuinely is not there. Nothing is being doubted about the observation.
- The **conclusion is not.** A project that declares no supported versions has said nothing about the version you are ruling out *and* nothing about the version you are standardizing on. The same evidence would "prove" both, so it settles neither.
The tag is `[ASSUMED]` because of the second point, not the first. It is not a complaint that the researcher was lazy.
## Tell the four cases apart before you do anything
Reaching for a probe when you did not need one is the common waste here, and treating a real declared constraint as an absence is the expensive mistake. Read which case you are in first:
| What RESEARCH.md shows | What it means | What to do |
|---|---|---|
| **No constraint declared** — no `python_requires`, no `engines`, no classifier for any version, no changelog entry | The project is silent. Silence binds nothing. | Probe it, or accept unlocked — below |
| **A constraint is declared and excludes you** — e.g. `requires-python = ">=3.9,<3.12"` and you want 3.14 | A real, positive, published constraint | Nothing to do. This is already `[VERIFIED]` and it is a genuine bound — honor it |
| **Docs state the incompatibility affirmatively** — "Python 3.14 is not supported" in the project's own documentation | A positive statement about the world | Nothing to do. This is `[CITED]` and it stands |
| **The lookup itself failed** — registry 5xx, package not found, tool errored | *No observation.* This is not the same as "no field declared" | Retry the lookup first. Do not treat a failed lookup as evidence of anything |
Row 2 is the one worth slowing down for. The rule targets metadata that is **missing**, never metadata that is merely **unfavorable** — so a declared upper bound is not weakened by any of this, and re-probing against it is wasted work.
## Probe it — the route to `[VERIFIED]`
For dependency-compatibility the probe is usually under ten lines: install it against the real target, exercise the capability the claim is about, print what happened.
```bash
# The real target — the interpreter/runtime the claim is about, not a proxy for it
uv run --python 3.14 --with ldap3 python -c "
import ldap3, importlib.metadata as md
print('ldap3', md.version('ldap3'))
conn = ldap3.Connection(ldap3.Server('ldaps://dir.example.internal', use_ssl=True), auto_bind=True)
print('bind', conn.bound)
print('search', conn.search('dc=example,dc=internal', '(objectClass=computer)', attributes=['cn']))
"
```
Then paste the **output** — not a summary of it, and not an assertion that you ran it — beside the claim in RESEARCH.md. The output is what makes the tag checkable; a claim that a probe was run is exactly as unfalsifiable as the absence it was meant to replace.
Three things decide whether the probe actually settled anything:
- **It must exercise the capability the claim is about.** `import ldap3` succeeding on 3.14 says nothing about whether `bind()` and `search()` work. A probe that only imports proves only that the import works.
- **A failure must be attributable to the incompatibility.** If the probe fails because a certificate is missing or the host is wrong, that is your environment failing, not the library. Fix the probe and rerun; do not bank the failure.
- **A probe that succeeds refutes the claim.** Do not soften it to "probably fine" — remove it. A disproved claim left in place as `[ASSUMED]` still steers the plan, and nobody looks again at something that already carries a hedge.
Record the result:
| Probe outcome | Tag | What the plan may do with it |
|---|---|---|
| Failed, attributably, output pasted | `[VERIFIED: probe]` | May lock a decision — the bound is real |
| Succeeded | *claim removed* | The premise is gone; do not bound on it |
| Failed for an unrelated reason | still `[ASSUMED]` | Fix the probe and rerun |
## Accept it unlocked — when you cannot probe
Sometimes the probe is genuinely unavailable: no interpreter for that version, no reachable target host, no network in this environment. That is a normal outcome and the rule is built for it.
Leave the claim `[ASSUMED]` and tell `/gsd-discuss-phase` to proceed without locking. The claim still reaches the planner and can still shape a plan — it simply cannot become a locked `CONTEXT.md` decision that downstream phases treat as settled. **A probe you cannot run costs you a confirmation checkpoint, not a blocked plan.**
What you should *not* do is confirm the claim at the checkpoint to make the prompt go away. Confirming is you asserting the premise on your own authority; it will be treated as settled by every phase after this one, and the checkpoint is the last place anybody will look at it.
## Cite it instead — the cheaper route
Before writing a probe, check whether the vendor states the incompatibility outright. An affirmative sentence in official documentation — a support matrix with your version marked unsupported, a release note saying support was dropped — is positive evidence and earns `[CITED: <url>]` with the sentence quoted. That is often a two-minute answer where the probe is a twenty-minute one.
An absence in the docs is not this. "The docs do not mention 3.14" is the same silence you started with.
## The mirror-image mistake
The rule cuts both ways, and the direction people forget is the optimistic one:
> ~~"No upper bound is declared, so any version works."~~
That is the same inference with the sign flipped, and it fails for the same reason. If you want to standardize *on* a version, the evidence for that is a declared constraint that covers it, a matching classifier, or a probe that succeeds — never the absence of a prohibition.
## What happens if you skip all of this
The failure this rule exists to prevent, from the report that prompted it ([#2951](https://github.com/open-gsd/gsd-core/issues/2951)): a missing `python_requires` was read as "no 3.14 support", locked as a decision, and turned into a `requires-python = ">=3.12,<3.14"` bound, a pinned `.python-version`, and a CI assertion. The claim was true and worthless — the same metadata was equally absent for 3.12 and 3.13, the versions being standardized on. Six pipeline stages and five review lanes passed it, because each one checked internal consistency and none re-derived the premise. A five-line probe against a live directory server disproved it after the downgrade had already been committed.
## Related
- [`/gsd-plan-phase`](../COMMANDS.md#gsd-plan-phase) — the two sibling provenance rules: package-name legitimacy, and in-repo value citation
- [Discuss a phase](discuss-a-phase.md) — where an `[ASSUMED]` claim reaches its confirmation checkpoint
- [gsd-phase-researcher](../AGENTS.md#gsd-phase-researcher) — the agent that applies the rule

View File

@@ -1,7 +1,7 @@
{ {
"version": 1, "version": 1,
"paths": { "paths": {
"gsd-phase-researcher.md": "#3409: guarded `cat \"$phase_dir\"/*-CONTEXT.md` against nullglob wiping the pattern to zero operands when no CONTEXT.md exists \u2014 a bare `cat` with no operands blocks reading stdin (hangs the agent) instead of the `2>/dev/null` guard ever firing, since a stalled read is not a failing exit. Now checks `${_CTX[0]}` is a real path before invoking cat. Growth is the array-guard idiom itself (+43 bytes).", "gsd-phase-researcher.md": "#3409: guarded `cat \"$phase_dir\"/*-CONTEXT.md` against nullglob wiping the pattern to zero operands when no CONTEXT.md exists \u2014 a bare `cat` with no operands blocks reading stdin (hangs the agent) instead of the `2>/dev/null` guard ever firing, since a stalled read is not a failing exit. Now checks `${_CTX[0]}` is a real path before invoking cat. Growth is the array-guard idiom itself (+43 bytes). \u2014 #2951 append (merged into this fragment because two ack sources may never name the same path; the #3409 entry above is at the base and therefore spent, and a second fragment naming this path is a hard failure in scripts/lint-emitted-drift-ack.cjs): +2352 bytes, 44250 -> 46602 (LARGE tier, cap 49152, 2550 bytes headroom). Adds the absent-evidence provenance rule to the claim-provenance section \u2014 the third member of the family alongside the package-name and in-repo-value rules. A compatibility claim resting on MISSING metadata (no python_requires, no engines field, no per-version classifier, no changelog entry, no matching support-matrix row) no longer earns [VERIFIED] however authoritative the source consulted; the only route from an absence to [VERIFIED] is a positive falsification attempt with its failing output pasted, and everything short of that is [ASSUMED], which the file already routes to \"needs user confirmation before becoming a locked decision\". Growth is inline prose in the agent body, deliberately NOT relocated into an eagerly @-imported reference, which ADR-1610 Decision 4 names as gaming the size proxy.",
"gsd-verifier.md": "#3409: same nullglob-hang fix as gsd-phase-researcher.md, applied to `cat \"$PHASE_DIR\"/*-VERIFICATION.md` in Step 0 \u2014 an absent VERIFICATION.md previously left a zero-operand `cat` blocking on stdin instead of falling through to first-verification mode. Growth is the array-guard idiom (+49 bytes). \u2014 #3206 append (merged into this fragment because two ack sources may never name the same path): +52 bytes, 49098 -> 49150 (2 under the LARGE cap). The growth is the literal fix for the term 5b used undefined: the compressed explicit-evidence definition inlined at 5b (+34 net on the rewritten line \u2014 the trailing honest-verifier cite there is dropped as superseded by the inline definition; honest-verifier.md stays cited at 5c) plus gsd-core/ path-prefix repairs on the two 404ing bare references/ cites at 5c (honest-verifier.md) and the MVP-mode section (verify-mvp-mode.md) (+9 each). Lazy extraction remains untakeable in this change: the large extractable blocks are content-pinned by tests that read the agent file directly (tests/verifier-behavior-unverified.test.cjs, tests/verification-overrides.test.cjs), so extraction is its own coordinated change.", "gsd-verifier.md": "#3409: same nullglob-hang fix as gsd-phase-researcher.md, applied to `cat \"$PHASE_DIR\"/*-VERIFICATION.md` in Step 0 \u2014 an absent VERIFICATION.md previously left a zero-operand `cat` blocking on stdin instead of falling through to first-verification mode. Growth is the array-guard idiom (+49 bytes). \u2014 #3206 append (merged into this fragment because two ack sources may never name the same path): +52 bytes, 49098 -> 49150 (2 under the LARGE cap). The growth is the literal fix for the term 5b used undefined: the compressed explicit-evidence definition inlined at 5b (+34 net on the rewritten line \u2014 the trailing honest-verifier cite there is dropped as superseded by the inline definition; honest-verifier.md stays cited at 5c) plus gsd-core/ path-prefix repairs on the two 404ing bare references/ cites at 5c (honest-verifier.md) and the MVP-mode section (verify-mvp-mode.md) (+9 each). Lazy extraction remains untakeable in this change: the large extractable blocks are content-pinned by tests that read the agent file directly (tests/verifier-behavior-unverified.test.cjs, tests/verification-overrides.test.cjs), so extraction is its own coordinated change.",
"complete-milestone.md": "#3409: guarded `cat .planning/phases/*-*/*-SUMMARY.md` \u2014 with `shopt -s nullglob` active in this block's preamble (#2962), zero matching phase summaries collapses the glob to nothing and a bare `cat` blocks reading stdin rather than producing empty output, wedging the milestone-completion review. Growth is the array-existence-check idiom (+73 bytes, two glob segments makes this longer than the single-glob sites). \u2014 #2142 append (merged into this fragment because two ack sources may never name the same path): +1605 bytes, 40498 -> 42103. The `archive_milestone` step now documents the opt-in `--archive-quick` quick-task archival flag (default OFF, deliberately NOT symmetrical with phase archival's default-ON posture), folds the AskUserQuestion decision for it into the SAME `milestone.complete` invocation (avoiding a redundant second call), and states the known bucket-all provenance limit.", "complete-milestone.md": "#3409: guarded `cat .planning/phases/*-*/*-SUMMARY.md` \u2014 with `shopt -s nullglob` active in this block's preamble (#2962), zero matching phase summaries collapses the glob to nothing and a bare `cat` blocks reading stdin rather than producing empty output, wedging the milestone-completion review. Growth is the array-existence-check idiom (+73 bytes, two glob segments makes this longer than the single-glob sites). \u2014 #2142 append (merged into this fragment because two ack sources may never name the same path): +1605 bytes, 40498 -> 42103. The `archive_milestone` step now documents the opt-in `--archive-quick` quick-task archival flag (default OFF, deliberately NOT symmetrical with phase archival's default-ON posture), folds the AskUserQuestion decision for it into the SAME `milestone.complete` invocation (avoiding a redundant second call), and states the known bucket-all provenance limit.",
"discuss-phase-assumptions.md": "#3409: replaced the unreachable `AUTO_MODE=$(gsd_run query check auto-mode --pick active 2>/dev/null || echo \"false\")` \u2014 `||` never fires because the query exits 0 with empty stdout when the field is absent, not a failure, so AUTO_MODE silently ended up empty rather than \"false\" \u2014 with a two-line capture-then-default (`AUTO_MODE=\"${AUTO_MODE:-false}\"`) that actually reaches the fallback. Growth is the extra default-assignment line (+19 bytes).", "discuss-phase-assumptions.md": "#3409: replaced the unreachable `AUTO_MODE=$(gsd_run query check auto-mode --pick active 2>/dev/null || echo \"false\")` \u2014 `||` never fires because the query exits 0 with empty stdout when the field is absent, not a failure, so AUTO_MODE silently ended up empty rather than \"false\" \u2014 with a two-line capture-then-default (`AUTO_MODE=\"${AUTO_MODE:-false}\"`) that actually reaches the fallback. Growth is the extra default-assignment line (+19 bytes).",

View File

@@ -627,3 +627,239 @@ describe('gsd-phase-researcher in-repo value provenance rule (#1699)', () => {
); );
}); });
}); });
// ---------------------------------------------------------------------------
// Absent-evidence provenance rule (#2951)
//
// The two sibling rules above classify WHERE you looked. Neither classifies
// whether what you saw supports the claim you drew from it. Consulting PyPI and
// finding no `python_requires` and no per-minor classifier is a tool-confirmed
// observation from an authoritative source — it earns `[VERIFIED: PyPI]` under a
// plain reading of the base taxonomy. The negative conclusion ("ldap3 does not
// support 3.14") then rides into a locked CONTEXT.md decision on a tag that was
// honestly applied to the lookup. The claim was true of 3.14 and equally true of
// 3.12 and 3.13 — the versions the plan was standardizing ON — so the same
// evidence "proved" both, and a wrong interpreter downgrade shipped.
//
// An absence can be verified; the tag did not distinguish a verified absence from
// a verified constraint. These assert the governed prose contract on the deployed
// agent definition.
// ---------------------------------------------------------------------------
describe('gsd-phase-researcher absent-evidence provenance rule (#2951)', () => {
const agentPath = path.join(ROOT, 'agents', 'gsd-phase-researcher.md');
const read = () => fs.readFileSync(agentPath, 'utf-8');
test('names the absent-metadata forms the rule governs', () => {
const content = read();
for (const term of ['`python_requires`', '`engines` field', 'per-version classifier', 'changelog entry', 'support matrix']) {
assert.ok(
content.includes(term),
`agent must name "${term}" as a governed form of absent metadata`
);
}
});
test('states that absence is silence about every value, not a constraint on one', () => {
const content = read();
assert.ok(
content.includes('Absence is silence about **every** value, not a constraint on one'),
'the core proposition must be stated, not paraphrased'
);
assert.match(
content,
/says nothing about the version you want \*and\* nothing about the version you are standardizing on/,
'the agent must spell out that the same absence "proves" both sides, which is why the ldap3 claim was worthless'
);
});
test('refuses the tag however authoritative the consulted source was', () => {
assert.ok(
read().includes('however authoritative the source you consulted'),
'source authority is what made the original claim pass review — it must be explicitly insufficient'
);
});
test("keys on the evidence, not the claim's wording", () => {
const content = read();
assert.ok(
content.includes('keys on the **evidence, not the wording**'),
'a rule keyed on surface polarity is evaded by rephrasing'
);
assert.ok(
content.includes('supports only up to 3.13'),
'the agent must name the positive rephrasing as resting on the identical absence'
);
});
test('rejects absence as evidence of support, not only of non-support', () => {
assert.match(
read(),
/an absence is equally not evidence that the target \*is\* supported/,
'the mirror-image error must be closed, or the rule licenses "no upper bound declared, so any version works"'
);
});
test('leaves a present declared constraint earning [VERIFIED]', () => {
const content = read();
assert.ok(
content.includes('A **present** constraint is the opposite case and is untouched'),
'the rule targets MISSING fields, never merely unfavorable ones'
);
assert.ok(
content.includes('`requires-python = ">=3.9,<3.12"` is a declared exclusion'),
'a concrete present-constraint example must show what still earns the tag'
);
});
test('leaves an affirmatively documented incompatibility on [CITED]', () => {
assert.match(
read(),
/documentation stating the incompatibility affirmatively \(`\[CITED: …\]`\)/,
'an affirmative statement about the world is not an absence and must keep its existing tag'
);
});
test('licenses a positive falsification attempt as the route to [VERIFIED]', () => {
const content = read();
// Pinned as ONE joined sentence, not as independent substrings: the rule's whole
// payoff is the TARGET of the route. A mutant swapping `[VERIFIED]` for `[CITED]`
// or `[ASSUMED]` inverts the rule while leaving every separate phrase intact.
assert.match(
content,
/The only route from an absence to `\[VERIFIED\]` is a \*\*positive falsification attempt\*\*/,
'the route and the tag it reaches must be asserted together, or inverting the tag survives'
);
assert.ok(
content.includes('run it against the real target'),
'the probe must exercise the real target, not a proxy'
);
});
test('makes the pasted failing output the artifact, not the claim of having run it', () => {
const content = read();
assert.ok(
content.includes('**paste the failing output**'),
'the output is the falsifiable artifact — mirrors the in-repo rule making the quote, not the citation, the artifact'
);
assert.ok(
content.includes('asserting that you ran it does not earn the tag'),
'an unpasted probe assertion is the box-ticking mode this rule exists to close'
);
});
test('rejects a failure not attributable to the incompatibility', () => {
assert.match(
read(),
// The parenthetical examples are illustrative; a copy-edit that swaps them must
// not break this. Only the substantive clause is pinned.
/a failure attributable to something else[^.]{0,80}is not a falsification/,
'any-failure-will-do turns the probe requirement into a formality'
);
});
test('refutes rather than downgrades a claim whose probe succeeds', () => {
assert.match(
read(),
/A probe that \*succeeds\* refutes the claim: drop it rather than downgrade it/,
'a disproved claim must leave RESEARCH.md entirely, not survive as [ASSUMED]'
);
});
test('separates a failed lookup from a declared absence', () => {
assert.match(
read(),
/When the lookup itself failed, report \*no observation\*, never a declared absence/,
'an unobserved field must not be laundered into a declared-absent field'
);
});
test('keeps [ASSUMED] available when a probe cannot be run', () => {
assert.match(
read(),
/costs a confirmation checkpoint, not a blocked plan/,
'a rule that demanded a feasible probe would become a gate that is routinely skipped'
);
});
test('sits inside the claim-provenance block ahead of the [ASSUMED] routing sentence', () => {
const content = read();
const taxonomy = content.indexOf('**Claim provenance:**');
const rule = content.indexOf('**Absent-evidence provenance rule:**');
const routing = content.indexOf('Claims tagged `[ASSUMED]` signal to the planner and discuss-phase');
assert.ok(taxonomy >= 0, 'the base claim-provenance taxonomy must still be present');
assert.ok(rule >= 0, 'the absent-evidence rule must be present');
assert.ok(routing >= 0, 'the [ASSUMED] routing sentence must still be present');
assert.ok(
taxonomy < rule && rule < routing,
'the rule must follow the taxonomy it constrains and precede the routing sentence it terminates in — '
+ 'otherwise "routes into the existing [ASSUMED] path" is not true of the deployed ordering'
);
});
test('distinguishes a bounding declaration from an allow-list that stops short', () => {
const content = read();
// The ldap3 class of case: classifiers present for some versions, absent for the
// target. Without this, the absence clause and the present-constraint carve-out
// give opposite verdicts on the same evidence and the rule cannot be applied.
assert.ok(
content.includes('whether the declaration bounds **every** value or only the ones it names'),
'the rule must give a decision procedure for present-list-vs-absent-entry, not two conflicting readings'
);
assert.match(
content,
/an enumerated allow-list that stops short of your target[\s\S]{0,400}is still a governed absence/,
'an allow-list missing the target version must stay a governed absence'
);
assert.ok(
content.includes('unless the project states the list is exhaustive'),
'the one condition that turns an allow-list into a real constraint must be named'
);
assert.match(
content,
/Reframing that silence as a positive finding[\s\S]{0,120}earns the same tag/,
'the positive-reframing evasion must be closed explicitly, not left to inference'
);
});
test('does not disturb the package name provenance rule', () => {
const content = read();
assert.ok(
content.includes('**Package name provenance rule:**'),
'the first sibling rule must survive unchanged'
);
assert.ok(
content.includes('a slopsquatted package also passes `npm view`'),
'package-legitimacy reasoning must remain intact'
);
});
test('does not disturb the in-repo value provenance rule', () => {
const content = read();
assert.ok(
content.includes('**In-repo value provenance rule:**'),
'the second sibling rule must survive unchanged'
);
assert.match(
content,
/opened the source-of-truth file with `Read` \*\*this session\*\*/,
'the same-session Read requirement must remain intact'
);
});
test('defines the rule and its core proposition once each (META.RULE.brief-no-paraphrase)', () => {
const content = read();
// Counting the heading alone would miss the drift mode this guard is named for: the
// same substance restated under a different heading. Pin the load-bearing sentence too.
assert.equal(
content.split('Absent-evidence provenance rule').length - 1,
1,
'the rule must be headed at exactly one site; a second copy is the prose-drift mode'
);
assert.equal(
content.split('Absence is silence about **every** value').length - 1,
1,
'the core proposition must appear once; a restatement elsewhere is the drift this guards'
);
});
});