fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in planner verify blocks (#1482)

* fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in verify blocks

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #1478/#1479/#1480 planner verify gate fix

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1478,#1479,#1480): fix test contract violations from new planner dimensions

gsd-planner.md exceeded both the planner-decomposition 48K char limit and
the reachability-check 50K char limit after the new HARD RULE blocks were
added inline. The full rule details already exist in planner-antipatterns.md
(added in the same PR); replace the verbose inline blocks with a single
@-reference pointer to the antipatterns file, reducing the file from 50981
to 49130 chars (under both limits).

Also regenerate tests/agent-size-baseline.json to reflect the new sizes of
gsd-planner.md and gsd-plan-checker.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-06-20 13:37:03 -04:00
committed by GitHub
parent 3a3b2135c2
commit fe0f9f904d
5 changed files with 94 additions and 2 deletions

View File

@@ -0,0 +1,8 @@
---
type: Fixed
pr: 1482
---
fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in planner verify blocks
<!-- docs-exempt: instruction-text updates to agents/gsd-planner.md, agents/gsd-plan-checker.md, and references/planner-antipatterns.md are themselves the documentation — no public user-facing API or CLI change -->

View File

@@ -647,6 +647,40 @@ issue:
fix_hint: "Add auth middleware pattern from PATTERNS.md ## Shared Patterns to plan"
```
## Dimension: Verify Command Format Sanity (#1478, #1479)
**Question:** Do `<verify>` commands use patterns that can actually match the tool's output? Are numeric counts measured? Are errors suppressed into comparison-feeding defaults?
**Red flags — BLOCKER:**
- `pnpm ls … | grep -E '^package'` — `^` anchor on tree-formatted package manager output (never matches tree-prefixed lines)
- Any verify block with `VAR=$(cmd 2>/dev/null || echo "0"); [ "$VAR" = ... ]` — swallowed error feeds passing comparison
- `|| true` or `|| :` as right-hand side of assignments that feed comparisons
**Red flags — WARNING:**
- Hard-coded count assertion (`grep '52 test files'`, `grep '714 passed'`) with no measurement provenance in the plan
**Process:**
1. For each `<automated>` block piping a package-manager list command into grep with a `^` anchor: BLOCKER.
2. For each `<automated>` block containing `2>/dev/null || echo` where the result feeds a `[ "$VAR" = ... ]` comparison: BLOCKER.
3. For each `<automated>` block asserting a specific numeric count not cited as measured in this plan: WARNING.
## Dimension: Numeric/Factual Claim Authority (#1480)
**Rule:** RESEARCH.md is produced at research time and may be stale. Numeric claims (test counts, file counts, version numbers) and factual state claims ("feature X is implemented") in RESEARCH.md may not reflect the current codebase. The plan may be more current. RESEARCH.md is authoritative for architectural decisions and constraints — not for measurements.
**Process when a plan's numeric/factual claim conflicts with RESEARCH.md:**
1. **Attempt live measurement first** with a targeted read-only command (e.g., `find . -name '*.test.*' | wc -l`). Run it. Use the result as ground truth:
- Measurement confirms plan → WARNING: RESEARCH.md is stale; recommend updating it.
- Measurement contradicts plan → BLOCKER: plan value is wrong; prescribe the measured value.
2. **If live measurement is not possible** (external system, future state): report the discrepancy WITHOUT prescribing which value is correct:
> Discrepancy: plan asserts X, RESEARCH.md asserts Y. Cannot determine ground truth without live measurement. Verify manually and update the stale artifact.
**NEVER** prescribe a specific value by assuming RESEARCH.md is authoritative for a numeric/factual claim.
**Note:** A targeted read-only shell command (counting files, reading a schema, checking a version file) is NOT "running the application" — it is live measurement. Such commands are permitted under this dimension even when the anti-pattern block says "DO NOT run the application."
</verification_dimensions>
<verification_process>

View File

@@ -200,6 +200,8 @@ Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Com
<region_scoped_negative_gate>
**Region-scoped negative gates (WARN, #968):** Region-scope a file-wide negative grep when a sibling task needs that construct elsewhere in the same file; `validate_plan` WARNS. See: @gsd-core/references/planner-antipatterns.md ("Region-Scoped Negative Gates").
**Verify-gate hygiene (#1478/#1479):** See @gsd-core/references/planner-antipatterns.md.
</region_scoped_negative_gate>
**<done>:** Acceptance criteria - measurable state of completion.

View File

@@ -174,3 +174,51 @@ If region-scoping is genuinely impractical and the file split is intentional, su
```
One marker per pattern. The marker exempts only the exact pattern it names. Prefer region-scoping over suppression.
## CLI Output Format Anchor Mismatch (#1478)
`pnpm ls vite | grep -E '^vite@7\.'` looks correct but silently fails. `pnpm ls` uses tree characters as line prefixes:
```
my-project@1.0.0
└── vite@7.3.5
```
Lines begin with `└──`, not `vite`. The `^` anchor matches line start, which is a tree character — the grep finds nothing.
**Bad:** `pnpm ls vite | grep -E '^vite@7\.'`
**Good:** `pnpm ls vite | grep -E 'vite@7\.'`
**Good (strict):** `pnpm ls vite | grep -E '(└|├)── vite@7\.'`
Same trap: `npm ls`, `yarn list`, `docker ps` column output, `kubectl get` table output.
## Fabricated Numeric Baselines (#1478)
Never emit `grep '714 tests'` or `grep '52 test files'` unless you ran the count command in this session. Model-recalled counts are stale from training.
**Bad:** `npm test 2>&1 | grep '714 passed'`
**Good:** `npm test 2>&1 | grep -E '[0-9]+ passed'` or just `npm test`
## Error-Suppressing Fallbacks in Verify Gates (#1479)
`2>/dev/null || echo "0"` in an assignment that feeds a comparison converts any failure into a passing gate that measures nothing.
**Bad — both sides default to "0" when files are missing:**
```bash
EN_KEYS=$(jq 'keys | length' i18n/en.json 2>/dev/null || echo "0")
DE_KEYS=$(jq 'keys | length' i18n/de.json 2>/dev/null || echo "0")
[ "$EN_KEYS" = "$DE_KEYS" ] && echo "ok"
```
If files don't exist (wrong path, etc.), both sides become `"0"`. Comparison passes. Gate certifies parity while measuring nothing.
**Good — let failure propagate:**
```bash
EN_KEYS=$(jq 'keys | length' src/i18n/en.json)
DE_KEYS=$(jq 'keys | length' src/i18n/de.json)
[ "$EN_KEYS" = "$DE_KEYS" ] && echo "ok"
```
**Good — explicit guard:**
```bash
test -f src/i18n/en.json && test -f src/i18n/de.json || { echo "missing input files"; exit 1; }
```
**When `|| echo "default"` is acceptable:** only when absence is semantically the default AND the result is NOT used in a comparison that should detect absence.

View File

@@ -22,8 +22,8 @@
"gsd-nyquist-auditor.md": 7255,
"gsd-pattern-mapper.md": 12487,
"gsd-phase-researcher.md": 40638,
"gsd-plan-checker.md": 42003,
"gsd-planner.md": 49216,
"gsd-plan-checker.md": 44646,
"gsd-planner.md": 49306,
"gsd-project-researcher.md": 22014,
"gsd-research-synthesizer.md": 13653,
"gsd-roadmapper.md": 21781,