Files
msd-core/gsd-core/references/planner-antipatterns.md
Dennis Alexis Valin Dittrich eedb6b5431 enhance(#4107): sequence external review after internal fixes (#4206)
* enhance(#4107): sequence external review after internal fixes

Teach the planner to finish internal review and accepted fixes before opening a PR known to trigger automatic external review. If an open-time property exists, re-check it immediately before opening with nothing intervening; post-open CI, review, changeset, and tracking work may follow.

Emitted-Drift-Ack-Growth: gsd-planner.md — issue #4107 adds the review-before-publish ordering rule

* chore(#4107): add PR #11 changeset

* chore(changeset): link upstream PR 4206

* fix(#4107): ground external-review terms and tighten ordering test

Addresses trek-e review on PR #4206:
- Ground 'known automatic external review' and 'open-time property' with
  concrete anchors (CodeRabbit App / .coderabbit.yaml, not-behind-base).
- Suffix the antipatterns heading with (#4107), matching sibling sections.
- Replace vacuous negative assertion with inverted-order fixtures that
  prove the ordering regexes reject bad phrasing, not just co-occurrence.

* fix(#4107): make directionality fixtures genuinely adversarial

agy (gemini-3.8-flash-high) adversarial review found the two negative
fixtures added in 584ec1cda were vacuous: they proved the ordering regexes
require certain keywords, not that they reject inverted order — the bad
strings simply omitted required tokens rather than reordering them.

- Rebuild both fixtures to contain every required token, reordered/negated,
  so a real reordering would still slip past a weaker regex.
- Drop the unsupported 'changeset' mention from the Wave 4+ antipatterns
  example — gsd-core/workflows/ship.md never references changeset work,
  so naming it here implied a step this rule doesn't actually govern.

* fix(#4107): make the full review-then-fix-then-open sequence explicit

CodeRabbit (fork PR #11) flagged that the planner prose only ordered
accepted fixes before PR open, without explicitly naming 'run internal
review' as its own earlier step, and that no fixture tested the planner
text's own wording for inversion (only the antipatterns example had one).

- Prose now reads 'run internal review and apply the accepted
  internal-review fixes before the final open'.
- Added a planner-text-specific inverted-order fixture alongside the
  existing antipatterns-example one.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 03:14:27 -04:00

12 KiB

Planner Anti-Patterns and Specificity Examples

Reference file for gsd-planner agent. Loaded on-demand via @ reference. For sub-200K context windows, this content is stripped from the agent prompt and available here for on-demand loading.

Checkpoint Anti-Patterns

Writing guidelines

DO: Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.

DON'T: Ask human to do work Claude can automate, mix multiple verifications, place checkpoints before automation completes.

Bad — Asking human to automate

<task type="checkpoint:human-action">
  <action>Deploy to Vercel</action>
  <instructions>Visit vercel.com, import repo, click deploy...</instructions>
</task>

Why bad: Vercel has a CLI. Claude should run vercel --yes. Never ask the user to do what Claude can automate via CLI/API.

Bad — Too many checkpoints

<task type="auto">Create schema</task>
<task type="checkpoint:human-verify">Check schema</task>
<task type="auto">Create API</task>
<task type="checkpoint:human-verify">Check API</task>

Why bad: Verification fatigue. Users should not be asked to verify every small step. Combine into one checkpoint at the end of meaningful work.

Good — Single verification checkpoint

<task type="auto">Create schema</task>
<task type="auto">Create API</task>
<task type="auto">Create UI</task>
<task type="checkpoint:human-verify">
  <what-built>Complete auth flow (schema + API + UI)</what-built>
  <how-to-verify>Test full flow: register, login, access protected page</how-to-verify>
</task>

Bad — Mixing checkpoints with implementation

A plan should not interleave multiple checkpoint types with implementation tasks. Checkpoints belong at natural verification boundaries, not scattered throughout.

Specificity Examples

TOO VAGUE JUST RIGHT
"Add authentication" "Add JWT auth with refresh rotation using jose library, store in httpOnly cookie, 15min access / 7day refresh"
"Create the API" "Create POST /api/projects endpoint accepting {name, description}, validates name length 3-50 chars, returns 201 with project object"
"Style the dashboard" "Add Tailwind classes to Dashboard.tsx: grid layout (3 cols on lg, 1 on mobile), card shadows, hover states on action buttons"
"Handle errors" "Wrap API calls in try/catch, return {error: string} on 4xx/5xx, show toast via sonner on client"
"Set up the database" "Add User and Project models to schema.prisma with UUID ids, email unique constraint, createdAt/updatedAt timestamps, run prisma db push"

Specificity test: Could a different Claude instance execute the task without asking clarifying questions? If not, add more detail.

Context Section Anti-Patterns

Bad — Reflexive SUMMARY chaining

<context>
@.planning/phases/01-foundation/01-01-SUMMARY.md
@.planning/phases/01-foundation/01-02-SUMMARY.md  <!-- Does Plan 02 actually need Plan 01's output? -->
@.planning/phases/01-foundation/01-03-SUMMARY.md  <!-- Chain grows, context bloats -->
</context>

Why bad: Plans are often independent. Reflexive chaining (02 refs 01, 03 refs 02...) wastes context. Only reference prior SUMMARY files when the plan genuinely uses types/exports from that prior plan or a decision from it affects the current plan.

Good — Selective context

<context>
@.planning/PROJECT.md
@.planning/STATE.md
@.planning/phases/01-foundation/01-01-SUMMARY.md  <!-- Uses User type defined in Plan 01 -->
</context>

Scope Reduction Anti-Patterns

Prohibited language in task actions:

  • "v1", "v2", "simplified version", "static for now", "hardcoded for now"
  • "future enhancement", "placeholder", "basic version", "minimal implementation"
  • "will be wired later", "dynamic in future phase", "skip for now"

If a decision from CONTEXT.md says "display cost calculated from billing table in impulses", the plan must deliver exactly that. Not "static label /min" as a "v1". If the phase is too complex, recommend a phase split instead of silently reducing scope.

Comment-Text Discipline (HARD GATE)

Enforced at plan-write time by verify.plan-structure (the validate_plan step). Issue #429.

When an <acceptance_criteria> or <verify> block uses a negative grep — grep -c 'LITERAL' file == 0, meaning "this literal must NOT appear in the file" — that same LITERAL must not appear verbatim anywhere in an <action> body. Verbatim code blocks, JSDoc samples, head-comment references, and "what NOT to do" illustrations get echoed into the file the executor writes, so the executor's commit-time gate fails on the comment text, not on a real code regression. The work is correct; the gate output is semantically wrong; the executor wastes cycles and learns to distrust the gate.

The gate: plan creation FAILS (error, valid: false) when a confidently-extracted (quoted) negative-grep literal also appears in an <action> block. When the grep literal is unquoted and cannot be extracted unambiguously, the gate WARNS instead of failing (so you still get the plan, with the risk surfaced).

Bad — JSDoc sample echoes the forbidden literal

<task>
  <action>
    Add a `?from=` query param to the share link. Do NOT reintroduce the old
    `?from=` referrer hack the JSDoc warned about.   <!-- echoes ?from= -->
  </action>
  <verify><automated>grep -c '?from=' src/animal-detail.tsx == 0</automated></verify>
</task>

Good — rephrase the comment by concept

<task>
  <action>
    Add the share-link query param. Do NOT reintroduce the legacy referrer hack.
  </action>
  <verify><automated>grep -c '?from=' src/animal-detail.tsx == 0</automated></verify>
</task>

Allowlist escape hatch

When the literal MUST appear in the plan body verbatim — e.g. the plan documents the test file that exercises the gate itself, or the literal is part of the verification command's own grep regex — add a marker on its own line so the gate skips that literal:

<!-- planner-discipline-allow: ?from= -->

One marker per literal. The marker exempts only the exact literal it names.

Region-Scoped Negative Gates

Surfaced at plan-write time by verify.plan-structure (the validate_plan step), WARN-level. Issue #968.

A negative grep — ! grep -Eq 'PAT' file or grep -c 'PAT' file == 0 — asserts a construct is absent. grep is file-scoped by nature: it has no notion of function or region. This breaks down when a phase splits one file across parallel tasks with legitimately opposite needs for the same construct in different regions:

  • Task A bans the construct file-wide — a synchronous factory must not block on a refresh: ! grep -Eq 'await .*refresh' app/page.py.
  • Task B legitimately requires it elsewhere in the same file — a post-reindex handler must await bridge.refresh() to repopulate state.

Both occurrences are real, correct production code in different functions of one file. A file-wide negative grep cannot say "absent in function X, present in function Y", so the two gates are mutually unsatisfiable with a direct call — the executor is pushed into an indirection whose only purpose is to relocate the matched string out of the file (pure gate-appeasement, zero behaviour change). This is distinct from Comment-Text Discipline (#429): there is no comment echo and no allowlist helps — the construct must genuinely be present in one region and absent in another.

The fix: region-scope the negative gate so "absent in region X" stops implying "absent file-wide."

Bad — file-wide ban unsatisfiable against a sibling's real code

<!-- Task A -->
<verify><automated>! grep -Eq 'await .*refresh' app/page.py</automated></verify>
<!-- Task B (same file, different function) -->
<action>Add a post-reindex handler in app/page.py that awaits bridge.refresh().</action>
<files>app/page.py</files>

Task B writing await bridge.refresh() trips Task A's file-wide gate, though both are correct.

Good — scope the gate to the factory region

<!-- Task A: ban only inside the synchronous factory, not the whole file -->
<verify><automated>! awk '/^def make_page/,/^def /' app/page.py | grep -Eq 'await .*refresh'</automated></verify>
<!-- or a fixed line range -->
<verify><automated>! sed -n '12,40p' app/page.py | grep -Eq 'await .*refresh'</automated></verify>

The factory region is asserted clean; the reindex handler elsewhere in the same file keeps its required await bridge.refresh(). Both gates pass with no code restructuring. Prefer an AST/structural check or a focused unit test where region extraction is fragile.

When the split is intentional and unavoidable

If region-scoping is genuinely impractical and the file split is intentional, suppress the warning with a marker naming the pattern:

<!-- planner-region-allow: await .*refresh -->

One marker per pattern. The marker exempts only the exact pattern it names. Prefer region-scoping over suppression.

CLI Output Format Anchor Mismatch (#1478)

pnpm ls vite | grep -E '^vite@7\.' looks correct but silently fails. pnpm ls uses tree characters as line prefixes:

my-project@1.0.0
└── vite@7.3.5

Lines begin with └──, not vite. The ^ anchor matches line start, which is a tree character — the grep finds nothing.

Bad: pnpm ls vite | grep -E '^vite@7\.' Good: pnpm ls vite | grep -E 'vite@7\.' Good (strict): pnpm ls vite | grep -E '(└|├)── vite@7\.'

Same trap: npm ls, yarn list, docker ps column output, kubectl get table output.

Fabricated Numeric Baselines (#1478)

Never emit grep '714 tests' or grep '52 test files' unless you ran the count command in this session. Model-recalled counts are stale from training.

Bad: npm test 2>&1 | grep '714 passed' Good: npm test 2>&1 | grep -E '[0-9]+ passed' or just npm test

Error-Suppressing Fallbacks in Verify Gates (#1479)

2>/dev/null || echo "0" in an assignment that feeds a comparison converts any failure into a passing gate that measures nothing.

Bad — both sides default to "0" when files are missing:

EN_KEYS=$(jq 'keys | length' i18n/en.json 2>/dev/null || echo "0")
DE_KEYS=$(jq 'keys | length' i18n/de.json 2>/dev/null || echo "0")
[ "$EN_KEYS" = "$DE_KEYS" ] && echo "ok"

If files don't exist (wrong path, etc.), both sides become "0". Comparison passes. Gate certifies parity while measuring nothing.

Good — let failure propagate:

EN_KEYS=$(jq 'keys | length' src/i18n/en.json)
DE_KEYS=$(jq 'keys | length' src/i18n/de.json)
[ "$EN_KEYS" = "$DE_KEYS" ] && echo "ok"

Good — explicit guard:

test -f src/i18n/en.json && test -f src/i18n/de.json || { echo "missing input files"; exit 1; }

When || echo "default" is acceptable: only when absence is semantically the default AND the result is NOT used in a comparison that should detect absence.

External Review Before PR Open (#4107)

Apply this ordering only when opening the PR is known to trigger automatic external review and the plan also has internal review lanes.

Bad:

Wave 1: Open PR; automatic external review starts
Wave 2: Run internal review
Wave 3: Apply accepted fixes

The external reviewer spends its first pass on a diff the plan already expects to change.

Good:

Wave 1: Run internal review
Wave 2: Apply accepted fixes
Wave 3: If applicable, re-check the open-time property; then immediately open PR
Wave 4+: Run post-open CI, external review, and tracking work

Nothing may intervene between an applicable re-check and the open. Post-open work may follow; "immediately" constrains only that gap. Opening-time properties do not justify an early PR.