Files
msd-core/gsd-core/templates/DEBUG.md
Tom Boucher 36a311c5bb enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests

Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking
(fast-check/Hypothesis, minimized seed, manual-minimization degradation), the
four oracle types (specified/derived/metamorphic/implicit with implicit flagged
weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the
equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in
(minimized seed + real oracle => the mutation guardrail bites).

Failing-first: reference, agent cross-refs, and template field do not yet exist.

* feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries)

Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First
Debugging (oracle classification + boundary neighbors):
- Shrinking: wrap an input-space failing input in a property (fast-check JS/TS,
  Hypothesis Python) and store the MINIMIZED counterexample as the regression
  seed; degrade to manual minimization when no PBT framework is present.
- Oracle classification: state specified / derived (contract/model) /
  metamorphic / implicit (crash, weakest) before writing the assertion; record
  under Resolution.oracle_type; never default to implicit silently.
- Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed
  defect's equivalence class.

Together they turn the regression test into a root-cause check — what the Phase
1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/
debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline +
install-parity goldens + AGENTS.md + DEBUG template updated.

* fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple)

- HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade-
  to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) —
  the gauntlet violation the sibling references already honored.
- Medium: test-provenance caveat (the failing input often comes from the bug
  report — author the generator from a sanitized description, cross-ref
  debugger-fix-acceptance.md).
- Medium: oracle scope note — the 4 types cover deterministic bugs; non-
  deterministic failures re-route to stability-stress per bug-taxonomy.
- Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient;
  boundary neighbors close the adjacent-input escape; the sufficient triple is
  seed+oracle+neighbors.
- Low: preserve the original noisy repro as a secondary reference; operationalize
  'equivalence class' (the predicate the fix draws). Nit: degradation reworded.

* chore(#1962): backfill changeset pr number (PR #2409)

---------

Co-authored-by: sim <sim@local>
2026-07-18 15:46:42 -04:00

5.9 KiB

Debug Template

Template for .planning/debug/[slug].md — active debug session tracking.


File Template

---
status: gathering | investigating | fixing | verifying | awaiting_human_verify | resolved
trigger: "[verbatim user input]"
created: [ISO timestamp]
updated: [ISO timestamp]
---

## Current Focus
<!-- OVERWRITE on each update - always reflects NOW -->

hypothesis: [current theory being tested]
test: [how testing it]
expecting: [what result means if true/false]
next_action: [immediate next step — be specific, not "continue investigating"]
bug_class: null  <!-- assigned at Phase 1.75 — bohrbug|heisenbug-mandelbug|concurrency — routes investigation technique (see gsd-core/references/debugger-bug-taxonomy.md) -->
reasoning_checkpoint: null  <!-- populated before every fix attempt — see structured_returns -->
tdd_checkpoint: null  <!-- populated when tdd_mode is active after root cause confirmed -->

## Symptoms
<!-- Written during gathering, then immutable -->

expected: [what should happen]
actual: [what actually happens]
errors: [error messages if any]
reproduction: [how to trigger]
started: [when it broke / always broken]

## Eliminated
<!-- APPEND only - prevents re-investigating after /clear -->

- hypothesis: [theory that was wrong]
  evidence: [what disproved it]
  timestamp: [when eliminated]

## Evidence
<!-- APPEND only - facts discovered during investigation -->

- timestamp: [when found]
  checked: [what was examined]
  found: [what was observed]
  implication: [what this means]

## Resolution
<!-- OVERWRITE as understanding evolves -->

root_cause: [empty until found — may hold one OR a small set of contributing causes when the AND-gate fires; see gsd-core/references/debugger-rca-branching.md]
fix: [empty until applied]
verification: [empty until verified — holds the nested per-signal fix-acceptance guardrail record (map shape) when active; see gsd-core/references/debugger-fix-acceptance.md]
oracle_type: [empty until the regression test is written — specified|derived|metamorphic|implicit; the assertion's oracle classification per gsd-core/references/debugger-repro-hardening.md]
files_changed: []

<section_rules>

Frontmatter (status, trigger, timestamps):

  • status: OVERWRITE - reflects current phase
  • trigger: IMMUTABLE - verbatim user input, never changes
  • created: IMMUTABLE - set once
  • updated: OVERWRITE - update on every change

Current Focus:

  • OVERWRITE entirely on each update
  • Always reflects what Claude is doing RIGHT NOW
  • If Claude reads this after /clear, it knows exactly where to resume
  • Fields: hypothesis, test, expecting, next_action, reasoning_checkpoint, tdd_checkpoint
  • next_action: must be concrete and actionable — bad: "continue investigating"; good: "Add logging at line 47 of auth.js to observe token value before jwt.verify()"
  • reasoning_checkpoint: OVERWRITE before every fix_and_verify — seven-field structured reasoning record (hypothesis, confirming_evidence, falsification_test, fix_rationale, blind_spots, candidate_causes, and_gate) — see gsd-debugger.md Structured Reasoning Checkpoint
  • tdd_checkpoint: OVERWRITE during TDD red/green phases — test file, name, status, failure output

Symptoms:

  • Written during initial gathering phase
  • IMMUTABLE after gathering complete
  • Reference point for what we're trying to fix
  • Fields: expected, actual, errors, reproduction, started

Eliminated:

  • APPEND only - never remove entries
  • Prevents re-investigating dead ends after context reset
  • Each entry: hypothesis, evidence that disproved it, timestamp
  • Critical for efficiency across /clear boundaries

Evidence:

  • APPEND only - never remove entries
  • Facts discovered during investigation
  • Each entry: timestamp, what checked, what found, implication
  • Builds the case for root cause

Resolution:

  • OVERWRITE as understanding evolves
  • May update multiple times as fixes are tried
  • Final state shows confirmed root cause and verified fix
  • Fields: root_cause, fix, verification, files_changed

</section_rules>

Creation: Immediately when /gsd:debug is called

  • Create file with trigger from user input
  • Set status to "gathering"
  • Current Focus: next_action = "gather symptoms"
  • Symptoms: empty, to be filled

During symptom gathering:

  • Update Symptoms section as user answers questions
  • Update Current Focus with each question
  • When complete: status → "investigating"

During investigation:

  • OVERWRITE Current Focus with each hypothesis
  • APPEND to Evidence with each finding
  • APPEND to Eliminated when hypothesis disproved
  • Update timestamp in frontmatter

During fixing:

  • status → "fixing"
  • Update Resolution.root_cause when confirmed
  • Update Resolution.fix when applied
  • Update Resolution.files_changed

During verification:

  • status → "verifying"
  • Update Resolution.verification with results
  • If verification fails: status → "investigating", try again

After self-verification passes:

  • status -> "awaiting_human_verify"
  • Request explicit user confirmation in a checkpoint
  • Do NOT move file to resolved yet

On resolution:

  • status → "resolved"
  • Move file to .planning/debug/resolved/ (only after user confirms fix)

<resume_behavior>

When Claude reads this file after /clear:

  1. Parse frontmatter → know status
  2. Read Current Focus → know exactly what was happening
  3. Read Eliminated → know what NOT to retry
  4. Read Evidence → know what's been learned
  5. Continue from next_action

The file IS the debugging brain. Claude should be able to resume perfectly from any interruption point.

</resume_behavior>

<size_constraint>

Keep debug files focused:

  • Evidence entries: 1-2 lines each, just the facts
  • Eliminated: brief - hypothesis + why it failed
  • No narrative prose - structured data only

If evidence grows very large (10+ entries), consider whether you're going in circles. Check Eliminated to ensure you're not re-treading.

</size_constraint>