Files
msd-core/gsd-core/references/debugger-techniques.md
Tom Boucher ed360cd99f chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point

Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers
only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim
into every runtime — and agent text is loaded into a subagent's context on every
dispatch.

The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op
for agents: agents never traverse that function. Agent content is read for
emission at five independent points, and the obvious chokepoint
`stageAgentsForProfile` short-circuits on the DEFAULT `full` profile
(`skills === '*'` returns the real unstaged directory), so a hook placed there is
dead code on most installs.

Composition now happens at two call sites instead of five parallel surfaces:
`stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind`
routed through it via an identity converter) and the inline agent loop in
bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` ->
`.windsurf/` regex can never reach inside a marker attribute — the ordering
#2930 established for workflows.

`installCodexConfig` was the fifth read point: Codex embeds each agent's prompt
into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it;
the exhaustive per-runtime emission sweep found it. That is why the new guard is
behavioral rather than structural — a sixth read point fails the sweep without
anyone remembering to extend a list.

tests/agent-fragments-emission.install.test.cjs spawns a real installer for every
runtime at every agent-bearing scope, derived from RUNTIME_META and the
capability registry at run time so a new runtime cannot be silently
under-covered. It asserts markers are absent AND the `when="always"` body is
retained, so marker-absence cannot be satisfied by dropping content. An
identity-composer negative control proves the assertion can fail.

Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope
paths; red before the wiring on claude(global+local), zcode(global+local),
kimi, codex and opencode.

Refs #2995

* chore(#2995): give the tightest agents headroom and correct the design lock

Epic #1671 Phase 6.4, second half.

`agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE
cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now
extract reference material to `gsd-core/references/` behind an @-reference — the
documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy:

  gsd-verifier  49,140 -> 46,371 B   headroom    12 -> 2,781
  gsd-debugger  57,197 -> 48,851 B   headroom   147 -> 8,493

Byte accounting proves no content was lost: the combined agent+reference delta
is exactly the new files' headers plus the agents' slim replacement blocks. Each
agent keeps its routing table and a one-line summary per entry, so it degrades
gracefully on a runtime that does not inline @-references.

`agents/gsd-planner.md` is untouched and still passes both char guards
(49,130 < 49,152); it needed no change, so it took none.

The other nine LARGE/XL agents carry NO gsd:section markers, and that is
deliberate, not deferred. `when=` selection is read from
gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives
from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key,
no per-agent init entry point. An agent atom therefore fails admission gate (2)
("a fact the init seam demonstrably computes at a real entry point") and would
evaluate false forever while looking like working gating. Marking agents would
manufacture exactly the silent-inertness rot the frozen vocabulary exists to
prevent.

ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage
found against what actually merged:

  - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR
    amendment, which that bullet's own rule forbids. Recorded now.
  - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to
    "the LARGE/XL rollout phase". Five shipped; this one is permanently
    rejected, and that disposition lived only in a merged PR body.
  - Phase 6.4's own finding: emission extends to agents/, gating does not.

CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still
listed the original 4-atom vocabulary and described when= as "not yet acted on",
and Section Manifest Module still described InvocationFacts as
{waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract.

Inventory manifest regenerated AFTER build:lib per the documented ordering
landmine; 19 install-tree fixtures pick up the two new references.

Refs #2995

* chore(#2995): correct the compose-site count and mark the raw stager

Self-review found two comment defects in the prior commit. The agentsKind
comment claimed composition lands at TWO call sites; it is three, since
installCodexConfig's per-agent .toml writer was added after that comment was
written. And stageAgentsForProfile is now production-dead — both callers route
through the composing stager — while staying exported and unit-tested, which
makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged
source directory under the default profile, so a future caller would silently
reintroduce the marker-shipping path. Its JSDoc now says so.

* test(#2995): guard the marker-documenting-doc class for agents

Widening the composer's scope to agents/ makes reachable the exact class #2930
narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an
unfenced example is indistinguishable from a real marker, so the composer drops
that line from the emitted artifact.

Three rows. A fenced example must compose byte-identically. No shipped agent may
carry a marker outside a fence — asserted by parsing every real agent and
requiring zero explicit sections, which is what makes the fence protection
load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED
marker IS parsed as a real marker, so if that ever stops being true the second
row is guarding nothing.

Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it
had no production caller, which is false — bin/install.js's _stageAgents still
calls it, and its consumers compose before writing. Corrected to state the
invariant instead. And a let/const nit in the emission sweep.

* fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture

The first remote run came back red with three failures. Both root causes were
mine.

1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally
   contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status
   vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent
   no longer had it.

   Byte accounting said no content was lost, and byte-wise that was true — but a
   contract required those tokens to live IN THE AGENT. That is ADR-1671:66's
   flexReserve floor stated concretely: a load-bearing fragment must not be
   trimmed out of its host, and "the bytes still exist somewhere" is not the
   test. The two status tables are restored to the agent and deliberately
   mirrored in the reference with a note saying so, so the procedure there still
   reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083,
   rather than the 2,781 the first attempt claimed.

2. Row 12b of the new marker-documentation guard asserted that an unfenced
   marker example parses as a real marker, and threw instead:
   "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The
   fixture had put the OPEN marker inline mid-sentence, so it was correctly not
   recognised as an open while the close, on its own line, was.

   That is a real refinement of the hazard this guard exists for: only a marker
   on its OWN line is mis-parsed — which is exactly how a documentation example
   is normally written. Row 12b now uses a whole-line marker, and a new row 12c
   pins the inline case as explicitly NOT a marker.

No test was weakened to accommodate the change; the change was corrected to
satisfy the tests.

Refs #2995

* chore(#2995): backfill changeset pr number to 3058

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:10:31 -04:00

9.3 KiB

Debugger technique catalog

Full technique bodies for agents/gsd-debugger.md, extracted per DEFECT.AGENT-FILE-SIZE-CAP-BREACH (issue #2995, epic #1671 Phase 6.4). The agent keeps each technique's name and routing entry; the step-by-step detail lives here.

Binary Search / Divide and Conquer

When: Large codebase, long execution path, many possible failure points.

How: Cut problem space in half repeatedly until you isolate the issue.

  1. Identify boundaries (where works, where fails)
  2. Add logging/testing at midpoint
  3. Determine which half contains the bug
  4. Repeat until you find exact line

Example: API returns wrong data

  • Test: Data leaves database correctly? YES
  • Test: Data reaches frontend correctly? NO
  • Test: Data leaves API route correctly? YES
  • Test: Data survives serialization? NO
  • Found: Bug in serialization layer (4 tests eliminated 90% of code)

Rubber Duck Debugging

When: Stuck, confused, mental model doesn't match reality.

How: Explain the problem out loud in complete detail.

Write or say:

  1. "The system should do X"
  2. "Instead it does Y"
  3. "I think this is because Z"
  4. "The code path is: A -> B -> C -> D"
  5. "I've verified that..." (list what you tested)
  6. "I'm assuming that..." (list assumptions)

Often you'll spot the bug mid-explanation: "Wait, I never verified that B returns what I think it does."

Delta Debugging

When: Large change set is suspected (many commits, a big refactor, or a complex feature that broke something). Also when "comment out everything" is too slow.

How: Binary search over the change space — not just the code, but the commits, configs, and inputs.

Over commits (use git bisect): Already covered under Git Bisect. But delta debugging extends it: after finding the breaking commit, delta-debug the commit itself — identify which of its N changed files/lines actually causes the failure.

Over code (systematic elimination):

  1. Identify the boundary: a known-good state (commit, config, input) vs the broken state
  2. List all differences between good and bad states
  3. Split the differences in half. Apply only half to the good state.
  4. If broken: bug is in the applied half. If not: bug is in the other half.
  5. Repeat until you have the minimal change set that causes the failure.

Over inputs:

  1. Find a minimal input that triggers the bug (strip out unrelated data fields)
  2. The minimal input reveals which code path is exercised

When to use:

  • "This worked yesterday, something changed" → delta debug commits
  • "Works with small data, fails with real data" → delta debug inputs
  • "Works without this config change, fails with it" → delta debug config diff

Example: 40-file commit introduces bug

Split into two 20-file halves.
Apply first 20: still works → bug in second half.
Split second half into 10+10.
Apply first 10: broken → bug in first 10.
... 6 splits later: single file isolated.

Minimal Reproduction

When: Complex system, many moving parts, unclear which part fails.

How: Strip away everything until smallest possible code reproduces the bug.

  1. Copy failing code to new file
  2. Remove one piece (dependency, function, feature)
  3. Test: Does it still reproduce? YES = keep removed. NO = put back.
  4. Repeat until bare minimum
  5. Bug is now obvious in stripped-down code
  6. Shrinking (input-space bugs) — when the bug triggers on a class of inputs, wrap it in a property (fast-check for JS/TS, Hypothesis for Python) and let the shrinker auto-minimize the counterexample; store the minimized input as the regression seed. See gsd-core/references/debugger-repro-hardening.md.

Example:

// Start: 500-line React component with 15 props, 8 hooks, 3 contexts
// End after stripping:
function MinimalRepro() {
  const [count, setCount] = useState(0);

  useEffect(() => {
    setCount(count + 1); // Bug: infinite loop, missing dependency array
  });

  return <div>{count}</div>;
}
// The bug was hidden in complexity. Minimal reproduction made it obvious.

Working Backwards

When: You know correct output, don't know why you're not getting it.

How: Start from desired end state, trace backwards.

  1. Define desired output precisely
  2. What function produces this output?
  3. Test that function with expected input - does it produce correct output?
    • YES: Bug is earlier (wrong input)
    • NO: Bug is here
  4. Repeat backwards through call stack
  5. Find divergence point (where expected vs actual first differ)

Example: UI shows "User not found" when user exists

Trace backwards:
1. UI displays: user.error → Is this the right value to display? YES
2. Component receives: user.error = "User not found" → Correct? NO, should be null
3. API returns: { error: "User not found" } → Why?
4. Database query: SELECT * FROM users WHERE id = 'undefined' → AH!
5. FOUND: User ID is 'undefined' (string) instead of a number

Differential Debugging

When: Something used to work and now doesn't. Works in one environment but not another.

Time-based (worked, now doesn't):

  • What changed in code since it worked?
  • What changed in environment? (Node version, OS, dependencies)
  • What changed in data?
  • What changed in configuration?

Environment-based (works in dev, fails in prod):

  • Configuration values
  • Environment variables
  • Network conditions (latency, reliability)
  • Data volume
  • Third-party service behavior

Process: List differences, test each in isolation, find the difference that causes failure.

Example: Works locally, fails in CI

Differences:
- Node version: Same ✓
- Environment variables: Same ✓
- Timezone: Different! ✗

Test: Set local timezone to UTC (like CI)
Result: Now fails locally too
FOUND: Date comparison logic assumes local timezone

Observability First

When: Always. Before making any fix.

Add visibility before changing behavior:

// Strategic logging (useful):
console.log('[handleSubmit] Input:', { email, password: '***' });
console.log('[handleSubmit] Validation result:', validationResult);
console.log('[handleSubmit] API response:', response);

// Assertion checks:
console.assert(user !== null, 'User is null!');
console.assert(user.id !== undefined, 'User ID is undefined!');

// Timing measurements:
console.time('Database query');
const result = await db.query(sql);
console.timeEnd('Database query');

// Stack traces at key points:
console.log('[updateUser] Called from:', new Error().stack);

Workflow: Add logging -> Run code -> Observe output -> Form hypothesis -> Then make changes.

Comment Out Everything

When: Many possible interactions, unclear which code causes issue.

How:

  1. Comment out everything in function/file
  2. Verify bug is gone
  3. Uncomment one piece at a time
  4. After each uncomment, test
  5. When bug returns, you found the culprit

Example: Some middleware breaks requests, but you have 8 middleware functions

app.use(helmet()); // Uncomment, test → works
app.use(cors()); // Uncomment, test → works
app.use(compression()); // Uncomment, test → works
app.use(bodyParser.json({ limit: '50mb' })); // Uncomment, test → BREAKS
// FOUND: Body size limit too high causes memory issues

Git Bisect

When: Feature worked in past, broke at unknown commit.

How: Binary search through git history.

git bisect start
git bisect bad              # Current commit is broken
git bisect good abc123      # This commit worked
# Git checks out middle commit
git bisect bad              # or good, based on testing
# Repeat until culprit found

100 commits between working and broken: ~7 tests to find exact breaking commit.

Follow the Indirection

When: Code constructs paths, URLs, keys, or references from variables — and the constructed value might not point where you expect.

The trap: You read code that builds a path like path.join(configDir, 'hooks') and assume it's correct because it looks reasonable. But you never verified that the constructed path matches where another part of the system actually writes/reads.

How:

  1. Find the code that produces the value (writer/installer/creator)
  2. Find the code that consumes the value (reader/checker/validator)
  3. Trace the actual resolved value in both — do they agree?
  4. Check every variable in the path construction — where does each come from? What's its actual value at runtime?

Common indirection bugs:

  • Path A writes to dir/sub/hooks/ but Path B checks dir/hooks/ (directory mismatch)
  • Config value comes from cache/template that wasn't updated
  • Variable is derived differently in two places (e.g., one adds a subdirectory, the other doesn't)
  • Template placeholder ({{VERSION}}) not substituted in all code paths

Example: Stale hook warning persists after update

Check code says:  hooksDir = path.join(configDir, 'hooks')
                  configDir = ~/.claude
                  → checks ~/.claude/hooks/

Installer says:   hooksDest = path.join(targetDir, 'hooks')
                  targetDir = ~/.claude/gsd-core
                  → writes to ~/.claude/gsd-core/hooks/

MISMATCH: Checker looks in wrong directory → hooks "not found" → reported as stale

The discipline: Never assume a constructed path is correct. Resolve it to its actual value and verify the other side agrees. When two systems share a resource (file, directory, key), trace the full path in both.