Files
msd-core/get-shit-done/references/context-budget.md
Tom Boucher 8c43ba7301 docs(#3025): MCP tool schema as a context-budget concern (#3032)
* docs(#3025): MCP tool schema as a context-budget concern

Adds documentation covering the largest GSD cost lever that GSD
itself does not own: MCP tool schema injection. Every enabled MCP
server adds its schema to every turn (often 20k+ tokens for
heavyweight servers like browser/playwright, mac-tools, etc.),
which can dwarf whatever `model_profile` tuning saves.

Two doc surfaces (per the bundle-docs-with-code skill depth gradient):

1. get-shit-done/references/context-budget.md
   - New "MCP Tool Schema Cost (Harness Concern)" section.
   - Explains schemas-per-turn cost framing.
   - Names enabledMcpjsonServers / disabledMcpjsonServers and
     .claude/settings.json explicitly.
   - Pre-phase audit checklist: browser/playwright, platform-specific,
     cross-project/stale, duplicate/shadow.
   - Explicit "GSD does not manage MCP enablement — harness concern"
     statement so users don't hunt for a GSD setting.
   - Links to Anthropic Claude Code MCP docs as canonical reference.
   - Notes compounding interaction with model_profile (additive levers).

2. docs/USER-GUIDE.md
   - New task-oriented "Trim MCP servers to reduce per-turn cost"
     section above "Using Non-Claude Runtimes".
   - Same checklist condensed.
   - Cross-link to context-budget.md for the full reference.

Tests:
- tests/feat-3025-mcp-token-budget-docs.test.cjs (12 cases) parses
  both docs into typed semantic-flag records and asserts behavioral
  invariants (mentions key, includes audit, names harness, etc.)
  rather than substring-matching prose. Adheres to CONTRIBUTING.md
  no-source-grep — section can be reworded freely as long as the
  required semantics survive.
- Markdownlint pre-flight tests (MD040 fence language, MD056 table
  column count) per the bundle-docs-with-code skill so CR can't
  ratchet on prose nitpicks across multiple review rounds.

Verification:
- 12/12 pass on regression test
- 6857/6857 full suite (12 net new)
- lint-no-source-grep clean (377 test files)

Companion to #3023 (per-phase-type model map) and #3024 (dynamic
routing). Together they cover the three biggest cost levers users
ask about; this issue covers the one GSD does not own.

Closes #3025

* docs(#3025): batch 3 CR fixes — pr id, relative link, named flag

CodeRabbit on PR #3032 (3 minor — 2 inline + 1 nitpick), all in one
push per the bundle-docs-with-code skill (avoid per-round nitpick
ratchet):

1. Inline (Minor) — .changeset/mcp-token-budget-docs.md:3
   `pr: TBD` → `pr: 3032` so changeset tooling can link the entry.

2. Inline (Minor) — docs/USER-GUIDE.md:1101
   Used a hardcoded `https://github.com/.../blob/main/...` URL for the
   cross-link to `context-budget.md`. Rest of USER-GUIDE.md uses
   relative links. Switched to `../get-shit-done/references/context-
   budget.md#mcp-tool-schema-cost-harness-concern` so feature-branch
   work shows the right content and rename-resilience is preserved.

3. Nitpick — tests/feat-3025-mcp-token-budget-docs.test.cjs:234
   The cross-link assertion used an inline `/context-budget/i.test(...)`
   while every other invariant in the file lived as a named flag in
   `parseMcpBudgetSection`. Per CONTRIBUTING.md no-source-grep, added
   `crossLinksContextBudget` to the parser and asserted on
   `parsed.crossLinksContextBudget` so the cross-link rule sits next
   to its siblings.

Verification:
- 12/12 pass on regression test (no count change; refactor only)
- No source code changes, only docs + tests

* test(#3025): strip inline markdown before phrase-match (CR nitpick)

CodeRabbit caught that the `explainsHarnessNotGsd` primary regex
branch couldn't match "GSD does **not** manage" in
context-budget.md because the markdown bold markers (`**`) sit
between contiguous words. The test passed today only via the
fallback `harness (concern|setting|controlled)` branch — the
primary branch was effectively dead code.

Fix: strip inline markdown emphasis (`**`, `*`, `~~`) and inline-
code backticks before any phrase-matching in `parseMcpBudgetSection`.
All seven flag computations now run against the stripped text so
markdown formatting can't silently invalidate any invariant.

Underscores are intentionally NOT stripped — `model_profile` and
other snake_case identifiers must survive intact for the
mentionsModelProfileInteraction check to find them.

Verification: 12/12 still pass; primary branches now fire on
real markdown content rather than relying on fallbacks.

* test(#3025): guard markdownlint tests against null section (CR nitpick)

CodeRabbit caught that the MD040 and MD056 markdownlint pre-flight
tests called `section.match(...)` and `section.split('\n')`
directly on the value returned by `extractSection`, which returns
null when no matching header is found. If the MCP section is ever
removed (regression), both tests would throw `TypeError: Cannot
read properties of null` instead of producing a clean assertion
failure naming the actual problem.

The semantic tests above are protected because parseMcpBudgetSection
short-circuits to a typed-falsy record on null input. The
markdownlint tests bypassed that guard since they need raw section
text, not parsed flags. Added `assert.ok(section, ...)` preconditions
to both so a missing section produces a meaningful failure message.

No content changes; defensive programming only.

Verification: 12/12 still pass.
2026-05-02 15:24:26 -04:00

6.1 KiB

Context Budget Rules

Standard rules for keeping orchestrator context lean. Reference this in workflows that spawn subagents or read significant content.

See also: references/universal-anti-patterns.md for the complete set of universal rules.


Universal Rules

Every workflow that spawns agents or reads significant content must follow these rules:

  1. Never read agent definition files (agents/*.md) -- subagent_type auto-loads them
  2. Never inline large files into subagent prompts -- tell agents to read files from disk instead
  3. Read depth scales with context window -- check context_window in .planning/config.json:
    • At < 500000 tokens (default 200k): read only frontmatter, status fields, or summaries. Never read full SUMMARY.md, VERIFICATION.md, or RESEARCH.md bodies.
    • At >= 500000 tokens (1M model): MAY read full subagent output bodies when the content is needed for inline presentation or decision-making. Still avoid unnecessary reads.
  4. Delegate heavy work to subagents -- the orchestrator routes, it doesn't execute
  5. Proactive warning: If you've already consumed significant context (large file reads, multiple subagent results), warn the user: "Context budget is getting heavy. Consider checkpointing progress."

Read Depth by Context Window

Context Window Subagent Output Reading SUMMARY.md VERIFICATION.md PLAN.md (other phases)
< 500k (200k model) Frontmatter only Frontmatter only Frontmatter only Current phase only
>= 500k (1M model) Full body permitted Full body permitted Full body permitted Current phase only

How to check: Read .planning/config.json and inspect context_window. If the field is absent, treat as 200k (conservative default).

Context Degradation Tiers

Monitor context usage and adjust behavior accordingly:

Tier Usage Behavior
PEAK 0-30% Full operations. Read bodies, spawn multiple agents, inline results.
GOOD 30-50% Normal operations. Prefer frontmatter reads, delegate aggressively.
DEGRADING 50-70% Economize. Frontmatter-only reads, minimal inlining, warn user about budget.
POOR 70%+ Emergency mode. Checkpoint progress immediately. No new reads unless critical.

Context Degradation Warning Signs

Quality degrades gradually before panic thresholds fire. Watch for these early signals:

  • Silent partial completion -- agent claims task is done but implementation is incomplete. Self-check catches file existence but not semantic completeness. Always verify agent output meets the plan's must_haves, not just that files exist.
  • Increasing vagueness -- agent starts using phrases like "appropriate handling" or "standard patterns" instead of specific code. This indicates context pressure even before budget warnings fire.
  • Skipped steps -- agent omits protocol steps it would normally follow. If an agent's success criteria has 8 items but it only reports 5, suspect context pressure.

When delegating to agents, the orchestrator cannot verify semantic correctness of agent output -- only structural completeness. This is a fundamental limitation. Mitigate with must_haves.truths and spot-check verification.

MCP Tool Schema Cost (Harness Concern)

Every enabled MCP server injects its tool schema into every turn, regardless of whether you call any of its tools. Heavyweight servers can cost 20k+ tokens per turn each — often dwarfing whatever GSD itself can save through model_profile tuning. This is a Claude Code harness concern, not a GSD concern: GSD does not manage MCP enablement. The toggle lives in .claude/settings.json under enabledMcpjsonServers and disabledMcpjsonServers.

Why this is the biggest cost lever you don't own

Tool schemas count against the same context budget as model context, prompts, and conversation history. If a project has 5 unused MCP servers averaging 5k tokens of schema each, every turn pays a 25k-token tax before the assistant reads a single project file. Trimming MCPs has a multiplier effect that compounds with whichever model_profile you've chosen — every-turn overhead drops regardless of which model is in use.

Pre-Phase MCP Audit

Before starting a long phase (especially /gsd-execute-phase, /gsd-plan-phase, or anything that fans out across many subagents), run this audit:

  • Browser / playwright tools enabled? If this phase has no UI work, disable them. They're among the heaviest per-turn schemas.
  • Platform-specific tools enabled? Mac-tools / Windows-tools / OS-specific helpers should be disabled when not actively needed for the phase at hand.
  • Cross-project / stale MCPs? Servers added for a different project that are still enabled here. These are often forgotten and pay a per-turn tax for zero benefit.
  • Duplicate or shadow servers? Two MCPs offering similar tools (e.g. two different filesystem helpers). Keep one.

Each item disabled removes its schema from every subsequent turn for the rest of the session.

How to toggle

The keys live in .claude/settings.json (project) or ~/.claude/settings.json (global) — not in .planning/config.json:

{
  "enabledMcpjsonServers": ["context7"],
  "disabledMcpjsonServers": ["playwright", "mac-tools"]
}

Either list works — enabledMcpjsonServers is an explicit allow-list, disabledMcpjsonServers is a block-list against the default. See the Claude Code MCP documentation for the canonical reference; this section just flags it as a context-budget lever GSD users routinely overlook.

Composition with model_profile

Trimming MCPs and tuning model_profile are independent levers that compound. Disabling a 25k-token MCP saves 25k per turn whether you're running quality (opus everywhere) or budget (sonnet/haiku); the savings are additive, not in lieu of model tuning. Don't pick one — do both, and audit MCPs first because the per-turn savings show up immediately and stack across every subagent the orchestrator spawns.