feat(3081): auto-trim review prompts for small-context model reviewers (#3708)
* feat(3081): auto-trim review prompts for small-context model reviewers Adds review.max_prompt_tokens and review.max_prompt_tokens_per_reviewer config keys. When configured, the /gsd-review workflow deterministically trims the assembled prompt before sending to each reviewer (drop CONTEXT → RESEARCH → REQUIREMENTS; head-shrink PROJECT.md; tail-truncate PLANs proportionally; reserve disclosure-note tokens upfront). Trim metadata is recorded in REVIEWS.md frontmatter. Reviewer is skipped with a warning if even the minimum review set exceeds the budget. Closes #3081 * fix(3081): register prompt-budget in SDK query registry and update inventory manifest review.md references `gsd-sdk query prompt-budget` at three call sites, but the command had no handler in the SDK registry — failing the registry-integration drift-guard test on all 6 CI matrix legs. Added a native TypeScript SDK handler (sdk/src/query/prompt-budget.ts) that ports the applyBudget logic from the CJS module, registered it in DOMAIN_STATIC_CATALOG, and regenerated docs/INVENTORY-MANIFEST.json to include the new cli_modules/prompt-budget.cjs entry. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(3081): bump ws to 8.20.1 and allowlist prompt-budget sibling pair Two additional CI failures after the registry fix: 1. ws moderate CVE (GHSA-58qx-3vcg-4xpx, uninitialized memory disclosure): The advisory covers ws >=8.0.0 <8.20.1. Both root and sdk/package.json pinned ^8.20.0 which resolved to 8.20.0. Bumped both to 8.20.1 to clear the npm audit drift-guard test (bug-3588-npm-audit-clean.test.cjs). 2. lint-shared-module-handsync detected the new prompt-budget.ts / prompt-budget.cjs sibling pair without an allowlist entry. Added a cooperatingSiblings entry to scripts/shared-module-handsync-allowlist.json with classification and justification matching the established pattern. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(3081): align prompt-budget skip semantics across CJS and SDK dispatch paths Replace brittle `[ $EXIT -eq 2 ]` guards with `[ $EXIT -ne 0 ]` in all three local-reviewer blocks (Ollama, LM Studio, llama.cpp) in workflows/review.md. Any non-zero exit from prompt-budget now triggers a skip with a descriptive warning — exit 2/11 prints "budget too small", any other non-zero prints "unexpected exit code". This ensures the SDK bridge dispatch path (exit 11 via GSDError(Blocked)) triggers the same skip as the CJS path (exit 2). The SDK handler (sdk/src/query/prompt-budget.ts) already writes both metadata and prompt files before throwing, so no change needed there. The Ollama block also gains the missing OLLAMA_SKIP guard so the reviewer invocation is actually skipped (previously the block only suppressed the OLLAMA_PROMPT_FILE update but still ran the curl invocation). SDK integration path (hardFailed via GSDError(Blocked) → exit 11) is covered by handler unit tests in tests/prompt-budget.test.cjs; no gsd-sdk-*.test.cjs exercising the full bridge dispatch for this command exists yet — that gap remains and is documented here. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix prompt-budget trim ordering and review guard follow-ups * perf: optimize prompt-budget and dedup reviewer trim workflow * fix(3708): drop source-grep theater tests to satisfy lint-no-source-grep All four test files added in commit 2df566ed were pure source-grep theater: they read .cjs / .ts / .md source files and asserted that specific string literals were present or absent. None exercised runtime behaviour. Deleted: - tests/gsd-tools-memory-optimizer.test.cjs — 7 includes() on gsd-tools.cjs - tests/prompt-budget-hotpath-optimizer.test.cjs — includes() on prompt-budget.cjs + .ts - tests/prompt-budget-io-optimizer.test.cjs — includes() on prompt-budget.ts + gsd-tools.cjs - tests/review-workflow-budget-dedup.test.cjs — includes() on review.md Behavioural coverage for the prompt-budget feature already exists in tests/prompt-budget.test.cjs and tests/prompt-budget-cli.test.cjs (also added by this PR). No replacement tests needed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(3708): correct budget-pressure threshold and minSet accounting Two bugs in applyBudget caused premature trimming and false hard-fails: 1. UNNEEDED_TRIM: budgetUnderPressure compared baseTokens against effectiveBudget - NOTE_RESERVE_TOKENS, triggering trim pressure 80 tokens before the budget was actually exceeded. Fix: compare against effectiveBudget directly; NOTE_RESERVE_TOKENS are still reserved in contentBudget once real pressure is confirmed. 2. FALSE_HARDFAIL: minSet included NOTE_RESERVE_TOKENS unconditionally, treating the note as mandatory even when no trim would occur and no note would be injected. Fix: exclude NOTE_RESERVE_TOKENS from minSet; a prompt that fits untrimmed needs no note and must not hard-fail. Both fixes applied in CJS and TypeScript implementations. Two regression tests added (cycles 11 and 12) that reproduce each case behaviorally. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -700,10 +700,16 @@ Configure per-CLI model selection for `/gsd-review`. When set, overrides the CLI
|
||||
| `review.models.lm_studio` | string | (server default) | Model name passed to LM Studio when `--lm-studio` reviewer is invoked. If unset, the first available model reported by the server is used. |
|
||||
| `review.models.llama_cpp` | string | (server default) | Model name passed to llama.cpp when `--llama-cpp` reviewer is invoked. If unset, the first model reported by `/v1/models` is used. |
|
||||
| `review.default_reviewers` | string[] \| null | (all detected reviewers) | Default reviewer subset for no-flag `/gsd-review`. Example: `["gemini","codex"]`. Explicit flags and `--all` override this setting. |
|
||||
| `review.max_prompt_tokens` | number\|null | null | Default maximum estimated tokens for the assembled review prompt. When set, the prompt is deterministically trimmed before being sent to each reviewer. Per-reviewer overrides via `review.max_prompt_tokens_per_reviewer` take precedence. null = no trim (current behavior). |
|
||||
| `review.max_prompt_tokens_per_reviewer` | object | {} | Per-reviewer token budget overrides. Keys are reviewer slugs (ollama, llama_cpp, lm_studio, gemini, claude, codex, opencode, qwen, cursor). Values override `review.max_prompt_tokens` for that reviewer. Recommended for local model servers. |
|
||||
| `review.ollama_host` | string | `http://localhost:11434` | Base URL of the Ollama server. Override when running Ollama on a non-default port or remote host: `gsd config-set review.ollama_host http://192.168.1.10:11434` |
|
||||
| `review.lm_studio_host` | string | `http://localhost:1234` | Base URL of the LM Studio local server. Override when using a non-default port. |
|
||||
| `review.llama_cpp_host` | string | `http://localhost:8080` | Base URL of the llama.cpp server (`llama-server`). Override when using a non-default port. |
|
||||
|
||||
### Prompt budgets for small-context reviewers
|
||||
|
||||
Local model servers (Ollama, llama.cpp, LM Studio) typically accept far fewer tokens than cloud APIs. Setting `review.max_prompt_tokens_per_reviewer` (or the global `review.max_prompt_tokens` fallback) triggers deterministic prompt trimming before the prompt is sent to that reviewer: CONTEXT is dropped first, then RESEARCH, then REQUIREMENTS; PROJECT.md is head-shrunk to the first 40 lines; PLANs are tail-truncated proportionally — instructions and roadmap are always preserved. When a reviewer is trimmed, a disclosure note is injected at the top of the prompt and trim metadata (budget, omitted sections, truncation percentage) is recorded in the REVIEWS.md frontmatter under `trimmed_reviewers`. If even the minimum review set (instructions + roadmap + plan stubs) exceeds the budget, the reviewer is skipped with a warning rather than sending a truncated prompt that would produce misleading feedback.
|
||||
|
||||
### Example
|
||||
|
||||
```json
|
||||
|
||||
@@ -1186,6 +1186,7 @@ When verification returns `human_needed`, items are persisted as a trackable HUM
|
||||
**User configuration note:**
|
||||
- Set `review.default_reviewers` in `.planning/config.json` (or via `gsd config-set`) to control no-flag `/gsd-review` fan-out.
|
||||
- Use `--all` for a full pre-merge sweep without changing project defaults.
|
||||
- For local model servers with small context windows, set `review.max_prompt_tokens_per_reviewer` to auto-trim prompts per reviewer — see [Prompt budgets for small-context reviewers](../docs/CONFIGURATION.md#prompt-budgets-for-small-context-reviewers) in CONFIGURATION.md.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"generated": "2026-05-17",
|
||||
"generated": "2026-05-18",
|
||||
"families": {
|
||||
"agents": [
|
||||
"gsd-advisor-researcher",
|
||||
@@ -302,6 +302,7 @@
|
||||
"profile-output.cjs",
|
||||
"profile-pipeline.cjs",
|
||||
"project-root.generated.cjs",
|
||||
"prompt-budget.cjs",
|
||||
"review-reviewer-selection.cjs",
|
||||
"roadmap-command-router.cjs",
|
||||
"roadmap.cjs",
|
||||
|
||||
@@ -361,7 +361,7 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
|
||||
|
||||
---
|
||||
|
||||
## CLI Modules (71 shipped)
|
||||
## CLI Modules (72 shipped)
|
||||
|
||||
Full listing: `get-shit-done/bin/lib/*.cjs`.
|
||||
|
||||
@@ -410,6 +410,7 @@ Full listing: `get-shit-done/bin/lib/*.cjs`.
|
||||
| `project-root.generated.cjs` | GENERATED — CJS artifact emitted from `sdk/src/project-root/index.ts` via `sdk/scripts/gen-project-root.mjs`; resolves a project root from a starting directory using four heuristics (own `.planning/` guard, `sub_repos` config, `multiRepo` flag, `.git` heuristic); do not edit directly |
|
||||
| `profile-output.cjs` | Profile rendering, USER-PROFILE.md and dev-preferences.md generation |
|
||||
| `profile-pipeline.cjs` | User behavioral profiling data pipeline, session file scanning |
|
||||
| `prompt-budget.cjs` | Pure token-budget accounting for review prompts — estimates tokens, applies deterministic trim priority (head-shrink PROJECT.md, proportional plan truncation, drop context/research/requirements, hard-fail guard), returns structured metadata for `review.max_prompt_tokens` (#3081) |
|
||||
| `review-reviewer-selection.cjs` | Reviewer selection/normalization helpers for `/gsd-review` default reviewer policy and precedence |
|
||||
| `roadmap-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools roadmap` |
|
||||
| `roadmap.cjs` | ROADMAP.md parsing, phase extraction, plan progress |
|
||||
|
||||
Reference in New Issue
Block a user