enhance(#4139): Phase 4 — measure the window instead of asserting it (#4502)

* enhance(#4404): add offline token benchmark for compact-content splits

ADR-4139 Decision 2 requires the finite-attention justification for
workflow.compact_content to be measured, not asserted. `npm run
benchmark:compact-content` computes, per registered spine/detail split
discovered under gsd-core/workflows/, the token count with the split
active (spine alone) vs inactive (spine + all detail parts read back
in), using gpt-tokenizer (pinned exact devDependency — Anthropic
publishes no tokenizer for Claude 3+, so every output surface labels
this a PROXY-TOKENIZER comparison: the on/off delta is exact under one
tokenizer applied identically to both sides, the absolute counts are
not Claude's real ones).

Reporting-only by design and verified so: --check diffs the live
recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json)
and prints drift, but never exits non-zero for a drifted or missing
baseline — the only thing allowed to fail this script is a genuine I/O
error reading a source .md file it's measuring. Not wired into lint:ci
or pretest.

Discovery is deliberately reimplemented rather than importing
tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a
scripts/ reporting tool from depending on a test-only module.

tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to
prove the benchmark makes no network call, monkeypatching http/https/
net/dns/fetch to throw rather than relying on sandboxing.

docs/CONFIGURATION.md documents the new benchmark against the
workflow.compact_content key to satisfy this repo's docs-required gate
for an Added-type changeset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4404): address orthogonal review findings on the token benchmark

Standards axis found two hard violations against documented rules:
- CLAUDE.md's Generative Fix Divergence rule requires a parity assertion
  for shared discovery logic maintained in two places. Added a test
  comparing benchmark-compact-content.cjs's own discoverRegisteredSplits
  against tests/helpers/compact-content-split.cjs's version on the real
  repo tree, so the two can never silently drift apart.
- The changeset body closed its bold span with a period and continued
  as a second sentence, instead of the canonical
  `**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents.

Spec axis found the "network disabled + identical output across two
runs" Done-when criterion was verified as two separate properties
(determinism tested without network denial, offline survival tested as
a single run) rather than as one combined property. Added a test that
runs the benchmark twice under the deny-network preload and asserts
byte-identical stdout.

Security axis found tests/fixtures/deny-network.cjs didn't patch
dns.promises (a separate binding from the callback dns API), tls.connect,
or http2.connect — inert today since nothing in the benchmark calls
them, but a silent gap in what the preload's own header claims to
guarantee. Patched all three.

CLAUDE.md's Property-Based Testing rule also requires a fast-check test
for budget-limit arithmetic; added one for computeAggregate's off/on
summation (true sum over N splits, never NaN/Infinity, never exceeds
100% when off >= on for every split).

Standards axis's remaining two findings (a Data Clumps observation on
the {offTokens, onTokens, reductionPct} triple, and mild duplication in
formatDriftReport's three line-formatters) are left as judgement calls:
introducing a named type for a 3-field local tuple, or a formatter
abstraction for three short lines, would be exactly the premature
abstraction CLAUDE.md's engineering guidance warns against for a script
this size.

All changes verified directly (parity logic, the fast-check property,
and the three newly-denied network surfaces actually throwing under the
preload) via node -e before committing; full npm run lint:ci passes
with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4404): backfill changeset pr number to 4502

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-09-07 19:19:03 -04:00
committed by GitHub
parent 8dcdcb253e
commit 93e141a006
8 changed files with 913 additions and 1 deletions

View File

@@ -517,7 +517,7 @@ All workflow toggles follow the **absent = enabled** pattern. If a key is missin
| `workflow.text_mode` | boolean | `false` | Replaces AskUserQuestion TUI menus with plain-text numbered lists. Required for Claude Code remote sessions (`/rc` mode) where TUI menus don't render. Can also be set per-session with `--text` flag on discuss-phase. Added in v1.28 |
| `workflow.use_worktrees` | boolean | `true` | When `false`, disables git worktree isolation for parallel execution. Users who prefer sequential execution or whose environment does not support worktrees can disable this. Added in v1.31. **Branch-divergence note:** when your branch has diverged from `origin/HEAD`, GSD auto-degrades to sequential and prints a warning. See [`worktree.baseRef`](#worktree-settings) to restore parallel execution on a diverged branch. **Per-runtime note:** whether this key can be honored depends on the runtime's declared `dispatch.isolation` capability, not on its name (#2584). Runtimes whose own harness isolates each executor (**Claude Code**, **Cursor**) run parallel worktrees natively; runtimes exposing a headless exec with an explicit working directory (**Codex**, **OpenCode**, **Kimi**, **Kimi Code**) get worktrees GSD itself creates and merges — where a dispatch site can only drive the harness model, those hosts degrade to sequential with a warning rather than aborting. Every other runtime declares no isolation primitive, and forcing `use_worktrees: true` there still fails closed before any executor dispatch. `/gsd-health` reports such a value as warning `W025` (#2486). **Default on a non-Claude install:** if a worktree-capable non-Claude host is not isolating as described above, check whether the install stamped this key's default to `false` and set an explicit `use_worktrees: true`. See [Executor isolation per runtime](#executor-isolation-per-runtime). |
| `workflow.agent_hint_routing` | boolean | `true` | Per-plan specialist executor routing (#1689). When `true`, a plan whose `agent_hint:` frontmatter names a subagent that resolves on the active runtime is dispatched to that specialist instead of `gsd-executor`. Default `true` — a no-op for plans without `agent_hint:`, so existing dispatch is unchanged. Set `false` to disable. See [PLAN.md `agent_hint`](reference/plan-md.md#per-plan-executor-routing). |
| `workflow.compact_content` | boolean | `false` | Compact content mode (#4139, [ADR-4139](adr/4139-compact-content-seam.md)). Per-project boolean selecting the terser form of GSD's own shipped prompt content (workflows, templates, agent-skill payloads). `/gsd-plan-phase` is the pilot workflow that actually branches on it today (#4402) — with the key off, its spine reads a deferred elaboration file back in before continuing (byte-identical instruction set to before); with it on, that read is skipped. The rest of the corpus does not branch on it yet — full coverage is later sub-issues (#4405+). |
| `workflow.compact_content` | boolean | `false` | Compact content mode (#4139, [ADR-4139](adr/4139-compact-content-seam.md)). Per-project boolean selecting the terser form of GSD's own shipped prompt content (workflows, templates, agent-skill payloads). `/gsd-plan-phase` is the pilot workflow that actually branches on it today (#4402) — with the key off, its spine reads a deferred elaboration file back in before continuing (byte-identical instruction set to before); with it on, that read is skipped. The rest of the corpus does not branch on it yet — full coverage is later sub-issues (#4405+). The eager-window token reduction each split actually achieves is measured, not asserted: `npm run benchmark:compact-content` reports per-split and aggregate on/off token counts (a proxy-tokenizer delta — Anthropic publishes no tokenizer for Claude 3+, so the comparison is exact under a pinned tokenizer even though the absolute counts are not Claude's real ones) against a committed baseline (`tests/fixtures/compact-content-benchmark-baseline.json`, #4404). Reporting-only — it never fails CI. |
| `workflow.worktree_skip_hooks` | boolean | `false` | When `true`, executor agents in worktree mode pass `--no-verify` (skipping pre-commit hooks) and post-wave hook validation runs against the merged result instead. Opt-in escape hatch for projects whose hooks cannot run in agent worktrees. Default `false` runs hooks on every commit (#2924). |
| `workflow.code_review` | boolean | `true` | Enable `/gsd-code-review` and `/gsd-code-review --fix` commands. When `false`, the commands exit with a configuration gate message. Added in v1.34 |
| `workflow.code_review_point` | string | `execute:post` | Loop point at which the code-review capability's step registers: `execute:post` reviews once, after every wave in a phase has landed (default — unchanged behavior); `execute:wave:post` reviews once per completed wave instead, scoped to what changed since the phase's prior review (the whole phase's diff on the first wave, each subsequent wave's own diff thereafter). Manual `/gsd-code-review <phase>` invocation is unaffected by this key — it is gated by `workflow.code_review` alone and runs regardless of which point is configured. `/gsd-autonomous` and `/gsd-quick` have no wave granularity of their own, so setting this to `execute:wave:post` means code review does not run automatically inside those two flows (consistent with how every other `execute:wave:post`-only capability already behaves for them). Added in #3661 |