* enhance(#4404): add offline token benchmark for compact-content splits
ADR-4139 Decision 2 requires the finite-attention justification for
workflow.compact_content to be measured, not asserted. `npm run
benchmark:compact-content` computes, per registered spine/detail split
discovered under gsd-core/workflows/, the token count with the split
active (spine alone) vs inactive (spine + all detail parts read back
in), using gpt-tokenizer (pinned exact devDependency — Anthropic
publishes no tokenizer for Claude 3+, so every output surface labels
this a PROXY-TOKENIZER comparison: the on/off delta is exact under one
tokenizer applied identically to both sides, the absolute counts are
not Claude's real ones).
Reporting-only by design and verified so: --check diffs the live
recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json)
and prints drift, but never exits non-zero for a drifted or missing
baseline — the only thing allowed to fail this script is a genuine I/O
error reading a source .md file it's measuring. Not wired into lint:ci
or pretest.
Discovery is deliberately reimplemented rather than importing
tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a
scripts/ reporting tool from depending on a test-only module.
tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to
prove the benchmark makes no network call, monkeypatching http/https/
net/dns/fetch to throw rather than relying on sandboxing.
docs/CONFIGURATION.md documents the new benchmark against the
workflow.compact_content key to satisfy this repo's docs-required gate
for an Added-type changeset.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(#4404): address orthogonal review findings on the token benchmark
Standards axis found two hard violations against documented rules:
- CLAUDE.md's Generative Fix Divergence rule requires a parity assertion
for shared discovery logic maintained in two places. Added a test
comparing benchmark-compact-content.cjs's own discoverRegisteredSplits
against tests/helpers/compact-content-split.cjs's version on the real
repo tree, so the two can never silently drift apart.
- The changeset body closed its bold span with a period and continued
as a second sentence, instead of the canonical
`**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents.
Spec axis found the "network disabled + identical output across two
runs" Done-when criterion was verified as two separate properties
(determinism tested without network denial, offline survival tested as
a single run) rather than as one combined property. Added a test that
runs the benchmark twice under the deny-network preload and asserts
byte-identical stdout.
Security axis found tests/fixtures/deny-network.cjs didn't patch
dns.promises (a separate binding from the callback dns API), tls.connect,
or http2.connect — inert today since nothing in the benchmark calls
them, but a silent gap in what the preload's own header claims to
guarantee. Patched all three.
CLAUDE.md's Property-Based Testing rule also requires a fast-check test
for budget-limit arithmetic; added one for computeAggregate's off/on
summation (true sum over N splits, never NaN/Infinity, never exceeds
100% when off >= on for every split).
Standards axis's remaining two findings (a Data Clumps observation on
the {offTokens, onTokens, reductionPct} triple, and mild duplication in
formatDriftReport's three line-formatters) are left as judgement calls:
introducing a named type for a 3-field local tuple, or a formatter
abstraction for three short lines, would be exactly the premature
abstraction CLAUDE.md's engineering guidance warns against for a script
this size.
All changes verified directly (parity logic, the fast-check property,
and the three newly-denied network surfaces actually throwing under the
preload) via node -e before committing; full npm run lint:ci passes
with the eslint cache cleared.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* docs(#4404): backfill changeset pr number to 4502
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>