Files
msd-core/scripts/benchmark-compact-content-variants.cjs
Tom Boucher a27cb6b2fa enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates

ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b
(gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4
(gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user:
two independent, complete files per covered path (canonical + .compact.md
sibling), with the gate picking which one gets Read at the call site. This is
a different shape from Phase 5's spine+detail partition, and is safe here
specifically because these files are already reached only by a runtime Read —
a missed Read already means zero overlay content today, with or without
workflow.compact_content, so selecting between two independently-complete
files introduces no new failure mode (documented in
gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section).

Disposition, after inspecting every candidate rather than trusting a byte-size
threshold (same rigor Phase 5 applied to review.md):

- Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing
  reference doc emitted verbatim, not orchestrator instruction). The other 9
  size-threshold candidates are dominated by fail-closed guards, exact CLI
  invocations, or output-format contracts (AskUserQuestion blocks) — recorded
  not-worth-compacting, same reasoning as Phase 5's review.md.
- Stream 4: a ground-truth reachability audit replaced the initial size-only
  candidate list. Two files (summary.md, user-setup.md) got compact variants;
  a third (spec.md) was drafted, then dropped after discovering its only two
  call sites are eager @-includes, not a runtime Read — stream-1 material
  hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md
  itself has 3 eager call sites and only 1 genuine runtime-Read call site
  (execute-plan.md); only that one was wired, so the compact variant's savings
  apply to the sequential single-plan execution path only.
- Discovered while auditing reachability: 12 gsd-core/templates/** files with
  zero references anywhere in workflow/agent/command prose, compiled source,
  or tests — dead scaffolding predating this phase. Deleted in this same PR
  per this repo's no-defer policy, after re-verifying against a computed
  path.join(...) pattern (not just a plain-string search) that nearly caused
  two genuinely load-bearing templates (user-profile.md, dev-preferences.md)
  to be misclassified as dead.

New checker (tests/helpers/compact-content-variant.cjs): registration,
reachability, protected-content-preserved, size-smaller — replacing Phase
3/5's disjointness/completeness checks, which assume a partition rather than
two deliberately-overlapping documents. The reachability check's own
"unprefixed match" guard had a real bug (rejected the repo's own
`~/.claude/gsd-core/...` convention), caught by running it against the
already-wired help/modes/full.compact.md pair rather than only synthetic
fixtures — fixed to anchor on the nearest `gsd-core` path segment instead.

Template consumer parity (tests/compact-content-template-variant-parity.test.cjs):
proves each compact variant's `## File Template` fenced block — the actual
output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed
against — is byte-identical to the canonical file, then runs the one real
deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent,
backing `gsd-tools uat classify-coverage`) against content built from that
shared contract.

Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs)
rather than extending the existing spine/detail one — different data shape,
and the existing script's own contract deliberately isolates it from a
test-only helper's shape changing.

Emitted-drift acknowledgement: not needed. Every changed/added path in this
diff is hand-authored and present in the diff itself, so diffEmitted's
attribution loop resolves `via` to the path's own source before reaching the
ack-lookup branch (same reasoning Phase 5 verified for its own diff).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* enhance(#4406): address code-review findings on the variant-swap gate

- docs/CONFIGURATION.md and gsd-core/references/planning-config.md's
  workflow.compact_content entries described only the spine+detail mechanism
  (Phase 5) and were missing this phase's variant-swap mechanism and its
  benchmark:compact-content-variants script entirely — required since this
  PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets"
  rule). Both now describe both mechanisms and which call sites are wired.
- Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved:
  a canonical file with zero <!-- gsd:protected --> blocks must be a
  no-op, not a violation — the only branch of that function the existing
  fixtures didn't exercise.
- Collapsed findCompactFiles/findMarkdownFiles in
  tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix
  helper — the two were identical recursive walks differing only in the
  extension predicate (minor Duplicated-Code finding).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification

gsd-test caught this, not static analysis: 10 real failures in
tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs,
and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install
path, which does
fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md'))
after copying gsd-core/templates/** into the target project, then merges it into both
.github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit
that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the
repo-root bin/install.js — a separately maintained installer bundle outside the
src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around
that read degrades to a silent skip rather than a crash when the template is missing,
which is why this surfaced only once the real E2E install test ran, not from any
static check.

Re-verified the remaining 11 deleted filenames against bin/install.js specifically
(plain substring and quoted-filename search) before trusting that list — all 11 have
zero hits there, confirmed dead by the same standard this one file failed.

Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to
reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md)
and the phase design doc from 12 to 11 deleted files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants
Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule

* docs(#4406): backfill changeset PR numbers

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion

CI's own full-test matrix (not gsd-test's matrix, which does not run this
check) caught 4 more false-positive dead-template classifications via
tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs
— a literal, word-boundary basename check across .github/workflows/,
gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every
file a PR deletes. It has no semantic awareness, so a deleted template's
basename colliding with something else entirely still fires:

- claude-md.md: gsd-core/templates/README.md had a stale table row claiming
  /gsd-profile reads this template to generate CLAUDE.md. Verified false (no
  code reads it anywhere, same search that already covered bin/install.js) —
  fixed the row to *(inline)*, matching every other command-generated
  artifact in that table. File stays deleted.
- codebase/testing.md: collided with docs/guides/testing.md, an illustrative
  example row in docs-update.md's sample output table (an unrelated real
  generated-docs path). Swapped the example topic to "contributing" — the
  row is illustrative, any topic works. File stays deleted.
- codebase/architecture.md, codebase/stack.md: collided with docs/reference/
  planning-artifacts.md's directory listing of a user's own generated
  .planning/codebase/architecture.md and stack.md output — the same
  semantic mismatch already investigated and dismissed as unrelated earlier
  in this phase's audit, now caught by a gate instead of judgment. That
  listing repeats across 5 locale copies of the doc.
- continue-here.md: collided with the real .continue-here.md pause-work
  artifact, referenced across 15+ locale and workflow files.

For the last two, the lint's own error message offers "restore the file or
update every consumer in the same commit." Rewording 15+ files across
languages I cannot verify translation quality for, to shave 2 already-tiny
templates that were merely presumed dead, is disproportionate to this PR's
actual scope — restored codebase/architecture.md, codebase/stack.md, and
continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates
section accordingly.

Final confirmed-dead set: claude-md.md, codebase/concerns.md,
codebase/conventions.md, codebase/integrations.md, codebase/structure.md,
codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down
from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next
node scripts/lint-removed-but-needed.cjs now passes clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes

* fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout

Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the
user asked to be actually fixed, not just re-run past: PR #4497 (landed
2026-09-07, one day before this PR's CI run) isolated
tests/codex-config.test.cjs into its own dedicated chunk because its
measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe
to share a chunk with any other file. That isolation was necessary but not
sufficient — even alone, with zero companion-file contention, the file's
real Windows execution time sits right at the 600s per-chunk ceiling. Two
independent CI runs on two unrelated PRs (this one and #4154) were both
killed within ~1.4s of the identical 600000ms mark — not random contention,
a deterministic near-miss the isolation fix couldn't address because it
never reduced the file's own cost, only removed the risk of a companion
file's cost stacking on top of it (which the PR #4497 comment explicitly
anticipated: "if a future profiling pass genuinely speeds up
codex-config.test.cjs itself, this isolation can be revisited").

The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79
describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760,
#3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more),
several of which are explicitly documented as "folded" in from separate
files that were never actually split back out ("Verified non-duplicate
against both the pre-existing target and the other three folded sources").

Split into 4 files by top-level AST statement boundaries (never a naive
column-0 regex — an early attempt at that overcounted 79 apparent
"describe(" matches when only 21 are genuinely top-level; the rest are
nested inside a handful of large folded-in blocks, which a regex can't tell
apart from real top-level statements). Verified lossless twice: the split
script asserts byte-for-byte reconstruction of every source character, and
independently, total test()/describe() call counts match exactly between
the original file and the sum across all 4 new files (433/79 both sides).
Each new file carries the complete original shared header (imports/helpers)
for safety; per-file unused-import warnings from that duplication are
resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard
form for an intentionally-unused destructured binding — never a bare `{
_foo }`, which would destructure a different, nonexistent property).

No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its
pinned test in tests/run-tests-harness.test.cjs: the file that keeps the
original name (tests/codex-config.test.cjs) is now only ~28% of the
original's size and safely isolated in its own chunk as before; the other
three new files re-enter normal weight-balanced packing, none individually
close to disproportionate. Confirmed no other file hardcodes the hardcoded
filename anywhere that would silently stop these tests from running (the
CI test-selection scripts determine scope algorithmically, not by literal
filename).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 14:31:04 -04:00

292 lines
12 KiB
JavaScript

#!/usr/bin/env node
'use strict';
/**
* Benchmarks the token-count effect of every registered variant-swap pair
* (ADR-4139, Phase 6 #4406 — `gsd-core/workflows/<name>/{modes,steps,templates}/*.compact.md`
* and `gsd-core/templates/**\/*.compact.md`). Sibling to
* `scripts/benchmark-compact-content.cjs` (Phase 3's spine/detail benchmark) rather than an
* extension of it — the two are different data shapes (a variant pair is two independent,
* deliberately-overlapping complete files; a spine/detail split is one document partitioned in
* two disjoint halves), and mixing them into one report would conflate an "off" total that means
* something different in each case.
*
* For every registered pair it reports the "off" token count (the canonical file — what
* `workflow.compact_content=false`, the default, pays at that call site) against the "on" token
* count (the `.compact.md` sibling — what `workflow.compact_content=true` pays once the gate
* resolves to it). See `gsd-core/references/compact-content-gate.md` §"Streams 1b and 4" for the
* resolution rule this measures.
*
* PROXY-TOKENIZER CAVEAT: same as the sibling script — `gpt-tokenizer` is a stand-in tokenizer;
* Anthropic publishes none for Claude 3+. The on/off comparison is exact under this one pinned
* tokenizer applied identically to both sides; the absolute counts are not Claude's real counts.
*
* Discovery is REIMPLEMENTED here rather than imported from
* `tests/helpers/compact-content-variant.cjs`, for the same reason
* `benchmark-compact-content.cjs` reimplements spine/detail discovery instead of importing it: a
* `scripts/` reporting tool depending on a test-only helper module inverts this repo's normal
* layering, and a test-only module changing shape should never be able to break a benchmark.
*
* Usage:
* node scripts/benchmark-compact-content-variants.cjs # print JSON to stdout
* node scripts/benchmark-compact-content-variants.cjs --write # write the committed baseline
* node scripts/benchmark-compact-content-variants.cjs --check # recompute, diff vs committed baseline
* node scripts/benchmark-compact-content-variants.cjs --check --baseline-path=<path>
*
* Same CRITICAL contract as the sibling script: this — and `--check` especially — MUST NEVER
* exit non-zero because a baseline is drifted, stale, or missing. Only a genuine I/O error
* reading a SOURCE file the benchmark measures may throw. This is a reporting instrument, never
* a gate.
*/
const fs = require('node:fs');
const path = require('node:path');
const { countTokens } = require('gpt-tokenizer');
const { runMain } = require('./lib/cli-exit.cjs');
const ROOT = path.resolve(__dirname, '..');
const VARIANT_ROOTS = [path.join(ROOT, 'gsd-core', 'workflows'), path.join(ROOT, 'gsd-core', 'templates')];
const BASELINE_PATH = path.join(ROOT, 'tests', 'fixtures', 'compact-content-variant-benchmark-baseline.json');
const COMPACT_SUFFIX = '.compact.md';
function getTokenizerVersion() {
const pkgPath = require.resolve('gpt-tokenizer/package.json');
const pkg = JSON.parse(fs.readFileSync(pkgPath, 'utf8'));
return pkg.version;
}
/**
* Discover every registered variant pair under `roots`: a `.compact.md` file
* with a same-directory, same-stem canonical `.md` sibling. Mirrors
* `tests/helpers/compact-content-variant.cjs`'s `discoverRegisteredVariants`
* in shape but is a from-scratch, self-contained implementation (see module
* header for why this is not a shared import). Skips an orphaned compact
* file with no canonical sibling — that is the guard's problem, not this
* benchmark's; a pair with no canonical baseline has no "off" number to
* report against.
*
* @param {string[]} [roots]
* @returns {Array<{name: string, canonicalPath: string, compactPath: string}>}
*/
function discoverRegisteredVariants(roots = VARIANT_ROOTS) {
const pairs = [];
function walk(dir) {
let entries;
try {
entries = fs.readdirSync(dir, { withFileTypes: true });
} catch {
return;
}
for (const entry of entries) {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) {
walk(full);
} else if (entry.isFile() && entry.name.endsWith(COMPACT_SUFFIX)) {
const stem = entry.name.slice(0, -COMPACT_SUFFIX.length);
const canonicalPath = path.join(dir, `${stem}.md`);
if (fs.existsSync(canonicalPath)) {
const name = path.relative(ROOT, canonicalPath).split(path.sep).join('/');
pairs.push({ name, canonicalPath, compactPath: full });
}
}
}
}
for (const root of roots) walk(root);
return pairs.sort((a, b) => a.name.localeCompare(b.name));
}
/**
* Compute the off/on/reduction numbers for ONE registered variant pair.
* Throws on a genuine read failure of a source file (the one thing allowed
* to throw, per the module-header CRITICAL note).
*
* @param {{canonicalPath: string, compactPath: string}} pair
* @returns {{offTokens: number, onTokens: number, reductionPct: number}}
*/
function computePairTokens(pair) {
const offTokens = countTokens(fs.readFileSync(pair.canonicalPath, 'utf8'));
const onTokens = countTokens(fs.readFileSync(pair.compactPath, 'utf8'));
const reductionPct = offTokens === 0 ? 0 : round2(((offTokens - onTokens) / offTokens) * 100);
return { offTokens, onTokens, reductionPct };
}
function round2(n) {
return Math.round(n * 100) / 100;
}
/**
* Aggregate per-pair numbers from SUMMED off/on totals, never averaged
* per-pair percentages. Reports `0` (not `NaN`) for zero registered pairs.
* @param {Record<string, {offTokens: number, onTokens: number}>} pairResults
*/
function computeAggregate(pairResults) {
let offTokens = 0;
let onTokens = 0;
for (const key of Object.keys(pairResults)) {
offTokens += pairResults[key].offTokens;
onTokens += pairResults[key].onTokens;
}
const reductionPct = offTokens === 0 ? 0 : round2(((offTokens - onTokens) / offTokens) * 100);
return { offTokens, onTokens, reductionPct };
}
const LABEL =
"PROXY-TOKENIZER DELTA — gpt-tokenizer is a stand-in; Anthropic publishes no tokenizer for Claude 3+. " +
"The on/off COMPARISON is exact under this pinned tokenizer; absolute counts are not Claude's real token counts.";
/**
* @param {string[]} [roots]
* @returns {object}
*/
function buildReport(roots = VARIANT_ROOTS) {
const pairs = discoverRegisteredVariants(roots);
const pairReports = {};
for (const pair of pairs) {
pairReports[pair.name] = computePairTokens(pair);
}
return {
schema_version: 1,
generated_by: 'scripts/benchmark-compact-content-variants.cjs',
tokenizer: { name: 'gpt-tokenizer', version: getTokenizerVersion() },
label: LABEL,
pairs: pairReports,
aggregate: computeAggregate(pairReports),
};
}
/**
* Format a human-readable drift summary between a (possibly missing/invalid)
* committed baseline and a freshly-computed live report. Never throws.
* @param {string} baselinePath
* @param {object} live
* @returns {string}
*/
function formatDriftReport(baselinePath, live) {
const lines = [];
let baseline = null;
let baselineReadError = null;
try {
const raw = fs.readFileSync(baselinePath, 'utf8');
try {
baseline = JSON.parse(raw);
} catch (parseErr) {
baselineReadError = `baseline at ${baselinePath} could not be parsed as JSON: ${parseErr.message}`;
}
} catch {
baselineReadError = `no baseline found at ${baselinePath}`;
}
if (baselineReadError) {
lines.push(`DRIFT: ${baselineReadError} — treating as fully drifted (this is reported, not an error).`);
lines.push('Live pairs:');
for (const name of Object.keys(live.pairs).sort()) {
const p = live.pairs[name];
lines.push(` + ${name}: off=${p.offTokens} on=${p.onTokens} reduction=${p.reductionPct}%`);
}
lines.push(
`Live aggregate: off=${live.aggregate.offTokens} on=${live.aggregate.onTokens} ` +
`reduction=${live.aggregate.reductionPct}%`,
);
return lines.join('\n');
}
const baselinePairs = (baseline && typeof baseline === 'object' && baseline.pairs) || {};
const livePairs = live.pairs;
const allNames = new Set([...Object.keys(baselinePairs), ...Object.keys(livePairs)]);
let anyDrift = false;
if (!baseline || typeof baseline.label !== 'string' || !baseline.label.includes('PROXY-TOKENIZER')) {
anyDrift = true;
lines.push('DRIFT: committed baseline is missing the required "PROXY-TOKENIZER" label.');
}
for (const name of [...allNames].sort()) {
const b = baselinePairs[name];
const l = livePairs[name];
if (!b) {
anyDrift = true;
lines.push(`DRIFT: pair "${name}" is new (not in committed baseline) — live off=${l.offTokens} on=${l.onTokens} reduction=${l.reductionPct}%`);
} else if (!l) {
anyDrift = true;
lines.push(`DRIFT: pair "${name}" was removed (present in committed baseline, not found live) — baseline off=${b.offTokens} on=${b.onTokens} reduction=${b.reductionPct}%`);
} else if (b.offTokens !== l.offTokens || b.onTokens !== l.onTokens || b.reductionPct !== l.reductionPct) {
anyDrift = true;
lines.push(
`DRIFT: pair "${name}": off ${b.offTokens} -> ${l.offTokens} (${l.offTokens - b.offTokens >= 0 ? '+' : ''}${l.offTokens - b.offTokens}), ` +
`on ${b.onTokens} -> ${l.onTokens} (${l.onTokens - b.onTokens >= 0 ? '+' : ''}${l.onTokens - b.onTokens}), ` +
`reduction ${b.reductionPct}% -> ${l.reductionPct}% (${round2(l.reductionPct - b.reductionPct) >= 0 ? '+' : ''}${round2(l.reductionPct - b.reductionPct)}pp)`,
);
}
}
const ba = (baseline && baseline.aggregate) || {};
const la = live.aggregate;
if (ba.offTokens !== la.offTokens || ba.onTokens !== la.onTokens || ba.reductionPct !== la.reductionPct) {
anyDrift = true;
lines.push(
`DRIFT: aggregate: off ${ba.offTokens} -> ${la.offTokens}, on ${ba.onTokens} -> ${la.onTokens}, ` +
`reduction ${ba.reductionPct}% -> ${la.reductionPct}%`,
);
}
if (!anyDrift) {
lines.push(`Baseline at ${baselinePath} is up to date with the live recompute.`);
} else {
lines.push('');
lines.push('Run `node scripts/benchmark-compact-content-variants.cjs --write` to refresh the committed baseline.');
lines.push('(This is a REPORT, not a gate — exiting 0 regardless of drift, per this script\'s own contract.)');
}
return lines.join('\n');
}
function parseArgs(argv) {
const opts = { write: false, check: false, baselinePath: BASELINE_PATH };
for (const arg of argv) {
if (arg === '--write') opts.write = true;
else if (arg === '--check') opts.check = true;
else if (arg.startsWith('--baseline-path=')) opts.baselinePath = arg.slice('--baseline-path='.length);
}
return opts;
}
function main() {
const opts = parseArgs(process.argv.slice(2));
if (opts.write) {
const report = buildReport();
fs.mkdirSync(path.dirname(BASELINE_PATH), { recursive: true });
fs.writeFileSync(BASELINE_PATH, JSON.stringify(report, null, 2) + '\n');
process.stdout.write(`Wrote ${BASELINE_PATH}\n`);
return;
}
if (opts.check) {
const live = buildReport();
process.stdout.write(formatDriftReport(opts.baselinePath, live) + '\n');
return;
}
process.stdout.write(JSON.stringify(buildReport(), null, 2) + '\n');
}
/* c8 ignore next 3 -- CLI entry guard; this repo measures coverage with c8, which does not honor istanbul pragmas */
if (require.main === module) {
runMain(main);
}
module.exports = {
discoverRegisteredVariants,
computePairTokens,
computeAggregate,
buildReport,
formatDriftReport,
getTokenizerVersion,
parseArgs,
LABEL,
BASELINE_PATH,
VARIANT_ROOTS,
};