* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills (src/init.cts) now selects between a canonical agents/<name>.md and a token-minimized agents/<name>.compact.md sibling based on workflow.compact_content, resolved in code (a real function call with a real exit code) rather than a prose config-get gate — the same precedent stream 1's spine/detail split established for a load-bearing seam, applied here because this seam already runs through TypeScript instead of an eager @-include. A missing compact sibling falls back to the canonical persona and discloses the fallback in the served payload itself (a leading HTML-comment provenance line), so the Done-when contract — compact when on, canonical when off, never silent or empty — holds even for an agent nobody has compacted yet. Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md), each an independent, complete rewrite (not an extraction — nothing is "moved" the way spine/detail moves text) that preserves frontmatter, every @-include, every output-format contract, and every guardrail verbatim while cutting restatement and verbose framing. Verified mechanically: every pair registers (a canonical sibling exists), every compact file is strictly smaller, and the full @-include set matches canonical's — including which references are standalone eager-load lines versus inline prose mentions, since demoting one to inline changes what the host actually substitutes. Traced the install path before writing any code (.gsd/phase/.../40-design.md): stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no stem filtering under the default full profile, so the new .compact.md files install for free with zero installer changes — matching issue #4407's stated scope. A tiered agent profile that doesn't stage a compact sibling degrades through the same fallback-with-provenance path already required for an unauthored one, so no installer change is needed there either. Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export (deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are reached by a generic code construction rather than a literal path in prose, and checkReachability's markdown-search shape has nothing to find there). Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new "#4407 compact payload selection" describe block spawns gsd_run agent-skills against real compact/canonical fixture pairs and asserts on the served payload, which can only pass if the seam genuinely wires through. Fixed a pre-existing test whose agents/*.md glob incidentally matched the new .compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's roster, both real, unrelated-to-content defects the new files' mere existence surfaced. Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token benchmark baseline (npm run benchmark:compact-content-variants --write). Closes #4407. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): apply orthogonal review findings from the compact-payload seam Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath) to collapse the duplicated read-and-empty-check shape between the compact and canonical branches in cmdAgentSkills, and updated the adjacent comment enumerating flat JSON extras to name agent_payload_variant alongside source/degraded (added by the prior commit, comment left stale). Security review and the Spec axis found no defects requiring a code change; their non-blocking observations (a pre-existing, unmodified path-construction pattern; the reasoned, documented substitution of a behavioral test for the literal reachability check) are recorded in .gsd/phase/enhance-4407-agent-skill-seam/60-review.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents Root-caused via a real gsd-test run (93 failures) rather than guessing which tests glob agents/ naively. Two classes of defect, both genuine: 1. Identity-roster confusion (11 files/areas): many tests and one production script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f => f.endsWith('.md'))`, which incidentally matched the new .compact.md variant siblings too — a compact file is a rendering of an EXISTING agent identity, not a new one. Fixed at the shared root (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests already consolidated on) and at each independent glob that didn't use it: agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix before checking XL/LARGE membership, so a compact file inherits its canonical sibling's tier instead of silently falling through to DEFAULT), agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual script, not just its test), codex-config.test.cjs (confirmed directly against generateCodexAgentToml that a compact role's derived sandbox_mode is byte-identical to its canonical sibling's before excluding it — not assumed), and copilot-install.test.cjs (two counts that legitimately DO need both files — an installed-file count and a full-conversion smoke test — fixed to expect 70, not stay pinned to 35). no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix: two compact files reproduce descriptive prose already allowlisted at their canonical file's line number; added matching entries at the compact files' own line numbers rather than excluding them from the scan (a genuine bare gsd-tools command-position bug in a compact file would be as real a defect as in canonical). 2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree run): six agents' compact renditions (gsd-debugger, gsd-executor, gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction — confirmed structural, not a compaction-quality gap: each is dominated by content this phase's own rules require verbatim (the ~2.6 KB gsd_run bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in every agent that calls gsd_run, output-format contracts, guardrails). ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing spot in cmdAgentSkills's single-file synchronous read. Removed these 6 compact files rather than ship an over-cap file or invent a multi-part read mechanism out of scope for this phase; recorded by name with the reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and 50-test-matrix.md, per #4407's own "or explicitly recorded as not worth covering" allowance. Their canonical personas are served correctly today via the fallback-with-disclosed-provenance path this phase's own Done-when #2 already requires — 29 of 35 agents now have a compact variant. Also fixes an unrelated, genuinely pre-existing defect this gsd-test run surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md | wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed: two meaning-preserving trims in the <success_criteria> block (a repeated parenthetical replaced with a same-exception reference; one redundant qualifier dropped) bring it to 40,940 bytes. Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6 now-orphaned roster rows removed alongside them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): make .compact.md-aware roster checks resilient to partial coverage Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception), breaking once 6 stems legitimately have none. - tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was picking up the "### Compact Payload Variants" subsection's rows as phantom/uncounted entries in the primary/advanced/inventory-only classification this test validates — a compact row documents an existing agent's alternate rendition and never gets its own AGENTS.md heading, so it was never meant to participate in that classification. Excluded at the parser, not per-assertion. - tests/copilot-install.test.cjs: the derived expected-file-list generator assumed every listAgentFiles() stem has a .compact.md source sibling; checks disk per stem now instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4407): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
299 lines
12 KiB
JavaScript
299 lines
12 KiB
JavaScript
#!/usr/bin/env node
|
|
'use strict';
|
|
|
|
/**
|
|
* Benchmarks the token-count effect of every registered variant-swap pair
|
|
* (ADR-4139, Phase 6 #4406 — `gsd-core/workflows/<name>/{modes,steps,templates}/*.compact.md`
|
|
* and `gsd-core/templates/**\/*.compact.md`). Sibling to
|
|
* `scripts/benchmark-compact-content.cjs` (Phase 3's spine/detail benchmark) rather than an
|
|
* extension of it — the two are different data shapes (a variant pair is two independent,
|
|
* deliberately-overlapping complete files; a spine/detail split is one document partitioned in
|
|
* two disjoint halves), and mixing them into one report would conflate an "off" total that means
|
|
* something different in each case.
|
|
*
|
|
* For every registered pair it reports the "off" token count (the canonical file — what
|
|
* `workflow.compact_content=false`, the default, pays at that call site) against the "on" token
|
|
* count (the `.compact.md` sibling — what `workflow.compact_content=true` pays once the gate
|
|
* resolves to it). See `gsd-core/references/compact-content-gate.md` §"Streams 1b and 4" for the
|
|
* resolution rule this measures.
|
|
*
|
|
* PROXY-TOKENIZER CAVEAT: same as the sibling script — `gpt-tokenizer` is a stand-in tokenizer;
|
|
* Anthropic publishes none for Claude 3+. The on/off comparison is exact under this one pinned
|
|
* tokenizer applied identically to both sides; the absolute counts are not Claude's real counts.
|
|
*
|
|
* Discovery is REIMPLEMENTED here rather than imported from
|
|
* `tests/helpers/compact-content-variant.cjs`, for the same reason
|
|
* `benchmark-compact-content.cjs` reimplements spine/detail discovery instead of importing it: a
|
|
* `scripts/` reporting tool depending on a test-only helper module inverts this repo's normal
|
|
* layering, and a test-only module changing shape should never be able to break a benchmark.
|
|
*
|
|
* Usage:
|
|
* node scripts/benchmark-compact-content-variants.cjs # print JSON to stdout
|
|
* node scripts/benchmark-compact-content-variants.cjs --write # write the committed baseline
|
|
* node scripts/benchmark-compact-content-variants.cjs --check # recompute, diff vs committed baseline
|
|
* node scripts/benchmark-compact-content-variants.cjs --check --baseline-path=<path>
|
|
*
|
|
* Same CRITICAL contract as the sibling script: this — and `--check` especially — MUST NEVER
|
|
* exit non-zero because a baseline is drifted, stale, or missing. Only a genuine I/O error
|
|
* reading a SOURCE file the benchmark measures may throw. This is a reporting instrument, never
|
|
* a gate.
|
|
*/
|
|
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const { countTokens } = require('gpt-tokenizer');
|
|
const { runMain } = require('./lib/cli-exit.cjs');
|
|
|
|
const ROOT = path.resolve(__dirname, '..');
|
|
// #4407: agents/ joins the scan. Reachability there is a code seam
|
|
// (cmdAgentSkills), not a markdown literal reference, but token accounting
|
|
// doesn't care how a pair is reached — only that it's registered.
|
|
const VARIANT_ROOTS = [
|
|
path.join(ROOT, 'gsd-core', 'workflows'),
|
|
path.join(ROOT, 'gsd-core', 'templates'),
|
|
path.join(ROOT, 'agents'),
|
|
];
|
|
const BASELINE_PATH = path.join(ROOT, 'tests', 'fixtures', 'compact-content-variant-benchmark-baseline.json');
|
|
const COMPACT_SUFFIX = '.compact.md';
|
|
|
|
function getTokenizerVersion() {
|
|
const pkgPath = require.resolve('gpt-tokenizer/package.json');
|
|
const pkg = JSON.parse(fs.readFileSync(pkgPath, 'utf8'));
|
|
return pkg.version;
|
|
}
|
|
|
|
/**
|
|
* Discover every registered variant pair under `roots`: a `.compact.md` file
|
|
* with a same-directory, same-stem canonical `.md` sibling. Mirrors
|
|
* `tests/helpers/compact-content-variant.cjs`'s `discoverRegisteredVariants`
|
|
* in shape but is a from-scratch, self-contained implementation (see module
|
|
* header for why this is not a shared import). Skips an orphaned compact
|
|
* file with no canonical sibling — that is the guard's problem, not this
|
|
* benchmark's; a pair with no canonical baseline has no "off" number to
|
|
* report against.
|
|
*
|
|
* @param {string[]} [roots]
|
|
* @returns {Array<{name: string, canonicalPath: string, compactPath: string}>}
|
|
*/
|
|
function discoverRegisteredVariants(roots = VARIANT_ROOTS) {
|
|
const pairs = [];
|
|
|
|
function walk(dir) {
|
|
let entries;
|
|
try {
|
|
entries = fs.readdirSync(dir, { withFileTypes: true });
|
|
} catch {
|
|
return;
|
|
}
|
|
for (const entry of entries) {
|
|
const full = path.join(dir, entry.name);
|
|
if (entry.isDirectory()) {
|
|
walk(full);
|
|
} else if (entry.isFile() && entry.name.endsWith(COMPACT_SUFFIX)) {
|
|
const stem = entry.name.slice(0, -COMPACT_SUFFIX.length);
|
|
const canonicalPath = path.join(dir, `${stem}.md`);
|
|
if (fs.existsSync(canonicalPath)) {
|
|
const name = path.relative(ROOT, canonicalPath).split(path.sep).join('/');
|
|
pairs.push({ name, canonicalPath, compactPath: full });
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
for (const root of roots) walk(root);
|
|
return pairs.sort((a, b) => a.name.localeCompare(b.name));
|
|
}
|
|
|
|
/**
|
|
* Compute the off/on/reduction numbers for ONE registered variant pair.
|
|
* Throws on a genuine read failure of a source file (the one thing allowed
|
|
* to throw, per the module-header CRITICAL note).
|
|
*
|
|
* @param {{canonicalPath: string, compactPath: string}} pair
|
|
* @returns {{offTokens: number, onTokens: number, reductionPct: number}}
|
|
*/
|
|
function computePairTokens(pair) {
|
|
const offTokens = countTokens(fs.readFileSync(pair.canonicalPath, 'utf8'));
|
|
const onTokens = countTokens(fs.readFileSync(pair.compactPath, 'utf8'));
|
|
const reductionPct = offTokens === 0 ? 0 : round2(((offTokens - onTokens) / offTokens) * 100);
|
|
return { offTokens, onTokens, reductionPct };
|
|
}
|
|
|
|
function round2(n) {
|
|
return Math.round(n * 100) / 100;
|
|
}
|
|
|
|
/**
|
|
* Aggregate per-pair numbers from SUMMED off/on totals, never averaged
|
|
* per-pair percentages. Reports `0` (not `NaN`) for zero registered pairs.
|
|
* @param {Record<string, {offTokens: number, onTokens: number}>} pairResults
|
|
*/
|
|
function computeAggregate(pairResults) {
|
|
let offTokens = 0;
|
|
let onTokens = 0;
|
|
for (const key of Object.keys(pairResults)) {
|
|
offTokens += pairResults[key].offTokens;
|
|
onTokens += pairResults[key].onTokens;
|
|
}
|
|
const reductionPct = offTokens === 0 ? 0 : round2(((offTokens - onTokens) / offTokens) * 100);
|
|
return { offTokens, onTokens, reductionPct };
|
|
}
|
|
|
|
const LABEL =
|
|
"PROXY-TOKENIZER DELTA — gpt-tokenizer is a stand-in; Anthropic publishes no tokenizer for Claude 3+. " +
|
|
"The on/off COMPARISON is exact under this pinned tokenizer; absolute counts are not Claude's real token counts.";
|
|
|
|
/**
|
|
* @param {string[]} [roots]
|
|
* @returns {object}
|
|
*/
|
|
function buildReport(roots = VARIANT_ROOTS) {
|
|
const pairs = discoverRegisteredVariants(roots);
|
|
const pairReports = {};
|
|
for (const pair of pairs) {
|
|
pairReports[pair.name] = computePairTokens(pair);
|
|
}
|
|
return {
|
|
schema_version: 1,
|
|
generated_by: 'scripts/benchmark-compact-content-variants.cjs',
|
|
tokenizer: { name: 'gpt-tokenizer', version: getTokenizerVersion() },
|
|
label: LABEL,
|
|
pairs: pairReports,
|
|
aggregate: computeAggregate(pairReports),
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Format a human-readable drift summary between a (possibly missing/invalid)
|
|
* committed baseline and a freshly-computed live report. Never throws.
|
|
* @param {string} baselinePath
|
|
* @param {object} live
|
|
* @returns {string}
|
|
*/
|
|
function formatDriftReport(baselinePath, live) {
|
|
const lines = [];
|
|
let baseline = null;
|
|
let baselineReadError = null;
|
|
try {
|
|
const raw = fs.readFileSync(baselinePath, 'utf8');
|
|
try {
|
|
baseline = JSON.parse(raw);
|
|
} catch (parseErr) {
|
|
baselineReadError = `baseline at ${baselinePath} could not be parsed as JSON: ${parseErr.message}`;
|
|
}
|
|
} catch {
|
|
baselineReadError = `no baseline found at ${baselinePath}`;
|
|
}
|
|
|
|
if (baselineReadError) {
|
|
lines.push(`DRIFT: ${baselineReadError} — treating as fully drifted (this is reported, not an error).`);
|
|
lines.push('Live pairs:');
|
|
for (const name of Object.keys(live.pairs).sort()) {
|
|
const p = live.pairs[name];
|
|
lines.push(` + ${name}: off=${p.offTokens} on=${p.onTokens} reduction=${p.reductionPct}%`);
|
|
}
|
|
lines.push(
|
|
`Live aggregate: off=${live.aggregate.offTokens} on=${live.aggregate.onTokens} ` +
|
|
`reduction=${live.aggregate.reductionPct}%`,
|
|
);
|
|
return lines.join('\n');
|
|
}
|
|
|
|
const baselinePairs = (baseline && typeof baseline === 'object' && baseline.pairs) || {};
|
|
const livePairs = live.pairs;
|
|
const allNames = new Set([...Object.keys(baselinePairs), ...Object.keys(livePairs)]);
|
|
let anyDrift = false;
|
|
|
|
if (!baseline || typeof baseline.label !== 'string' || !baseline.label.includes('PROXY-TOKENIZER')) {
|
|
anyDrift = true;
|
|
lines.push('DRIFT: committed baseline is missing the required "PROXY-TOKENIZER" label.');
|
|
}
|
|
|
|
for (const name of [...allNames].sort()) {
|
|
const b = baselinePairs[name];
|
|
const l = livePairs[name];
|
|
if (!b) {
|
|
anyDrift = true;
|
|
lines.push(`DRIFT: pair "${name}" is new (not in committed baseline) — live off=${l.offTokens} on=${l.onTokens} reduction=${l.reductionPct}%`);
|
|
} else if (!l) {
|
|
anyDrift = true;
|
|
lines.push(`DRIFT: pair "${name}" was removed (present in committed baseline, not found live) — baseline off=${b.offTokens} on=${b.onTokens} reduction=${b.reductionPct}%`);
|
|
} else if (b.offTokens !== l.offTokens || b.onTokens !== l.onTokens || b.reductionPct !== l.reductionPct) {
|
|
anyDrift = true;
|
|
lines.push(
|
|
`DRIFT: pair "${name}": off ${b.offTokens} -> ${l.offTokens} (${l.offTokens - b.offTokens >= 0 ? '+' : ''}${l.offTokens - b.offTokens}), ` +
|
|
`on ${b.onTokens} -> ${l.onTokens} (${l.onTokens - b.onTokens >= 0 ? '+' : ''}${l.onTokens - b.onTokens}), ` +
|
|
`reduction ${b.reductionPct}% -> ${l.reductionPct}% (${round2(l.reductionPct - b.reductionPct) >= 0 ? '+' : ''}${round2(l.reductionPct - b.reductionPct)}pp)`,
|
|
);
|
|
}
|
|
}
|
|
|
|
const ba = (baseline && baseline.aggregate) || {};
|
|
const la = live.aggregate;
|
|
if (ba.offTokens !== la.offTokens || ba.onTokens !== la.onTokens || ba.reductionPct !== la.reductionPct) {
|
|
anyDrift = true;
|
|
lines.push(
|
|
`DRIFT: aggregate: off ${ba.offTokens} -> ${la.offTokens}, on ${ba.onTokens} -> ${la.onTokens}, ` +
|
|
`reduction ${ba.reductionPct}% -> ${la.reductionPct}%`,
|
|
);
|
|
}
|
|
|
|
if (!anyDrift) {
|
|
lines.push(`Baseline at ${baselinePath} is up to date with the live recompute.`);
|
|
} else {
|
|
lines.push('');
|
|
lines.push('Run `node scripts/benchmark-compact-content-variants.cjs --write` to refresh the committed baseline.');
|
|
lines.push('(This is a REPORT, not a gate — exiting 0 regardless of drift, per this script\'s own contract.)');
|
|
}
|
|
return lines.join('\n');
|
|
}
|
|
|
|
function parseArgs(argv) {
|
|
const opts = { write: false, check: false, baselinePath: BASELINE_PATH };
|
|
for (const arg of argv) {
|
|
if (arg === '--write') opts.write = true;
|
|
else if (arg === '--check') opts.check = true;
|
|
else if (arg.startsWith('--baseline-path=')) opts.baselinePath = arg.slice('--baseline-path='.length);
|
|
}
|
|
return opts;
|
|
}
|
|
|
|
function main() {
|
|
const opts = parseArgs(process.argv.slice(2));
|
|
|
|
if (opts.write) {
|
|
const report = buildReport();
|
|
fs.mkdirSync(path.dirname(BASELINE_PATH), { recursive: true });
|
|
fs.writeFileSync(BASELINE_PATH, JSON.stringify(report, null, 2) + '\n');
|
|
process.stdout.write(`Wrote ${BASELINE_PATH}\n`);
|
|
return;
|
|
}
|
|
|
|
if (opts.check) {
|
|
const live = buildReport();
|
|
process.stdout.write(formatDriftReport(opts.baselinePath, live) + '\n');
|
|
return;
|
|
}
|
|
|
|
process.stdout.write(JSON.stringify(buildReport(), null, 2) + '\n');
|
|
}
|
|
|
|
/* c8 ignore next 3 -- CLI entry guard; this repo measures coverage with c8, which does not honor istanbul pragmas */
|
|
if (require.main === module) {
|
|
runMain(main);
|
|
}
|
|
|
|
module.exports = {
|
|
discoverRegisteredVariants,
|
|
computePairTokens,
|
|
computeAggregate,
|
|
buildReport,
|
|
formatDriftReport,
|
|
getTokenizerVersion,
|
|
parseArgs,
|
|
LABEL,
|
|
BASELINE_PATH,
|
|
VARIANT_ROOTS,
|
|
};
|