* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills (src/init.cts) now selects between a canonical agents/<name>.md and a token-minimized agents/<name>.compact.md sibling based on workflow.compact_content, resolved in code (a real function call with a real exit code) rather than a prose config-get gate — the same precedent stream 1's spine/detail split established for a load-bearing seam, applied here because this seam already runs through TypeScript instead of an eager @-include. A missing compact sibling falls back to the canonical persona and discloses the fallback in the served payload itself (a leading HTML-comment provenance line), so the Done-when contract — compact when on, canonical when off, never silent or empty — holds even for an agent nobody has compacted yet. Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md), each an independent, complete rewrite (not an extraction — nothing is "moved" the way spine/detail moves text) that preserves frontmatter, every @-include, every output-format contract, and every guardrail verbatim while cutting restatement and verbose framing. Verified mechanically: every pair registers (a canonical sibling exists), every compact file is strictly smaller, and the full @-include set matches canonical's — including which references are standalone eager-load lines versus inline prose mentions, since demoting one to inline changes what the host actually substitutes. Traced the install path before writing any code (.gsd/phase/.../40-design.md): stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no stem filtering under the default full profile, so the new .compact.md files install for free with zero installer changes — matching issue #4407's stated scope. A tiered agent profile that doesn't stage a compact sibling degrades through the same fallback-with-provenance path already required for an unauthored one, so no installer change is needed there either. Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export (deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are reached by a generic code construction rather than a literal path in prose, and checkReachability's markdown-search shape has nothing to find there). Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new "#4407 compact payload selection" describe block spawns gsd_run agent-skills against real compact/canonical fixture pairs and asserts on the served payload, which can only pass if the seam genuinely wires through. Fixed a pre-existing test whose agents/*.md glob incidentally matched the new .compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's roster, both real, unrelated-to-content defects the new files' mere existence surfaced. Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token benchmark baseline (npm run benchmark:compact-content-variants --write). Closes #4407. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): apply orthogonal review findings from the compact-payload seam Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath) to collapse the duplicated read-and-empty-check shape between the compact and canonical branches in cmdAgentSkills, and updated the adjacent comment enumerating flat JSON extras to name agent_payload_variant alongside source/degraded (added by the prior commit, comment left stale). Security review and the Spec axis found no defects requiring a code change; their non-blocking observations (a pre-existing, unmodified path-construction pattern; the reasoned, documented substitution of a behavioral test for the literal reachability check) are recorded in .gsd/phase/enhance-4407-agent-skill-seam/60-review.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents Root-caused via a real gsd-test run (93 failures) rather than guessing which tests glob agents/ naively. Two classes of defect, both genuine: 1. Identity-roster confusion (11 files/areas): many tests and one production script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f => f.endsWith('.md'))`, which incidentally matched the new .compact.md variant siblings too — a compact file is a rendering of an EXISTING agent identity, not a new one. Fixed at the shared root (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests already consolidated on) and at each independent glob that didn't use it: agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix before checking XL/LARGE membership, so a compact file inherits its canonical sibling's tier instead of silently falling through to DEFAULT), agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual script, not just its test), codex-config.test.cjs (confirmed directly against generateCodexAgentToml that a compact role's derived sandbox_mode is byte-identical to its canonical sibling's before excluding it — not assumed), and copilot-install.test.cjs (two counts that legitimately DO need both files — an installed-file count and a full-conversion smoke test — fixed to expect 70, not stay pinned to 35). no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix: two compact files reproduce descriptive prose already allowlisted at their canonical file's line number; added matching entries at the compact files' own line numbers rather than excluding them from the scan (a genuine bare gsd-tools command-position bug in a compact file would be as real a defect as in canonical). 2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree run): six agents' compact renditions (gsd-debugger, gsd-executor, gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction — confirmed structural, not a compaction-quality gap: each is dominated by content this phase's own rules require verbatim (the ~2.6 KB gsd_run bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in every agent that calls gsd_run, output-format contracts, guardrails). ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing spot in cmdAgentSkills's single-file synchronous read. Removed these 6 compact files rather than ship an over-cap file or invent a multi-part read mechanism out of scope for this phase; recorded by name with the reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and 50-test-matrix.md, per #4407's own "or explicitly recorded as not worth covering" allowance. Their canonical personas are served correctly today via the fallback-with-disclosed-provenance path this phase's own Done-when #2 already requires — 29 of 35 agents now have a compact variant. Also fixes an unrelated, genuinely pre-existing defect this gsd-test run surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md | wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed: two meaning-preserving trims in the <success_criteria> block (a repeated parenthetical replaced with a same-exception reference; one redundant qualifier dropped) bring it to 40,940 bytes. Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6 now-orphaned roster rows removed alongside them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): make .compact.md-aware roster checks resilient to partial coverage Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception), breaking once 6 stems legitimately have none. - tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was picking up the "### Compact Payload Variants" subsection's rows as phantom/uncounted entries in the primary/advanced/inventory-only classification this test validates — a compact row documents an existing agent's alternate rendition and never gets its own AGENTS.md heading, so it was never meant to participate in that classification. Excluded at the parser, not per-assertion. - tests/copilot-install.test.cjs: the derived expected-file-list generator assumed every listAgentFiles() stem has a .compact.md source sibling; checks disk per stem now instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4407): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
331 lines
15 KiB
JavaScript
331 lines
15 KiB
JavaScript
// Workflow .md / agent .md / command .md / reference .md files — their text
|
|
// IS what the runtime loads. Testing text content tests the deployed contract.
|
|
// Per CONTRIBUTING.md exception matrix.
|
|
|
|
/**
|
|
* Agent size budget (measured in BYTES — see #717).
|
|
*
|
|
* Agent definitions in `agents/gsd-*.md` are loaded verbatim into the agent's
|
|
* context on every subagent dispatch. Unbounded growth is paid on every call
|
|
* across every workflow.
|
|
*
|
|
* ## Enforcement model (issue #1074)
|
|
*
|
|
* Mirrors tests/workflow-size-budget.test.cjs — two complementary guards, no
|
|
* tier-max ceiling:
|
|
*
|
|
* 1. Per-agent baseline (the anti-creep): every agent is pinned to its exact
|
|
* byte size in `tests/agent-size-baseline.json`. Any growth fails with the
|
|
* file and delta; `npm run size:baseline` records a deliberate change as a
|
|
* reviewable one-line diff. This replaced the tier-max tighten-only ratchet
|
|
* (which only bound the single largest agent per tier).
|
|
*
|
|
* 2. Tier hard caps (the outer bound): XL/LARGE/DEFAULT absolute red lines
|
|
* with real headroom, never raised in normal work. Crossing one means
|
|
* extracting shared boilerplate to `gsd-core/references/`, not a +N bump.
|
|
* A net-new agent is DEFAULT-tier, so the DEFAULT cap already bounds it —
|
|
* no separate new-file cap is needed (DEFAULT is already small).
|
|
*
|
|
* Tiers:
|
|
* - XL : top-level orchestrators that own end-to-end rubrics
|
|
* - LARGE : multi-phase operators with branching workflows
|
|
* - DEFAULT : focused single-purpose agents
|
|
*
|
|
* See:
|
|
* - https://github.com/open-gsd/gsd-core/issues/1074 (per-file baseline + hard caps)
|
|
* - https://github.com/open-gsd/gsd-core/issues/717 (bytes, not lines)
|
|
* - https://github.com/open-gsd/gsd-core/issues/683 (LF-normalized byte count)
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('fs');
|
|
const os = require('node:os');
|
|
const path = require('path');
|
|
const {
|
|
lfByteCount,
|
|
measureMdFiles,
|
|
MARGIN_RATIO,
|
|
marginFor,
|
|
buildHeadroomRows,
|
|
formatHeadroomTable,
|
|
buildHeadroomSummaryMarkdown,
|
|
appendHeadroomStepSummary,
|
|
} = require('../scripts/workflow-size.cjs');
|
|
const { cleanup } = require('./helpers.cjs');
|
|
|
|
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
|
|
const isGsdAgent = (f) => f.startsWith('gsd-');
|
|
|
|
// Tier HARD CAPS (#1074, bytes) — absolute red lines, not high-water-hugging
|
|
// ceilings. Day-to-day creep is caught per-agent by the baseline guard below;
|
|
// these sit above each tier's current high-water with real headroom:
|
|
// XL 56 KiB
|
|
// LARGE 48 KiB
|
|
// DEFAULT 24 KiB
|
|
//
|
|
// #4261: the per-tier high-water marks that used to be written out here went
|
|
// stale — the LARGE line claimed "gsd-executor 42,342 → ~6.8 KB headroom"
|
|
// while the real high-water had reached 99.6% of the cap, so the comment
|
|
// documenting the margin was itself the reason nobody noticed the margin was
|
|
// gone. Hand-maintained measurements of a moving tree do not survive; the
|
|
// headroom census below emits the live numbers on every run instead.
|
|
const XL_CAP = 57344; // 56 KiB
|
|
const LARGE_CAP = 49152; // 48 KiB
|
|
const DEFAULT_CAP = 24576; // 24 KiB
|
|
|
|
const XL_AGENTS = new Set([
|
|
'gsd-debugger',
|
|
'gsd-planner',
|
|
]);
|
|
|
|
const LARGE_AGENTS = new Set([
|
|
'gsd-phase-researcher',
|
|
'gsd-verifier',
|
|
'gsd-doc-writer',
|
|
'gsd-plan-checker',
|
|
'gsd-executor',
|
|
'gsd-code-fixer',
|
|
'gsd-codebase-mapper',
|
|
'gsd-project-researcher',
|
|
'gsd-roadmapper',
|
|
]);
|
|
|
|
const ALL_AGENTS = fs.readdirSync(AGENTS_DIR)
|
|
.filter(f => isGsdAgent(f) && f.endsWith('.md'))
|
|
.map(f => f.replace('.md', ''));
|
|
|
|
// #4407: a `.compact.md` variant is a terser rewrite of its canonical sibling,
|
|
// not a new agent — it belongs to the SAME complexity tier. Stripping the
|
|
// `.compact` suffix before lookup lets e.g. `gsd-planner.compact` inherit
|
|
// `gsd-planner`'s XL tier instead of silently falling through to DEFAULT,
|
|
// which would apply a cap sized for a "focused single-purpose agent" to a
|
|
// compacted rewrite of an XL top-level orchestrator.
|
|
function canonicalStem(agent) {
|
|
return agent.endsWith('.compact') ? agent.slice(0, -'.compact'.length) : agent;
|
|
}
|
|
|
|
function capFor(agent) {
|
|
const stem = canonicalStem(agent);
|
|
if (XL_AGENTS.has(stem)) return { tier: 'XL', cap: XL_CAP };
|
|
if (LARGE_AGENTS.has(stem)) return { tier: 'LARGE', cap: LARGE_CAP };
|
|
return { tier: 'DEFAULT', cap: DEFAULT_CAP };
|
|
}
|
|
|
|
describe('SIZE: agent tier hard caps (issue #1074)', () => {
|
|
// Absolute outer bound per tier. A cap is NOT raised when an agent approaches
|
|
// it — crossing it means extract shared boilerplate to gsd-core/references/.
|
|
for (const agent of ALL_AGENTS) {
|
|
const { tier, cap } = capFor(agent);
|
|
test(`${agent} (${tier}) stays within the ${tier} hard cap (${cap} bytes)`, () => {
|
|
const bytes = lfByteCount(path.join(AGENTS_DIR, agent + '.md'));
|
|
assert.ok(
|
|
bytes <= cap,
|
|
`${agent}.md is ${bytes} bytes — exceeds the ${tier} hard cap of ${cap}. ` +
|
|
`This cap is a red line, NOT a budget to raise: extract shared boilerplate ` +
|
|
`to gsd-core/references/ and load it lazily.`
|
|
);
|
|
});
|
|
}
|
|
});
|
|
|
|
describe('SIZE: agent hard-cap boundary fixtures (#1074 — negative proof)', () => {
|
|
// The per-tier loop above only iterates the real, fully-compliant agent
|
|
// corpus, so its `bytes <= cap` failure branch never executes. Exercise that
|
|
// exact comparison on synthetic files measured at cap-1 / cap / cap+1 (the
|
|
// limit boundary — RULESET.TESTS.boundary-coverage.fixtures) through the SAME
|
|
// lfByteCount path the guard uses, so a future threshold or operator edit
|
|
// cannot silently neuter a cap (RULESET.TESTS.regression-must-fail-first).
|
|
test('cap comparison fires at the limit boundary for every tier', () => {
|
|
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-'));
|
|
try {
|
|
// ASCII 'a' is 1 byte/char and has no CRLF, so lfByteCount == length.
|
|
const measureAt = (n) => {
|
|
const p = path.join(tmp, `fixture-${n}.md`);
|
|
fs.writeFileSync(p, 'a'.repeat(n));
|
|
return lfByteCount(p);
|
|
};
|
|
for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) {
|
|
assert.equal(measureAt(cap - 1) <= cap, true, `${cap - 1} must be within cap ${cap}`);
|
|
assert.equal(measureAt(cap) <= cap, true, `${cap} (exactly at cap) must be within cap ${cap}`);
|
|
assert.equal(measureAt(cap + 1) <= cap, false, `${cap + 1} must exceed cap ${cap}`);
|
|
}
|
|
} finally {
|
|
cleanup(tmp);
|
|
}
|
|
});
|
|
});
|
|
|
|
// A prior "SIZE: per-agent baseline (issue #1074)" describe block lived here,
|
|
// asserting every agent's exact byte count against the committed
|
|
// `tests/agent-size-baseline.json` snapshot. #2724 (ADR-2719 Phase 4) deletes that
|
|
// snapshot: it was a pure function of the source tree, and its purpose — "growth
|
|
// must be noticed and justified" — is now served by the same differential machine
|
|
// that replaced the golden-install-parity fixtures (tests/emitted-attribution.test.cjs's
|
|
// real-tree test, via `emitted-diff.cjs`'s size ratchet: growth is reported with its
|
|
// exact byte delta and requires an entry in tests/emitted-drift-ack.json, ADR-2719 §4 /
|
|
// must-have 6). The tier hard caps above are unaffected — they are independent of the
|
|
// deleted baseline and remain the outer bound.
|
|
|
|
// ─── #4261: headroom visibility + reserved margin ──────────────────────────
|
|
//
|
|
// The hard caps above are red lines and this changes none of them. What was
|
|
// missing is everything BELOW the red line: a passing run said nothing, so a
|
|
// contributor sitting at 99.6% of a cap and one at 60% got identical
|
|
// feedback — green — and the density that produces merge-time collisions was
|
|
// invisible to the people creating it.
|
|
//
|
|
// Two levels, matching the shape `execute-phase.md` already has by hand (a
|
|
// hard ceiling plus a lower margin "so minor future edits don't re-trip the
|
|
// gate"), which until now was the only capped file with one:
|
|
//
|
|
// 1. the census, printed every run, green or not
|
|
// 2. the reserved margin, which REPORTS rather than fails
|
|
//
|
|
// The margin deliberately does not fail. A cap breach is a red line; a file
|
|
// at 96% is not broken, it is a file whose next contributor should know to
|
|
// extract before adding. Failing there would turn a warning into a second
|
|
// red line and force exactly the +N bumps this policy forbids.
|
|
const AGENT_HEADROOM_ROWS = buildHeadroomRows(
|
|
measureMdFiles(AGENTS_DIR, isGsdAgent),
|
|
capFor,
|
|
);
|
|
|
|
describe('SIZE: agent headroom census (issue #4261)', () => {
|
|
test('reports every agent\'s remaining bytes, and never fails for it', (t) => {
|
|
for (const line of formatHeadroomTable(AGENT_HEADROOM_ROWS)) t.diagnostic(line);
|
|
const pressured = AGENT_HEADROOM_ROWS.filter((r) => r.overMargin);
|
|
t.diagnostic(
|
|
`agents: ${AGENT_HEADROOM_ROWS.length} | over the ${Math.round(MARGIN_RATIO * 100)}% margin: ${pressured.length}`,
|
|
);
|
|
appendHeadroomStepSummary('Agent size headroom', AGENT_HEADROOM_ROWS);
|
|
|
|
// The census is reporting, not a gate — the only thing asserted is that it
|
|
// measured the corpus at all. A census that silently went empty (a moved
|
|
// directory, a broken predicate) would otherwise read as good news.
|
|
assert.equal(AGENT_HEADROOM_ROWS.length, ALL_AGENTS.length);
|
|
});
|
|
|
|
test('names the agents inside the reserved margin', (t) => {
|
|
for (const r of AGENT_HEADROOM_ROWS.filter((row) => row.overMargin)) {
|
|
t.diagnostic(
|
|
`RESERVED MARGIN: ${r.name}.md is ${r.bytes} bytes — ${r.headroom} under the ${r.tier} cap ` +
|
|
`(${r.usedPct.toFixed(1)}%), past the ${r.margin}-byte margin. The cap is not moving: ` +
|
|
`extract shared boilerplate to gsd-core/references/ and load it lazily before adding more.`,
|
|
);
|
|
}
|
|
// Intentionally no assertion on the COUNT. Pinning "3 agents are over the
|
|
// margin" would make this a baseline that every extraction has to update,
|
|
// which is the maintenance burden #2724 removed when it deleted the
|
|
// per-file size snapshot. The hard caps stay the only failing gate.
|
|
});
|
|
});
|
|
|
|
describe('SIZE: reserved-margin boundary fixtures (#4261 — negative proof)', () => {
|
|
// The margin loop above reports whatever the real corpus happens to be, so
|
|
// its comparison branch is not exercised by construction — the same gap the
|
|
// hard-cap fixtures above exist to close. Pin marginFor and the > operator
|
|
// at the boundary so a future ratio or operator edit cannot quietly widen
|
|
// the margin to nothing (RULESET.TESTS.boundary-coverage.fixtures).
|
|
test('marginFor sits strictly below its cap and fires at the boundary', () => {
|
|
for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) {
|
|
const margin = marginFor(cap);
|
|
assert.ok(margin < cap, `margin ${margin} must sit below cap ${cap}`);
|
|
assert.equal(margin > cap * MARGIN_RATIO - 1, true, `margin ${margin} must track the ratio`);
|
|
const rows = buildHeadroomRows({
|
|
'below.md': margin - 1,
|
|
'exact.md': margin,
|
|
'above.md': margin + 1,
|
|
}, () => ({ tier: 'FIXTURE', cap }));
|
|
const byName = new Map(rows.map((row) => [row.name, row]));
|
|
assert.equal(byName.get('below').overMargin, false, 'margin - 1 is NOT over it');
|
|
assert.equal(byName.get('exact').overMargin, false, 'exactly at the margin is NOT over it');
|
|
assert.equal(byName.get('above').overMargin, true, 'margin + 1 IS over it');
|
|
}
|
|
});
|
|
|
|
test('marginFor floors, so a tiny cap can never produce a margin at the cap', () => {
|
|
// Rounding here would put the margin ON the cap for small caps, making the
|
|
// warning fire only when the hard gate already had.
|
|
assert.equal(marginFor(1), 0);
|
|
assert.equal(marginFor(20), 19);
|
|
assert.ok(marginFor(20) < 20);
|
|
});
|
|
});
|
|
|
|
describe('SIZE: headroom job summary (#4261)', () => {
|
|
// The census is only useful if it reaches a human. The diagnostics above go
|
|
// to the test log; this is the copy that lands on the PR's checks page,
|
|
// which is where a reviewer actually looks.
|
|
const row = (over) => ({
|
|
name: 'gsd-example', tier: 'LARGE', bytes: over ? 49000 : 10000,
|
|
cap: 49152, margin: 46694, headroom: over ? 152 : 39152,
|
|
usedPct: over ? 99.7 : 20.3, overMargin: over,
|
|
});
|
|
|
|
test('a clean corpus still reports how many files it measured', () => {
|
|
const md = buildHeadroomSummaryMarkdown('Agent size headroom', [row(false), row(false)]);
|
|
// "no table" must be distinguishable from "nothing ran".
|
|
assert.match(md, /All 2 files are under the 95% reserved margin\./);
|
|
assert.doesNotMatch(md, /\| file \|/);
|
|
});
|
|
|
|
test('a pressured corpus renders one table row per file over the margin', () => {
|
|
const md = buildHeadroomSummaryMarkdown('Agent size headroom', [row(true), row(false)]);
|
|
assert.match(md, /\*\*1 of 2\*\* files are over the 95% reserved margin\./);
|
|
assert.match(md, /\| `gsd-example` \| LARGE \| 49000 \| 49152 \| 152 \| 99\.7% \|/);
|
|
});
|
|
|
|
test('writes to GITHUB_STEP_SUMMARY when set, and is a no-op when it is not', () => {
|
|
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-summary-'));
|
|
try {
|
|
const file = path.join(tmp, 'summary.md');
|
|
assert.equal(appendHeadroomStepSummary('T', [row(true)], { GITHUB_STEP_SUMMARY: file }), true);
|
|
assert.match(fs.readFileSync(file, 'utf8'), /### T/);
|
|
assert.equal(appendHeadroomStepSummary('T', [row(true)], {}), false);
|
|
} finally {
|
|
cleanup(tmp);
|
|
}
|
|
});
|
|
|
|
test('an unwritable summary path reports but does not fail the run', () => {
|
|
// Reporting must never be able to red a suite that is otherwise green —
|
|
// the whole point of this block is that it is additive.
|
|
const unwritable = path.join(os.tmpdir(), 'agent-size-nope', 'nested', 'summary.md');
|
|
assert.equal(appendHeadroomStepSummary('T', [row(true)], { GITHUB_STEP_SUMMARY: unwritable }), false);
|
|
});
|
|
});
|
|
|
|
describe('SIZE: every agent is classified', () => {
|
|
test('every agent falls in exactly one tier', () => {
|
|
for (const agent of ALL_AGENTS) {
|
|
const inXL = XL_AGENTS.has(agent);
|
|
const inLarge = LARGE_AGENTS.has(agent);
|
|
assert.ok(
|
|
!(inXL && inLarge),
|
|
`${agent} is in both XL_AGENTS and LARGE_AGENTS — pick one`
|
|
);
|
|
}
|
|
});
|
|
|
|
test('every named XL agent exists', () => {
|
|
for (const agent of XL_AGENTS) {
|
|
const filePath = path.join(AGENTS_DIR, agent + '.md');
|
|
assert.ok(
|
|
fs.existsSync(filePath),
|
|
`XL_AGENTS references ${agent}.md which does not exist — clean up the set`
|
|
);
|
|
}
|
|
});
|
|
|
|
test('every named LARGE agent exists', () => {
|
|
for (const agent of LARGE_AGENTS) {
|
|
const filePath = path.join(AGENTS_DIR, agent + '.md');
|
|
assert.ok(
|
|
fs.existsSync(filePath),
|
|
`LARGE_AGENTS references ${agent}.md which does not exist — clean up the set`
|
|
);
|
|
}
|
|
});
|
|
});
|