diff --git a/CONTEXT.md b/CONTEXT.md index b11ebd623..330972b03 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -268,6 +268,7 @@ The canonical lint infrastructure adopted in ADR 452 (`docs/adr/452-eslint-lint- `RULESET.WORKFLOW_MARKDOWN.FENCES=preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)` `RULESET.WORKFLOW_SIZE_BUDGET=workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = per-file baseline (PRIMARY anti-creep: tests/workflow-size-baseline.json pins each file's exact size) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + new-file cap (un-baselined files <32768, the Codex anchor) + discuss-phase<32000; a file that grew fails the baseline guard — fix with `npm run size:baseline`, commit the one-line diff, and justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump` +`RULESET.AGENT_SIZE_BUDGET=agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = per-file baseline (PRIMARY anti-creep: tests/agent-size-baseline.json pins each agents/gsd-*.md exact byte size) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). One 'npm run size:baseline' regenerates BOTH workflow and agent baselines via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter. A grown agent fails the baseline guard — regenerate + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes` `RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name` `RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/bug-3135-capture-backlog-workflow.test.cjs; INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill` `RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets` diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index cd7a5b0f4..2594adee7 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -856,9 +856,10 @@ gsd-core/ If you legitimately grow or shrink a workflow file, run `npm run size:baseline` to update the snapshot and justify any growth in your PR (or extract content - lazily). Full how-to + reference in - docs/TESTING-SUITES.md (Workflow size budget); see - issue #1074. + lazily). The same guard covers agent files + (agents/gsd-*.md). Full how-to + reference in + docs/TESTING-SUITES.md (Workflow & agent size + budget); see issue #1074. references/ — Reference documentation (.md) templates/ — File templates agents/ — Agent definitions (.md) — CANONICAL SOURCE diff --git a/docs/TESTING-SUITES.md b/docs/TESTING-SUITES.md index 35e546f7e..c026b61bc 100644 --- a/docs/TESTING-SUITES.md +++ b/docs/TESTING-SUITES.md @@ -59,17 +59,21 @@ the sanctioned layout (see the #443 strategy below), not a one-off regression pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way `bug-*` did, extend the ratchet's regex and regenerate the allowlist. -## Workflow size budget +## Workflow & agent size budget > Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074). > Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717); > LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683). -Workflow files (`gsd-core/workflows/*.md`) ship in the installed runtime and are -loaded into the agent's context, so their byte size is a real cost. The size -guard in `tests/workflow-size-budget.test.cjs` keeps that cost from creeping up -invisibly. It is an **anti-creep ratchet**, sibling to the regression-name -ratchet above — three layers, ordered from day-to-day to last-resort: +Workflow files (`gsd-core/workflows/*.md`) and agent files (`agents/gsd-*.md`) +both ship in the installed runtime and are loaded into context — workflows on +every command, agents on every subagent dispatch — so their byte size is a real +cost. Two sibling guards (`tests/workflow-size-budget.test.cjs` and +`tests/agent-size-budget.test.cjs`) keep that cost from creeping up invisibly, +sharing one byte-counter (`measureMdFiles`) and one `npm run size:baseline` +command that regenerates **both** snapshots. Each is an **anti-creep ratchet**, +sibling to the regression-name ratchet above — three layers (workflows), ordered +from day-to-day to last-resort: | Layer | What it does | Where | |---|---|---| @@ -80,9 +84,18 @@ ratchet above — three layers, ordered from day-to-day to last-resort: `discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes (issue [#2551](https://github.com/open-gsd/gsd-core/issues/2551)). -### How-to: a workflow grew and CI is red +**Agents** (`tests/agent-size-budget.test.cjs`) use the same per-agent baseline +(`tests/agent-size-baseline.json`) + loose tier hard caps — `XL ≤ 57344` / +`LARGE ≤ 49152` / `DEFAULT ≤ 24576` bytes. There is no new-agent cap: a net-new +agent is DEFAULT-tier and already bounded by the DEFAULT cap. (This is distinct +from the separate 45 KB-*char* extraction-evidence threshold on `gsd-planner` +enforced by `tests/planner-decomposition.test.cjs` — that one proves mode +sections were extracted; this one bounds total agent bytes.) -The baseline guard reports the file and the byte delta. To resolve: +### How-to: a workflow or agent grew and CI is red + +The baseline guard reports the file and the byte delta (the same flow for both +the workflow and agent guards). To resolve: 1. **Regenerate the snapshot** and inspect the one-line diff: ```bash @@ -93,9 +106,10 @@ The baseline guard reports the file and the byte delta. To resolve: the committed baseline diff is the review record that the larger size was a deliberate, seen decision, not silent drift. 3. **Or shrink it instead of baselining.** Prefer extraction when the growth is - incidental: move per-mode bodies to `workflows//modes/`, templates to - `workflows//templates/`, and shared prose to `gsd-core/references/` — - then load them **LAZILY**. Do *not* convert them to eager `@-required_reading` + incidental: for a workflow, move per-mode bodies to `workflows//modes/`, + templates to `workflows//templates/`, and shared prose to + `gsd-core/references/`; for an agent, lift shared boilerplate into + `gsd-core/references/` and `@`-reference it — then load it **LAZILY**. Do *not* convert them to eager `@-required_reading` includes: that shrinks the file's bytes without shrinking loaded context, so it games the guard while making the real cost worse. See `workflows/discuss-phase/` for the progressive-disclosure pattern. @@ -107,10 +121,12 @@ that is the signal to extract, per step 3. | Artifact | Role | |---|---| -| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) plus workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guard and the generator so they can never measure or enumerate differently. | -| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates `tests/workflow-size-baseline.json` — sorted keys, trailing newline, idempotent. | -| `tests/workflow-size-baseline.json` | The committed per-file snapshot (one entry per workflow). | -| `tests/workflow-size-budget.test.cjs` | The three guards above, plus the `discuss-phase` progressive-disclosure checks. | +| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) + generic `measureMdFiles(dir, predicate)` (backs both workflows and agents) + workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guards and the generator so they can never measure differently. | +| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates **both** `tests/workflow-size-baseline.json` and `tests/agent-size-baseline.json` — sorted keys, trailing newline, idempotent. | +| `tests/workflow-size-baseline.json` | The committed per-workflow snapshot (one entry per workflow). | +| `tests/agent-size-baseline.json` | The committed per-agent snapshot (one entry per `gsd-*` agent). | +| `tests/workflow-size-budget.test.cjs` | The three workflow guards above, plus the `discuss-phase` progressive-disclosure checks. | +| `tests/agent-size-budget.test.cjs` | The per-agent baseline + tier hard-cap guards (the agent analog). | ## Running suites locally diff --git a/scripts/update-size-baseline.cjs b/scripts/update-size-baseline.cjs index 6dd494d36..1ff6b423d 100644 --- a/scripts/update-size-baseline.cjs +++ b/scripts/update-size-baseline.cjs @@ -16,9 +16,12 @@ const fs = require('fs'); const path = require('path'); -const { measureWorkflows, WORKFLOWS_DIR } = require('./workflow-size.cjs'); +const { measureMdFiles, WORKFLOWS_DIR } = require('./workflow-size.cjs'); const BASELINE_PATH = path.join(__dirname, '..', 'tests', 'workflow-size-baseline.json'); +const AGENTS_DIR = path.join(__dirname, '..', 'agents'); +const AGENT_BASELINE_PATH = path.join(__dirname, '..', 'tests', 'agent-size-baseline.json'); +const isGsdAgent = (f) => f.startsWith('gsd-'); /** * Serialize a size map to the on-disk baseline format: keys sorted, 2-space @@ -34,25 +37,32 @@ function serializeBaseline(sizes) { } /** - * Write the baseline file from the measured workflow sizes. + * Write a baseline file from the measured `.md` sizes in a directory. * * @param {object} [opts] - * @param {string} [opts.dir] - Workflows dir to measure (default canonical). - * @param {string} [opts.outPath] - Baseline file to write (default canonical). + * @param {string} [opts.dir] - Directory to measure (default: workflows). + * @param {string} [opts.outPath] - Baseline file to write (default: workflow). + * @param {function(string): boolean} [opts.predicate] - Filename filter. * @returns {{ outPath: string, count: number, content: string }} */ -function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH } = {}) { - const sizes = measureWorkflows(dir); +function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH, predicate } = {}) { + const sizes = measureMdFiles(dir, predicate); const content = serializeBaseline(sizes); fs.writeFileSync(outPath, content); return { outPath, count: Object.keys(sizes).length, content }; } if (require.main === module) { - const { outPath, count } = generateBaseline(); - process.stdout.write( - `Wrote ${count} workflow sizes to ${path.relative(process.cwd(), outPath)}\n` - ); + const targets = [ + { label: 'workflow', dir: WORKFLOWS_DIR, outPath: BASELINE_PATH }, + { label: 'agent', dir: AGENTS_DIR, outPath: AGENT_BASELINE_PATH, predicate: isGsdAgent }, + ]; + for (const t of targets) { + const { outPath, count } = generateBaseline(t); + process.stdout.write( + `Wrote ${count} ${t.label} sizes to ${path.relative(process.cwd(), outPath)}\n` + ); + } } -module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH }; +module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH, AGENT_BASELINE_PATH }; diff --git a/scripts/workflow-size.cjs b/scripts/workflow-size.cjs index 9a2f23e31..990359d74 100644 --- a/scripts/workflow-size.cjs +++ b/scripts/workflow-size.cjs @@ -51,18 +51,40 @@ function listWorkflowStems(dir = WORKFLOWS_DIR) { } /** - * Measure every top-level workflow file, keyed by filename (`.md`). + * Measure every top-level `.md` file in `dir`, keyed by filename, byte sizes. + * Generic over directory and an optional filename predicate — used for both + * workflows (`gsd-core/workflows/*.md`) and agents (`agents/gsd-*.md`) so the + * size guards and the baseline generator share one measurement path (#1074). + * Non-recursive by design. * - * @param {string} [dir] - Workflows directory (defaults to the canonical one). - * @returns {Object} Map of `.md` → LF byte size, with - * keys inserted in sorted order. + * @param {string} dir - Directory to scan. + * @param {function(string): boolean} [predicate] - Filename filter (default: all `.md`). + * @returns {Object} Map of filename → LF byte size, keys sorted. */ -function measureWorkflows(dir = WORKFLOWS_DIR) { +function measureMdFiles(dir, predicate = () => true) { const out = {}; - for (const stem of listWorkflowStems(dir)) { - out[`${stem}.md`] = lfByteCount(path.join(dir, `${stem}.md`)); - } + const names = fs + .readdirSync(dir) + .filter((f) => f.endsWith('.md') && predicate(f)) + .sort(); + for (const name of names) out[name] = lfByteCount(path.join(dir, name)); return out; } -module.exports = { WORKFLOWS_DIR, lfByteCount, listWorkflowStems, measureWorkflows }; +/** + * Measure every top-level workflow file, keyed by filename (`.md`). + * + * @param {string} [dir] - Workflows directory (defaults to the canonical one). + * @returns {Object} Map of `.md` → LF byte size, sorted. + */ +function measureWorkflows(dir = WORKFLOWS_DIR) { + return measureMdFiles(dir); +} + +module.exports = { + WORKFLOWS_DIR, + lfByteCount, + listWorkflowStems, + measureMdFiles, + measureWorkflows, +}; diff --git a/tests/agent-size-baseline.json b/tests/agent-size-baseline.json new file mode 100644 index 000000000..30b0b83a2 --- /dev/null +++ b/tests/agent-size-baseline.json @@ -0,0 +1,35 @@ +{ + "gsd-advisor-researcher.md": 4536, + "gsd-ai-researcher.md": 5851, + "gsd-assumptions-analyzer.md": 4489, + "gsd-code-fixer.md": 36499, + "gsd-code-reviewer.md": 16773, + "gsd-codebase-mapper.md": 21388, + "gsd-debug-session-manager.md": 14159, + "gsd-debugger.md": 51213, + "gsd-doc-classifier.md": 7629, + "gsd-doc-synthesizer.md": 9722, + "gsd-doc-verifier.md": 12403, + "gsd-doc-writer.md": 38827, + "gsd-domain-researcher.md": 6938, + "gsd-eval-auditor.md": 7754, + "gsd-eval-planner.md": 7008, + "gsd-executor.md": 42512, + "gsd-framework-selector.md": 6778, + "gsd-integration-checker.md": 15141, + "gsd-intel-updater.md": 18122, + "gsd-nyquist-auditor.md": 7245, + "gsd-pattern-mapper.md": 12487, + "gsd-phase-researcher.md": 40611, + "gsd-plan-checker.md": 41996, + "gsd-planner.md": 48885, + "gsd-project-researcher.md": 21987, + "gsd-research-synthesizer.md": 13646, + "gsd-roadmapper.md": 19652, + "gsd-security-auditor.md": 6216, + "gsd-ui-auditor.md": 17152, + "gsd-ui-checker.md": 11081, + "gsd-ui-researcher.md": 19265, + "gsd-user-profiler.md": 8516, + "gsd-verifier.md": 41285 +} diff --git a/tests/agent-size-budget.test.cjs b/tests/agent-size-budget.test.cjs index 3dfddc173..6e440cebe 100644 --- a/tests/agent-size-budget.test.cjs +++ b/tests/agent-size-budget.test.cjs @@ -4,48 +4,62 @@ // Per CONTRIBUTING.md exception matrix. /** - * Agent size budget. + * Agent size budget (measured in BYTES — see #717). * - * Agent definitions in `agents/gsd-*.md` are loaded verbatim into Claude's + * Agent definitions in `agents/gsd-*.md` are loaded verbatim into the agent's * context on every subagent dispatch. Unbounded growth is paid on every call * across every workflow. * - * Budgets are tiered to reflect the intent of each agent class: + * ## Enforcement model (issue #1074) + * + * Mirrors tests/workflow-size-budget.test.cjs — two complementary guards, no + * tier-max ceiling: + * + * 1. Per-agent baseline (the anti-creep): every agent is pinned to its exact + * byte size in `tests/agent-size-baseline.json`. Any growth fails with the + * file and delta; `npm run size:baseline` records a deliberate change as a + * reviewable one-line diff. This replaced the tier-max tighten-only ratchet + * (which only bound the single largest agent per tier). + * + * 2. Tier hard caps (the outer bound): XL/LARGE/DEFAULT absolute red lines + * with real headroom, never raised in normal work. Crossing one means + * extracting shared boilerplate to `gsd-core/references/`, not a +N bump. + * A net-new agent is DEFAULT-tier, so the DEFAULT cap already bounds it — + * no separate new-file cap is needed (DEFAULT is already small). + * + * Tiers: * - XL : top-level orchestrators that own end-to-end rubrics * - LARGE : multi-phase operators with branching workflows * - DEFAULT : focused single-purpose agents * - * Raising a budget is a deliberate choice — adjust the constant, write a - * rationale in the PR, and make sure the bloat is not duplicated content - * that belongs in `gsd-core/references/`. - * - * Tighten-only invariant (issue #597): ceilings track the tier high-water mark - * within GRACE lines. Budgets may only decrease, never silently creep upward. - * The assertTightCeiling() call below enforces this automatically. - * - * See: https://github.com/open-gsd/gsd-core/issues/2361 + * See: + * - https://github.com/open-gsd/gsd-core/issues/1074 (per-file baseline + hard caps) + * - https://github.com/open-gsd/gsd-core/issues/717 (bytes, not lines) + * - https://github.com/open-gsd/gsd-core/issues/683 (LF-normalized byte count) */ const { test, describe } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('fs'); +const os = require('node:os'); const path = require('path'); -const { assertTightCeiling } = require('../scripts/lib/allowlist-ratchet.cjs'); +const { assertFileBaseline } = require('../scripts/lib/allowlist-ratchet.cjs'); +const { lfByteCount, measureMdFiles } = require('../scripts/workflow-size.cjs'); +const { cleanup } = require('./helpers.cjs'); const AGENTS_DIR = path.join(__dirname, '..', 'agents'); +const BASELINE_PATH = path.join(__dirname, 'agent-size-baseline.json'); +const isGsdAgent = (f) => f.startsWith('gsd-'); -// Ceilings tightened to actualMax + GRACE per the ratchet-down rule (#597). -// XL ceiling lowered from 1600 → 1512 (actualMax=1452, gsd-debugger). -const XL_BUDGET = 1512; -// LARGE ceiling kept at 1000 (actualMax=978, slack=22 ≤ GRACE=60). -const LARGE_BUDGET = 1000; -// DEFAULT ceiling kept at 500 (actualMax=495, slack=5 ≤ GRACE=60). -const DEFAULT_BUDGET = 500; - -// Grace band: maximum allowed slack (ceiling − actualMax) before a ceiling is -// considered too loose. 60 lines gives one reasonable screen of breathing room -// without permitting gross inflation. -const GRACE = 60; +// Tier HARD CAPS (#1074, bytes) — absolute red lines, not high-water-hugging +// ceilings. Day-to-day creep is caught per-agent by the baseline guard below; +// these sit above each tier's current high-water with real headroom: +// XL 56 KiB — high-water gsd-debugger 51,043 → ~6.3 KB headroom +// LARGE 48 KiB — high-water gsd-executor 42,342 → ~6.8 KB headroom +// DEFAULT 24 KiB — high-water gsd-ui-researcher 19,095 → ~5.5 KB headroom +const XL_CAP = 57344; // 56 KiB +const LARGE_CAP = 49152; // 48 KiB +const DEFAULT_CAP = 24576; // 24 KiB const XL_AGENTS = new Set([ 'gsd-debugger', @@ -65,64 +79,76 @@ const LARGE_AGENTS = new Set([ ]); const ALL_AGENTS = fs.readdirSync(AGENTS_DIR) - .filter(f => f.startsWith('gsd-') && f.endsWith('.md')) + .filter(f => isGsdAgent(f) && f.endsWith('.md')) .map(f => f.replace('.md', '')); -function budgetFor(agent) { - if (XL_AGENTS.has(agent)) return { tier: 'XL', limit: XL_BUDGET }; - if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', limit: LARGE_BUDGET }; - return { tier: 'DEFAULT', limit: DEFAULT_BUDGET }; +function capFor(agent) { + if (XL_AGENTS.has(agent)) return { tier: 'XL', cap: XL_CAP }; + if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', cap: LARGE_CAP }; + return { tier: 'DEFAULT', cap: DEFAULT_CAP }; } -function lineCount(filePath) { - const content = fs.readFileSync(filePath, 'utf-8'); - if (content.length === 0) return 0; - const trailingNewline = content.endsWith('\n') ? 1 : 0; - return content.split('\n').length - trailingNewline; -} - -describe('SIZE: agent line-count budget', () => { +describe('SIZE: agent tier hard caps (issue #1074)', () => { + // Absolute outer bound per tier. A cap is NOT raised when an agent approaches + // it — crossing it means extract shared boilerplate to gsd-core/references/. for (const agent of ALL_AGENTS) { - const { tier, limit } = budgetFor(agent); - test(`${agent} (${tier}) stays under ${limit} lines`, () => { - const filePath = path.join(AGENTS_DIR, agent + '.md'); - const lines = lineCount(filePath); + const { tier, cap } = capFor(agent); + test(`${agent} (${tier}) stays within the ${tier} hard cap (${cap} bytes)`, () => { + const bytes = lfByteCount(path.join(AGENTS_DIR, agent + '.md')); assert.ok( - lines <= limit, - `${agent}.md has ${lines} lines — exceeds ${tier} budget of ${limit}. ` + - `Extract shared boilerplate to gsd-core/references/ or raise the budget ` + - `in tests/agent-size-budget.test.cjs with a rationale.` + bytes <= cap, + `${agent}.md is ${bytes} bytes — exceeds the ${tier} hard cap of ${cap}. ` + + `This cap is a red line, NOT a budget to raise: extract shared boilerplate ` + + `to gsd-core/references/ and load it lazily.` ); }); } }); -describe('SIZE: tier anti-creep (tighten-only ceilings, issue #597)', () => { - // For each tier, compute the high-water mark across all files in that tier - // and assert the ceiling stays tight. Prevents budgets from silently drifting - // upward: ceiling − actualMax must not exceed GRACE. - test('XL tier: ceiling tracks high-water mark within GRACE', () => { - const values = ALL_AGENTS - .filter(a => XL_AGENTS.has(a)) - .map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); - const actualMax = Math.max(...values); - assertTightCeiling({ label: 'XL', actualMax, ceiling: XL_BUDGET, grace: GRACE, fail: assert.fail }); +describe('SIZE: agent hard-cap boundary fixtures (#1074 — negative proof)', () => { + // The per-tier loop above only iterates the real, fully-compliant agent + // corpus, so its `bytes <= cap` failure branch never executes. Exercise that + // exact comparison on synthetic files measured at cap-1 / cap / cap+1 (the + // limit boundary — RULESET.TESTS.boundary-coverage.fixtures) through the SAME + // lfByteCount path the guard uses, so a future threshold or operator edit + // cannot silently neuter a cap (RULESET.TESTS.regression-must-fail-first). + test('cap comparison fires at the limit boundary for every tier', () => { + const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-')); + try { + // ASCII 'a' is 1 byte/char and has no CRLF, so lfByteCount == length. + const measureAt = (n) => { + const p = path.join(tmp, `fixture-${n}.md`); + fs.writeFileSync(p, 'a'.repeat(n)); + return lfByteCount(p); + }; + for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) { + assert.equal(measureAt(cap - 1) <= cap, true, `${cap - 1} must be within cap ${cap}`); + assert.equal(measureAt(cap) <= cap, true, `${cap} (exactly at cap) must be within cap ${cap}`); + assert.equal(measureAt(cap + 1) <= cap, false, `${cap + 1} must exceed cap ${cap}`); + } + } finally { + cleanup(tmp); + } }); +}); - test('LARGE tier: ceiling tracks high-water mark within GRACE', () => { - const values = ALL_AGENTS - .filter(a => LARGE_AGENTS.has(a)) - .map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); - const actualMax = Math.max(...values); - assertTightCeiling({ label: 'LARGE', actualMax, ceiling: LARGE_BUDGET, grace: GRACE, fail: assert.fail }); - }); - - test('DEFAULT tier: ceiling tracks high-water mark within GRACE', () => { - const values = ALL_AGENTS - .filter(a => !XL_AGENTS.has(a) && !LARGE_AGENTS.has(a)) - .map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); - const actualMax = Math.max(...values); - assertTightCeiling({ label: 'DEFAULT', actualMax, ceiling: DEFAULT_BUDGET, grace: GRACE, fail: assert.fail }); +describe('SIZE: per-agent baseline (issue #1074)', () => { + // Per-agent exact-size ratchet — the primary anti-creep guard. Guards EVERY + // agent by name against tests/agent-size-baseline.json. Growth fails with the + // file and delta; shrinkage fails as a stale snapshot. Fix: `npm run + // size:baseline` plus a PR justification for genuine growth (or extraction). + test('every agent matches its committed baseline', () => { + const baseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8')); + const current = measureMdFiles(AGENTS_DIR, isGsdAgent); + assertFileBaseline({ + label: 'agent-size', + current, + baseline, + fail: assert.fail, + updateHint: + 'Run `npm run size:baseline` to update tests/agent-size-baseline.json, ' + + 'then justify any growth in your PR (or extract shared boilerplate to gsd-core/references/).', + }); }); }); diff --git a/tests/update-size-baseline.test.cjs b/tests/update-size-baseline.test.cjs index 26feaacea..1522fc57f 100644 --- a/tests/update-size-baseline.test.cjs +++ b/tests/update-size-baseline.test.cjs @@ -124,4 +124,21 @@ describe('generateBaseline', () => { /ENOENT/ ); }); + + test('predicate filters which files are baselined (agent path)', () => { + const agentDir = path.join(dir, 'agents'); + fs.mkdirSync(agentDir); + fs.writeFileSync(path.join(agentDir, 'gsd-one.md'), 'a\n'); + fs.writeFileSync(path.join(agentDir, 'gsd-two.md'), 'bb\n'); + fs.writeFileSync(path.join(agentDir, 'README.md'), 'not an agent\n'); + + const result = generateBaseline({ + dir: agentDir, + outPath, + predicate: (f) => f.startsWith('gsd-'), + }); + assert.strictEqual(result.count, 2, 'only gsd-*.md files should be baselined'); + const written = JSON.parse(fs.readFileSync(outPath, 'utf-8')); + assert.deepStrictEqual(Object.keys(written), ['gsd-one.md', 'gsd-two.md']); + }); });