test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)

Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the
same assertTightCeiling tier ratchet but was still line-based (never rebased in
#717). Completes the migration — the last part of the #1074 epic.

- Rebase agent sizing from lines to LF-normalized bytes (#717/#683).
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests);
  add a per-agent baseline (tests/agent-size-baseline.json) as the primary
  anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB),
  each above its tier high-water with real headroom. No separate new-file cap:
  a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap.
- Keep the agent-classification tests verbatim.
- scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate)
  (workflows + agents share one byte-measurement path); measureWorkflows now
  delegates to it.
- scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates
  BOTH the workflow and agent baselines (gsd-* filter for agents).

Rebased onto next after PR 2/3 (#1096) merged: replicate the
scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across
the generator and the agent test's require; regenerate the agent baseline
against current agents (a uniform +170 B preamble drift on all 33 since
authoring).

Addresses the #1097 review (trek-e):
- BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md.
  Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline,
  dual size:baseline, shared measureMdFiles seam) and disambiguates it from the
  separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two
  purposes).
- Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section
  is in next, fold in the agent coverage here (renamed to "Workflow & agent
  size budget"): agent caps + per-agent baseline + the how-to + reference rows,
  and the disambiguation from the 45K-char guard.
- Minor (negative proof): add a boundary-fixture test exercising the hard-cap
  comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future
  threshold/operator edit can't silently neuter a cap.
- Nit: align the tier test name wording ("stays within") with the <= operator.

Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B;
XL hard cap catches 57,516 > 57,344 with the baseline current.

Closes #1095 (PR 3/3 child); landing this completes the #1074 epic.
This commit is contained in:
Rezolv
2026-06-12 09:58:44 -04:00
committed by GitHub
parent b055ca14e2
commit e4f0910d62
8 changed files with 236 additions and 108 deletions

View File

@@ -268,6 +268,7 @@ The canonical lint infrastructure adopted in ADR 452 (`docs/adr/452-eslint-lint-
`RULESET.WORKFLOW_MARKDOWN.FENCES=preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)` `RULESET.WORKFLOW_MARKDOWN.FENCES=preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)`
`RULESET.WORKFLOW_SIZE_BUDGET=workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = per-file baseline (PRIMARY anti-creep: tests/workflow-size-baseline.json pins each file's exact size) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + new-file cap (un-baselined files <32768, the Codex anchor) + discuss-phase<32000; a file that grew fails the baseline guard — fix with `npm run size:baseline`, commit the one-line diff, and justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump` `RULESET.WORKFLOW_SIZE_BUDGET=workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = per-file baseline (PRIMARY anti-creep: tests/workflow-size-baseline.json pins each file's exact size) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + new-file cap (un-baselined files <32768, the Codex anchor) + discuss-phase<32000; a file that grew fails the baseline guard — fix with `npm run size:baseline`, commit the one-line diff, and justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump`
`RULESET.AGENT_SIZE_BUDGET=agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = per-file baseline (PRIMARY anti-creep: tests/agent-size-baseline.json pins each agents/gsd-*.md exact byte size) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). One 'npm run size:baseline' regenerates BOTH workflow and agent baselines via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter. A grown agent fails the baseline guard — regenerate + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes`
`RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; <step name="..."> XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name` `RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; <step name="..."> XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name`
`RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/bug-3135-capture-backlog-workflow.test.cjs; INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill` `RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/bug-3135-capture-backlog-workflow.test.cjs; INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill`
`RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets` `RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets`

View File

@@ -856,9 +856,10 @@ gsd-core/
If you legitimately grow or shrink a workflow file, If you legitimately grow or shrink a workflow file,
run `npm run size:baseline` to update the snapshot and run `npm run size:baseline` to update the snapshot and
justify any growth in your PR (or extract content justify any growth in your PR (or extract content
lazily). Full how-to + reference in lazily). The same guard covers agent files
docs/TESTING-SUITES.md (Workflow size budget); see (agents/gsd-*.md). Full how-to + reference in
issue #1074. docs/TESTING-SUITES.md (Workflow & agent size
budget); see issue #1074.
references/ — Reference documentation (.md) references/ — Reference documentation (.md)
templates/ — File templates templates/ — File templates
agents/ — Agent definitions (.md) — CANONICAL SOURCE agents/ — Agent definitions (.md) — CANONICAL SOURCE

View File

@@ -59,17 +59,21 @@ the sanctioned layout (see the #443 strategy below), not a one-off regression
pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way
`bug-*` did, extend the ratchet's regex and regenerate the allowlist. `bug-*` did, extend the ratchet's regex and regenerate the allowlist.
## Workflow size budget ## Workflow & agent size budget
> Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074). > Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074).
> Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717); > Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717);
> LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683). > LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683).
Workflow files (`gsd-core/workflows/*.md`) ship in the installed runtime and are Workflow files (`gsd-core/workflows/*.md`) and agent files (`agents/gsd-*.md`)
loaded into the agent's context, so their byte size is a real cost. The size both ship in the installed runtime and are loaded into context — workflows on
guard in `tests/workflow-size-budget.test.cjs` keeps that cost from creeping up every command, agents on every subagent dispatch — so their byte size is a real
invisibly. It is an **anti-creep ratchet**, sibling to the regression-name cost. Two sibling guards (`tests/workflow-size-budget.test.cjs` and
ratchet above — three layers, ordered from day-to-day to last-resort: `tests/agent-size-budget.test.cjs`) keep that cost from creeping up invisibly,
sharing one byte-counter (`measureMdFiles`) and one `npm run size:baseline`
command that regenerates **both** snapshots. Each is an **anti-creep ratchet**,
sibling to the regression-name ratchet above — three layers (workflows), ordered
from day-to-day to last-resort:
| Layer | What it does | Where | | Layer | What it does | Where |
|---|---|---| |---|---|---|
@@ -80,9 +84,18 @@ ratchet above — three layers, ordered from day-to-day to last-resort:
`discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes `discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes
(issue [#2551](https://github.com/open-gsd/gsd-core/issues/2551)). (issue [#2551](https://github.com/open-gsd/gsd-core/issues/2551)).
### How-to: a workflow grew and CI is red **Agents** (`tests/agent-size-budget.test.cjs`) use the same per-agent baseline
(`tests/agent-size-baseline.json`) + loose tier hard caps — `XL ≤ 57344` /
`LARGE ≤ 49152` / `DEFAULT ≤ 24576` bytes. There is no new-agent cap: a net-new
agent is DEFAULT-tier and already bounded by the DEFAULT cap. (This is distinct
from the separate 45 KB-*char* extraction-evidence threshold on `gsd-planner`
enforced by `tests/planner-decomposition.test.cjs` — that one proves mode
sections were extracted; this one bounds total agent bytes.)
The baseline guard reports the file and the byte delta. To resolve: ### How-to: a workflow or agent grew and CI is red
The baseline guard reports the file and the byte delta (the same flow for both
the workflow and agent guards). To resolve:
1. **Regenerate the snapshot** and inspect the one-line diff: 1. **Regenerate the snapshot** and inspect the one-line diff:
```bash ```bash
@@ -93,9 +106,10 @@ The baseline guard reports the file and the byte delta. To resolve:
the committed baseline diff is the review record that the larger size was a the committed baseline diff is the review record that the larger size was a
deliberate, seen decision, not silent drift. deliberate, seen decision, not silent drift.
3. **Or shrink it instead of baselining.** Prefer extraction when the growth is 3. **Or shrink it instead of baselining.** Prefer extraction when the growth is
incidental: move per-mode bodies to `workflows/<name>/modes/`, templates to incidental: for a workflow, move per-mode bodies to `workflows/<name>/modes/`,
`workflows/<name>/templates/`, and shared prose to `gsd-core/references/` — templates to `workflows/<name>/templates/`, and shared prose to
then load them **LAZILY**. Do *not* convert them to eager `@-required_reading` `gsd-core/references/`; for an agent, lift shared boilerplate into
`gsd-core/references/` and `@`-reference it — then load it **LAZILY**. Do *not* convert them to eager `@-required_reading`
includes: that shrinks the file's bytes without shrinking loaded context, so includes: that shrinks the file's bytes without shrinking loaded context, so
it games the guard while making the real cost worse. See it games the guard while making the real cost worse. See
`workflows/discuss-phase/` for the progressive-disclosure pattern. `workflows/discuss-phase/` for the progressive-disclosure pattern.
@@ -107,10 +121,12 @@ that is the signal to extract, per step 3.
| Artifact | Role | | Artifact | Role |
|---|---| |---|---|
| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) plus workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guard and the generator so they can never measure or enumerate differently. | | `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) + generic `measureMdFiles(dir, predicate)` (backs both workflows and agents) + workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guards and the generator so they can never measure differently. |
| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates `tests/workflow-size-baseline.json` — sorted keys, trailing newline, idempotent. | | `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates **both** `tests/workflow-size-baseline.json` and `tests/agent-size-baseline.json` — sorted keys, trailing newline, idempotent. |
| `tests/workflow-size-baseline.json` | The committed per-file snapshot (one entry per workflow). | | `tests/workflow-size-baseline.json` | The committed per-workflow snapshot (one entry per workflow). |
| `tests/workflow-size-budget.test.cjs` | The three guards above, plus the `discuss-phase` progressive-disclosure checks. | | `tests/agent-size-baseline.json` | The committed per-agent snapshot (one entry per `gsd-*` agent). |
| `tests/workflow-size-budget.test.cjs` | The three workflow guards above, plus the `discuss-phase` progressive-disclosure checks. |
| `tests/agent-size-budget.test.cjs` | The per-agent baseline + tier hard-cap guards (the agent analog). |
## Running suites locally ## Running suites locally

View File

@@ -16,9 +16,12 @@
const fs = require('fs'); const fs = require('fs');
const path = require('path'); const path = require('path');
const { measureWorkflows, WORKFLOWS_DIR } = require('./workflow-size.cjs'); const { measureMdFiles, WORKFLOWS_DIR } = require('./workflow-size.cjs');
const BASELINE_PATH = path.join(__dirname, '..', 'tests', 'workflow-size-baseline.json'); const BASELINE_PATH = path.join(__dirname, '..', 'tests', 'workflow-size-baseline.json');
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
const AGENT_BASELINE_PATH = path.join(__dirname, '..', 'tests', 'agent-size-baseline.json');
const isGsdAgent = (f) => f.startsWith('gsd-');
/** /**
* Serialize a size map to the on-disk baseline format: keys sorted, 2-space * Serialize a size map to the on-disk baseline format: keys sorted, 2-space
@@ -34,25 +37,32 @@ function serializeBaseline(sizes) {
} }
/** /**
* Write the baseline file from the measured workflow sizes. * Write a baseline file from the measured `.md` sizes in a directory.
* *
* @param {object} [opts] * @param {object} [opts]
* @param {string} [opts.dir] - Workflows dir to measure (default canonical). * @param {string} [opts.dir] - Directory to measure (default: workflows).
* @param {string} [opts.outPath] - Baseline file to write (default canonical). * @param {string} [opts.outPath] - Baseline file to write (default: workflow).
* @param {function(string): boolean} [opts.predicate] - Filename filter.
* @returns {{ outPath: string, count: number, content: string }} * @returns {{ outPath: string, count: number, content: string }}
*/ */
function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH } = {}) { function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH, predicate } = {}) {
const sizes = measureWorkflows(dir); const sizes = measureMdFiles(dir, predicate);
const content = serializeBaseline(sizes); const content = serializeBaseline(sizes);
fs.writeFileSync(outPath, content); fs.writeFileSync(outPath, content);
return { outPath, count: Object.keys(sizes).length, content }; return { outPath, count: Object.keys(sizes).length, content };
} }
if (require.main === module) { if (require.main === module) {
const { outPath, count } = generateBaseline(); const targets = [
process.stdout.write( { label: 'workflow', dir: WORKFLOWS_DIR, outPath: BASELINE_PATH },
`Wrote ${count} workflow sizes to ${path.relative(process.cwd(), outPath)}\n` { label: 'agent', dir: AGENTS_DIR, outPath: AGENT_BASELINE_PATH, predicate: isGsdAgent },
); ];
for (const t of targets) {
const { outPath, count } = generateBaseline(t);
process.stdout.write(
`Wrote ${count} ${t.label} sizes to ${path.relative(process.cwd(), outPath)}\n`
);
}
} }
module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH }; module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH, AGENT_BASELINE_PATH };

View File

@@ -51,18 +51,40 @@ function listWorkflowStems(dir = WORKFLOWS_DIR) {
} }
/** /**
* Measure every top-level workflow file, keyed by filename (`<stem>.md`). * Measure every top-level `.md` file in `dir`, keyed by filename, byte sizes.
* Generic over directory and an optional filename predicate — used for both
* workflows (`gsd-core/workflows/*.md`) and agents (`agents/gsd-*.md`) so the
* size guards and the baseline generator share one measurement path (#1074).
* Non-recursive by design.
* *
* @param {string} [dir] - Workflows directory (defaults to the canonical one). * @param {string} dir - Directory to scan.
* @returns {Object<string, number>} Map of `<stem>.md` → LF byte size, with * @param {function(string): boolean} [predicate] - Filename filter (default: all `.md`).
* keys inserted in sorted order. * @returns {Object<string, number>} Map of filename → LF byte size, keys sorted.
*/ */
function measureWorkflows(dir = WORKFLOWS_DIR) { function measureMdFiles(dir, predicate = () => true) {
const out = {}; const out = {};
for (const stem of listWorkflowStems(dir)) { const names = fs
out[`${stem}.md`] = lfByteCount(path.join(dir, `${stem}.md`)); .readdirSync(dir)
} .filter((f) => f.endsWith('.md') && predicate(f))
.sort();
for (const name of names) out[name] = lfByteCount(path.join(dir, name));
return out; return out;
} }
module.exports = { WORKFLOWS_DIR, lfByteCount, listWorkflowStems, measureWorkflows }; /**
* Measure every top-level workflow file, keyed by filename (`<stem>.md`).
*
* @param {string} [dir] - Workflows directory (defaults to the canonical one).
* @returns {Object<string, number>} Map of `<stem>.md` → LF byte size, sorted.
*/
function measureWorkflows(dir = WORKFLOWS_DIR) {
return measureMdFiles(dir);
}
module.exports = {
WORKFLOWS_DIR,
lfByteCount,
listWorkflowStems,
measureMdFiles,
measureWorkflows,
};

View File

@@ -0,0 +1,35 @@
{
"gsd-advisor-researcher.md": 4536,
"gsd-ai-researcher.md": 5851,
"gsd-assumptions-analyzer.md": 4489,
"gsd-code-fixer.md": 36499,
"gsd-code-reviewer.md": 16773,
"gsd-codebase-mapper.md": 21388,
"gsd-debug-session-manager.md": 14159,
"gsd-debugger.md": 51213,
"gsd-doc-classifier.md": 7629,
"gsd-doc-synthesizer.md": 9722,
"gsd-doc-verifier.md": 12403,
"gsd-doc-writer.md": 38827,
"gsd-domain-researcher.md": 6938,
"gsd-eval-auditor.md": 7754,
"gsd-eval-planner.md": 7008,
"gsd-executor.md": 42512,
"gsd-framework-selector.md": 6778,
"gsd-integration-checker.md": 15141,
"gsd-intel-updater.md": 18122,
"gsd-nyquist-auditor.md": 7245,
"gsd-pattern-mapper.md": 12487,
"gsd-phase-researcher.md": 40611,
"gsd-plan-checker.md": 41996,
"gsd-planner.md": 48885,
"gsd-project-researcher.md": 21987,
"gsd-research-synthesizer.md": 13646,
"gsd-roadmapper.md": 19652,
"gsd-security-auditor.md": 6216,
"gsd-ui-auditor.md": 17152,
"gsd-ui-checker.md": 11081,
"gsd-ui-researcher.md": 19265,
"gsd-user-profiler.md": 8516,
"gsd-verifier.md": 41285
}

View File

@@ -4,48 +4,62 @@
// Per CONTRIBUTING.md exception matrix. // Per CONTRIBUTING.md exception matrix.
/** /**
* Agent size budget. * Agent size budget (measured in BYTES — see #717).
* *
* Agent definitions in `agents/gsd-*.md` are loaded verbatim into Claude's * Agent definitions in `agents/gsd-*.md` are loaded verbatim into the agent's
* context on every subagent dispatch. Unbounded growth is paid on every call * context on every subagent dispatch. Unbounded growth is paid on every call
* across every workflow. * across every workflow.
* *
* Budgets are tiered to reflect the intent of each agent class: * ## Enforcement model (issue #1074)
*
* Mirrors tests/workflow-size-budget.test.cjs — two complementary guards, no
* tier-max ceiling:
*
* 1. Per-agent baseline (the anti-creep): every agent is pinned to its exact
* byte size in `tests/agent-size-baseline.json`. Any growth fails with the
* file and delta; `npm run size:baseline` records a deliberate change as a
* reviewable one-line diff. This replaced the tier-max tighten-only ratchet
* (which only bound the single largest agent per tier).
*
* 2. Tier hard caps (the outer bound): XL/LARGE/DEFAULT absolute red lines
* with real headroom, never raised in normal work. Crossing one means
* extracting shared boilerplate to `gsd-core/references/`, not a +N bump.
* A net-new agent is DEFAULT-tier, so the DEFAULT cap already bounds it —
* no separate new-file cap is needed (DEFAULT is already small).
*
* Tiers:
* - XL : top-level orchestrators that own end-to-end rubrics * - XL : top-level orchestrators that own end-to-end rubrics
* - LARGE : multi-phase operators with branching workflows * - LARGE : multi-phase operators with branching workflows
* - DEFAULT : focused single-purpose agents * - DEFAULT : focused single-purpose agents
* *
* Raising a budget is a deliberate choice — adjust the constant, write a * See:
* rationale in the PR, and make sure the bloat is not duplicated content * - https://github.com/open-gsd/gsd-core/issues/1074 (per-file baseline + hard caps)
* that belongs in `gsd-core/references/`. * - https://github.com/open-gsd/gsd-core/issues/717 (bytes, not lines)
* * - https://github.com/open-gsd/gsd-core/issues/683 (LF-normalized byte count)
* Tighten-only invariant (issue #597): ceilings track the tier high-water mark
* within GRACE lines. Budgets may only decrease, never silently creep upward.
* The assertTightCeiling() call below enforces this automatically.
*
* See: https://github.com/open-gsd/gsd-core/issues/2361
*/ */
const { test, describe } = require('node:test'); const { test, describe } = require('node:test');
const assert = require('node:assert/strict'); const assert = require('node:assert/strict');
const fs = require('fs'); const fs = require('fs');
const os = require('node:os');
const path = require('path'); const path = require('path');
const { assertTightCeiling } = require('../scripts/lib/allowlist-ratchet.cjs'); const { assertFileBaseline } = require('../scripts/lib/allowlist-ratchet.cjs');
const { lfByteCount, measureMdFiles } = require('../scripts/workflow-size.cjs');
const { cleanup } = require('./helpers.cjs');
const AGENTS_DIR = path.join(__dirname, '..', 'agents'); const AGENTS_DIR = path.join(__dirname, '..', 'agents');
const BASELINE_PATH = path.join(__dirname, 'agent-size-baseline.json');
const isGsdAgent = (f) => f.startsWith('gsd-');
// Ceilings tightened to actualMax + GRACE per the ratchet-down rule (#597). // Tier HARD CAPS (#1074, bytes) — absolute red lines, not high-water-hugging
// XL ceiling lowered from 1600 → 1512 (actualMax=1452, gsd-debugger). // ceilings. Day-to-day creep is caught per-agent by the baseline guard below;
const XL_BUDGET = 1512; // these sit above each tier's current high-water with real headroom:
// LARGE ceiling kept at 1000 (actualMax=978, slack=22 ≤ GRACE=60). // XL 56 KiB — high-water gsd-debugger 51,043 → ~6.3 KB headroom
const LARGE_BUDGET = 1000; // LARGE 48 KiB — high-water gsd-executor 42,342 → ~6.8 KB headroom
// DEFAULT ceiling kept at 500 (actualMax=495, slack=5 ≤ GRACE=60). // DEFAULT 24 KiB — high-water gsd-ui-researcher 19,095 → ~5.5 KB headroom
const DEFAULT_BUDGET = 500; const XL_CAP = 57344; // 56 KiB
const LARGE_CAP = 49152; // 48 KiB
// Grace band: maximum allowed slack (ceiling − actualMax) before a ceiling is const DEFAULT_CAP = 24576; // 24 KiB
// considered too loose. 60 lines gives one reasonable screen of breathing room
// without permitting gross inflation.
const GRACE = 60;
const XL_AGENTS = new Set([ const XL_AGENTS = new Set([
'gsd-debugger', 'gsd-debugger',
@@ -65,64 +79,76 @@ const LARGE_AGENTS = new Set([
]); ]);
const ALL_AGENTS = fs.readdirSync(AGENTS_DIR) const ALL_AGENTS = fs.readdirSync(AGENTS_DIR)
.filter(f => f.startsWith('gsd-') && f.endsWith('.md')) .filter(f => isGsdAgent(f) && f.endsWith('.md'))
.map(f => f.replace('.md', '')); .map(f => f.replace('.md', ''));
function budgetFor(agent) { function capFor(agent) {
if (XL_AGENTS.has(agent)) return { tier: 'XL', limit: XL_BUDGET }; if (XL_AGENTS.has(agent)) return { tier: 'XL', cap: XL_CAP };
if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', limit: LARGE_BUDGET }; if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', cap: LARGE_CAP };
return { tier: 'DEFAULT', limit: DEFAULT_BUDGET }; return { tier: 'DEFAULT', cap: DEFAULT_CAP };
} }
function lineCount(filePath) { describe('SIZE: agent tier hard caps (issue #1074)', () => {
const content = fs.readFileSync(filePath, 'utf-8'); // Absolute outer bound per tier. A cap is NOT raised when an agent approaches
if (content.length === 0) return 0; // it — crossing it means extract shared boilerplate to gsd-core/references/.
const trailingNewline = content.endsWith('\n') ? 1 : 0;
return content.split('\n').length - trailingNewline;
}
describe('SIZE: agent line-count budget', () => {
for (const agent of ALL_AGENTS) { for (const agent of ALL_AGENTS) {
const { tier, limit } = budgetFor(agent); const { tier, cap } = capFor(agent);
test(`${agent} (${tier}) stays under ${limit} lines`, () => { test(`${agent} (${tier}) stays within the ${tier} hard cap (${cap} bytes)`, () => {
const filePath = path.join(AGENTS_DIR, agent + '.md'); const bytes = lfByteCount(path.join(AGENTS_DIR, agent + '.md'));
const lines = lineCount(filePath);
assert.ok( assert.ok(
lines <= limit, bytes <= cap,
`${agent}.md has ${lines} lines — exceeds ${tier} budget of ${limit}. ` + `${agent}.md is ${bytes} bytes — exceeds the ${tier} hard cap of ${cap}. ` +
`Extract shared boilerplate to gsd-core/references/ or raise the budget ` + `This cap is a red line, NOT a budget to raise: extract shared boilerplate ` +
`in tests/agent-size-budget.test.cjs with a rationale.` `to gsd-core/references/ and load it lazily.`
); );
}); });
} }
}); });
describe('SIZE: tier anti-creep (tighten-only ceilings, issue #597)', () => { describe('SIZE: agent hard-cap boundary fixtures (#1074 — negative proof)', () => {
// For each tier, compute the high-water mark across all files in that tier // The per-tier loop above only iterates the real, fully-compliant agent
// and assert the ceiling stays tight. Prevents budgets from silently drifting // corpus, so its `bytes <= cap` failure branch never executes. Exercise that
// upward: ceiling − actualMax must not exceed GRACE. // exact comparison on synthetic files measured at cap-1 / cap / cap+1 (the
test('XL tier: ceiling tracks high-water mark within GRACE', () => { // limit boundary — RULESET.TESTS.boundary-coverage.fixtures) through the SAME
const values = ALL_AGENTS // lfByteCount path the guard uses, so a future threshold or operator edit
.filter(a => XL_AGENTS.has(a)) // cannot silently neuter a cap (RULESET.TESTS.regression-must-fail-first).
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); test('cap comparison fires at the limit boundary for every tier', () => {
const actualMax = Math.max(...values); const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-'));
assertTightCeiling({ label: 'XL', actualMax, ceiling: XL_BUDGET, grace: GRACE, fail: assert.fail }); try {
// ASCII 'a' is 1 byte/char and has no CRLF, so lfByteCount == length.
const measureAt = (n) => {
const p = path.join(tmp, `fixture-${n}.md`);
fs.writeFileSync(p, 'a'.repeat(n));
return lfByteCount(p);
};
for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) {
assert.equal(measureAt(cap - 1) <= cap, true, `${cap - 1} must be within cap ${cap}`);
assert.equal(measureAt(cap) <= cap, true, `${cap} (exactly at cap) must be within cap ${cap}`);
assert.equal(measureAt(cap + 1) <= cap, false, `${cap + 1} must exceed cap ${cap}`);
}
} finally {
cleanup(tmp);
}
}); });
});
test('LARGE tier: ceiling tracks high-water mark within GRACE', () => { describe('SIZE: per-agent baseline (issue #1074)', () => {
const values = ALL_AGENTS // Per-agent exact-size ratchet — the primary anti-creep guard. Guards EVERY
.filter(a => LARGE_AGENTS.has(a)) // agent by name against tests/agent-size-baseline.json. Growth fails with the
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); // file and delta; shrinkage fails as a stale snapshot. Fix: `npm run
const actualMax = Math.max(...values); // size:baseline` plus a PR justification for genuine growth (or extraction).
assertTightCeiling({ label: 'LARGE', actualMax, ceiling: LARGE_BUDGET, grace: GRACE, fail: assert.fail }); test('every agent matches its committed baseline', () => {
}); const baseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
const current = measureMdFiles(AGENTS_DIR, isGsdAgent);
test('DEFAULT tier: ceiling tracks high-water mark within GRACE', () => { assertFileBaseline({
const values = ALL_AGENTS label: 'agent-size',
.filter(a => !XL_AGENTS.has(a) && !LARGE_AGENTS.has(a)) current,
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md'))); baseline,
const actualMax = Math.max(...values); fail: assert.fail,
assertTightCeiling({ label: 'DEFAULT', actualMax, ceiling: DEFAULT_BUDGET, grace: GRACE, fail: assert.fail }); updateHint:
'Run `npm run size:baseline` to update tests/agent-size-baseline.json, ' +
'then justify any growth in your PR (or extract shared boilerplate to gsd-core/references/).',
});
}); });
}); });

View File

@@ -124,4 +124,21 @@ describe('generateBaseline', () => {
/ENOENT/ /ENOENT/
); );
}); });
test('predicate filters which files are baselined (agent path)', () => {
const agentDir = path.join(dir, 'agents');
fs.mkdirSync(agentDir);
fs.writeFileSync(path.join(agentDir, 'gsd-one.md'), 'a\n');
fs.writeFileSync(path.join(agentDir, 'gsd-two.md'), 'bb\n');
fs.writeFileSync(path.join(agentDir, 'README.md'), 'not an agent\n');
const result = generateBaseline({
dir: agentDir,
outPath,
predicate: (f) => f.startsWith('gsd-'),
});
assert.strictEqual(result.count, 2, 'only gsd-*.md files should be baselined');
const written = JSON.parse(fs.readFileSync(outPath, 'utf-8'));
assert.deepStrictEqual(Object.keys(written), ['gsd-one.md', 'gsd-two.md']);
});
}); });