Files
msd-core/scripts/workflow-size.cjs
Rezolv e4f0910d62 test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the
same assertTightCeiling tier ratchet but was still line-based (never rebased in
#717). Completes the migration — the last part of the #1074 epic.

- Rebase agent sizing from lines to LF-normalized bytes (#717/#683).
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests);
  add a per-agent baseline (tests/agent-size-baseline.json) as the primary
  anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB),
  each above its tier high-water with real headroom. No separate new-file cap:
  a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap.
- Keep the agent-classification tests verbatim.
- scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate)
  (workflows + agents share one byte-measurement path); measureWorkflows now
  delegates to it.
- scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates
  BOTH the workflow and agent baselines (gsd-* filter for agents).

Rebased onto next after PR 2/3 (#1096) merged: replicate the
scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across
the generator and the agent test's require; regenerate the agent baseline
against current agents (a uniform +170 B preamble drift on all 33 since
authoring).

Addresses the #1097 review (trek-e):
- BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md.
  Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline,
  dual size:baseline, shared measureMdFiles seam) and disambiguates it from the
  separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two
  purposes).
- Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section
  is in next, fold in the agent coverage here (renamed to "Workflow & agent
  size budget"): agent caps + per-agent baseline + the how-to + reference rows,
  and the disambiguation from the 45K-char guard.
- Minor (negative proof): add a boundary-fixture test exercising the hard-cap
  comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future
  threshold/operator edit can't silently neuter a cap.
- Nit: align the tier test name wording ("stays within") with the <= operator.

Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B;
XL hard cap catches 57,516 > 57,344 with the baseline current.

Closes #1095 (PR 3/3 child); landing this completes the #1074 epic.
2026-06-12 09:58:44 -04:00

91 lines
3.2 KiB
JavaScript

'use strict';
/**
* @file workflow-size.cjs
*
* Single source of truth for measuring workflow `.md` file sizes in bytes.
*
* Shared by `tests/workflow-size-budget.test.cjs` (the CI guard) and
* `scripts/update-size-baseline.cjs` (the baseline generator) so the two can
* never disagree on HOW a file is measured. A divergence between the generator
* and the guard would silently mis-record the baseline (issue #1074).
*/
const fs = require('fs');
const path = require('path');
const WORKFLOWS_DIR = path.join(__dirname, '..', 'gsd-core', 'workflows');
/**
* Byte size of a file, counted as on an LF (Unix) checkout.
*
* The size budget is calibrated against `wc -c` on a Unix (LF) checkout, but
* these `.md` files have no `eol=lf` in `.gitattributes`, so Windows checks
* them out as CRLF. Counting raw on-disk bytes there adds one byte per line,
* a Windows-only false positive that diverges from the LF calibration basis
* (issue #683). Stripping CR yields the same LF byte count on every platform.
* This is still a raw byte count (not a trailing-newline-stripping line count).
*
* @param {string} filePath - Absolute or relative path to the file.
* @returns {number} LF-normalized byte length.
*/
function lfByteCount(filePath) {
const content = fs.readFileSync(filePath, 'utf-8');
return Buffer.byteLength(content.replace(/\r\n/g, '\n'), 'utf-8');
}
/**
* List top-level workflow stems (filenames without the `.md` extension), sorted.
* Non-recursive by design: per-mode bodies under `workflows/<name>/modes/` and
* templates are NOT measured — only the always-loaded top-level workflows.
*
* @param {string} [dir] - Workflows directory (defaults to the canonical one).
* @returns {string[]} Sorted stems, e.g. `['autonomous', 'plan-phase', ...]`.
*/
function listWorkflowStems(dir = WORKFLOWS_DIR) {
return fs
.readdirSync(dir)
.filter((f) => f.endsWith('.md'))
.map((f) => f.replace(/\.md$/, ''))
.sort();
}
/**
* Measure every top-level `.md` file in `dir`, keyed by filename, byte sizes.
* Generic over directory and an optional filename predicate — used for both
* workflows (`gsd-core/workflows/*.md`) and agents (`agents/gsd-*.md`) so the
* size guards and the baseline generator share one measurement path (#1074).
* Non-recursive by design.
*
* @param {string} dir - Directory to scan.
* @param {function(string): boolean} [predicate] - Filename filter (default: all `.md`).
* @returns {Object<string, number>} Map of filename → LF byte size, keys sorted.
*/
function measureMdFiles(dir, predicate = () => true) {
const out = {};
const names = fs
.readdirSync(dir)
.filter((f) => f.endsWith('.md') && predicate(f))
.sort();
for (const name of names) out[name] = lfByteCount(path.join(dir, name));
return out;
}
/**
* Measure every top-level workflow file, keyed by filename (`<stem>.md`).
*
* @param {string} [dir] - Workflows directory (defaults to the canonical one).
* @returns {Object<string, number>} Map of `<stem>.md` → LF byte size, sorted.
*/
function measureWorkflows(dir = WORKFLOWS_DIR) {
return measureMdFiles(dir);
}
module.exports = {
WORKFLOWS_DIR,
lfByteCount,
listWorkflowStems,
measureMdFiles,
measureWorkflows,
};