Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the same assertTightCeiling tier ratchet but was still line-based (never rebased in #717). Completes the migration — the last part of the #1074 epic. - Rebase agent sizing from lines to LF-normalized bytes (#717/#683). - Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests); add a per-agent baseline (tests/agent-size-baseline.json) as the primary anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB), each above its tier high-water with real headroom. No separate new-file cap: a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap. - Keep the agent-classification tests verbatim. - scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate) (workflows + agents share one byte-measurement path); measureWorkflows now delegates to it. - scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates BOTH the workflow and agent baselines (gsd-* filter for agents). Rebased onto next after PR 2/3 (#1096) merged: replicate the scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across the generator and the agent test's require; regenerate the agent baseline against current agents (a uniform +170 B preamble drift on all 33 since authoring). Addresses the #1097 review (trek-e): - BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md. Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline, dual size:baseline, shared measureMdFiles seam) and disambiguates it from the separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two purposes). - Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section is in next, fold in the agent coverage here (renamed to "Workflow & agent size budget"): agent caps + per-agent baseline + the how-to + reference rows, and the disambiguation from the 45K-char guard. - Minor (negative proof): add a boundary-fixture test exercising the hard-cap comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future threshold/operator edit can't silently neuter a cap. - Nit: align the tier test name wording ("stays within") with the <= operator. Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B; XL hard cap catches 57,516 > 57,344 with the baseline current. Closes #1095 (PR 3/3 child); landing this completes the #1074 epic.
This commit is contained in:
@@ -268,6 +268,7 @@ The canonical lint infrastructure adopted in ADR 452 (`docs/adr/452-eslint-lint-
|
|||||||
|
|
||||||
`RULESET.WORKFLOW_MARKDOWN.FENCES=preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)`
|
`RULESET.WORKFLOW_MARKDOWN.FENCES=preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)`
|
||||||
`RULESET.WORKFLOW_SIZE_BUDGET=workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = per-file baseline (PRIMARY anti-creep: tests/workflow-size-baseline.json pins each file's exact size) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + new-file cap (un-baselined files <32768, the Codex anchor) + discuss-phase<32000; a file that grew fails the baseline guard — fix with `npm run size:baseline`, commit the one-line diff, and justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump`
|
`RULESET.WORKFLOW_SIZE_BUDGET=workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = per-file baseline (PRIMARY anti-creep: tests/workflow-size-baseline.json pins each file's exact size) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + new-file cap (un-baselined files <32768, the Codex anchor) + discuss-phase<32000; a file that grew fails the baseline guard — fix with `npm run size:baseline`, commit the one-line diff, and justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump`
|
||||||
|
`RULESET.AGENT_SIZE_BUDGET=agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = per-file baseline (PRIMARY anti-creep: tests/agent-size-baseline.json pins each agents/gsd-*.md exact byte size) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). One 'npm run size:baseline' regenerates BOTH workflow and agent baselines via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter. A grown agent fails the baseline guard — regenerate + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes`
|
||||||
`RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; <step name="..."> XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name`
|
`RULESET.WORKFLOW_FILE_NAMES=workflow files use hyphens; <step name="..."> XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name`
|
||||||
`RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/bug-3135-capture-backlog-workflow.test.cjs; INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill`
|
`RULESET.WORKFLOW_EXECUTION_CONTEXT=@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/bug-3135-capture-backlog-workflow.test.cjs; INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; "Invoked by" attribution must move when a flag absorbs a micro-skill`
|
||||||
`RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets`
|
`RULESET.WORKFLOW_EXECUTE_END_TO_END=ADR-0002 standard for single-workflow commands is "Execute end-to-end." (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses "execute the X workflow end-to-end." in routing bullets`
|
||||||
|
|||||||
@@ -856,9 +856,10 @@ gsd-core/
|
|||||||
If you legitimately grow or shrink a workflow file,
|
If you legitimately grow or shrink a workflow file,
|
||||||
run `npm run size:baseline` to update the snapshot and
|
run `npm run size:baseline` to update the snapshot and
|
||||||
justify any growth in your PR (or extract content
|
justify any growth in your PR (or extract content
|
||||||
lazily). Full how-to + reference in
|
lazily). The same guard covers agent files
|
||||||
docs/TESTING-SUITES.md (Workflow size budget); see
|
(agents/gsd-*.md). Full how-to + reference in
|
||||||
issue #1074.
|
docs/TESTING-SUITES.md (Workflow & agent size
|
||||||
|
budget); see issue #1074.
|
||||||
references/ — Reference documentation (.md)
|
references/ — Reference documentation (.md)
|
||||||
templates/ — File templates
|
templates/ — File templates
|
||||||
agents/ — Agent definitions (.md) — CANONICAL SOURCE
|
agents/ — Agent definitions (.md) — CANONICAL SOURCE
|
||||||
|
|||||||
@@ -59,17 +59,21 @@ the sanctioned layout (see the #443 strategy below), not a one-off regression
|
|||||||
pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way
|
pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way
|
||||||
`bug-*` did, extend the ratchet's regex and regenerate the allowlist.
|
`bug-*` did, extend the ratchet's regex and regenerate the allowlist.
|
||||||
|
|
||||||
## Workflow size budget
|
## Workflow & agent size budget
|
||||||
|
|
||||||
> Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074).
|
> Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074).
|
||||||
> Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717);
|
> Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717);
|
||||||
> LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683).
|
> LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683).
|
||||||
|
|
||||||
Workflow files (`gsd-core/workflows/*.md`) ship in the installed runtime and are
|
Workflow files (`gsd-core/workflows/*.md`) and agent files (`agents/gsd-*.md`)
|
||||||
loaded into the agent's context, so their byte size is a real cost. The size
|
both ship in the installed runtime and are loaded into context — workflows on
|
||||||
guard in `tests/workflow-size-budget.test.cjs` keeps that cost from creeping up
|
every command, agents on every subagent dispatch — so their byte size is a real
|
||||||
invisibly. It is an **anti-creep ratchet**, sibling to the regression-name
|
cost. Two sibling guards (`tests/workflow-size-budget.test.cjs` and
|
||||||
ratchet above — three layers, ordered from day-to-day to last-resort:
|
`tests/agent-size-budget.test.cjs`) keep that cost from creeping up invisibly,
|
||||||
|
sharing one byte-counter (`measureMdFiles`) and one `npm run size:baseline`
|
||||||
|
command that regenerates **both** snapshots. Each is an **anti-creep ratchet**,
|
||||||
|
sibling to the regression-name ratchet above — three layers (workflows), ordered
|
||||||
|
from day-to-day to last-resort:
|
||||||
|
|
||||||
| Layer | What it does | Where |
|
| Layer | What it does | Where |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -80,9 +84,18 @@ ratchet above — three layers, ordered from day-to-day to last-resort:
|
|||||||
`discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes
|
`discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes
|
||||||
(issue [#2551](https://github.com/open-gsd/gsd-core/issues/2551)).
|
(issue [#2551](https://github.com/open-gsd/gsd-core/issues/2551)).
|
||||||
|
|
||||||
### How-to: a workflow grew and CI is red
|
**Agents** (`tests/agent-size-budget.test.cjs`) use the same per-agent baseline
|
||||||
|
(`tests/agent-size-baseline.json`) + loose tier hard caps — `XL ≤ 57344` /
|
||||||
|
`LARGE ≤ 49152` / `DEFAULT ≤ 24576` bytes. There is no new-agent cap: a net-new
|
||||||
|
agent is DEFAULT-tier and already bounded by the DEFAULT cap. (This is distinct
|
||||||
|
from the separate 45 KB-*char* extraction-evidence threshold on `gsd-planner`
|
||||||
|
enforced by `tests/planner-decomposition.test.cjs` — that one proves mode
|
||||||
|
sections were extracted; this one bounds total agent bytes.)
|
||||||
|
|
||||||
The baseline guard reports the file and the byte delta. To resolve:
|
### How-to: a workflow or agent grew and CI is red
|
||||||
|
|
||||||
|
The baseline guard reports the file and the byte delta (the same flow for both
|
||||||
|
the workflow and agent guards). To resolve:
|
||||||
|
|
||||||
1. **Regenerate the snapshot** and inspect the one-line diff:
|
1. **Regenerate the snapshot** and inspect the one-line diff:
|
||||||
```bash
|
```bash
|
||||||
@@ -93,9 +106,10 @@ The baseline guard reports the file and the byte delta. To resolve:
|
|||||||
the committed baseline diff is the review record that the larger size was a
|
the committed baseline diff is the review record that the larger size was a
|
||||||
deliberate, seen decision, not silent drift.
|
deliberate, seen decision, not silent drift.
|
||||||
3. **Or shrink it instead of baselining.** Prefer extraction when the growth is
|
3. **Or shrink it instead of baselining.** Prefer extraction when the growth is
|
||||||
incidental: move per-mode bodies to `workflows/<name>/modes/`, templates to
|
incidental: for a workflow, move per-mode bodies to `workflows/<name>/modes/`,
|
||||||
`workflows/<name>/templates/`, and shared prose to `gsd-core/references/` —
|
templates to `workflows/<name>/templates/`, and shared prose to
|
||||||
then load them **LAZILY**. Do *not* convert them to eager `@-required_reading`
|
`gsd-core/references/`; for an agent, lift shared boilerplate into
|
||||||
|
`gsd-core/references/` and `@`-reference it — then load it **LAZILY**. Do *not* convert them to eager `@-required_reading`
|
||||||
includes: that shrinks the file's bytes without shrinking loaded context, so
|
includes: that shrinks the file's bytes without shrinking loaded context, so
|
||||||
it games the guard while making the real cost worse. See
|
it games the guard while making the real cost worse. See
|
||||||
`workflows/discuss-phase/` for the progressive-disclosure pattern.
|
`workflows/discuss-phase/` for the progressive-disclosure pattern.
|
||||||
@@ -107,10 +121,12 @@ that is the signal to extract, per step 3.
|
|||||||
|
|
||||||
| Artifact | Role |
|
| Artifact | Role |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) plus workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guard and the generator so they can never measure or enumerate differently. |
|
| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) + generic `measureMdFiles(dir, predicate)` (backs both workflows and agents) + workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guards and the generator so they can never measure differently. |
|
||||||
| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates `tests/workflow-size-baseline.json` — sorted keys, trailing newline, idempotent. |
|
| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates **both** `tests/workflow-size-baseline.json` and `tests/agent-size-baseline.json` — sorted keys, trailing newline, idempotent. |
|
||||||
| `tests/workflow-size-baseline.json` | The committed per-file snapshot (one entry per workflow). |
|
| `tests/workflow-size-baseline.json` | The committed per-workflow snapshot (one entry per workflow). |
|
||||||
| `tests/workflow-size-budget.test.cjs` | The three guards above, plus the `discuss-phase` progressive-disclosure checks. |
|
| `tests/agent-size-baseline.json` | The committed per-agent snapshot (one entry per `gsd-*` agent). |
|
||||||
|
| `tests/workflow-size-budget.test.cjs` | The three workflow guards above, plus the `discuss-phase` progressive-disclosure checks. |
|
||||||
|
| `tests/agent-size-budget.test.cjs` | The per-agent baseline + tier hard-cap guards (the agent analog). |
|
||||||
|
|
||||||
## Running suites locally
|
## Running suites locally
|
||||||
|
|
||||||
|
|||||||
@@ -16,9 +16,12 @@
|
|||||||
|
|
||||||
const fs = require('fs');
|
const fs = require('fs');
|
||||||
const path = require('path');
|
const path = require('path');
|
||||||
const { measureWorkflows, WORKFLOWS_DIR } = require('./workflow-size.cjs');
|
const { measureMdFiles, WORKFLOWS_DIR } = require('./workflow-size.cjs');
|
||||||
|
|
||||||
const BASELINE_PATH = path.join(__dirname, '..', 'tests', 'workflow-size-baseline.json');
|
const BASELINE_PATH = path.join(__dirname, '..', 'tests', 'workflow-size-baseline.json');
|
||||||
|
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
|
||||||
|
const AGENT_BASELINE_PATH = path.join(__dirname, '..', 'tests', 'agent-size-baseline.json');
|
||||||
|
const isGsdAgent = (f) => f.startsWith('gsd-');
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Serialize a size map to the on-disk baseline format: keys sorted, 2-space
|
* Serialize a size map to the on-disk baseline format: keys sorted, 2-space
|
||||||
@@ -34,25 +37,32 @@ function serializeBaseline(sizes) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Write the baseline file from the measured workflow sizes.
|
* Write a baseline file from the measured `.md` sizes in a directory.
|
||||||
*
|
*
|
||||||
* @param {object} [opts]
|
* @param {object} [opts]
|
||||||
* @param {string} [opts.dir] - Workflows dir to measure (default canonical).
|
* @param {string} [opts.dir] - Directory to measure (default: workflows).
|
||||||
* @param {string} [opts.outPath] - Baseline file to write (default canonical).
|
* @param {string} [opts.outPath] - Baseline file to write (default: workflow).
|
||||||
|
* @param {function(string): boolean} [opts.predicate] - Filename filter.
|
||||||
* @returns {{ outPath: string, count: number, content: string }}
|
* @returns {{ outPath: string, count: number, content: string }}
|
||||||
*/
|
*/
|
||||||
function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH } = {}) {
|
function generateBaseline({ dir = WORKFLOWS_DIR, outPath = BASELINE_PATH, predicate } = {}) {
|
||||||
const sizes = measureWorkflows(dir);
|
const sizes = measureMdFiles(dir, predicate);
|
||||||
const content = serializeBaseline(sizes);
|
const content = serializeBaseline(sizes);
|
||||||
fs.writeFileSync(outPath, content);
|
fs.writeFileSync(outPath, content);
|
||||||
return { outPath, count: Object.keys(sizes).length, content };
|
return { outPath, count: Object.keys(sizes).length, content };
|
||||||
}
|
}
|
||||||
|
|
||||||
if (require.main === module) {
|
if (require.main === module) {
|
||||||
const { outPath, count } = generateBaseline();
|
const targets = [
|
||||||
process.stdout.write(
|
{ label: 'workflow', dir: WORKFLOWS_DIR, outPath: BASELINE_PATH },
|
||||||
`Wrote ${count} workflow sizes to ${path.relative(process.cwd(), outPath)}\n`
|
{ label: 'agent', dir: AGENTS_DIR, outPath: AGENT_BASELINE_PATH, predicate: isGsdAgent },
|
||||||
);
|
];
|
||||||
|
for (const t of targets) {
|
||||||
|
const { outPath, count } = generateBaseline(t);
|
||||||
|
process.stdout.write(
|
||||||
|
`Wrote ${count} ${t.label} sizes to ${path.relative(process.cwd(), outPath)}\n`
|
||||||
|
);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH };
|
module.exports = { generateBaseline, serializeBaseline, BASELINE_PATH, AGENT_BASELINE_PATH };
|
||||||
|
|||||||
@@ -51,18 +51,40 @@ function listWorkflowStems(dir = WORKFLOWS_DIR) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Measure every top-level workflow file, keyed by filename (`<stem>.md`).
|
* Measure every top-level `.md` file in `dir`, keyed by filename, byte sizes.
|
||||||
|
* Generic over directory and an optional filename predicate — used for both
|
||||||
|
* workflows (`gsd-core/workflows/*.md`) and agents (`agents/gsd-*.md`) so the
|
||||||
|
* size guards and the baseline generator share one measurement path (#1074).
|
||||||
|
* Non-recursive by design.
|
||||||
*
|
*
|
||||||
* @param {string} [dir] - Workflows directory (defaults to the canonical one).
|
* @param {string} dir - Directory to scan.
|
||||||
* @returns {Object<string, number>} Map of `<stem>.md` → LF byte size, with
|
* @param {function(string): boolean} [predicate] - Filename filter (default: all `.md`).
|
||||||
* keys inserted in sorted order.
|
* @returns {Object<string, number>} Map of filename → LF byte size, keys sorted.
|
||||||
*/
|
*/
|
||||||
function measureWorkflows(dir = WORKFLOWS_DIR) {
|
function measureMdFiles(dir, predicate = () => true) {
|
||||||
const out = {};
|
const out = {};
|
||||||
for (const stem of listWorkflowStems(dir)) {
|
const names = fs
|
||||||
out[`${stem}.md`] = lfByteCount(path.join(dir, `${stem}.md`));
|
.readdirSync(dir)
|
||||||
}
|
.filter((f) => f.endsWith('.md') && predicate(f))
|
||||||
|
.sort();
|
||||||
|
for (const name of names) out[name] = lfByteCount(path.join(dir, name));
|
||||||
return out;
|
return out;
|
||||||
}
|
}
|
||||||
|
|
||||||
module.exports = { WORKFLOWS_DIR, lfByteCount, listWorkflowStems, measureWorkflows };
|
/**
|
||||||
|
* Measure every top-level workflow file, keyed by filename (`<stem>.md`).
|
||||||
|
*
|
||||||
|
* @param {string} [dir] - Workflows directory (defaults to the canonical one).
|
||||||
|
* @returns {Object<string, number>} Map of `<stem>.md` → LF byte size, sorted.
|
||||||
|
*/
|
||||||
|
function measureWorkflows(dir = WORKFLOWS_DIR) {
|
||||||
|
return measureMdFiles(dir);
|
||||||
|
}
|
||||||
|
|
||||||
|
module.exports = {
|
||||||
|
WORKFLOWS_DIR,
|
||||||
|
lfByteCount,
|
||||||
|
listWorkflowStems,
|
||||||
|
measureMdFiles,
|
||||||
|
measureWorkflows,
|
||||||
|
};
|
||||||
|
|||||||
35
tests/agent-size-baseline.json
Normal file
35
tests/agent-size-baseline.json
Normal file
@@ -0,0 +1,35 @@
|
|||||||
|
{
|
||||||
|
"gsd-advisor-researcher.md": 4536,
|
||||||
|
"gsd-ai-researcher.md": 5851,
|
||||||
|
"gsd-assumptions-analyzer.md": 4489,
|
||||||
|
"gsd-code-fixer.md": 36499,
|
||||||
|
"gsd-code-reviewer.md": 16773,
|
||||||
|
"gsd-codebase-mapper.md": 21388,
|
||||||
|
"gsd-debug-session-manager.md": 14159,
|
||||||
|
"gsd-debugger.md": 51213,
|
||||||
|
"gsd-doc-classifier.md": 7629,
|
||||||
|
"gsd-doc-synthesizer.md": 9722,
|
||||||
|
"gsd-doc-verifier.md": 12403,
|
||||||
|
"gsd-doc-writer.md": 38827,
|
||||||
|
"gsd-domain-researcher.md": 6938,
|
||||||
|
"gsd-eval-auditor.md": 7754,
|
||||||
|
"gsd-eval-planner.md": 7008,
|
||||||
|
"gsd-executor.md": 42512,
|
||||||
|
"gsd-framework-selector.md": 6778,
|
||||||
|
"gsd-integration-checker.md": 15141,
|
||||||
|
"gsd-intel-updater.md": 18122,
|
||||||
|
"gsd-nyquist-auditor.md": 7245,
|
||||||
|
"gsd-pattern-mapper.md": 12487,
|
||||||
|
"gsd-phase-researcher.md": 40611,
|
||||||
|
"gsd-plan-checker.md": 41996,
|
||||||
|
"gsd-planner.md": 48885,
|
||||||
|
"gsd-project-researcher.md": 21987,
|
||||||
|
"gsd-research-synthesizer.md": 13646,
|
||||||
|
"gsd-roadmapper.md": 19652,
|
||||||
|
"gsd-security-auditor.md": 6216,
|
||||||
|
"gsd-ui-auditor.md": 17152,
|
||||||
|
"gsd-ui-checker.md": 11081,
|
||||||
|
"gsd-ui-researcher.md": 19265,
|
||||||
|
"gsd-user-profiler.md": 8516,
|
||||||
|
"gsd-verifier.md": 41285
|
||||||
|
}
|
||||||
@@ -4,48 +4,62 @@
|
|||||||
// Per CONTRIBUTING.md exception matrix.
|
// Per CONTRIBUTING.md exception matrix.
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Agent size budget.
|
* Agent size budget (measured in BYTES — see #717).
|
||||||
*
|
*
|
||||||
* Agent definitions in `agents/gsd-*.md` are loaded verbatim into Claude's
|
* Agent definitions in `agents/gsd-*.md` are loaded verbatim into the agent's
|
||||||
* context on every subagent dispatch. Unbounded growth is paid on every call
|
* context on every subagent dispatch. Unbounded growth is paid on every call
|
||||||
* across every workflow.
|
* across every workflow.
|
||||||
*
|
*
|
||||||
* Budgets are tiered to reflect the intent of each agent class:
|
* ## Enforcement model (issue #1074)
|
||||||
|
*
|
||||||
|
* Mirrors tests/workflow-size-budget.test.cjs — two complementary guards, no
|
||||||
|
* tier-max ceiling:
|
||||||
|
*
|
||||||
|
* 1. Per-agent baseline (the anti-creep): every agent is pinned to its exact
|
||||||
|
* byte size in `tests/agent-size-baseline.json`. Any growth fails with the
|
||||||
|
* file and delta; `npm run size:baseline` records a deliberate change as a
|
||||||
|
* reviewable one-line diff. This replaced the tier-max tighten-only ratchet
|
||||||
|
* (which only bound the single largest agent per tier).
|
||||||
|
*
|
||||||
|
* 2. Tier hard caps (the outer bound): XL/LARGE/DEFAULT absolute red lines
|
||||||
|
* with real headroom, never raised in normal work. Crossing one means
|
||||||
|
* extracting shared boilerplate to `gsd-core/references/`, not a +N bump.
|
||||||
|
* A net-new agent is DEFAULT-tier, so the DEFAULT cap already bounds it —
|
||||||
|
* no separate new-file cap is needed (DEFAULT is already small).
|
||||||
|
*
|
||||||
|
* Tiers:
|
||||||
* - XL : top-level orchestrators that own end-to-end rubrics
|
* - XL : top-level orchestrators that own end-to-end rubrics
|
||||||
* - LARGE : multi-phase operators with branching workflows
|
* - LARGE : multi-phase operators with branching workflows
|
||||||
* - DEFAULT : focused single-purpose agents
|
* - DEFAULT : focused single-purpose agents
|
||||||
*
|
*
|
||||||
* Raising a budget is a deliberate choice — adjust the constant, write a
|
* See:
|
||||||
* rationale in the PR, and make sure the bloat is not duplicated content
|
* - https://github.com/open-gsd/gsd-core/issues/1074 (per-file baseline + hard caps)
|
||||||
* that belongs in `gsd-core/references/`.
|
* - https://github.com/open-gsd/gsd-core/issues/717 (bytes, not lines)
|
||||||
*
|
* - https://github.com/open-gsd/gsd-core/issues/683 (LF-normalized byte count)
|
||||||
* Tighten-only invariant (issue #597): ceilings track the tier high-water mark
|
|
||||||
* within GRACE lines. Budgets may only decrease, never silently creep upward.
|
|
||||||
* The assertTightCeiling() call below enforces this automatically.
|
|
||||||
*
|
|
||||||
* See: https://github.com/open-gsd/gsd-core/issues/2361
|
|
||||||
*/
|
*/
|
||||||
|
|
||||||
const { test, describe } = require('node:test');
|
const { test, describe } = require('node:test');
|
||||||
const assert = require('node:assert/strict');
|
const assert = require('node:assert/strict');
|
||||||
const fs = require('fs');
|
const fs = require('fs');
|
||||||
|
const os = require('node:os');
|
||||||
const path = require('path');
|
const path = require('path');
|
||||||
const { assertTightCeiling } = require('../scripts/lib/allowlist-ratchet.cjs');
|
const { assertFileBaseline } = require('../scripts/lib/allowlist-ratchet.cjs');
|
||||||
|
const { lfByteCount, measureMdFiles } = require('../scripts/workflow-size.cjs');
|
||||||
|
const { cleanup } = require('./helpers.cjs');
|
||||||
|
|
||||||
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
|
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
|
||||||
|
const BASELINE_PATH = path.join(__dirname, 'agent-size-baseline.json');
|
||||||
|
const isGsdAgent = (f) => f.startsWith('gsd-');
|
||||||
|
|
||||||
// Ceilings tightened to actualMax + GRACE per the ratchet-down rule (#597).
|
// Tier HARD CAPS (#1074, bytes) — absolute red lines, not high-water-hugging
|
||||||
// XL ceiling lowered from 1600 → 1512 (actualMax=1452, gsd-debugger).
|
// ceilings. Day-to-day creep is caught per-agent by the baseline guard below;
|
||||||
const XL_BUDGET = 1512;
|
// these sit above each tier's current high-water with real headroom:
|
||||||
// LARGE ceiling kept at 1000 (actualMax=978, slack=22 ≤ GRACE=60).
|
// XL 56 KiB — high-water gsd-debugger 51,043 → ~6.3 KB headroom
|
||||||
const LARGE_BUDGET = 1000;
|
// LARGE 48 KiB — high-water gsd-executor 42,342 → ~6.8 KB headroom
|
||||||
// DEFAULT ceiling kept at 500 (actualMax=495, slack=5 ≤ GRACE=60).
|
// DEFAULT 24 KiB — high-water gsd-ui-researcher 19,095 → ~5.5 KB headroom
|
||||||
const DEFAULT_BUDGET = 500;
|
const XL_CAP = 57344; // 56 KiB
|
||||||
|
const LARGE_CAP = 49152; // 48 KiB
|
||||||
// Grace band: maximum allowed slack (ceiling − actualMax) before a ceiling is
|
const DEFAULT_CAP = 24576; // 24 KiB
|
||||||
// considered too loose. 60 lines gives one reasonable screen of breathing room
|
|
||||||
// without permitting gross inflation.
|
|
||||||
const GRACE = 60;
|
|
||||||
|
|
||||||
const XL_AGENTS = new Set([
|
const XL_AGENTS = new Set([
|
||||||
'gsd-debugger',
|
'gsd-debugger',
|
||||||
@@ -65,64 +79,76 @@ const LARGE_AGENTS = new Set([
|
|||||||
]);
|
]);
|
||||||
|
|
||||||
const ALL_AGENTS = fs.readdirSync(AGENTS_DIR)
|
const ALL_AGENTS = fs.readdirSync(AGENTS_DIR)
|
||||||
.filter(f => f.startsWith('gsd-') && f.endsWith('.md'))
|
.filter(f => isGsdAgent(f) && f.endsWith('.md'))
|
||||||
.map(f => f.replace('.md', ''));
|
.map(f => f.replace('.md', ''));
|
||||||
|
|
||||||
function budgetFor(agent) {
|
function capFor(agent) {
|
||||||
if (XL_AGENTS.has(agent)) return { tier: 'XL', limit: XL_BUDGET };
|
if (XL_AGENTS.has(agent)) return { tier: 'XL', cap: XL_CAP };
|
||||||
if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', limit: LARGE_BUDGET };
|
if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', cap: LARGE_CAP };
|
||||||
return { tier: 'DEFAULT', limit: DEFAULT_BUDGET };
|
return { tier: 'DEFAULT', cap: DEFAULT_CAP };
|
||||||
}
|
}
|
||||||
|
|
||||||
function lineCount(filePath) {
|
describe('SIZE: agent tier hard caps (issue #1074)', () => {
|
||||||
const content = fs.readFileSync(filePath, 'utf-8');
|
// Absolute outer bound per tier. A cap is NOT raised when an agent approaches
|
||||||
if (content.length === 0) return 0;
|
// it — crossing it means extract shared boilerplate to gsd-core/references/.
|
||||||
const trailingNewline = content.endsWith('\n') ? 1 : 0;
|
|
||||||
return content.split('\n').length - trailingNewline;
|
|
||||||
}
|
|
||||||
|
|
||||||
describe('SIZE: agent line-count budget', () => {
|
|
||||||
for (const agent of ALL_AGENTS) {
|
for (const agent of ALL_AGENTS) {
|
||||||
const { tier, limit } = budgetFor(agent);
|
const { tier, cap } = capFor(agent);
|
||||||
test(`${agent} (${tier}) stays under ${limit} lines`, () => {
|
test(`${agent} (${tier}) stays within the ${tier} hard cap (${cap} bytes)`, () => {
|
||||||
const filePath = path.join(AGENTS_DIR, agent + '.md');
|
const bytes = lfByteCount(path.join(AGENTS_DIR, agent + '.md'));
|
||||||
const lines = lineCount(filePath);
|
|
||||||
assert.ok(
|
assert.ok(
|
||||||
lines <= limit,
|
bytes <= cap,
|
||||||
`${agent}.md has ${lines} lines — exceeds ${tier} budget of ${limit}. ` +
|
`${agent}.md is ${bytes} bytes — exceeds the ${tier} hard cap of ${cap}. ` +
|
||||||
`Extract shared boilerplate to gsd-core/references/ or raise the budget ` +
|
`This cap is a red line, NOT a budget to raise: extract shared boilerplate ` +
|
||||||
`in tests/agent-size-budget.test.cjs with a rationale.`
|
`to gsd-core/references/ and load it lazily.`
|
||||||
);
|
);
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
describe('SIZE: tier anti-creep (tighten-only ceilings, issue #597)', () => {
|
describe('SIZE: agent hard-cap boundary fixtures (#1074 — negative proof)', () => {
|
||||||
// For each tier, compute the high-water mark across all files in that tier
|
// The per-tier loop above only iterates the real, fully-compliant agent
|
||||||
// and assert the ceiling stays tight. Prevents budgets from silently drifting
|
// corpus, so its `bytes <= cap` failure branch never executes. Exercise that
|
||||||
// upward: ceiling − actualMax must not exceed GRACE.
|
// exact comparison on synthetic files measured at cap-1 / cap / cap+1 (the
|
||||||
test('XL tier: ceiling tracks high-water mark within GRACE', () => {
|
// limit boundary — RULESET.TESTS.boundary-coverage.fixtures) through the SAME
|
||||||
const values = ALL_AGENTS
|
// lfByteCount path the guard uses, so a future threshold or operator edit
|
||||||
.filter(a => XL_AGENTS.has(a))
|
// cannot silently neuter a cap (RULESET.TESTS.regression-must-fail-first).
|
||||||
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md')));
|
test('cap comparison fires at the limit boundary for every tier', () => {
|
||||||
const actualMax = Math.max(...values);
|
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-'));
|
||||||
assertTightCeiling({ label: 'XL', actualMax, ceiling: XL_BUDGET, grace: GRACE, fail: assert.fail });
|
try {
|
||||||
|
// ASCII 'a' is 1 byte/char and has no CRLF, so lfByteCount == length.
|
||||||
|
const measureAt = (n) => {
|
||||||
|
const p = path.join(tmp, `fixture-${n}.md`);
|
||||||
|
fs.writeFileSync(p, 'a'.repeat(n));
|
||||||
|
return lfByteCount(p);
|
||||||
|
};
|
||||||
|
for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) {
|
||||||
|
assert.equal(measureAt(cap - 1) <= cap, true, `${cap - 1} must be within cap ${cap}`);
|
||||||
|
assert.equal(measureAt(cap) <= cap, true, `${cap} (exactly at cap) must be within cap ${cap}`);
|
||||||
|
assert.equal(measureAt(cap + 1) <= cap, false, `${cap + 1} must exceed cap ${cap}`);
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
cleanup(tmp);
|
||||||
|
}
|
||||||
});
|
});
|
||||||
|
});
|
||||||
|
|
||||||
test('LARGE tier: ceiling tracks high-water mark within GRACE', () => {
|
describe('SIZE: per-agent baseline (issue #1074)', () => {
|
||||||
const values = ALL_AGENTS
|
// Per-agent exact-size ratchet — the primary anti-creep guard. Guards EVERY
|
||||||
.filter(a => LARGE_AGENTS.has(a))
|
// agent by name against tests/agent-size-baseline.json. Growth fails with the
|
||||||
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md')));
|
// file and delta; shrinkage fails as a stale snapshot. Fix: `npm run
|
||||||
const actualMax = Math.max(...values);
|
// size:baseline` plus a PR justification for genuine growth (or extraction).
|
||||||
assertTightCeiling({ label: 'LARGE', actualMax, ceiling: LARGE_BUDGET, grace: GRACE, fail: assert.fail });
|
test('every agent matches its committed baseline', () => {
|
||||||
});
|
const baseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||||||
|
const current = measureMdFiles(AGENTS_DIR, isGsdAgent);
|
||||||
test('DEFAULT tier: ceiling tracks high-water mark within GRACE', () => {
|
assertFileBaseline({
|
||||||
const values = ALL_AGENTS
|
label: 'agent-size',
|
||||||
.filter(a => !XL_AGENTS.has(a) && !LARGE_AGENTS.has(a))
|
current,
|
||||||
.map(a => lineCount(path.join(AGENTS_DIR, a + '.md')));
|
baseline,
|
||||||
const actualMax = Math.max(...values);
|
fail: assert.fail,
|
||||||
assertTightCeiling({ label: 'DEFAULT', actualMax, ceiling: DEFAULT_BUDGET, grace: GRACE, fail: assert.fail });
|
updateHint:
|
||||||
|
'Run `npm run size:baseline` to update tests/agent-size-baseline.json, ' +
|
||||||
|
'then justify any growth in your PR (or extract shared boilerplate to gsd-core/references/).',
|
||||||
|
});
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
|||||||
@@ -124,4 +124,21 @@ describe('generateBaseline', () => {
|
|||||||
/ENOENT/
|
/ENOENT/
|
||||||
);
|
);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test('predicate filters which files are baselined (agent path)', () => {
|
||||||
|
const agentDir = path.join(dir, 'agents');
|
||||||
|
fs.mkdirSync(agentDir);
|
||||||
|
fs.writeFileSync(path.join(agentDir, 'gsd-one.md'), 'a\n');
|
||||||
|
fs.writeFileSync(path.join(agentDir, 'gsd-two.md'), 'bb\n');
|
||||||
|
fs.writeFileSync(path.join(agentDir, 'README.md'), 'not an agent\n');
|
||||||
|
|
||||||
|
const result = generateBaseline({
|
||||||
|
dir: agentDir,
|
||||||
|
outPath,
|
||||||
|
predicate: (f) => f.startsWith('gsd-'),
|
||||||
|
});
|
||||||
|
assert.strictEqual(result.count, 2, 'only gsd-*.md files should be baselined');
|
||||||
|
const written = JSON.parse(fs.readFileSync(outPath, 'utf-8'));
|
||||||
|
assert.deepStrictEqual(Object.keys(written), ['gsd-one.md', 'gsd-two.md']);
|
||||||
|
});
|
||||||
});
|
});
|
||||||
|
|||||||
Reference in New Issue
Block a user