* test(#2724): delete golden-install-parity fixtures, test, and generator Removes the 19 committed path->hash manifests, the two per-file size baselines, tests/golden-install-parity.test.cjs, and scripts/gen-golden-install-parity-zcode.cjs. These were pure functions of the source tree (ADR-2719); the differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the sole gate for emitted-artifact propagation. tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs are unchanged (ADR-2719 section 7 exception). Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's existence guard, .gitattributes, package.json scripts, the emitted-provenance totality guard's IO, the differential check's baseline acquisition, CI wiring to publish/restore the baseline artifact, and docs. * refactor(#2724): make the differential attribution check self-sufficient Three fixes required to delete the golden fixtures without breaking CI: - scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs from the three rules that named it. #2759's missingRuleTestFiles guard hard-throws at module load if a rule names a test file absent from disk, which would break the changes job on every PR the moment the fixture-deletion commit landed. - tests/helpers/emitted-provenance.cjs: loadManifests() read the committed golden fixture directory. With that directory deleted at every future ref, this would throw at module load forever, taking the Phase 2 totality guard down with it. Rebuilt from real installer spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest), the same shape emitted-runtime.cjs's currentManifests() already uses. - tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs: the real-tree test's baseline acquisition swaps from baselineManifestsAtRef(base) (git show at a ref that no longer carries fixtures) to resolveBaseline()'s documented precedence: env, then the on-disk cache, then an in-job build. The build fallback (buildBaselineAtRef, new) checks out base into a throwaway git worktree and runs the new scripts/gen-emitted-baseline.cjs there -- no npm ci needed, since bin/install.js and the test helper shells are Node-builtins-only. That script also publishes the baseline artifact from CI's push-to-next job (wired in a follow-up commit). * refactor(#2724): retire the merge-driver bridge and per-file size baselines The Phase 1 bridge (#2721) is retired now that the artifacts it guarded are deleted: scripts/git-merge-regen-driver.cjs, its test, and the 'setup:merge-driver' npm script are removed, and the .gitattributes merge=gsd-regen/linguist-generated block for the three deleted-path globs is dropped. tests/fixtures/install-tree/*.json keeps its normal merge behavior, unchanged (ADR-2719 section 7). scripts/update-size-baseline.cjs and its test are removed: their sole purpose was regenerating tests/workflow-size-baseline.json and tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm script and its step in 'regen:derived' go with it. The per-file baseline describe blocks in tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs are removed for the same reason; the independent loose-tier hard caps are untouched. The differential attribution check's size ratchet (tests/emitted-diff.cjs, already shipped in #2723) is the replacement anti-creep mechanism. 'npm run gen:golden' is replaced by 'npm run gen:install-tree', which keeps regenerating tests/fixtures/install-tree/*.json (the one artifact family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's error messages point at the new command name. tests/golden-parity-single-source.test.cjs's anti-divergence guard (#2266) is retargeted from the two deleted golden-parity consumers to their two replacements (tests/helpers/emitted-runtime.cjs and tests/helpers/emitted-provenance.cjs), which import buildParityManifest the same way — the divergence risk the guard exists for is unchanged. Also wires CI: a new publish-emitted-baseline job runs scripts/gen-emitted-baseline.cjs after a push to next and caches the result keyed on the sha; the test and test-full jobs restore that cache on pull_request events, keyed on the PR's base sha, and export GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree test to pick up. * docs(#2724): flip ADR-2719 to Accepted and update contributor docs Status: Proposed -> Accepted. Regenerated docs/adr/README.md index. CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET. EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET. AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry) no longer point at the deleted golden-install-parity fixtures, size baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver / git-merge-regen-driver.cjs bridge. Editing shipped content now requires zero manual fixture regeneration, documented against the differential attribution check instead of the deleted commands. * docs(#2724): add changeset for removed golden-parity commands * fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs token in CONTEXT.md resolves to a real file. The RULESET. EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks, which the checker reads as a live reference, not historical prose. * test(#2724): retarget ci-test-scope tests off the deleted golden test tests/ci-test-scope.test.cjs asserted specific RULES entries select tests/golden-install-parity.test.cjs, and that every rule selecting it also selects both emitted gates. Both premises broke when the golden test was deleted (#2724): the deleted filename never re-appears in targeted_tests, and there was no longer a third file for the gates to travel alongside. Retargeted the two selection describe blocks to assert tests/emitted-provenance.test.cjs directly (the drift guard the golden gate's rules were retargeted to), and simplified the third block to assert the two emitted gates always travel together, without reference to the golden filename. * docs(#2724): repoint two contributor how-to guides at the differential check Both guides told contributors to regenerate a baseline against tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed at the differential attribution check (tests/emitted-attribution.test.cjs, ADR-2719), which needs no manual regeneration step. * fix(#2724): repair phase6-capstone-conformance's deleted-baseline read An independent orthogonal review caught a real regression this branch introduced into a test file the branch's diff never touched: tests/phase6-capstone-conformance.test.cjs read tests/workflow-size-baseline.json (deleted earlier in this branch) with no fallback, so the whole suite would throw ENOENT the moment this branch landed. The test's actual intent — prove the host-loop workflow files are real, tracked, non-empty docs — is preserved by asserting the live byte count via the same shared counter (scripts/workflow-size.cjs) the size guards already use, instead of a committed snapshot. Also, from the same review: a stale doc comment in scripts/workflow-size.cjs still named the deleted scripts/update-size-baseline.cjs as a consumer, and buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left two fs.rmSync calls unguarded against masking the primary result/error, inconsistent with the try/catch already wrapping the git cleanup beside them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef explaining why it (and its siblings) are kept despite having no production caller post-cutover — they still answer real questions about refs that predate the cutover. * fix(#2724): repair three real regressions found by remote verification 1. tests/emitted-provenance.test.cjs's two hostile-input tests (non-object manifest, unreadable fixture) drove loadManifests(tmp) and monkeypatched fs.readFileSync, both premised on the deleted fixture-directory read this branch already replaced with real installer spawns -- the negative assertions silently stopped firing. loadManifests() now accepts injected {families, install, build, clean} (defaulting to production values), giving the tests a real seam to drive a bad build result and a build failure through the ACTUAL loader instead of a reimplementation, and added coverage that clean() still runs on both paths. 2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE' steps hardcoded shell: bash, which is wrong on windows-latest (native pwsh) and on test-full's macos-latest legs (native zsh per that job's own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning .test.cjs) caught it. Replaced the inline bash script with scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a bare 'node <path>' command line has no shell-specific syntax, so it runs correctly under bash, zsh, and pwsh without a shell override. tests/phase6-capstone-conformance.test.cjs's deleted-baseline read (caught by the same remote run, at a commit prior to this one) was already fixed in d0c3b1242 and is not touched here; verified still passing after these changes. * fix(#2724): revive ADR-1610's new-file size cap inside the differential An isolated review caught a real regression: deleting tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP (ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor) with no successor. tests/helpers/emitted-diff.cjs's size ratchet already 'continue's past any file absent from sizeBaseline -- exactly the files this cap exists to bound -- so a brand-new workflow file sized 32,769-40,960 bytes passed CI clean and shipped, then risked silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted and never referenced anywhere in this branch. Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline, name) signal the growth check already computes -- 'new' is exactly 'present in sizeCurrent, absent from sizeBaseline'. Not ack-able, matching the tier hard caps it sits beside: the fix is extraction, not an acknowledgment entry. Documented, disclosed narrowing: the pure differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering (tests/workflow-size-budget.test.cjs's classification), so a legitimately large new file must extract rather than tier in, one release earlier than an existing file would need to. ADR-1610 itself is left unamended -- this restores its decision rather than re-litigating it. Also fixes a stale comment plus a redundant real 19-installer-spawn assertion left over from the pre-injection-seam version of tests/emitted-provenance.test.cjs's build-failure test, and annotates 3 of 4 stale golden-fixture citations in docs/reference/host-integration-capability-matrix.md as superseded (the 4th is an accurate historical PR narrative, left alone). * fix(#2724): repair three red CI defects on the golden-fixture cutover Windows-only provenance false attribution (defect A): the `hooks-built` provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are Windows-only installer output (ensureCodexHooksJsonSessionStart / ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so the self-attribution resolved to a path that exists on no platform. Only windows-latest ever emits the key, so this only failed there. Fixed by special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated rule) — a dedicated rule would match zero paths, and therefore report as a dead rule, on every non-Windows lane of the same totality guard. `sources` already supported per-match functions; `transforms` is extended to support the same shape so the attribution can vary by match within one rule. Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef` ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but that script is new in this PR and therefore absent at any base ref that predates it — every call failed closed with "Cannot find module". Fixed by running the PR checkout's own generator against the worktree via a new `--dir` parameter, decoupling "which copy of the script runs" from "which tree it measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override, threaded down to `runMinimalInstall`'s new `installScript` override). This is not just a bootstrap fix: a differential needs ONE measurement schema applied to both sides, or the two stop being comparable the moment that schema evolves — running each side's own copy would silently reintroduce that risk. Verified locally end-to-end against real origin/next: resolves a valid {version, sha, manifests, sizes} artifact with the correct sha and no leaked worktree. Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let docs-lint evaluate the fragment for the first time; it already passes (docs/TESTING-SUITES.md and friends already document the removed scripts). Also fixed while in this file: an eslint no-unused-vars warning surfaced by the changed lint run (unused `cleanup` import in tests/emitted-provenance.test.cjs). Added regression coverage for both A and B: a cross-platform spot-check that drives the real hooks-built rule against `.cmd` keys directly (not through a real Windows install), and a real-tree test that drives buildBaselineAtRef against a base ref verified (via git cat-file) to lack the generator, both skipping honestly rather than false-passing when their precondition does not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test Two isolated-review findings on PR #2767: - `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the wrapped `hooks/<name>.js` script, asserting a byte-provenance link that does not exist — traced against buildCodexHookWindowsShimIR (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that same file) flows into the .cmd bytes, never its content. Point `sources` at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention used elsewhere in the table. Since `sources` is checked before `transforms` in the differential, the wrong mapping silently excused any .cmd byte movement caused by editing the wrapped .js file. - The `buildBaselineAtRef` regression test skipped unless a resolvable base ref still lacked scripts/gen-emitted-baseline.cjs — true only until this PR merges, after which every base ref carries the file and the test skips forever with zero ongoing coverage. Rebuilt hermetically: synthesize the missing-generator condition in-place via git plumbing (a throwaway commit, child of HEAD, with just that one file removed from a scratch index), never touching the real working tree, HEAD, or index, and never depending on ambient history or remotes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path The runner container mounts the repo at a path owned by a different uid than the process running the suite, so git's dubious-ownership protection refuses every git operation there. GitHub Actions never hits this because actions/checkout registers the workspace as safe automatically; this runner's container does not. buildBaselineAtRef is the production build-fallback the sole remaining emitted gate depends on (resolveBaseline's in-job-build leg), not just a test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs) that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's worktree add/remove/prune, and the hermetic regression test added in the prior commit — funnels through, plus gen-emitted-baseline.cjs's own rev-parse (now reusing that same wrapper instead of a second execFileSync, so the fix has one source of truth). Each call declares -c safe.directory=<the exact directory it already operates on>, never the * wildcard. Audited every other helper on this surface (emitted-diff.cjs, emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them shell out to git at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
176 lines
7.2 KiB
JavaScript
176 lines
7.2 KiB
JavaScript
// allow-test-rule: source-text-is-the-product
|
|
// Workflow .md / agent .md / command .md / reference .md files — their text
|
|
// IS what the runtime loads. Testing text content tests the deployed contract.
|
|
// Per CONTRIBUTING.md exception matrix.
|
|
|
|
/**
|
|
* Agent size budget (measured in BYTES — see #717).
|
|
*
|
|
* Agent definitions in `agents/gsd-*.md` are loaded verbatim into the agent's
|
|
* context on every subagent dispatch. Unbounded growth is paid on every call
|
|
* across every workflow.
|
|
*
|
|
* ## Enforcement model (issue #1074)
|
|
*
|
|
* Mirrors tests/workflow-size-budget.test.cjs — two complementary guards, no
|
|
* tier-max ceiling:
|
|
*
|
|
* 1. Per-agent baseline (the anti-creep): every agent is pinned to its exact
|
|
* byte size in `tests/agent-size-baseline.json`. Any growth fails with the
|
|
* file and delta; `npm run size:baseline` records a deliberate change as a
|
|
* reviewable one-line diff. This replaced the tier-max tighten-only ratchet
|
|
* (which only bound the single largest agent per tier).
|
|
*
|
|
* 2. Tier hard caps (the outer bound): XL/LARGE/DEFAULT absolute red lines
|
|
* with real headroom, never raised in normal work. Crossing one means
|
|
* extracting shared boilerplate to `gsd-core/references/`, not a +N bump.
|
|
* A net-new agent is DEFAULT-tier, so the DEFAULT cap already bounds it —
|
|
* no separate new-file cap is needed (DEFAULT is already small).
|
|
*
|
|
* Tiers:
|
|
* - XL : top-level orchestrators that own end-to-end rubrics
|
|
* - LARGE : multi-phase operators with branching workflows
|
|
* - DEFAULT : focused single-purpose agents
|
|
*
|
|
* See:
|
|
* - https://github.com/open-gsd/gsd-core/issues/1074 (per-file baseline + hard caps)
|
|
* - https://github.com/open-gsd/gsd-core/issues/717 (bytes, not lines)
|
|
* - https://github.com/open-gsd/gsd-core/issues/683 (LF-normalized byte count)
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('fs');
|
|
const os = require('node:os');
|
|
const path = require('path');
|
|
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
|
const { cleanup } = require('./helpers.cjs');
|
|
|
|
const AGENTS_DIR = path.join(__dirname, '..', 'agents');
|
|
const isGsdAgent = (f) => f.startsWith('gsd-');
|
|
|
|
// Tier HARD CAPS (#1074, bytes) — absolute red lines, not high-water-hugging
|
|
// ceilings. Day-to-day creep is caught per-agent by the baseline guard below;
|
|
// these sit above each tier's current high-water with real headroom:
|
|
// XL 56 KiB — high-water gsd-debugger 51,043 → ~6.3 KB headroom
|
|
// LARGE 48 KiB — high-water gsd-executor 42,342 → ~6.8 KB headroom
|
|
// DEFAULT 24 KiB — high-water gsd-ui-researcher 19,095 → ~5.5 KB headroom
|
|
const XL_CAP = 57344; // 56 KiB
|
|
const LARGE_CAP = 49152; // 48 KiB
|
|
const DEFAULT_CAP = 24576; // 24 KiB
|
|
|
|
const XL_AGENTS = new Set([
|
|
'gsd-debugger',
|
|
'gsd-planner',
|
|
]);
|
|
|
|
const LARGE_AGENTS = new Set([
|
|
'gsd-phase-researcher',
|
|
'gsd-verifier',
|
|
'gsd-doc-writer',
|
|
'gsd-plan-checker',
|
|
'gsd-executor',
|
|
'gsd-code-fixer',
|
|
'gsd-codebase-mapper',
|
|
'gsd-project-researcher',
|
|
'gsd-roadmapper',
|
|
]);
|
|
|
|
const ALL_AGENTS = fs.readdirSync(AGENTS_DIR)
|
|
.filter(f => isGsdAgent(f) && f.endsWith('.md'))
|
|
.map(f => f.replace('.md', ''));
|
|
|
|
function capFor(agent) {
|
|
if (XL_AGENTS.has(agent)) return { tier: 'XL', cap: XL_CAP };
|
|
if (LARGE_AGENTS.has(agent)) return { tier: 'LARGE', cap: LARGE_CAP };
|
|
return { tier: 'DEFAULT', cap: DEFAULT_CAP };
|
|
}
|
|
|
|
describe('SIZE: agent tier hard caps (issue #1074)', () => {
|
|
// Absolute outer bound per tier. A cap is NOT raised when an agent approaches
|
|
// it — crossing it means extract shared boilerplate to gsd-core/references/.
|
|
for (const agent of ALL_AGENTS) {
|
|
const { tier, cap } = capFor(agent);
|
|
test(`${agent} (${tier}) stays within the ${tier} hard cap (${cap} bytes)`, () => {
|
|
const bytes = lfByteCount(path.join(AGENTS_DIR, agent + '.md'));
|
|
assert.ok(
|
|
bytes <= cap,
|
|
`${agent}.md is ${bytes} bytes — exceeds the ${tier} hard cap of ${cap}. ` +
|
|
`This cap is a red line, NOT a budget to raise: extract shared boilerplate ` +
|
|
`to gsd-core/references/ and load it lazily.`
|
|
);
|
|
});
|
|
}
|
|
});
|
|
|
|
describe('SIZE: agent hard-cap boundary fixtures (#1074 — negative proof)', () => {
|
|
// The per-tier loop above only iterates the real, fully-compliant agent
|
|
// corpus, so its `bytes <= cap` failure branch never executes. Exercise that
|
|
// exact comparison on synthetic files measured at cap-1 / cap / cap+1 (the
|
|
// limit boundary — RULESET.TESTS.boundary-coverage.fixtures) through the SAME
|
|
// lfByteCount path the guard uses, so a future threshold or operator edit
|
|
// cannot silently neuter a cap (RULESET.TESTS.regression-must-fail-first).
|
|
test('cap comparison fires at the limit boundary for every tier', () => {
|
|
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-size-'));
|
|
try {
|
|
// ASCII 'a' is 1 byte/char and has no CRLF, so lfByteCount == length.
|
|
const measureAt = (n) => {
|
|
const p = path.join(tmp, `fixture-${n}.md`);
|
|
fs.writeFileSync(p, 'a'.repeat(n));
|
|
return lfByteCount(p);
|
|
};
|
|
for (const cap of [DEFAULT_CAP, LARGE_CAP, XL_CAP]) {
|
|
assert.equal(measureAt(cap - 1) <= cap, true, `${cap - 1} must be within cap ${cap}`);
|
|
assert.equal(measureAt(cap) <= cap, true, `${cap} (exactly at cap) must be within cap ${cap}`);
|
|
assert.equal(measureAt(cap + 1) <= cap, false, `${cap + 1} must exceed cap ${cap}`);
|
|
}
|
|
} finally {
|
|
cleanup(tmp);
|
|
}
|
|
});
|
|
});
|
|
|
|
// A prior "SIZE: per-agent baseline (issue #1074)" describe block lived here,
|
|
// asserting every agent's exact byte count against the committed
|
|
// `tests/agent-size-baseline.json` snapshot. #2724 (ADR-2719 Phase 4) deletes that
|
|
// snapshot: it was a pure function of the source tree, and its purpose — "growth
|
|
// must be noticed and justified" — is now served by the same differential machine
|
|
// that replaced the golden-install-parity fixtures (tests/emitted-attribution.test.cjs's
|
|
// real-tree test, via `emitted-diff.cjs`'s size ratchet: growth is reported with its
|
|
// exact byte delta and requires an entry in tests/emitted-drift-ack.json, ADR-2719 §4 /
|
|
// must-have 6). The tier hard caps above are unaffected — they are independent of the
|
|
// deleted baseline and remain the outer bound.
|
|
|
|
describe('SIZE: every agent is classified', () => {
|
|
test('every agent falls in exactly one tier', () => {
|
|
for (const agent of ALL_AGENTS) {
|
|
const inXL = XL_AGENTS.has(agent);
|
|
const inLarge = LARGE_AGENTS.has(agent);
|
|
assert.ok(
|
|
!(inXL && inLarge),
|
|
`${agent} is in both XL_AGENTS and LARGE_AGENTS — pick one`
|
|
);
|
|
}
|
|
});
|
|
|
|
test('every named XL agent exists', () => {
|
|
for (const agent of XL_AGENTS) {
|
|
const filePath = path.join(AGENTS_DIR, agent + '.md');
|
|
assert.ok(
|
|
fs.existsSync(filePath),
|
|
`XL_AGENTS references ${agent}.md which does not exist — clean up the set`
|
|
);
|
|
}
|
|
});
|
|
|
|
test('every named LARGE agent exists', () => {
|
|
for (const agent of LARGE_AGENTS) {
|
|
const filePath = path.join(AGENTS_DIR, agent + '.md');
|
|
assert.ok(
|
|
fs.existsSync(filePath),
|
|
`LARGE_AGENTS references ${agent}.md which does not exist — clean up the set`
|
|
);
|
|
}
|
|
});
|
|
});
|