* test(#2724): delete golden-install-parity fixtures, test, and generator Removes the 19 committed path->hash manifests, the two per-file size baselines, tests/golden-install-parity.test.cjs, and scripts/gen-golden-install-parity-zcode.cjs. These were pure functions of the source tree (ADR-2719); the differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the sole gate for emitted-artifact propagation. tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs are unchanged (ADR-2719 section 7 exception). Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's existence guard, .gitattributes, package.json scripts, the emitted-provenance totality guard's IO, the differential check's baseline acquisition, CI wiring to publish/restore the baseline artifact, and docs. * refactor(#2724): make the differential attribution check self-sufficient Three fixes required to delete the golden fixtures without breaking CI: - scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs from the three rules that named it. #2759's missingRuleTestFiles guard hard-throws at module load if a rule names a test file absent from disk, which would break the changes job on every PR the moment the fixture-deletion commit landed. - tests/helpers/emitted-provenance.cjs: loadManifests() read the committed golden fixture directory. With that directory deleted at every future ref, this would throw at module load forever, taking the Phase 2 totality guard down with it. Rebuilt from real installer spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest), the same shape emitted-runtime.cjs's currentManifests() already uses. - tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs: the real-tree test's baseline acquisition swaps from baselineManifestsAtRef(base) (git show at a ref that no longer carries fixtures) to resolveBaseline()'s documented precedence: env, then the on-disk cache, then an in-job build. The build fallback (buildBaselineAtRef, new) checks out base into a throwaway git worktree and runs the new scripts/gen-emitted-baseline.cjs there -- no npm ci needed, since bin/install.js and the test helper shells are Node-builtins-only. That script also publishes the baseline artifact from CI's push-to-next job (wired in a follow-up commit). * refactor(#2724): retire the merge-driver bridge and per-file size baselines The Phase 1 bridge (#2721) is retired now that the artifacts it guarded are deleted: scripts/git-merge-regen-driver.cjs, its test, and the 'setup:merge-driver' npm script are removed, and the .gitattributes merge=gsd-regen/linguist-generated block for the three deleted-path globs is dropped. tests/fixtures/install-tree/*.json keeps its normal merge behavior, unchanged (ADR-2719 section 7). scripts/update-size-baseline.cjs and its test are removed: their sole purpose was regenerating tests/workflow-size-baseline.json and tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm script and its step in 'regen:derived' go with it. The per-file baseline describe blocks in tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs are removed for the same reason; the independent loose-tier hard caps are untouched. The differential attribution check's size ratchet (tests/emitted-diff.cjs, already shipped in #2723) is the replacement anti-creep mechanism. 'npm run gen:golden' is replaced by 'npm run gen:install-tree', which keeps regenerating tests/fixtures/install-tree/*.json (the one artifact family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's error messages point at the new command name. tests/golden-parity-single-source.test.cjs's anti-divergence guard (#2266) is retargeted from the two deleted golden-parity consumers to their two replacements (tests/helpers/emitted-runtime.cjs and tests/helpers/emitted-provenance.cjs), which import buildParityManifest the same way — the divergence risk the guard exists for is unchanged. Also wires CI: a new publish-emitted-baseline job runs scripts/gen-emitted-baseline.cjs after a push to next and caches the result keyed on the sha; the test and test-full jobs restore that cache on pull_request events, keyed on the PR's base sha, and export GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree test to pick up. * docs(#2724): flip ADR-2719 to Accepted and update contributor docs Status: Proposed -> Accepted. Regenerated docs/adr/README.md index. CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET. EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET. AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry) no longer point at the deleted golden-install-parity fixtures, size baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver / git-merge-regen-driver.cjs bridge. Editing shipped content now requires zero manual fixture regeneration, documented against the differential attribution check instead of the deleted commands. * docs(#2724): add changeset for removed golden-parity commands * fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs token in CONTEXT.md resolves to a real file. The RULESET. EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks, which the checker reads as a live reference, not historical prose. * test(#2724): retarget ci-test-scope tests off the deleted golden test tests/ci-test-scope.test.cjs asserted specific RULES entries select tests/golden-install-parity.test.cjs, and that every rule selecting it also selects both emitted gates. Both premises broke when the golden test was deleted (#2724): the deleted filename never re-appears in targeted_tests, and there was no longer a third file for the gates to travel alongside. Retargeted the two selection describe blocks to assert tests/emitted-provenance.test.cjs directly (the drift guard the golden gate's rules were retargeted to), and simplified the third block to assert the two emitted gates always travel together, without reference to the golden filename. * docs(#2724): repoint two contributor how-to guides at the differential check Both guides told contributors to regenerate a baseline against tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed at the differential attribution check (tests/emitted-attribution.test.cjs, ADR-2719), which needs no manual regeneration step. * fix(#2724): repair phase6-capstone-conformance's deleted-baseline read An independent orthogonal review caught a real regression this branch introduced into a test file the branch's diff never touched: tests/phase6-capstone-conformance.test.cjs read tests/workflow-size-baseline.json (deleted earlier in this branch) with no fallback, so the whole suite would throw ENOENT the moment this branch landed. The test's actual intent — prove the host-loop workflow files are real, tracked, non-empty docs — is preserved by asserting the live byte count via the same shared counter (scripts/workflow-size.cjs) the size guards already use, instead of a committed snapshot. Also, from the same review: a stale doc comment in scripts/workflow-size.cjs still named the deleted scripts/update-size-baseline.cjs as a consumer, and buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left two fs.rmSync calls unguarded against masking the primary result/error, inconsistent with the try/catch already wrapping the git cleanup beside them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef explaining why it (and its siblings) are kept despite having no production caller post-cutover — they still answer real questions about refs that predate the cutover. * fix(#2724): repair three real regressions found by remote verification 1. tests/emitted-provenance.test.cjs's two hostile-input tests (non-object manifest, unreadable fixture) drove loadManifests(tmp) and monkeypatched fs.readFileSync, both premised on the deleted fixture-directory read this branch already replaced with real installer spawns -- the negative assertions silently stopped firing. loadManifests() now accepts injected {families, install, build, clean} (defaulting to production values), giving the tests a real seam to drive a bad build result and a build failure through the ACTUAL loader instead of a reimplementation, and added coverage that clean() still runs on both paths. 2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE' steps hardcoded shell: bash, which is wrong on windows-latest (native pwsh) and on test-full's macos-latest legs (native zsh per that job's own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning .test.cjs) caught it. Replaced the inline bash script with scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a bare 'node <path>' command line has no shell-specific syntax, so it runs correctly under bash, zsh, and pwsh without a shell override. tests/phase6-capstone-conformance.test.cjs's deleted-baseline read (caught by the same remote run, at a commit prior to this one) was already fixed in d0c3b1242 and is not touched here; verified still passing after these changes. * fix(#2724): revive ADR-1610's new-file size cap inside the differential An isolated review caught a real regression: deleting tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP (ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor) with no successor. tests/helpers/emitted-diff.cjs's size ratchet already 'continue's past any file absent from sizeBaseline -- exactly the files this cap exists to bound -- so a brand-new workflow file sized 32,769-40,960 bytes passed CI clean and shipped, then risked silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted and never referenced anywhere in this branch. Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline, name) signal the growth check already computes -- 'new' is exactly 'present in sizeCurrent, absent from sizeBaseline'. Not ack-able, matching the tier hard caps it sits beside: the fix is extraction, not an acknowledgment entry. Documented, disclosed narrowing: the pure differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering (tests/workflow-size-budget.test.cjs's classification), so a legitimately large new file must extract rather than tier in, one release earlier than an existing file would need to. ADR-1610 itself is left unamended -- this restores its decision rather than re-litigating it. Also fixes a stale comment plus a redundant real 19-installer-spawn assertion left over from the pre-injection-seam version of tests/emitted-provenance.test.cjs's build-failure test, and annotates 3 of 4 stale golden-fixture citations in docs/reference/host-integration-capability-matrix.md as superseded (the 4th is an accurate historical PR narrative, left alone). * fix(#2724): repair three red CI defects on the golden-fixture cutover Windows-only provenance false attribution (defect A): the `hooks-built` provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are Windows-only installer output (ensureCodexHooksJsonSessionStart / ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so the self-attribution resolved to a path that exists on no platform. Only windows-latest ever emits the key, so this only failed there. Fixed by special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated rule) — a dedicated rule would match zero paths, and therefore report as a dead rule, on every non-Windows lane of the same totality guard. `sources` already supported per-match functions; `transforms` is extended to support the same shape so the attribution can vary by match within one rule. Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef` ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but that script is new in this PR and therefore absent at any base ref that predates it — every call failed closed with "Cannot find module". Fixed by running the PR checkout's own generator against the worktree via a new `--dir` parameter, decoupling "which copy of the script runs" from "which tree it measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override, threaded down to `runMinimalInstall`'s new `installScript` override). This is not just a bootstrap fix: a differential needs ONE measurement schema applied to both sides, or the two stop being comparable the moment that schema evolves — running each side's own copy would silently reintroduce that risk. Verified locally end-to-end against real origin/next: resolves a valid {version, sha, manifests, sizes} artifact with the correct sha and no leaked worktree. Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let docs-lint evaluate the fragment for the first time; it already passes (docs/TESTING-SUITES.md and friends already document the removed scripts). Also fixed while in this file: an eslint no-unused-vars warning surfaced by the changed lint run (unused `cleanup` import in tests/emitted-provenance.test.cjs). Added regression coverage for both A and B: a cross-platform spot-check that drives the real hooks-built rule against `.cmd` keys directly (not through a real Windows install), and a real-tree test that drives buildBaselineAtRef against a base ref verified (via git cat-file) to lack the generator, both skipping honestly rather than false-passing when their precondition does not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test Two isolated-review findings on PR #2767: - `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the wrapped `hooks/<name>.js` script, asserting a byte-provenance link that does not exist — traced against buildCodexHookWindowsShimIR (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that same file) flows into the .cmd bytes, never its content. Point `sources` at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention used elsewhere in the table. Since `sources` is checked before `transforms` in the differential, the wrong mapping silently excused any .cmd byte movement caused by editing the wrapped .js file. - The `buildBaselineAtRef` regression test skipped unless a resolvable base ref still lacked scripts/gen-emitted-baseline.cjs — true only until this PR merges, after which every base ref carries the file and the test skips forever with zero ongoing coverage. Rebuilt hermetically: synthesize the missing-generator condition in-place via git plumbing (a throwaway commit, child of HEAD, with just that one file removed from a scratch index), never touching the real working tree, HEAD, or index, and never depending on ambient history or remotes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path The runner container mounts the repo at a path owned by a different uid than the process running the suite, so git's dubious-ownership protection refuses every git operation there. GitHub Actions never hits this because actions/checkout registers the workspace as safe automatically; this runner's container does not. buildBaselineAtRef is the production build-fallback the sole remaining emitted gate depends on (resolveBaseline's in-job-build leg), not just a test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs) that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's worktree add/remove/prune, and the hermetic regression test added in the prior commit — funnels through, plus gen-emitted-baseline.cjs's own rev-parse (now reusing that same wrapper instead of a second execFileSync, so the fix has one source of truth). Each call declares -c safe.directory=<the exact directory it already operates on>, never the * wildcard. Audited every other helper on this surface (emitted-diff.cjs, emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them shell out to git at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
380 lines
18 KiB
JavaScript
380 lines
18 KiB
JavaScript
// allow-test-rule: source-text-is-the-product
|
|
'use strict';
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const { execFileSync } = require('node:child_process');
|
|
const { cleanup } = require('./helpers.cjs');
|
|
|
|
const ROOT = path.join(__dirname, '..');
|
|
const { HOST_LOOP_FILES, scanWiredPoints } = require('../scripts/gen-loop-host-contract.cjs');
|
|
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
|
|
|
const CORE_SUBSTRATE_TERMS = [
|
|
'Verification substrate',
|
|
'verifier↔predicate contract',
|
|
'Probe Core Module',
|
|
'Edge Probe Module',
|
|
];
|
|
|
|
const registry = require('../gsd-core/bin/lib/capability-registry.cjs');
|
|
const { isCentralConfigKey } = require('../gsd-core/bin/lib/config-schema.cjs');
|
|
|
|
function readRepoFile(relativePath) {
|
|
return fs.readFileSync(path.join(ROOT, relativePath), 'utf8');
|
|
}
|
|
|
|
function escapeRegExp(value) {
|
|
return value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
|
}
|
|
|
|
function activeWhenKeys() {
|
|
const keys = new Set();
|
|
for (const cap of Object.values(registry.capabilities)) {
|
|
for (const group of ['steps', 'gates', 'contributions']) {
|
|
for (const hook of cap[group] || []) {
|
|
if (hook.when) keys.add(hook.when);
|
|
}
|
|
}
|
|
}
|
|
return [...keys].sort();
|
|
}
|
|
|
|
describe('ADR-857 Phase 6 capstone conformance (#1139)', () => {
|
|
test('first-party optional feature capabilities are declared in the generated registry', () => {
|
|
const expectedFeatureCapabilities = [
|
|
'ai-integration',
|
|
'audit',
|
|
'code-review',
|
|
'graphify',
|
|
'intel',
|
|
'nyquist',
|
|
'pattern-mapper',
|
|
'research',
|
|
'security',
|
|
'ui',
|
|
];
|
|
|
|
for (const capId of expectedFeatureCapabilities) {
|
|
assert.equal(registry.capabilities[capId]?.role, 'feature', `${capId} must be a feature Capability`);
|
|
}
|
|
});
|
|
|
|
test('core verification substrate is documented as deliberately not capability-owned', () => {
|
|
const context = readRepoFile('CONTEXT.md');
|
|
for (const term of CORE_SUBSTRATE_TERMS) {
|
|
assert.match(context, new RegExp(escapeRegExp(term)), `${term} must be documented in CONTEXT.md`);
|
|
}
|
|
});
|
|
|
|
test('host loop files do not read capability hook activation keys directly', () => {
|
|
const forbiddenKeys = activeWhenKeys();
|
|
assert.ok(forbiddenKeys.length > 0, 'registry must expose hook activation keys');
|
|
|
|
for (const relativePath of HOST_LOOP_FILES) {
|
|
const content = readRepoFile(relativePath);
|
|
for (const key of forbiddenKeys) {
|
|
assert.doesNotMatch(
|
|
content,
|
|
new RegExp(`\\bconfig-get\\s+${escapeRegExp(key)}\\b`),
|
|
`${relativePath} must resolve ${key} through Capability hooks/state, not direct config-get`,
|
|
);
|
|
}
|
|
}
|
|
});
|
|
|
|
test('capability-owned config keys are not reintroduced into the central schema', () => {
|
|
for (const key of Object.keys(registry.configKeys).sort()) {
|
|
assert.equal(
|
|
isCentralConfigKey(key),
|
|
false,
|
|
`${key} is owned by capability ${registry.configKeys[key]} and must stay out of central config schema`,
|
|
);
|
|
}
|
|
});
|
|
|
|
test('host loop workflow files have a measurable, non-empty byte size', () => {
|
|
// Was asserted against the committed tests/workflow-size-baseline.json snapshot;
|
|
// #2724 (ADR-2719 Phase 4) deletes that file — the differential attribution
|
|
// check's size ratchet (tests/emitted-attribution.test.cjs) is the replacement
|
|
// anti-creep mechanism, but this test's actual intent was narrower: prove these
|
|
// host-loop files are real, tracked, non-empty workflow docs. Asserting the
|
|
// live byte count via the same shared counter the size guards use preserves
|
|
// that intent without depending on a committed snapshot.
|
|
for (const relativePath of HOST_LOOP_FILES) {
|
|
const fileName = path.basename(relativePath);
|
|
const bytes = lfByteCount(path.join(ROOT, relativePath));
|
|
assert.ok(bytes > 0, `${fileName} must be a non-empty workflow file`);
|
|
}
|
|
});
|
|
|
|
// ─── Phase-6 conformance: RED BY DESIGN until phase 6 is actually complete ──────
|
|
//
|
|
// #1139 closed (via #1158) with a green "capstone conformance gate" while the
|
|
// ADR-857 phase-6 acceptance criteria were unmet — a false green. The three
|
|
// tests below assert the real criteria with NO paper-over allowlist, so the
|
|
// gate stays RED until the work lands. Green here must mean "phase 6 conformant,"
|
|
// not "no new regression." Fixes tracked in #1167 / #1168 / #1169.
|
|
|
|
test('every declared capability hook point has a render-hooks call site in the host loop (#1168)', () => {
|
|
// No allowlist: every point a capability declares a hook at MUST have a
|
|
// `render-hooks` call site in the host loop, or those hooks can never fire.
|
|
const declaredPoints = new Set();
|
|
for (const cap of Object.values(registry.capabilities)) {
|
|
for (const group of ['steps', 'gates', 'contributions']) {
|
|
for (const hook of cap[group] || []) {
|
|
if (hook.point) declaredPoints.add(hook.point);
|
|
}
|
|
}
|
|
}
|
|
|
|
// Scan only the host loop files (a `render-hooks` mention in a non-host
|
|
// workflow must not mask a lost host call site).
|
|
const callSites = new Set();
|
|
for (const relativePath of HOST_LOOP_FILES) {
|
|
const content = readRepoFile(relativePath);
|
|
for (const pt of scanWiredPoints(content)) callSites.add(pt);
|
|
}
|
|
|
|
const orphaned = [...declaredPoints].sort().filter((p) => !callSites.has(p));
|
|
assert.deepEqual(
|
|
orphaned, [],
|
|
`ADR-857 phase 6 is NOT complete: capability hooks declare these extension points ` +
|
|
`but no host-loop workflow calls \`gsd_run loop render-hooks <point>\`, so the hooks ` +
|
|
`can never fire: ${orphaned.join(', ')}. Wire each call site (#1167/#1169).`,
|
|
);
|
|
});
|
|
|
|
test('all ADR-857-named optional features are real Capabilities, not empty stubs (#1169)', () => {
|
|
// ADR-857 §53 + Decision 7 enumerate these optional, non-loop modules as
|
|
// Capabilities. "Migrated" means the feature OWNS its behavior: hook-based
|
|
// features (tdd/schema-gate/drift/gap-analysis) must declare >=1 hook;
|
|
// command-family features (profile-pipeline) must declare a command family.
|
|
// A registration-only stub (role:feature but no hooks/commands) games this
|
|
// gate while the logic stays welded into the loop — rejected here.
|
|
const REQUIRED = ['tdd', 'schema-gate', 'drift', 'gap-analysis', 'profile-pipeline'];
|
|
const problems = [];
|
|
for (const id of REQUIRED) {
|
|
const cap = registry.capabilities[id];
|
|
if (!cap) { problems.push(`${id}: not registered`); continue; }
|
|
if (cap.role !== 'feature') { problems.push(`${id}: role="${cap.role}", must be "feature"`); continue; }
|
|
const hookCount = (cap.steps?.length || 0) + (cap.contributions?.length || 0) + (cap.gates?.length || 0);
|
|
const isCommandFamily = (cap.commands?.length || 0) > 0;
|
|
if (hookCount === 0 && !isCommandFamily) {
|
|
problems.push(`${id}: EMPTY STUB (no hooks, no command family) — inline logic was not migrated; declare the real hooks/commands and remove the inline branch`);
|
|
}
|
|
}
|
|
assert.deepEqual(
|
|
problems, [],
|
|
`ADR-857 phase 6 is NOT complete:\n ${problems.join('\n ')}\n` +
|
|
`Each feature must OWN its behavior via hooks or a command family — not exist as a registration-only stub (#1169).`,
|
|
);
|
|
});
|
|
|
|
test('host loop reads no capability-owned config key inline (#1169)', () => {
|
|
// Phase 6 requires the loop to resolve capability behavior via render-hooks,
|
|
// not by reading capability-owned keys directly. Any inline `config-get` of a
|
|
// registry-owned key is an incomplete migration (the loop still owns the
|
|
// feature's params).
|
|
const leaks = [];
|
|
for (const relativePath of HOST_LOOP_FILES) {
|
|
const content = readRepoFile(relativePath);
|
|
for (const key of Object.keys(registry.configKeys)) {
|
|
if (new RegExp(`\\bconfig-get\\s+${escapeRegExp(key)}\\b`).test(content)) {
|
|
leaks.push(`${path.basename(relativePath)} → ${key} (owned by ${registry.configKeys[key]})`);
|
|
}
|
|
}
|
|
}
|
|
leaks.sort();
|
|
assert.deepEqual(
|
|
leaks, [],
|
|
`ADR-857 phase 6 is NOT complete: the host loop reads capability-owned config keys ` +
|
|
`inline:\n ${leaks.join('\n ')}\nThe owning capability must render/consume these (#1169).`,
|
|
);
|
|
});
|
|
|
|
test('host loop bodies are materially smaller than the pre-phase-6 baseline (#1168)', () => {
|
|
// #1139 AC: plan-phase.md / execute-phase.md must shrink as optional features
|
|
// extract to capabilities. Frozen pre-phase-6 sizes (LF bytes); the files must
|
|
// drop strictly below these. This also defeats double-run gaming — declaring a
|
|
// hook while leaving the inline block keeps the file from shrinking -> red.
|
|
//
|
|
// #1298: the execute-phase.md ceiling was raised from 93166 to accommodate
|
|
// wiring the mandatory `worktree record-agent` writer verb into the per-agent
|
|
// wave-manifest append. That verb is privileged host machinery (ADR-857
|
|
// Decision #1) — NOT the optional-feature inline logic this budget ratchets
|
|
// toward capabilities — so its footprint legitimately raises the host-loop
|
|
// ceiling rather than signalling an un-extracted optional feature.
|
|
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
|
const PRE_PHASE6 = { 'plan-phase.md': 94519, 'execute-phase.md': 93600 };
|
|
const notShrunk = [];
|
|
for (const [file, frozen] of Object.entries(PRE_PHASE6)) {
|
|
const now = lfByteCount(path.join(ROOT, 'gsd-core', 'workflows', file));
|
|
if (now >= frozen) notShrunk.push(`${file}: ${now} bytes (must be < pre-phase-6 ${frozen})`);
|
|
}
|
|
assert.deepEqual(
|
|
notShrunk, [],
|
|
`ADR-857 phase 6 is NOT complete: host loop bodies have not shrunk — the optional ` +
|
|
`feature logic has not actually been extracted:\n ${notShrunk.join('\n ')}`,
|
|
);
|
|
});
|
|
|
|
describe('ADR-857 phase 6 — capabilities must not bake install paths into the registry', () => {
|
|
// Matches GSD install paths that LEAK when copied verbatim to non-Claude runtimes.
|
|
// (~/.claude/projects is a legit runtime feature and is intentionally NOT matched.)
|
|
const LEAK = /\.claude[/\\](?:gsd-core|commands|agents|hooks)\b/;
|
|
|
|
test('no capability source (capability.json or fragment) embeds a ~/.claude install path', () => {
|
|
const capsDir = path.join(__dirname, '..', 'capabilities');
|
|
const offenders = [];
|
|
for (const id of fs.readdirSync(capsDir)) {
|
|
const dir = path.join(capsDir, id);
|
|
if (!fs.statSync(dir).isDirectory()) continue;
|
|
const cj = path.join(dir, 'capability.json');
|
|
if (fs.existsSync(cj) && LEAK.test(fs.readFileSync(cj, 'utf8'))) {
|
|
offenders.push(`capabilities/${id}/capability.json`);
|
|
}
|
|
const fragDir = path.join(dir, 'fragments');
|
|
if (fs.existsSync(fragDir)) {
|
|
for (const f of fs.readdirSync(fragDir)) {
|
|
if (LEAK.test(fs.readFileSync(path.join(fragDir, f), 'utf8'))) {
|
|
offenders.push(`capabilities/${id}/fragments/${f}`);
|
|
}
|
|
}
|
|
}
|
|
}
|
|
assert.deepEqual(offenders, [],
|
|
`capability sources embed ~/.claude install paths — these leak into the verbatim-copied capability-registry.cjs on non-Claude runtimes. Make the fragment path-free. Offenders: ${offenders.join(', ')}`);
|
|
});
|
|
|
|
test('generated capability-registry.cjs contains no ~/.claude install path', () => {
|
|
const reg = fs.readFileSync(path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'capability-registry.cjs'), 'utf8');
|
|
const leakLines = reg.split(/\r?\n/).map((l, i) => [i + 1, l]).filter(([, l]) => LEAK.test(l)).map(([n]) => n);
|
|
assert.deepEqual(leakLines, [],
|
|
`capability-registry.cjs leaks ~/.claude install paths at line(s) ${leakLines.join(', ')} — the registry is copied verbatim to non-Claude runtimes (only workflow .md files are path-converted at install). Make the source capability fragment path-free.`);
|
|
});
|
|
});
|
|
|
|
test('every plan:pre planner contribution is injected generically (not per-capId hardcode)', () => {
|
|
// FIX C regression guard: plan-phase.md must inject planner contributions
|
|
// generically (by into == "planner") rather than only injecting a single
|
|
// hardcoded capId (e.g. "tdd"). A generic injection ensures any active
|
|
// plan:pre contribution with into=="planner" reaches the planner — including
|
|
// tdd, schema-gate, and security contributions.
|
|
//
|
|
// Heuristic: the planner prompt section must reference injecting where
|
|
// into == "planner" (or iterate contributions), AND must NOT rely solely
|
|
// on a single capId == "tdd" injection as the only planner contribution
|
|
// delivery mechanism.
|
|
const planPhase = readRepoFile('gsd-core/workflows/plan-phase.md');
|
|
|
|
// The file must contain a generic reference to into == "planner" contribution injection.
|
|
assert.match(
|
|
planPhase,
|
|
/into\s*==\s*["']planner["']/,
|
|
'plan-phase.md must inject planner contributions generically via into == "planner" ' +
|
|
'(not just a single hardcoded capId). Fix C regression: all active planner contributions must reach the planner.',
|
|
);
|
|
|
|
// Verify the file does NOT rely SOLELY on a hardcoded capId == "tdd" injection
|
|
// for the planner contribution. If only a tdd-specific injection exists (old form),
|
|
// the schema-gate and security contributions are silently dropped.
|
|
// We check: every occurrence of 'capId == "tdd"' contribution injection must be
|
|
// accompanied somewhere by a generic into=="planner" dispatch (already verified above).
|
|
// Additionally, the old exact tdd-only injection prose must not be the only delivery.
|
|
const onlyTddInjection = /\bRead from `PLAN_PRE_HOOKS_JSON` where `kind == "contribution"` and `capId == "tdd"`\b/;
|
|
// If the old tdd-only prose still exists WITHOUT the generic into=="planner" prose,
|
|
// that's a regression. Since we already asserted into=="planner" exists, we just
|
|
// confirm the tdd-only prose is no longer the sole injection mechanism.
|
|
if (onlyTddInjection.test(planPhase)) {
|
|
// Old prose still present: acceptable only if generic prose is ALSO present (already asserted).
|
|
// Verify the into=="planner" injection appears NEAR the planner prompt (within 5000 chars of it).
|
|
const plannerPromptIdx = planPhase.indexOf('into == "planner"');
|
|
assert.ok(
|
|
plannerPromptIdx >= 0,
|
|
'plan-phase.md has tdd-only injection prose but no generic into=="planner" injection. ' +
|
|
'Remove the tdd-only injection and replace with generic contribution dispatch.',
|
|
);
|
|
}
|
|
});
|
|
|
|
test('every declared gate check.query returns a uniform boolean `block` field', () => {
|
|
// FIX A regression guard: every gate check command must return a top-level
|
|
// boolean `block` field so the host-loop dispatch can read a single consistent
|
|
// field regardless of which capability owns the gate.
|
|
//
|
|
// For each unique check.query declared in the registry's gate hooks, invoke
|
|
// the check command against a temp directory and assert the JSON output
|
|
// contains `block` as a boolean. Uses a minimal temp dir so the command
|
|
// returns quickly without real project state.
|
|
const os = require('node:os');
|
|
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gate-block-contract-'));
|
|
|
|
// Collect unique gate check.queries from the registry
|
|
const queries = new Set();
|
|
for (const cap of Object.values(registry.capabilities)) {
|
|
for (const gate of cap.gates || []) {
|
|
if (gate.check && gate.check.query) queries.add(gate.check.query);
|
|
}
|
|
}
|
|
assert.ok(queries.size > 0, 'Registry must declare at least one gate check.query');
|
|
|
|
const gsdTools = path.join(ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
|
const failures = [];
|
|
|
|
for (const query of [...queries].sort()) {
|
|
let rawOut = '';
|
|
try {
|
|
// Invoke with --raw (the real dispatch form used by the host loop).
|
|
// Most commands accept a phase number and return valid JSON even when
|
|
// no real project state exists.
|
|
rawOut = execFileSync(
|
|
process.execPath,
|
|
[gsdTools, 'check', query, '1', '--raw'],
|
|
{ cwd: tmpDir, encoding: 'utf-8', timeout: 10000 },
|
|
);
|
|
const parsed = JSON.parse(rawOut.trim());
|
|
if (typeof parsed.block !== 'boolean') {
|
|
failures.push(
|
|
`check ${query}: returned JSON without a boolean \`block\` field ` +
|
|
`(got: ${JSON.stringify(parsed.block)}, type: ${typeof parsed.block}). ` +
|
|
`Add \`block\` to the command's output per the uniform gate contract.`,
|
|
);
|
|
}
|
|
} catch (err) {
|
|
// If it threw because the command required a different arg shape, try with a path
|
|
try {
|
|
rawOut = execFileSync(
|
|
process.execPath,
|
|
[gsdTools, 'check', query, tmpDir, '--raw'],
|
|
{ cwd: tmpDir, encoding: 'utf-8', timeout: 10000 },
|
|
);
|
|
const parsed = JSON.parse(rawOut.trim());
|
|
if (typeof parsed.block !== 'boolean') {
|
|
failures.push(
|
|
`check ${query}: returned JSON without a boolean \`block\` field ` +
|
|
`(got: ${JSON.stringify(parsed.block)}, type: ${typeof parsed.block}).`,
|
|
);
|
|
}
|
|
} catch (err2) {
|
|
failures.push(
|
|
`check ${query}: command failed or returned non-JSON output. ` +
|
|
`Error: ${err2 instanceof Error ? err2.message : String(err2)}. ` +
|
|
`Stdout: ${rawOut.slice(0, 200)}`,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
// Clean up temp dir
|
|
cleanup(tmpDir);
|
|
|
|
assert.deepEqual(
|
|
failures, [],
|
|
`Gate check commands must all return a top-level boolean \`block\` field:\n ${failures.join('\n ')}`,
|
|
);
|
|
});
|
|
});
|