Files
msd-core/tests/phase6-capstone-conformance.test.cjs
Tom Boucher 1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00

380 lines
18 KiB
JavaScript

// allow-test-rule: source-text-is-the-product
'use strict';
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const { cleanup } = require('./helpers.cjs');
const ROOT = path.join(__dirname, '..');
const { HOST_LOOP_FILES, scanWiredPoints } = require('../scripts/gen-loop-host-contract.cjs');
const { lfByteCount } = require('../scripts/workflow-size.cjs');
const CORE_SUBSTRATE_TERMS = [
'Verification substrate',
'verifier↔predicate contract',
'Probe Core Module',
'Edge Probe Module',
];
const registry = require('../gsd-core/bin/lib/capability-registry.cjs');
const { isCentralConfigKey } = require('../gsd-core/bin/lib/config-schema.cjs');
function readRepoFile(relativePath) {
return fs.readFileSync(path.join(ROOT, relativePath), 'utf8');
}
function escapeRegExp(value) {
return value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}
function activeWhenKeys() {
const keys = new Set();
for (const cap of Object.values(registry.capabilities)) {
for (const group of ['steps', 'gates', 'contributions']) {
for (const hook of cap[group] || []) {
if (hook.when) keys.add(hook.when);
}
}
}
return [...keys].sort();
}
describe('ADR-857 Phase 6 capstone conformance (#1139)', () => {
test('first-party optional feature capabilities are declared in the generated registry', () => {
const expectedFeatureCapabilities = [
'ai-integration',
'audit',
'code-review',
'graphify',
'intel',
'nyquist',
'pattern-mapper',
'research',
'security',
'ui',
];
for (const capId of expectedFeatureCapabilities) {
assert.equal(registry.capabilities[capId]?.role, 'feature', `${capId} must be a feature Capability`);
}
});
test('core verification substrate is documented as deliberately not capability-owned', () => {
const context = readRepoFile('CONTEXT.md');
for (const term of CORE_SUBSTRATE_TERMS) {
assert.match(context, new RegExp(escapeRegExp(term)), `${term} must be documented in CONTEXT.md`);
}
});
test('host loop files do not read capability hook activation keys directly', () => {
const forbiddenKeys = activeWhenKeys();
assert.ok(forbiddenKeys.length > 0, 'registry must expose hook activation keys');
for (const relativePath of HOST_LOOP_FILES) {
const content = readRepoFile(relativePath);
for (const key of forbiddenKeys) {
assert.doesNotMatch(
content,
new RegExp(`\\bconfig-get\\s+${escapeRegExp(key)}\\b`),
`${relativePath} must resolve ${key} through Capability hooks/state, not direct config-get`,
);
}
}
});
test('capability-owned config keys are not reintroduced into the central schema', () => {
for (const key of Object.keys(registry.configKeys).sort()) {
assert.equal(
isCentralConfigKey(key),
false,
`${key} is owned by capability ${registry.configKeys[key]} and must stay out of central config schema`,
);
}
});
test('host loop workflow files have a measurable, non-empty byte size', () => {
// Was asserted against the committed tests/workflow-size-baseline.json snapshot;
// #2724 (ADR-2719 Phase 4) deletes that file — the differential attribution
// check's size ratchet (tests/emitted-attribution.test.cjs) is the replacement
// anti-creep mechanism, but this test's actual intent was narrower: prove these
// host-loop files are real, tracked, non-empty workflow docs. Asserting the
// live byte count via the same shared counter the size guards use preserves
// that intent without depending on a committed snapshot.
for (const relativePath of HOST_LOOP_FILES) {
const fileName = path.basename(relativePath);
const bytes = lfByteCount(path.join(ROOT, relativePath));
assert.ok(bytes > 0, `${fileName} must be a non-empty workflow file`);
}
});
// ─── Phase-6 conformance: RED BY DESIGN until phase 6 is actually complete ──────
//
// #1139 closed (via #1158) with a green "capstone conformance gate" while the
// ADR-857 phase-6 acceptance criteria were unmet — a false green. The three
// tests below assert the real criteria with NO paper-over allowlist, so the
// gate stays RED until the work lands. Green here must mean "phase 6 conformant,"
// not "no new regression." Fixes tracked in #1167 / #1168 / #1169.
test('every declared capability hook point has a render-hooks call site in the host loop (#1168)', () => {
// No allowlist: every point a capability declares a hook at MUST have a
// `render-hooks` call site in the host loop, or those hooks can never fire.
const declaredPoints = new Set();
for (const cap of Object.values(registry.capabilities)) {
for (const group of ['steps', 'gates', 'contributions']) {
for (const hook of cap[group] || []) {
if (hook.point) declaredPoints.add(hook.point);
}
}
}
// Scan only the host loop files (a `render-hooks` mention in a non-host
// workflow must not mask a lost host call site).
const callSites = new Set();
for (const relativePath of HOST_LOOP_FILES) {
const content = readRepoFile(relativePath);
for (const pt of scanWiredPoints(content)) callSites.add(pt);
}
const orphaned = [...declaredPoints].sort().filter((p) => !callSites.has(p));
assert.deepEqual(
orphaned, [],
`ADR-857 phase 6 is NOT complete: capability hooks declare these extension points ` +
`but no host-loop workflow calls \`gsd_run loop render-hooks <point>\`, so the hooks ` +
`can never fire: ${orphaned.join(', ')}. Wire each call site (#1167/#1169).`,
);
});
test('all ADR-857-named optional features are real Capabilities, not empty stubs (#1169)', () => {
// ADR-857 §53 + Decision 7 enumerate these optional, non-loop modules as
// Capabilities. "Migrated" means the feature OWNS its behavior: hook-based
// features (tdd/schema-gate/drift/gap-analysis) must declare >=1 hook;
// command-family features (profile-pipeline) must declare a command family.
// A registration-only stub (role:feature but no hooks/commands) games this
// gate while the logic stays welded into the loop — rejected here.
const REQUIRED = ['tdd', 'schema-gate', 'drift', 'gap-analysis', 'profile-pipeline'];
const problems = [];
for (const id of REQUIRED) {
const cap = registry.capabilities[id];
if (!cap) { problems.push(`${id}: not registered`); continue; }
if (cap.role !== 'feature') { problems.push(`${id}: role="${cap.role}", must be "feature"`); continue; }
const hookCount = (cap.steps?.length || 0) + (cap.contributions?.length || 0) + (cap.gates?.length || 0);
const isCommandFamily = (cap.commands?.length || 0) > 0;
if (hookCount === 0 && !isCommandFamily) {
problems.push(`${id}: EMPTY STUB (no hooks, no command family) — inline logic was not migrated; declare the real hooks/commands and remove the inline branch`);
}
}
assert.deepEqual(
problems, [],
`ADR-857 phase 6 is NOT complete:\n ${problems.join('\n ')}\n` +
`Each feature must OWN its behavior via hooks or a command family — not exist as a registration-only stub (#1169).`,
);
});
test('host loop reads no capability-owned config key inline (#1169)', () => {
// Phase 6 requires the loop to resolve capability behavior via render-hooks,
// not by reading capability-owned keys directly. Any inline `config-get` of a
// registry-owned key is an incomplete migration (the loop still owns the
// feature's params).
const leaks = [];
for (const relativePath of HOST_LOOP_FILES) {
const content = readRepoFile(relativePath);
for (const key of Object.keys(registry.configKeys)) {
if (new RegExp(`\\bconfig-get\\s+${escapeRegExp(key)}\\b`).test(content)) {
leaks.push(`${path.basename(relativePath)} → ${key} (owned by ${registry.configKeys[key]})`);
}
}
}
leaks.sort();
assert.deepEqual(
leaks, [],
`ADR-857 phase 6 is NOT complete: the host loop reads capability-owned config keys ` +
`inline:\n ${leaks.join('\n ')}\nThe owning capability must render/consume these (#1169).`,
);
});
test('host loop bodies are materially smaller than the pre-phase-6 baseline (#1168)', () => {
// #1139 AC: plan-phase.md / execute-phase.md must shrink as optional features
// extract to capabilities. Frozen pre-phase-6 sizes (LF bytes); the files must
// drop strictly below these. This also defeats double-run gaming — declaring a
// hook while leaving the inline block keeps the file from shrinking -> red.
//
// #1298: the execute-phase.md ceiling was raised from 93166 to accommodate
// wiring the mandatory `worktree record-agent` writer verb into the per-agent
// wave-manifest append. That verb is privileged host machinery (ADR-857
// Decision #1) — NOT the optional-feature inline logic this budget ratchets
// toward capabilities — so its footprint legitimately raises the host-loop
// ceiling rather than signalling an un-extracted optional feature.
const { lfByteCount } = require('../scripts/workflow-size.cjs');
const PRE_PHASE6 = { 'plan-phase.md': 94519, 'execute-phase.md': 93600 };
const notShrunk = [];
for (const [file, frozen] of Object.entries(PRE_PHASE6)) {
const now = lfByteCount(path.join(ROOT, 'gsd-core', 'workflows', file));
if (now >= frozen) notShrunk.push(`${file}: ${now} bytes (must be < pre-phase-6 ${frozen})`);
}
assert.deepEqual(
notShrunk, [],
`ADR-857 phase 6 is NOT complete: host loop bodies have not shrunk — the optional ` +
`feature logic has not actually been extracted:\n ${notShrunk.join('\n ')}`,
);
});
describe('ADR-857 phase 6 — capabilities must not bake install paths into the registry', () => {
// Matches GSD install paths that LEAK when copied verbatim to non-Claude runtimes.
// (~/.claude/projects is a legit runtime feature and is intentionally NOT matched.)
const LEAK = /\.claude[/\\](?:gsd-core|commands|agents|hooks)\b/;
test('no capability source (capability.json or fragment) embeds a ~/.claude install path', () => {
const capsDir = path.join(__dirname, '..', 'capabilities');
const offenders = [];
for (const id of fs.readdirSync(capsDir)) {
const dir = path.join(capsDir, id);
if (!fs.statSync(dir).isDirectory()) continue;
const cj = path.join(dir, 'capability.json');
if (fs.existsSync(cj) && LEAK.test(fs.readFileSync(cj, 'utf8'))) {
offenders.push(`capabilities/${id}/capability.json`);
}
const fragDir = path.join(dir, 'fragments');
if (fs.existsSync(fragDir)) {
for (const f of fs.readdirSync(fragDir)) {
if (LEAK.test(fs.readFileSync(path.join(fragDir, f), 'utf8'))) {
offenders.push(`capabilities/${id}/fragments/${f}`);
}
}
}
}
assert.deepEqual(offenders, [],
`capability sources embed ~/.claude install paths — these leak into the verbatim-copied capability-registry.cjs on non-Claude runtimes. Make the fragment path-free. Offenders: ${offenders.join(', ')}`);
});
test('generated capability-registry.cjs contains no ~/.claude install path', () => {
const reg = fs.readFileSync(path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'capability-registry.cjs'), 'utf8');
const leakLines = reg.split(/\r?\n/).map((l, i) => [i + 1, l]).filter(([, l]) => LEAK.test(l)).map(([n]) => n);
assert.deepEqual(leakLines, [],
`capability-registry.cjs leaks ~/.claude install paths at line(s) ${leakLines.join(', ')} — the registry is copied verbatim to non-Claude runtimes (only workflow .md files are path-converted at install). Make the source capability fragment path-free.`);
});
});
test('every plan:pre planner contribution is injected generically (not per-capId hardcode)', () => {
// FIX C regression guard: plan-phase.md must inject planner contributions
// generically (by into == "planner") rather than only injecting a single
// hardcoded capId (e.g. "tdd"). A generic injection ensures any active
// plan:pre contribution with into=="planner" reaches the planner — including
// tdd, schema-gate, and security contributions.
//
// Heuristic: the planner prompt section must reference injecting where
// into == "planner" (or iterate contributions), AND must NOT rely solely
// on a single capId == "tdd" injection as the only planner contribution
// delivery mechanism.
const planPhase = readRepoFile('gsd-core/workflows/plan-phase.md');
// The file must contain a generic reference to into == "planner" contribution injection.
assert.match(
planPhase,
/into\s*==\s*["']planner["']/,
'plan-phase.md must inject planner contributions generically via into == "planner" ' +
'(not just a single hardcoded capId). Fix C regression: all active planner contributions must reach the planner.',
);
// Verify the file does NOT rely SOLELY on a hardcoded capId == "tdd" injection
// for the planner contribution. If only a tdd-specific injection exists (old form),
// the schema-gate and security contributions are silently dropped.
// We check: every occurrence of 'capId == "tdd"' contribution injection must be
// accompanied somewhere by a generic into=="planner" dispatch (already verified above).
// Additionally, the old exact tdd-only injection prose must not be the only delivery.
const onlyTddInjection = /\bRead from `PLAN_PRE_HOOKS_JSON` where `kind == "contribution"` and `capId == "tdd"`\b/;
// If the old tdd-only prose still exists WITHOUT the generic into=="planner" prose,
// that's a regression. Since we already asserted into=="planner" exists, we just
// confirm the tdd-only prose is no longer the sole injection mechanism.
if (onlyTddInjection.test(planPhase)) {
// Old prose still present: acceptable only if generic prose is ALSO present (already asserted).
// Verify the into=="planner" injection appears NEAR the planner prompt (within 5000 chars of it).
const plannerPromptIdx = planPhase.indexOf('into == "planner"');
assert.ok(
plannerPromptIdx >= 0,
'plan-phase.md has tdd-only injection prose but no generic into=="planner" injection. ' +
'Remove the tdd-only injection and replace with generic contribution dispatch.',
);
}
});
test('every declared gate check.query returns a uniform boolean `block` field', () => {
// FIX A regression guard: every gate check command must return a top-level
// boolean `block` field so the host-loop dispatch can read a single consistent
// field regardless of which capability owns the gate.
//
// For each unique check.query declared in the registry's gate hooks, invoke
// the check command against a temp directory and assert the JSON output
// contains `block` as a boolean. Uses a minimal temp dir so the command
// returns quickly without real project state.
const os = require('node:os');
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gate-block-contract-'));
// Collect unique gate check.queries from the registry
const queries = new Set();
for (const cap of Object.values(registry.capabilities)) {
for (const gate of cap.gates || []) {
if (gate.check && gate.check.query) queries.add(gate.check.query);
}
}
assert.ok(queries.size > 0, 'Registry must declare at least one gate check.query');
const gsdTools = path.join(ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
const failures = [];
for (const query of [...queries].sort()) {
let rawOut = '';
try {
// Invoke with --raw (the real dispatch form used by the host loop).
// Most commands accept a phase number and return valid JSON even when
// no real project state exists.
rawOut = execFileSync(
process.execPath,
[gsdTools, 'check', query, '1', '--raw'],
{ cwd: tmpDir, encoding: 'utf-8', timeout: 10000 },
);
const parsed = JSON.parse(rawOut.trim());
if (typeof parsed.block !== 'boolean') {
failures.push(
`check ${query}: returned JSON without a boolean \`block\` field ` +
`(got: ${JSON.stringify(parsed.block)}, type: ${typeof parsed.block}). ` +
`Add \`block\` to the command's output per the uniform gate contract.`,
);
}
} catch (err) {
// If it threw because the command required a different arg shape, try with a path
try {
rawOut = execFileSync(
process.execPath,
[gsdTools, 'check', query, tmpDir, '--raw'],
{ cwd: tmpDir, encoding: 'utf-8', timeout: 10000 },
);
const parsed = JSON.parse(rawOut.trim());
if (typeof parsed.block !== 'boolean') {
failures.push(
`check ${query}: returned JSON without a boolean \`block\` field ` +
`(got: ${JSON.stringify(parsed.block)}, type: ${typeof parsed.block}).`,
);
}
} catch (err2) {
failures.push(
`check ${query}: command failed or returned non-JSON output. ` +
`Error: ${err2 instanceof Error ? err2.message : String(err2)}. ` +
`Stdout: ${rawOut.slice(0, 200)}`,
);
}
}
}
// Clean up temp dir
cleanup(tmpDir);
assert.deepEqual(
failures, [],
`Gate check commands must all return a top-level boolean \`block\` field:\n ${failures.join('\n ')}`,
);
});
});