* test(#3333): fold the runtime & install surface fix-* cluster — Wave 1 Folds 11 legacy tests/fix-*.test.cjs regression files into their module's main test suite: 6 folded into existing suites (host-integration-descriptors, effort-surface-axis, trae-imperative-reference, hermes-skills-migration, gsd-agent-isolation-guard), 5 renamed to become the module's sole suite (cursor-hook-workspace-roots, cursor-subagent-isolation, lint-compiled-artifact-sync, hooks-commonjs-marker, shared-hooks-dir-resolution). All 195 test() blocks preserved with zero drops; lint-test-file-count.cjs and eslint remain clean. No production code changed. Wave 1 of 7 in #3315 (H3 of epic #3053). * test(#3333): replace try/finally with t.after() in isolation-guard tests CONTRIBUTING.md bans try/finally inside test bodies (masks failures, not an approved pattern). The fold in the prior commit carried 27 instances forward verbatim from the deleted fix-3045-dispatch-isolation-resolver.test.cjs into an otherwise-clean file. Converts each to the approved per-test t.after() cleanup pattern — same cleanup call, registered instead of finally-wrapped. No assertion, fixture, or test-name change; test( count unchanged at 50. Found by the Standards review pass on Wave 1 (#3333, H3 of epic #3053). * fix(#3333): restore raw NUL byte mangled by the fold in hermes-skills-migration.test.cjs The prior fold commit copied fix-2284-hermes-agent-delegate-task-projection's "collision-robust" test via a text-based Read/Write pipeline, which silently turned a raw NUL byte (0x00) embedded in two string literals into a regular space character. That corrupted the test's actual purpose (proving a NUL byte survives a string-rewrite operation untouched) and produced a genuine gsd-test failure: `24 !== 1` for `out.split(' ').length`, because splitting on a space finds every space in the sentence instead of the single NUL byte the test meant to isolate. Root-caused by diffing the raw bytes (via `git cat-file blob` + `cat -v`) between the pre-fold source and the folded target — confirmed exactly two bytes differ. Restored via a byte-precise patch (latin1 round-trip) touching only those two lines; test( count and every other byte unchanged. * fix(#3333): use \x00 escape sequence instead of a raw NUL byte in test fixture The prior commit restored a byte-exact raw NUL byte matching the original fix-2284 source, and the production function (applyClaudeCodeBrandSwap) was confirmed correct in a standalone repro. But the same raw byte still failed through gsd-test's remote pipeline. Root cause is upstream of gsd-core: some step in that transfer path does not carry a raw 0x00 byte through untouched. A raw embedded NUL byte was never necessary here — `\x00` as a 4-character escape sequence in the source text produces the identical runtime character (U+0000) without ever putting a raw byte in the tracked file, sidestepping any byte-oriented transfer step. Applied at both call sites (the fixture string and the split() delimiter). No behavior change; test( count unchanged at 76. * fix(#3333): harden copyWithPathReplacement against a source file vanishing mid-copy (TOCTOU) Surfaced by this PR's own gsd-test run: tests/install-minimal-hooks.test.cjs and tests/opencode-command-dir-plural.test.cjs intermittently crashed with ENOENT reading gsd-core/workflows/zzz-e5-drift-fixture.md. Root cause is unrelated to test-file consolidation — tests/planning-prompt-drift.test.cjs writes that fixture directly into the real, shared gsd-core/workflows/ tree (main() hardcodes its scan root to the real repo) and deletes it in t.after(); copyWithPathReplacement's readdirSync-then-read loop has no protection against the listed file vanishing before it gets there, so a concurrently-running install path can crash entirely on what is otherwise a completely benign race. Fixed by skipping (not crashing on) a listed entry that no longer exists by the time the loop reaches it. Added a regression test that deterministically reproduces the race (readdirSync snapshot still lists the file; it is deleted immediately after) and proves both outcomes: no throw, and the vanished entry's destination is never partially written. Per CLAUDE.md's no-defer rule, a defect surfaced while verifying this PR is fixed inline rather than deferred — this overrides one-concern-per-PR. * fix(#3333): fix third NUL-byte-mangled occurrence missed by prior fix passes The fold originally mangled three raw-NUL-byte occurrences to spaces, not two — the earlier byte-restore and escape-sequence commits both only targeted the fixture string and the split() delimiter, missing out.includes('[ ]') a few lines below (should read out.includes('[\x00]')). A remote gsd-test run kept failing on this exact assertion even after both prior fixes, which is what surfaced the miss. Verified via a standalone repro using the file's real (not retyped) fixture content: all six assertions in the collision-robust test now pass. Zero raw NUL bytes remain in the file; test( count unchanged at 76. * chore(#3333): add changeset for the copyWithPathReplacement TOCTOU fix Fixed-type fragment for the production defect fixed inline in this PR (bin/install.js's copyWithPathReplacement). Exempt from docs/ requirements per CONTRIBUTING.md (only Added/Changed/Deprecated/Removed require it). * chore(#3333): backfill changeset PR number (pr:0 -> pr:3341) --------- Co-authored-by: sim <sim@local>
503 lines
22 KiB
JavaScript
503 lines
22 KiB
JavaScript
// allow-test-rule: source-text-is-the-product (see #2481, #2615)
|
|
// The final describe block asserts on gsd-core/workflows/review.md's text. A
|
|
// workflow .md IS what the runtime loads — its literal command lines are the
|
|
// deployed contract, and there is no runtime seam that executes review.md here.
|
|
// The #2615 matrix-parity block below is the same kind of contract assertion:
|
|
// docs/reference/host-integration-capability-matrix.md IS the cited source of
|
|
// truth for every descriptor axis (ADR-1239), so asserting a shipped axis
|
|
// value appears there and matches is a contract assertion, not a source grep.
|
|
// Every other block in this file is behavioral (CLI + module surface).
|
|
|
|
/**
|
|
* #2481 — ADR-1239 `effortSurface` axis + ADR-443 path (a).
|
|
*
|
|
* Before this change effort reached a runtime only through install-time channels
|
|
* (EFFORT_RENDERING's `frontmatter`/`api`), so a reviewer CLI spawned as a
|
|
* subprocess silently inherited whatever effort sat in the user's own global CLI
|
|
* config. These tests pin the invocation-time channel: the negotiated axis that
|
|
* decides WHETHER effort is deliverable, the renderer that knows the syntax, and
|
|
* the live orchestration path that carries it.
|
|
*/
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const fc = require('fast-check');
|
|
|
|
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
|
|
|
const REPO_ROOT = path.resolve(__dirname, '..');
|
|
const {
|
|
renderEffortArgv,
|
|
EFFORT_ARGV,
|
|
} = require(path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib', 'model-catalog.cjs'));
|
|
const {
|
|
HOST_INTEGRATION_AXES,
|
|
negotiateHostCapabilities,
|
|
degradationFor,
|
|
} = require(path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib', 'host-integration.cjs'));
|
|
const {
|
|
_HOST_INTEGRATION_VOCAB,
|
|
validateRuntimeBody,
|
|
} = require(path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib', 'capability-validator.cjs'));
|
|
const registry = require(path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs'));
|
|
|
|
// #2615: the host-integration capability matrix, normalized so CRLF checkouts
|
|
// (Windows autocrlf) don't break the row regexes below.
|
|
const MATRIX = path.join(REPO_ROOT, 'docs', 'reference', 'host-integration-capability-matrix.md');
|
|
const MATRIX_TEXT = fs.readFileSync(MATRIX, 'utf-8').replace(/\r\n/g, '\n');
|
|
|
|
/** Extract a `## <host>` section body, stopping at the next top-level host heading. */
|
|
function matrixSection(host) {
|
|
const start = MATRIX_TEXT.indexOf(`\n## ${host}\n`);
|
|
if (start === -1) return null;
|
|
const rest = MATRIX_TEXT.slice(start + 1);
|
|
const end = rest.indexOf('\n## ');
|
|
return end === -1 ? rest : rest.slice(0, end);
|
|
}
|
|
|
|
/** Read the value cell of a `| <axis> | <value> | …` row. */
|
|
function matrixAxisValue(body, axis) {
|
|
const row = body.split(/\r?\n/).find((l) => l.startsWith(`| ${axis} |`));
|
|
return row ? row.split('|')[2].trim() : null;
|
|
}
|
|
|
|
const MATRIX_RUNTIMES = Object.keys(registry.runtimes).filter(
|
|
(id) => registry.runtimes[id]?.runtime?.hostIntegration,
|
|
);
|
|
|
|
/**
|
|
* A real shipped descriptor with one hostIntegration axis stripped.
|
|
*
|
|
* Deriving the fixture from a descriptor this gate did not author satisfies the
|
|
* fixture-provenance rule (#2371) — a hand-built body would only ever encode the
|
|
* author's mental model of a valid descriptor, which is how the required-axis
|
|
* defect reached the runner in the first place.
|
|
*/
|
|
function shippedDescriptorWithout(axis) {
|
|
const cap = JSON.parse(
|
|
fs.readFileSync(path.join(REPO_ROOT, 'capabilities', 'vscode', 'capability.json'), 'utf8'),
|
|
);
|
|
delete cap.runtime.hostIntegration[axis];
|
|
return cap;
|
|
}
|
|
|
|
/** Write a project whose effort cascade resolves to a known universal value. */
|
|
function projectWithEffort(effort) {
|
|
const dir = createTempProject();
|
|
fs.writeFileSync(
|
|
path.join(dir, '.planning', 'config.json'),
|
|
JSON.stringify({ effort: { default: effort } }, null, 2),
|
|
);
|
|
return dir;
|
|
}
|
|
|
|
describe('#2481 effortSurface — closed vocabulary', () => {
|
|
test('is exactly argv|none — no config-file member', () => {
|
|
// Gemini CLI was the only host with a config-file effort surface and was
|
|
// removed as a sunset runtime (8f2ebbe9b / #1928 / PR #1996). A member no
|
|
// supported host can claim would invite guessed descriptor values.
|
|
assert.deepEqual([...HOST_INTEGRATION_AXES.effortSurface], ['argv', 'none']);
|
|
});
|
|
|
|
test('engine vocabulary and validator mirror agree (parity guard)', () => {
|
|
assert.deepEqual(
|
|
[...HOST_INTEGRATION_AXES.effortSurface],
|
|
[..._HOST_INTEGRATION_VOCAB.effortSurface],
|
|
);
|
|
});
|
|
|
|
test('undocumented is NOT a vocabulary member — it is the corpus sentinel', () => {
|
|
assert.ok(!HOST_INTEGRATION_AXES.effortSurface.includes('undocumented'));
|
|
});
|
|
});
|
|
|
|
describe('#2481 effortSurface — negotiation fails closed', () => {
|
|
const cases = [
|
|
['argv declared', 'argv', 'argv'],
|
|
['undocumented sentinel', 'undocumented', 'none'],
|
|
['retired value (config-file)', 'config-file', 'none'],
|
|
['unknown/future value', 'quantum-telepathy', 'none'],
|
|
['empty string', '', 'none'],
|
|
['none declared', 'none', 'none'],
|
|
];
|
|
for (const [label, declared, expected] of cases) {
|
|
test(`${label} -> ${expected}`, () => {
|
|
const r = negotiateHostCapabilities({ protocolVersion: 1, modelMode: 'active', effortSurface: declared });
|
|
assert.equal(r.effective.effortSurface, expected);
|
|
});
|
|
}
|
|
|
|
test('axis omitted entirely -> safe floor, and the omission is warned', () => {
|
|
const r = negotiateHostCapabilities({ protocolVersion: 1, modelMode: 'active' });
|
|
assert.equal(r.effective.effortSurface, 'none');
|
|
assert.ok(r.warnings.some((w) => String(w).includes('effortSurface')));
|
|
});
|
|
|
|
test('a descriptor that omits the axis entirely still validates clean', () => {
|
|
// The axis was added after descriptors existed. Requiring it would invalidate
|
|
// every pre-existing descriptor — including third-party ones — and break the
|
|
// "purely additive" property ADR-1239 promises for external descriptors.
|
|
// Regression guard: 48 suites failed across both node lanes when it was required.
|
|
const cap = shippedDescriptorWithout('effortSurface');
|
|
const errors = validateRuntimeBody(cap);
|
|
assert.deepEqual(
|
|
errors, [],
|
|
`a descriptor without effortSurface must validate clean, got: ${JSON.stringify(errors)}`,
|
|
);
|
|
});
|
|
|
|
test('optional does not mean unvalidated — a present bad value is still rejected', () => {
|
|
const cap = shippedDescriptorWithout('effortSurface');
|
|
cap.runtime.hostIntegration.effortSurface = 'config-file'; // retired value
|
|
const errors = validateRuntimeBody(cap);
|
|
assert.ok(
|
|
errors.some((e) => String(e).includes('effortSurface')),
|
|
`a present invalid value must error, got: ${JSON.stringify(errors)}`,
|
|
);
|
|
});
|
|
|
|
test('an undeclared axis is never invented from a profile baseline', () => {
|
|
// Regression guard for the failure this axis was designed against: a
|
|
// programmatic-cli host must not inherit `argv` merely by being programmatic.
|
|
const r = negotiateHostCapabilities({
|
|
protocolVersion: 1,
|
|
embeddingMode: 'imperative',
|
|
commandSurface: 'slash-file',
|
|
modelMode: 'active',
|
|
});
|
|
assert.equal(r.effective.effortSurface, 'none');
|
|
});
|
|
});
|
|
|
|
describe('#2481 effortSurface — is its own axis, not folded into the model point', () => {
|
|
test('the model interface point still grades on modelMode alone', () => {
|
|
// Deliberate: modelMode has graded interface point 3 since Phase A. Widening
|
|
// it to also mean "delivers effort" would silently redefine that contract for
|
|
// every existing consumer. Effort is read from effective.effortSurface.
|
|
assert.equal(degradationFor('model', { modelMode: 'active' }).level, 'full');
|
|
assert.equal(degradationFor('model', { modelMode: 'passive' }).level, 'degraded');
|
|
});
|
|
|
|
test('declaring an effort surface does not change the model point', () => {
|
|
for (const es of ['argv', 'none', 'undocumented', undefined]) {
|
|
assert.equal(degradationFor('model', { modelMode: 'active', effortSurface: es }).level, 'full');
|
|
}
|
|
});
|
|
|
|
test('degradationFor never throws on a malformed axes object', () => {
|
|
for (const axes of [{}, { modelMode: null }, { effortSurface: 42 }, { modelMode: 'active', effortSurface: [] }]) {
|
|
assert.ok(['full', 'degraded', 'absent'].includes(degradationFor('model', axes).level));
|
|
}
|
|
});
|
|
});
|
|
|
|
describe('#2481 renderEffortArgv — per-host syntax and clamping', () => {
|
|
test('claude renders --effort', () => {
|
|
assert.deepEqual(renderEffortArgv('claude', 'xhigh', 'argv').argv, ['--effort', 'xhigh']);
|
|
});
|
|
|
|
test('opencode renders --variant', () => {
|
|
assert.deepEqual(renderEffortArgv('opencode', 'high', 'argv').argv, ['--variant', 'high']);
|
|
});
|
|
|
|
test('codex renders the generic -c config override, not a dedicated flag', () => {
|
|
// codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI flag
|
|
// (config.toml key only), so -c key=value is the only argv route.
|
|
assert.deepEqual(
|
|
renderEffortArgv('codex', 'high', 'argv').argv,
|
|
['-c', 'model_reasoning_effort=high'],
|
|
);
|
|
});
|
|
|
|
test('clamps the provider-unique tail levels', () => {
|
|
// claude has no `minimal`; codex has no `max`.
|
|
assert.deepEqual(renderEffortArgv('claude', 'minimal', 'argv').argv, ['--effort', 'low']);
|
|
assert.deepEqual(renderEffortArgv('codex', 'max', 'argv').argv, ['-c', 'model_reasoning_effort=xhigh']);
|
|
});
|
|
|
|
test('emits nothing when the surface is not argv', () => {
|
|
for (const surface of ['none', 'undocumented', 'config-file', '', null, undefined]) {
|
|
assert.deepEqual(renderEffortArgv('claude', 'xhigh', surface).argv, []);
|
|
}
|
|
});
|
|
|
|
test('emits nothing for a host with no known syntax', () => {
|
|
for (const host of ['gemini', 'cursor', 'zcode', '']) {
|
|
assert.deepEqual(renderEffortArgv(host, 'high', 'argv').argv, []);
|
|
}
|
|
});
|
|
|
|
test('inherited Object members are not mistaken for host specs', () => {
|
|
// Regression guard: a bare EFFORT_ARGV[host] lookup resolves these to
|
|
// inherited members — truthy, but with no clamp/render — so a hostile host
|
|
// id from an untrusted descriptor threw instead of degrading.
|
|
for (const host of ['__proto__', 'constructor', 'toString', 'hasOwnProperty', 'valueOf']) {
|
|
assert.deepEqual(
|
|
renderEffortArgv(host, 'high', 'argv').argv, [],
|
|
`${host} must degrade to no argument, not throw`,
|
|
);
|
|
}
|
|
});
|
|
|
|
test('emits nothing for a missing or unrecognised effort level', () => {
|
|
for (const level of ['', 'bogus', 'HIGH', ' high', null, undefined, 42]) {
|
|
assert.deepEqual(renderEffortArgv('claude', level, 'argv').argv, []);
|
|
}
|
|
});
|
|
|
|
test('property: a rendered level is always inside that host\'s supported set', () => {
|
|
const hosts = Object.keys(EFFORT_ARGV);
|
|
const levels = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max'];
|
|
fc.assert(
|
|
fc.property(fc.constantFrom(...hosts), fc.constantFrom(...levels), (host, level) => {
|
|
const r = renderEffortArgv(host, level, 'argv');
|
|
if (r.argv.length === 0) return true;
|
|
return EFFORT_ARGV[host].supported.has(r.value);
|
|
}),
|
|
{ numRuns: 200, seed: 2481 },
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('#2481 live path — resolve-execution carries invocation-time effort', () => {
|
|
test('--host renders the argument for an argv host', (t) => {
|
|
const dir = projectWithEffort('xhigh');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const r = runGsdTools('query resolve-execution gsd-planner --host claude', dir);
|
|
assert.ok(r.success, `resolve-execution failed: ${r.error}`);
|
|
const out = JSON.parse(r.output);
|
|
assert.equal(out.effort, 'xhigh');
|
|
assert.equal(out.effort_surface, 'argv');
|
|
assert.deepEqual(out.effort_argv, ['--effort', 'xhigh']);
|
|
assert.equal(out.effort_argv_string, '--effort xhigh');
|
|
});
|
|
|
|
test('--host on a host without a documented surface renders nothing', (t) => {
|
|
const dir = projectWithEffort('xhigh');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const out = JSON.parse(runGsdTools('query resolve-execution gsd-planner --host cursor', dir).output);
|
|
assert.equal(out.effort_surface, 'none');
|
|
assert.deepEqual(out.effort_argv, []);
|
|
assert.equal(out.effort_argv_string, '');
|
|
});
|
|
|
|
test('an unknown host degrades closed rather than erroring', (t) => {
|
|
const dir = projectWithEffort('high');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const r = runGsdTools('query resolve-execution gsd-planner --host not-a-real-host', dir);
|
|
assert.ok(r.success, 'an unknown host must degrade, not fail');
|
|
const out = JSON.parse(r.output);
|
|
assert.equal(out.effort_surface, 'none');
|
|
assert.deepEqual(out.effort_argv, []);
|
|
});
|
|
|
|
test('omitting --host leaves the JSON contract untouched', (t) => {
|
|
const dir = projectWithEffort('high');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const out = JSON.parse(runGsdTools('query resolve-execution gsd-planner', dir).output);
|
|
for (const k of ['host', 'effort_surface', 'effort_argv', 'effort_argv_string', 'effort_argv_value']) {
|
|
assert.ok(!(k in out), `--host absent must not add "${k}" to the contract`);
|
|
}
|
|
});
|
|
|
|
test('--host requires a value', (t) => {
|
|
const dir = projectWithEffort('high');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const r = runGsdTools('query resolve-execution gsd-planner --host', dir);
|
|
assert.ok(!r.success, 'a valueless --host must be a usage error');
|
|
});
|
|
|
|
test('a shell-metacharacter host is not interpolated, just unmatched', (t) => {
|
|
const dir = projectWithEffort('high');
|
|
t.after(() => cleanup(dir));
|
|
|
|
const r = runGsdTools('query resolve-execution gsd-planner --host "claude; touch pwned"', dir);
|
|
assert.ok(r.success);
|
|
assert.deepEqual(JSON.parse(r.output).effort_argv, []);
|
|
assert.ok(!fs.existsSync(path.join(dir, 'pwned')), 'no shell interpolation of the host value');
|
|
});
|
|
});
|
|
|
|
describe('#2481 — the escalation surface renders argv (CLI-level, not a workflow claim)', () => {
|
|
// NAMING IS DELIBERATE. This exercises `resolve-execution --attempt` directly,
|
|
// which is the CLI surface ADR-443's blocker explicitly EXCLUDES when it asks
|
|
// for "a real caller outside src/commands.cts's CLI surface and tests". It
|
|
// proves the escalation ladder still renders a host argument; it does NOT
|
|
// prove any workflow escalates. The live workflow caller for Decision item 6
|
|
// is #2296's gsd-core/references/execute-phase-quota-recovery.md, asserted
|
|
// separately below.
|
|
test('--attempt walks the effort ladder above the configured default', (t) => {
|
|
const dir = createTempProject();
|
|
t.after(() => cleanup(dir));
|
|
fs.writeFileSync(
|
|
path.join(dir, '.planning', 'config.json'),
|
|
JSON.stringify({
|
|
effort: { default: 'low' },
|
|
dynamic_routing: { enabled: true, escalate_on_failure: true, max_escalations: 3 },
|
|
}, null, 2),
|
|
);
|
|
|
|
const at = (n) => JSON.parse(
|
|
runGsdTools(`query resolve-execution gsd-planner --host claude --attempt ${n}`, dir).output,
|
|
);
|
|
|
|
// attempt 0 is the un-escalated baseline; a later attempt must not be lower.
|
|
const base = at(0);
|
|
const later = at(2);
|
|
const RANK = { minimal: 0, low: 1, medium: 2, high: 3, xhigh: 4, max: 5 };
|
|
assert.equal(base.effort, 'low');
|
|
assert.ok(
|
|
RANK[later.effort] >= RANK[base.effort],
|
|
`escalation must not lower effort: attempt0=${base.effort} attempt2=${later.effort}`,
|
|
);
|
|
// Whatever the ladder resolved, it must still reach the host as an argument.
|
|
assert.equal(later.effort_surface, 'argv');
|
|
assert.deepEqual(later.effort_argv, ['--effort', later.effort]);
|
|
});
|
|
|
|
test('a negative --attempt is rejected', (t) => {
|
|
const dir = projectWithEffort('high');
|
|
t.after(() => cleanup(dir));
|
|
assert.ok(!runGsdTools('query resolve-execution gsd-planner --attempt -1', dir).success);
|
|
});
|
|
});
|
|
|
|
describe('#2481 — ADR-443 mechanism callers, as they actually exist', () => {
|
|
const quotaRecovery = fs.readFileSync(
|
|
path.join(REPO_ROOT, 'gsd-core', 'references', 'execute-phase-quota-recovery.md'),
|
|
'utf8',
|
|
);
|
|
const executePhase = fs.readFileSync(
|
|
path.join(REPO_ROOT, 'gsd-core', 'workflows', 'execute-phase.md'),
|
|
'utf8',
|
|
);
|
|
|
|
test('Decision item 6 (escalation) has its live caller — from #2296, not this change', () => {
|
|
assert.match(
|
|
quotaRecovery, /resolve-execution\s+gsd-executor\s+--attempt/,
|
|
'execute-phase-quota-recovery.md must invoke resolve-execution with --attempt (#2296)',
|
|
);
|
|
assert.ok(
|
|
executePhase.includes('references/execute-phase-quota-recovery.md'),
|
|
'that reference must be @-included into execute-phase.md, or it is not a live caller',
|
|
);
|
|
});
|
|
|
|
test('Decision item 1 (invocation override) still has NO live caller', () => {
|
|
// Guards the corrected ADR-443 claim. If someone later wires --effort into a
|
|
// workflow, this fails and the ADR status text must be revisited — that is
|
|
// the point: the ADR must not silently drift back to being wrong.
|
|
const dirs = ['gsd-core/workflows', 'gsd-core/references', 'agents', 'commands'];
|
|
const hits = [];
|
|
const walk = (d) => {
|
|
const abs = path.join(REPO_ROOT, d);
|
|
if (!fs.existsSync(abs)) return;
|
|
for (const e of fs.readdirSync(abs, { withFileTypes: true })) {
|
|
const full = path.join(abs, e.name);
|
|
if (e.isDirectory()) walk(path.relative(REPO_ROOT, full));
|
|
else if (e.name.endsWith('.md') && /resolve-execution[^\r\n]*--effort\s/.test(fs.readFileSync(full, 'utf8'))) {
|
|
hits.push(path.relative(REPO_ROOT, full));
|
|
}
|
|
}
|
|
};
|
|
dirs.forEach(walk);
|
|
assert.deepEqual(
|
|
hits, [],
|
|
`ADR-443 records Decision item 1 as having no live caller; found: ${JSON.stringify(hits)}. ` +
|
|
'Update the ADR-443 amendment before adding one.',
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('#2481 review workflow resolves effort per reviewer', () => {
|
|
test('shipped orchestration invokes resolve-execution — the grep ADR-443 said returned zero hits', () => {
|
|
// Phase 5b (#2799) moved the call out of review.md's per-lane bash and into the review-lane
|
|
// route, which resolves effort once per selected lane through the SAME surface. ADR-443's
|
|
// invariant is about shipped orchestration calling resolve-execution at all, not about which
|
|
// file it lives in — so the assertion follows the call rather than pinning the old location.
|
|
const toolsSrc = fs.readFileSync(
|
|
path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs'), 'utf-8',
|
|
);
|
|
assert.ok(
|
|
toolsSrc.includes('resolve-execution'),
|
|
'ADR-443 blocks on no shipped orchestration calling resolve-execution',
|
|
);
|
|
});
|
|
|
|
test('each argv-effort reviewer places effort in its resolved command line', () => {
|
|
// Stronger than the old shell-variable check: this asserts the effort actually lands in the
|
|
// argv AT THE POSITION the lane declares, which a `$VAR` substring never proved. Lanes whose
|
|
// effortChannel is not `argv` must receive nothing.
|
|
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
|
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
|
const EFFORT = ['--effort', 'high'];
|
|
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
|
const r = resolveLanePlan({
|
|
lane, configGet: () => undefined, runDir: '/run', repoRoot: '/repo', effortArgs: EFFORT,
|
|
});
|
|
assert.equal(r.ok, true, `${lane.slug} failed to resolve`);
|
|
const carries = r.plan.argv.includes('--effort');
|
|
assert.equal(
|
|
carries, lane.invoke.effortChannel === 'argv',
|
|
`${lane.slug}: effortChannel=${lane.invoke.effortChannel} but argv ${carries ? 'carries' : 'omits'} effort`,
|
|
);
|
|
}
|
|
// The three lanes ADR-1239 #2481 named must still be the argv-effort set.
|
|
const argvEffort = REVIEWER_LANES
|
|
.filter((l) => l.transport === 'spawn' && l.invoke.effortChannel === 'argv')
|
|
.map((l) => l.slug).sort();
|
|
assert.deepStrictEqual(argvEffort, ['claude', 'codex', 'opencode']);
|
|
});
|
|
});
|
|
|
|
describe('#2615: the matrix documents the effortSurface axis', () => {
|
|
test('the axes legend defines effortSurface and its vocabulary', () => {
|
|
const legendRow = MATRIX_TEXT.split(/\r?\n/).find((l) => l.startsWith('| `effortSurface` |'));
|
|
assert.ok(legendRow, 'the axes legend must define effortSurface (#2615)');
|
|
for (const member of ['`argv`', '`none`', '`undocumented`']) {
|
|
assert.ok(legendRow.includes(member),
|
|
`the legend must document the ${member} vocabulary member (#2615)`);
|
|
}
|
|
});
|
|
|
|
test('there is at least one runtime to check', () => {
|
|
// Guards the loops below against silently asserting nothing.
|
|
assert.ok(MATRIX_RUNTIMES.length >= 18, `expected the full runtime corpus, got ${MATRIX_RUNTIMES.length}`);
|
|
});
|
|
|
|
for (const id of MATRIX_RUNTIMES) {
|
|
describe(`runtime: ${id}`, () => {
|
|
test('has a matrix section', () => {
|
|
assert.ok(matrixSection(id), `${id}: every installed runtime needs a matrix section (ADR-1239)`);
|
|
});
|
|
|
|
test('documents effortSurface, and the value matches the descriptor', () => {
|
|
const body = matrixSection(id);
|
|
assert.ok(body, `${id}: missing matrix section`);
|
|
|
|
const documented = matrixAxisValue(body, 'effortSurface');
|
|
assert.ok(documented, `${id}: the matrix must carry an effortSurface row (#2615)`);
|
|
|
|
const declared = registry.runtimes[id].runtime.hostIntegration.effortSurface;
|
|
if (declared === undefined) {
|
|
// kimi-code declares no value: its mechanism (`/effort`) is interactive-only
|
|
// and neither `argv` nor `none` describes it. The matrix must say so rather
|
|
// than invent a value.
|
|
assert.match(documented, /not declared/i,
|
|
`${id}: an absent descriptor value must be documented as absent, not guessed (#2615)`);
|
|
} else {
|
|
assert.equal(documented, declared,
|
|
`${id}: the matrix effortSurface value must match the shipped descriptor`);
|
|
}
|
|
});
|
|
});
|
|
}
|
|
});
|