Files
msd-core/tests/vscode-lm-tools.test.cjs
Tom Boucher aceea3ce4a refactor(#3217): withhold a percentage when its scope is not complete (#3318)
* wip(#3217): rule-4 scope withholding — parked, two open findings

Implemented but NOT shippable. An isolated review found buildStateFrontmatter
still hardcodes SCOPE.COMPLETE, so state json reports percent 0 where roadmap
analyze, stats and query progress all correctly report null on the same disk
state - rule 4 reintroduced at a site this phase claims to close. Also: roadmap
analyze emits scope complete beside progress_percent null with nothing
explaining it.

Parked to build Phase 4 (#3186) first, which is unblocked. Findings recorded in
.gsd/phase/refactor-3217-completion-ratio-scoping/60-review.json.

* fix(#3217): withhold the sync percentage on a non-complete scope

The parked blocker is fixed - buildStateFrontmatter no longer hardcodes
SCOPE.COMPLETE, and the prose Progress fallback is gated too, which was a second
leak found while tracing the first. roadmap analyze exposes progress_scope so a
consumer can tell WHY a percentage is absent from the JSON alone.

Then a residual gap was reproduced rather than assumed. cmdStateSync carried the
same hardcode behind a written reason claiming it did not reproduce. It did: on a
TRUNCATED window and on UNSCOPED row 4, state sync wrote Progress 0 percent to 100
percent while state json, roadmap analyze, stats and query progress all withheld -
and it persisted a self-contradictory file, body claiming 100 percent while its own
frontmatter correctly omitted percent.

The excuse was also wrong. syncRoadmapRaw is already parsed in that function and is
exactly what produces a real scope, so there was a scope to pass. Threaded through
listMilestonePhaseDirs; a non-complete scope now skips the write with a reason in
changes. milestoneBounded stays as the orthogonal 1761 guard for row 5.

Second time this epic a does-not-reproduce claim was too generous. Recorded in
ADR Amendment 8 as a correction rather than a quiet rewrite.

Verified on the remote runner.

* test(#3217): give the withholding fixtures a resolvable scope

40 matrix failures, all fixture drift - no code regression. My own hypothesis
that this was over-withholding was wrong and is recorded as such: the worry case,
a plain ROADMAP with Phase entries and no version heading, resolves to complete
exactly as ADR 7.1 says it should.

The real causes were two fixture shapes. Most had no ROADMAP.md at all, which is
unreadable via a pre-existing graceful path, and asserted a numeric percent. The
five vscode, pi-extension, mcp-server and shell-projection failures were that
shape - bare temp dirs using progress json as a reachability proxy while
asserting typeof percent is number, which under rule 4 is now null.

The rest had a version token in a title or heading with no STATE.md milestone
pointer to resolve it, which is classification row 4, versioned but unresolved,
so withholding is correct per the contract.

Verified on the remote runner.

* test(#3217): make the LM-tools reachability tests dispatch against their fixture

The gsd_progress reachability test was never testing its fixture. invoke()
resolves cwd from vscode.workspace.workspaceFolders by design (the real
LanguageModelToolInvocationOptions has no cwd field, per the 2103 fix in
extension.js), the mock had no workspace at all, and the test passed a cwd option
nothing reads - so it dispatched against the repo working directory. Writing a
ROADMAP into the temp dir had no effect. Rule 4 only made it visible.

Fixed by mocking workspaceFolders. The two siblings in the same file carried the
identical dead cwd and were dispatching against the repo too; they were not
failing only because their assertions did not touch scope-dependent output. Both
now use their own fixture with assertions unchanged - the no-planning fallback
paths already satisfy them honestly.

Re-scanned the other five reachability files: no further instances. They thread
cwd into parameters that genuinely read it, not through an options shape that
ignores it.

Verified on the remote runner.

* chore(#3217): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3318 exists.

* ci(#3217): give the coverage merge enough heap for the merged shards

The coverage gate OOMed at exit 134. c8 report merges three shard artifacts,
roughly 358MB of V8 dumps in coverage/tmp, and died holding their per-file
position maps at the ~4GB default heap. Verified as this branch's delta rather
than pre-existing: the same job succeeded on next at 14:18, after phases 4 and 5
merged.

Both coverage-gate steps get the bump because both re-slice the same merged data.
8192 doubles what failed and leaves headroom on a 16GB ubuntu runner, matching
the idiom the shard step already uses at 6144.

This is a memory bound, not a change to what is measured. No threshold was
touched. The test file was checked for gratuitous subprocess spawning and is
already reasonable at 43 spawns, each a distinct fixture-by-surface pairing.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-10 11:59:51 -04:00

184 lines
9.0 KiB
JavaScript

'use strict';
/**
* VS Code Language Model Tools test — #2103 UPGRADE 1.
*
* Proves the GSD extension's Language Model Tools are keystone-WIRED:
* 1. contributes.languageModelTools is present in package.json and its
* `name` entries match the runtime registration names (extension.js's
* LM_TOOLS / browser.js's LM_TOOL_NAMES) exactly — a mismatch here would
* mean VS Code rejects the tool registration at activation.
* 2. registerLanguageModelTools() calls vscode.lm.registerTool for every
* manifest entry (mock vscode.lm — no real VS Code host available in CI).
* 3. A registered tool's invoke() dispatches through the SAME shared
* dispatchGsdCommand as gsd.invoke (desktop) and returns REAL output —
* the "user can invoke X" proof, matching the reachability-test pattern.
* 4. The desktop and web entries register the identical tool NAME set (only
* the invoke() behavior differs — real dispatch vs. honest web-mode message).
*/
const { test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const pkg = require('../vscode/package.json');
const extension = require('../vscode/extension.js');
const browser = require('../vscode/browser.js');
const { createTempDir, cleanup } = require('./helpers.cjs');
class FakeTextPart {
constructor(text) { this.text = text; }
}
class FakeToolResult {
constructor(parts) { this.parts = parts; }
}
// `workspaceCwd`: createLanguageModelTool's invoke() resolves cwd via
// resolveWorkspaceCwd(vscode) — vscode.workspace.workspaceFolders[0].uri.fsPath
// — NOT via a `cwd` field on the invoke() options object (LanguageModelToolInvocationOptions
// has no such field on the real API; see extension.js's #2103 FIX comment on
// createLanguageModelTool). A test that wants a real dispatch to run against a
// fixture directory must mock workspace.workspaceFolders here, not pass
// `{ cwd }` in the options object handed to invoke() (that field is simply
// never read).
function mockVscodeLm(workspaceCwd) {
const registered = [];
return {
lm: {
registerTool(name, impl) {
registered.push({ name, impl });
return { dispose() {} };
},
},
workspace: workspaceCwd
? { workspaceFolders: [{ uri: { fsPath: workspaceCwd } }] }
: undefined,
LanguageModelTextPart: FakeTextPart,
LanguageModelToolResult: FakeToolResult,
registered,
};
}
test('package.json contributes.languageModelTools is present with well-formed entries', () => {
assert.ok(pkg.contributes && Array.isArray(pkg.contributes.languageModelTools),
'contributes.languageModelTools must be an array');
assert.ok(pkg.contributes.languageModelTools.length > 0, 'must declare at least one tool');
for (const tool of pkg.contributes.languageModelTools) {
assert.equal(typeof tool.name, 'string');
assert.ok(tool.name.length > 0);
assert.equal(typeof tool.toolReferenceName, 'string');
assert.equal(typeof tool.displayName, 'string');
assert.equal(typeof tool.modelDescription, 'string');
assert.equal(typeof tool.userDescription, 'string');
assert.equal(tool.canBeReferencedInPrompt, true);
assert.ok(Array.isArray(tool.tags));
assert.equal(typeof tool.inputSchema, 'object');
}
});
test('manifest tool names exactly match extension.js LM_TOOLS registration names', () => {
const manifestNames = pkg.contributes.languageModelTools.map((t) => t.name).sort();
const runtimeNames = extension.LM_TOOLS.map((t) => t.name).sort();
assert.deepEqual(runtimeNames, manifestNames,
'a mismatch here means VS Code would reject the runtime registerTool call against the manifest');
});
test('manifest tool names exactly match browser.js LM_TOOL_NAMES (web entry registers the same surface)', () => {
const manifestNames = pkg.contributes.languageModelTools.map((t) => t.name).sort();
assert.deepEqual([...browser.LM_TOOL_NAMES].sort(), manifestNames);
});
test('REACHABILITY (desktop): registerLanguageModelTools registers every manifest tool via vscode.lm.registerTool', () => {
const mock = mockVscodeLm();
const context = { subscriptions: [] };
const count = extension.registerLanguageModelTools(mock, context);
assert.equal(count, pkg.contributes.languageModelTools.length);
assert.equal(mock.registered.length, pkg.contributes.languageModelTools.length);
assert.equal(context.subscriptions.length, pkg.contributes.languageModelTools.length);
assert.deepEqual(mock.registered.map((r) => r.name).sort(), extension.LM_TOOLS.map((t) => t.name).sort());
});
test('REACHABILITY (desktop): gsd_progress tool.invoke() dispatches through the hub and returns REAL output', async () => {
const dir = createTempDir();
try {
// #3217 (ADR-3180 §7.6 rule 4): a free-form ROADMAP.md (no version
// token) is COMPLETE scope for windowing (§7.1) — without this, a
// bare temp dir has no ROADMAP.md at all (UNREADABLE) and `percent`
// is withheld (null), breaking this reachability proxy.
fs.mkdirSync(path.join(dir, '.planning'), { recursive: true });
fs.writeFileSync(path.join(dir, '.planning', 'ROADMAP.md'), '# Roadmap\n');
// The LM tool's invoke() resolves cwd via vscode.workspace.workspaceFolders
// (resolveWorkspaceCwd), not via an options.cwd field — pass the fixture
// dir through the mock's workspace so dispatch actually runs against it.
const mock = mockVscodeLm(dir);
extension.registerLanguageModelTools(mock, { subscriptions: [] });
const progressTool = mock.registered.find((r) => r.name === 'gsd_progress');
assert.ok(progressTool, 'gsd_progress must be registered');
const result = await progressTool.impl.invoke({ input: {} }, {});
assert.ok(result instanceof FakeToolResult, 'invoke must return a LanguageModelToolResult');
assert.ok(Array.isArray(result.parts) && result.parts.length === 1);
assert.ok(result.parts[0] instanceof FakeTextPart, 'result part must be a LanguageModelTextPart');
const parsed = JSON.parse(result.parts[0].text);
assert.equal(typeof parsed.percent, 'number', 'the real progress command ran (engine reached, not a stub)');
} finally {
cleanup(dir);
}
});
test('REACHABILITY (desktop): gsd_plan_phase tool.invoke() forwards the "phase" input through dispatch (real, not UnknownCommand)', async () => {
const dir = createTempDir();
try {
// The LM tool's invoke() resolves cwd via vscode.workspace.workspaceFolders
// (resolveWorkspaceCwd), not via an options.cwd field — pass the fixture
// dir through the mock's workspace so dispatch actually runs against it.
const mock = mockVscodeLm(dir);
extension.registerLanguageModelTools(mock, { subscriptions: [] });
const planPhaseTool = mock.registered.find((r) => r.name === 'gsd_plan_phase');
assert.ok(planPhaseTool);
const result = await planPhaseTool.impl.invoke({ input: { phase: 'nonexistent-phase-8675309' } }, {});
const parsed = JSON.parse(result.parts[0].text);
// Real dispatch reaches gsd-tools.cjs and returns a structured "phase not
// found" response (proves the engine was reached) — not the manifest's
// own family/subcommand rejected as unknown.
assert.equal(parsed.phase, 'nonexistent-phase-8675309');
assert.ok('error' in parsed || 'plans' in parsed, 'expected a real phase-plan-index response shape');
} finally {
cleanup(dir);
}
});
test('gsd_workstreams tool.invoke() dispatches through the hub and returns REAL output', async () => {
const dir = createTempDir();
try {
// The LM tool's invoke() resolves cwd via vscode.workspace.workspaceFolders
// (resolveWorkspaceCwd), not via an options.cwd field — pass the fixture
// dir through the mock's workspace so dispatch actually runs against it.
const mock = mockVscodeLm(dir);
extension.registerLanguageModelTools(mock, { subscriptions: [] });
const wsTool = mock.registered.find((r) => r.name === 'gsd_workstreams');
const result = await wsTool.impl.invoke({ input: {} }, {});
const parsed = JSON.parse(result.parts[0].text);
assert.ok('workstreams' in parsed || 'mode' in parsed, 'expected a real workstream list response shape');
} finally {
cleanup(dir);
}
});
test('registerLanguageModelTools fails soft (returns 0, does not throw) when vscode.lm is absent', () => {
assert.doesNotThrow(() => {
const count = extension.registerLanguageModelTools({}, { subscriptions: [] });
assert.equal(count, 0);
});
});
test('WEB MODE: browser.js registerLanguageModelTools registers the same names but invoke() returns an honest web-mode message (no engine dispatch)', async () => {
const mock = mockVscodeLm();
const count = browser.registerLanguageModelTools(mock, { subscriptions: [] });
assert.equal(count, browser.LM_TOOL_NAMES.length);
const progressTool = mock.registered.find((r) => r.name === 'gsd_progress');
const result = await progressTool.impl.invoke({ input: {} }, {});
assert.match(result.parts[0].text, /web mode/i);
assert.match(result.parts[0].text, /MCP server/);
});