* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local>
187 lines
7.2 KiB
JavaScript
187 lines
7.2 KiB
JavaScript
/**
|
|
* Tests for gsd-workflow-guard.js PreToolUse hook.
|
|
*
|
|
* #2304 — Kimi tool vocabulary engages the guard: Kimi CLI registers this
|
|
* guard with matcher 'Shell|WriteFile|StrReplaceFile' and forwards its own
|
|
* tool vocabulary (tool_name 'Shell', possibly module-qualified). kimi-cli's
|
|
* Shell.Params names its field `command` (src/kimi_cli/tools/shell/
|
|
* __init__.py), same as Claude's Bash, so only the tool name needs
|
|
* normalization. Pre-fix the guard's Bash branch never matched on Kimi and
|
|
* the force-add block was silently dormant.
|
|
*/
|
|
|
|
'use strict';
|
|
|
|
process.env.GSD_TEST_MODE = '1';
|
|
|
|
const { test, describe, before, after } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const os = require('node:os');
|
|
const path = require('node:path');
|
|
const { runHook: runHookSeam } = require('./helpers/process-seam.cjs');
|
|
const { throwIfFailed } = require('./helpers/git-fixture.cjs');
|
|
|
|
const { cleanup } = require('./helpers.cjs');
|
|
|
|
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-workflow-guard.js');
|
|
|
|
function runHook(payload, timeoutMs = 5000) {
|
|
const input = JSON.stringify(payload);
|
|
const r = runHookSeam(HOOK_PATH, [], { input, timeoutMs });
|
|
if (r.exitCode === 0) {
|
|
return { exitCode: 0, stdout: r.stdout.trim(), stderr: '' };
|
|
}
|
|
return {
|
|
exitCode: r.exitCode ?? 1,
|
|
stdout: r.stdout.trim(),
|
|
stderr: r.stderr.trim(),
|
|
};
|
|
}
|
|
|
|
describe('#2304: Kimi tool vocabulary engages the workflow guard', () => {
|
|
// A repo on a worktree-agent-* branch with the guard enabled: the one
|
|
// state where the Bash branch produces an observable block, so a dormant
|
|
// guard (silent exit 0) is distinguishable from a working one (exit 2).
|
|
let repoDir;
|
|
|
|
before(() => {
|
|
repoDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-workflow-guard-'));
|
|
const initResult = runHookSeam(
|
|
'-c',
|
|
['git init -q -b worktree-agent-test && git config user.email t@t && git config user.name t'],
|
|
{ interpreter: 'bash', cwd: repoDir },
|
|
);
|
|
throwIfFailed(initResult, 'bash -c <git init/config for gsd-workflow-guard fixture>');
|
|
fs.mkdirSync(path.join(repoDir, '.planning'));
|
|
fs.writeFileSync(
|
|
path.join(repoDir, '.planning', 'config.json'),
|
|
JSON.stringify({ hooks: { workflow_guard: true } })
|
|
);
|
|
});
|
|
|
|
after(() => {
|
|
cleanup(repoDir);
|
|
});
|
|
|
|
test('Shell force-add on a worktree-agent branch is blocked like Bash', () => {
|
|
const r = runHook({
|
|
tool_name: 'Shell',
|
|
tool_input: { command: 'git add -f secrets.env' },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 2, 'Kimi Shell should reach the Bash branch and block');
|
|
const output = JSON.parse(r.stdout);
|
|
assert.equal(
|
|
output.code,
|
|
'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN',
|
|
'block payload should carry the force-add code'
|
|
);
|
|
assert.ok(
|
|
r.stderr.includes('must not run git add -f'),
|
|
'reason must reach stderr — that is what Kimi feeds back to the model on exit 2'
|
|
);
|
|
});
|
|
|
|
test('module-qualified kimi_cli.tools.shell:Shell is recognized', () => {
|
|
const r = runHook({
|
|
tool_name: 'kimi_cli.tools.shell:Shell',
|
|
tool_input: { command: 'git add --force secrets.env' },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 2);
|
|
assert.equal(JSON.parse(r.stdout).code, 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN');
|
|
});
|
|
|
|
test('benign Shell command passes through', () => {
|
|
const r = runHook({
|
|
tool_name: 'Shell',
|
|
tool_input: { command: 'git status' },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 0);
|
|
assert.equal(r.stdout, '');
|
|
});
|
|
|
|
test('Bash (Claude vocabulary) still blocks — normalization is additive', () => {
|
|
const r = runHook({
|
|
tool_name: 'Bash',
|
|
tool_input: { command: 'git add -f secrets.env' },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 2);
|
|
assert.equal(JSON.parse(r.stdout).code, 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN');
|
|
});
|
|
|
|
test('WriteFile outside .planning/ gets the workflow advisory like Write', () => {
|
|
const r = runHook({
|
|
tool_name: 'WriteFile',
|
|
tool_input: { path: path.join(repoDir, 'src', 'app.js'), content: 'x' },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 0);
|
|
const output = JSON.parse(r.stdout);
|
|
assert.ok(
|
|
output.hookSpecificOutput?.additionalContext?.includes('WORKFLOW ADVISORY'),
|
|
'Kimi WriteFile should reach the write branch and emit the advisory'
|
|
);
|
|
});
|
|
|
|
test('StrReplaceFile editing .planning/ passes silently', () => {
|
|
const r = runHook({
|
|
tool_name: 'StrReplaceFile',
|
|
tool_input: { path: path.join(repoDir, '.planning', 'notes.md'), edit: { old: 'a', new: 'b' } },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 0);
|
|
assert.equal(r.stdout, '');
|
|
});
|
|
|
|
// #2547 — normalizeKimiPayload rebuilt old_string/new_string with
|
|
// `String(e.old ?? '')`. `??` guards the value, not the dereference, so a
|
|
// NULLISH entry threw a TypeError at the top of the handler, before the Bash
|
|
// branch ran. The outer `catch { process.exit(0) }` swallowed it, so a Shell
|
|
// payload carrying a spurious malformed `edit` field walked straight past the
|
|
// force-add hard block. The `edit` field is never read on the Bash path — it
|
|
// only has to be present to trigger the crash, which is what makes this
|
|
// reachable from a command that has nothing to do with editing.
|
|
//
|
|
// The boundary is nullish specifically: `('x').old` is a legal property read
|
|
// yielding undefined, so a string entry never threw. The `null entry` case is
|
|
// the regression (exits 0 against pre-fix code); the rest are controls.
|
|
describe('#2547: a spurious malformed edit field does not disarm the force-add block', () => {
|
|
for (const [label, edit] of [
|
|
['null entry (the #2547 bypass)', [null]],
|
|
// `{"toString": null}` is valid JSON whose coercion throws "Cannot
|
|
// convert object to primitive value" — the same crash-to-allow reached
|
|
// through String() rather than through the property read.
|
|
['non-coercible old (the #2547 String() bypass)', [{ old: { toString: null }, new: 'x' }]],
|
|
['non-coercible new (the #2547 String() bypass)', [{ old: 'x', new: { toString: null } }]],
|
|
['string entry (control — never threw)', ['nope']],
|
|
['bare null, not a list (control — normalizes to no edits)', null],
|
|
]) {
|
|
test(`force-add still blocks with a spurious edit field (${label})`, () => {
|
|
const r = runHook({
|
|
tool_name: 'Shell',
|
|
tool_input: { command: 'git add -f secrets.env', edit },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 2,
|
|
`a spurious malformed edit field (${label}) must not downgrade the force-add ` +
|
|
`block to a silent allow. Got exit ${r.exitCode}. stderr: ${r.stderr}`);
|
|
assert.equal(JSON.parse(r.stdout).code, 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN');
|
|
});
|
|
}
|
|
|
|
test('benign command with a malformed edit field still passes (no over-block)', () => {
|
|
const r = runHook({
|
|
tool_name: 'Shell',
|
|
tool_input: { command: 'git status', edit: [null] },
|
|
cwd: repoDir,
|
|
});
|
|
assert.equal(r.exitCode, 0, `benign command must stay allowed. stderr: ${r.stderr}`);
|
|
assert.equal(r.stdout, '');
|
|
});
|
|
});
|
|
});
|