Files
msd-core/tests/read-guard.test.cjs
0xdhx a8b40fa53f fix(#2547): fail closed on crashing and path-shadowing Kimi payloads (#2595)
* fix(#2547): fail closed on a malformed Kimi edit list in normalizeKimiPayload

`normalizeKimiPayload` rebuilt old_string/new_string with
`String(e.old ?? '')`. `??` guards the value, not the dereference, so a
nullish entry in a Kimi `edit` list threw a TypeError at the top of the
handler, before any tool dispatch. Each guard's outer
`catch { process.exit(0) }` swallowed that crash and emitted the same exit
code as "nothing to report" — turning a should-BLOCK call into a silent
allow.

Two hard blocks were bypassable:

  * gsd-worktree-path-guard's cross-git-root write block (#260) — a
    StrReplaceFile write whose path resolves to a different git root is
    correctly blocked with a well-formed edit list, and silently allowed
    with `edit: [null]`.
  * gsd-workflow-guard's force-add block on agent-* branches — a Shell
    payload carrying a spurious `edit: [null]` field walks past it. The
    Bash path never reads `edit`; the field only has to be present to
    trigger the crash.

Fixed with `e?.old` / `e?.new`, landed identically in all five copies so
tests/kimi-guard-normalization-parity.test.cjs's byte-identity assertion
still holds.

The crash boundary is nullish specifically, not "non-object": `('x').old`
and `(7).old` are legal reads yielding undefined, so string/number entries
never threw. Both are kept as controls proving the fix did not change
their behaviour.

Regression coverage is folded into the owning suites per CONTRIBUTING.md
(no new bug-* files). Negative-controlled: the nullish cases exit 0
against pre-fix guards and exit 2 after, with positive controls (the
equivalent well-formed payload blocks) and negative controls (in-worktree
writes and benign commands still pass) alongside.

Refs #2547

* test(#2547): exercise the production Kimi payload shape in read-guard tests

The `#2304: Kimi tool vocabulary engages the read guard` cases send
payloads with no `session_id`, and runHook injects none. A live Kimi turn
always carries one — kimi-cli's hooks/events.py `_base()` sets it
unconditionally, and soul/kimisoul.py calls `set_session_id()` at the top
of every turn before tool dispatch, so the ContextVar's `default=""` never
reaches a tool call.

gsd-read-guard treats any non-empty `data.session_id` as "Claude Code
already enforces read-before-edit, skip" (#2520). So the advisory those
tests assert fires only for a shape production never sends: the tests were
green, and the guard was dormant on Kimi. A sibling #2520 case in the same
file asserts the skip when `session_id` IS present — both passed, and the
production shape hits the skip.

Two changes, test-validity only:

  * Retitle the #2304 block to say what it proves — the tool VOCABULARY is
    normalized through to the Write/Edit branch — with a comment warning
    not to read it as production evidence.
  * Add a #2547 block asserting behaviour against the production shape
    (session_id populated), including a case that pins the delta directly:
    the same payload fires without session_id and is silent with it.

The #2547 block characterizes a known gap; it does not endorse it.
Redesigning how the guard discriminates runtimes is explicitly out of
scope for #2547. If a later change makes the advisory fire on Kimi these
tests are supposed to fail — update them then rather than dropping the
coverage.

Refs #2547

* docs(#2547): scope the Kimi guard-engagement claim to what Kimi enforces

#2518 engaged the guards' Kimi matchers and the release notes describe the
result as "All seven guard hooks now engage on Kimi", singling out the
prompt-injection read scanner as "the security-relevant guard" taken "from
silently dormant to engaged". That is not achievable for the scanner at
the emit layer.

gsd-read-injection-scanner.js is a PostToolUse hook, and kimi-cli's
dispatch never inspects PostToolUse hook results: src/kimi_cli/soul/
toolset.py awaits PreToolUse and honours `result.action == "block"`, but
fires PostToolUse via asyncio.create_task() and returns the ToolResult
without awaiting it — the done_callback only retrieves the task's own
exception. So no output shape the scanner emits can block or flag a Kimi
tool call, and `security.injection_blocking` cannot take effect there.
Reshaping the scanner's output would not change this; the enforcement gap
is in kimi-cli's PostToolUse handling, which is out of scope here.

This corrects the claim rather than the code — there is no gsd-core emit
fix that would make it true:

  * .changeset/2304-kimi-guard-tool-name.md — the fragment is unreleased,
    so it would otherwise ship this as a CHANGELOG security claim.
    Headline narrowed to "normalize Kimi's payload shape" and a scope
    paragraph added naming what actually blocks on Kimi (the two
    PreToolUse blocks) versus what cannot.
  * docs/migration/kimi-to-kimi-code.md — the scanner was listed under
    "Every GSD `PreToolUse` guard"; it is PostToolUse. Corrected, and the
    "What about the dormant guards?" section now splits enforceable from
    not-enforceable instead of saying Phase 0 "fixed all seven".
  * hooks/gsd-read-injection-scanner.js — the same scope note in the
    file's own Kimi rationale comment, where the next contributor to touch
    the normalization will actually read it. Comment only; the shared
    normalization block is untouched and byte-identity still holds.

Refs #2547

* chore(#2547): regenerate golden install-parity fixtures for the guard fix

The golden install-parity fixtures record a content hash per installed
file, so changing the five guard hooks changes their hashes across every
runtime's fixture. Regenerated with the full sweep (build, gen:golden,
size:baseline) rather than a single generator — running gen:golden alone
leaves tests/workflow-size-baseline.json stale and loses CI jobs to a
regeneration that looked complete.

The size baselines came out unchanged (no workflow or agent bodies
touched) and the hash delta is confined to exactly the five guards:
gsd-prompt-guard, gsd-read-guard, gsd-read-injection-scanner,
gsd-workflow-guard, gsd-worktree-path-guard.

Refs #2547

* fix(#2547): guard the String() coercion in normalizeKimiPayload too

Found by adversarial review of the first commit, then reproduced against
pristine next: `e?.old` closes the nullish dereference but leaves a second
route to the same crash-to-allow.

`{"toString": null}` is valid JSON, and coercing it throws
`TypeError: Cannot convert object to primitive value` — so an edit entry
that IS a well-formed object still crashes normalization, still lands in
the outer `catch { process.exit(0) }`, and still downgrades a should-BLOCK
call to a silent allow. Confirmed on both hard blocks:

  {"tool_name":"Shell","tool_input":{
     "command":"git add -f secret.env",
     "edit":[{"old":{"toString":null},"new":"x"}]}}      -> exit 0 (was)

  {"tool_name":"StrReplaceFile","tool_input":{
     "path":"<main-repo>/src/index.ts",
     "edit":[{"old":{"toString":null},"new":"x"}]}}      -> exit 0 (was)

Both exit 2 now.

The coercion is wrapped rather than replaced with a `typeof === 'string'`
test on purpose. Degrading only the non-coercible entry keeps
stringification identical for every value that CAN coerce — numbers,
arrays, plain objects — which matters because gsd-prompt-guard scans
new_string for injection patterns, and a `typeof` test would silently stop
scanning content that reaches that scan today (e.g. `new: ["ignore all
previous instructions"]` currently stringifies and is scanned). Verified:
zero behaviour change across string, number, bool, null, array-of-strings,
nested array, plain object and `__proto__`-keyed input; only the throwing
case changes, from crash to ''.

Regression cases are negative-controlled against the previous commit: the
four new coercion-trap tests fail with only the `e?.old` fix in place and
pass with this one.

Refs #2547

* chore(#2547): cover the String() coercion vector in the changeset

The release note described only the nullish-dereference route. Both routes
reach the same fail-open, so both belong in the changelog entry, along with
why the coercion is wrapped rather than type-tested.

Refs #2547

* chore(#2547): point the changeset fragment at the real PR number

The fragment has to exist before `gh pr create` runs, so it carried the
issue number as a placeholder. Corrected to 2595 now that the PR is open.

Refs #2547

* fix(#2547): make Kimi's `path` authoritative over a model-supplied `file_path`

normalizeKimiPayload copied Kimi's `path` into `file_path` only when
`file_path === undefined`, so any `file_path` the model chose to include won
outright. Every guard reads `file_path`; kimi-cli executes on `path`. The guard
therefore inspected one file while the write landed on another.

This bypass needs no crash. A payload pairing a cross-root `path` with a
spurious `file_path: ""` left gsd-worktree-path-guard reading an empty string
and exiting 0, while the identical write without the extra key blocked — the
same cross-root write the #260 block exists to catch. The shadowing also
preserved a non-string `file_path` (`[]`), which threw inside that guard's
path.isAbsolute() and reached its outer `catch { process.exit(0) }`: the same
crash-to-allow the rest of #2547 closes, reached through the guard's own read
rather than through normalization.

Reachability is not speculative. kimi-cli's soul/toolset.py json-parses the
model's raw tool arguments and passes that dict verbatim as tool_input to
PreToolUse, performing typed validation only later inside tool.call() — after
the hook has already decided. So the model controls extra keys in tool_input at
the moment the guard runs. kimi-cli's file tools carry no `file_path` field at
all (src/kimi_cli/tools/file/write.py, replace.py), so a `file_path` in a Kimi
payload is always model-supplied.

`path` now wins outright. Overwriting can only ever narrow what a guard inspects
to the path that will actually be written, so it cannot under-block.
Normalization returns early for non-Kimi tool names, so the native Claude Code
contract (file_path governs) is untouched.

Landed identically across all five inlined copies; the byte-identity assertion
in tests/kimi-guard-normalization-parity.test.cjs enforces that.

* test(#2547): cover the file_path-shadowing bypass in the #260 guard suite

Four cases, each exiting 0 (bypass) against the pre-fix guards: a spurious
empty-string file_path, an in-worktree decoy file_path, and non-string
file_path values (array and object) that additionally crashed
path.isAbsolute() into the outer catch.

Two controls that are not bypass cases and matter as much:

  - an in-worktree write carrying a cross-root DECOY file_path must still exit
    0. Pre-fix this blocked, because the decoy won; the guard now follows the
    path kimi-cli executes on in both directions, so the fix narrows what is
    inspected without over-blocking.

  - a native Claude Edit (no `path` field) must still block on file_path alone.
    normalizeKimiPayload returns early for non-Kimi tool names, and this pins
    that the non-Kimi contract did not move. It passes both pre- and post-fix
    by design.

Negative-controlled: run against the pre-fix hooks, the four bypass cases and
the decoy control fail, and the native-Claude control passes.

* test(#2547): back the totality claim with property tests over fc.anything()

This PR claims the fix "makes normalization total over the inputs JSON can
express" — a for-all guarantee — while the tests backing it are example-based,
each shape added reactively after a crash was found by hand (the String()
coercion trap was itself found by adversarial review after the first commit
shipped). Example-based tests cannot substantiate a for-all claim; they record
the counterexamples someone happened to think of.

Four properties over fc.anything(), which is exactly the JSON-expressible
domain the claim names:

  (a) totality over any tool_input
  (b) totality over any edit list — the crash surface both #2547 fixes targeted
  (c) `path` always wins over any model-supplied `file_path` (the review blocker
      invariant: a guard reading file_path can never be aimed at a file other
      than the one kimi-cli writes)
  (d) a non-Kimi tool_name passes through untouched — the native Claude contract

normalizeKimiPayload is inlined per hook with no runtime binding, so there is
nothing to require. The block is extracted from hook source and evaluated via
the SAME extraction contract kimi-guard-normalization-parity.test.cjs uses, so
a source edit that breaks one breaks both instead of silently testing a stale
block. An extraction floor test fails loudly if the extraction yields a no-op.

Non-vacuous, and checked rather than assumed: against pristine pre-#2547 `next`,
(a), (b) and (c) all FAIL and (d) passes. (a) needed the fix that makes it
meaningful — a bare fc.anything() for tool_input passed even against the live
defect, because arbitrary generation essentially never invents the `edit` key
the crash lives behind, so the generator is biased onto the keys normalization
actually reads and unioned back with unbiased input.

* chore(#2547): cover the shadowing vector in the changeset and regen goldens

Golden install-parity churn is hash-only, on exactly the five hook files this
round changed. gsd-phase-boundary.sh is deliberately unchanged.

* test(#2547): make the property test able to kill the coercion mutant

Review Major 1: the generative test added to stop the NEXT counterexample
could not kill the one it was written for. Reproduced the reviewer's matrix
independently — against the shipped generator, a mutant reverting `editText`
to the unguarded `String(v ?? '')` passed all four properties.

Cause, confirmed by measurement: the edit-array ENTRIES were bare
`fc.anything()`, which essentially never invents an `old`/`new` key, so
`e?.old` was always undefined and `String(undefined ?? '')` never coerced
anything. That is the same vacuity the file's own comment describes one level
up, reproduced one level down.

The review's prescribed fix — bias the entry onto `{old, new}` — is necessary
but NOT sufficient, and this is the part worth recording: measured over 20,000
draws, bare `fc.anything()` yields a non-coercible value 3 times (0.015%). At
`numRuns: 200` an `old` key holding a hostile value essentially never
co-occurs, and the mutant survives the entry bias too. Both levels need
biasing — the entry onto the keys normalization reads, and the VALUE onto the
shape that actually throws.

`{"toString": <non-function>}` is that shape and stays inside the
"JSON-expressible" domain the claim names (JSON.parse produces it verbatim);
`fc.anything({withNullPrototype: true})` would also kill the mutant but widens
the domain past what the PR asserts, so it is not used.

Verified: M1 now dies at every seed tried (1/7/42/99/4242/31337, failing
within 3-31 cases) while HEAD stays green at all of them.

Also closes three coverage gaps the review listed as nits — properties (e)
totality over any JSON value as the WHOLE payload, (f) the tool_output →
tool_response mapping (including that an existing tool_response is not
clobbered), and (g) an empty edit list reconstructing nothing.

Property (e) required a one-line fix in the normalizer itself: `JSON.parse
('null')` is null, and null/primitive payloads threw on the `data.tool_name`
read — falsifying the "total over the inputs JSON can express" claim. Harmless
in practice (the throw landed in the same fail-open catch as the exit 0 it now
takes deliberately), but the claim should be true as stated. Landed
byte-identically across all five copies.

Refs #2547

* fix(#2547): make old_string/new_string authoritative, not merely defaulted

Review Major 2: the shadowing class the prior round's BLOCKER closed for
`file_path` survived one field over. `old_string`/`new_string` were still
reconstructed only `if (input.<field> === undefined)`, so a model-supplied
value won.

The argument for making `path` authoritative applies verbatim here.
kimi-cli's StrReplaceFile schema is `path` + `edit` only
(src/kimi_cli/tools/file/replace.py @ 4a550ef) and carries no
`old_string`/`new_string` at all, so either key appearing in a Kimi payload is
always model-supplied — exactly like `file_path`.

Verified end-to-end against the reviewer's payload: a cross-root write
carrying `new_string: ""` alongside an injected `edit[].new` left
gsd-prompt-guard reading '' and returning at its `if (!content)` guard, so the
injection advisory never fired and the reconstructed content was never
scanned. `new_string: null` behaved identically. Negative-controlled: both
produce empty output against pre-fix source and fire the advisory after.

Chose unconditional reconstruction over the offered `typeof` alternative
deliberately. A type test closes `""`/`null` but leaves the interesting case
open — a benign NON-EMPTY decoy (`new_string: "chore: tidy"`) shadows just as
effectively and passes any type test. The new suite includes that case
specifically; it is what discriminates between the two candidate fixes.

Also pins the kimi-cli SHA in the authoritative-path comment, as requested —
it cited file names with no version while the issue pins 4a550ef.

Landed byte-identically across all five inlined copies; the parity test's
byte-identity assertion holds.

Refs #2547

* fix(#2547): close the non-string file_path crash-to-allow at every read site

Review Major 3: the crash-to-allow was closed only as a side effect of `path`
masking the bad value, while the changeset read as though it were closed
outright. Confirmed both of the review's reachability claims: `[]`/`{}`/`42`
are truthy, survive the `!rawFilePath` early-out, and throw inside
path.isAbsolute() into the outer `catch { process.exit(0) }`; and normalization
returns early for native Claude Code payloads (KIMI_TOOL_NAMES has no 'Edit'
entry), so `{"tool_name":"Edit","tool_input":{"file_path":[]}}` reached it
untouched — this guard's original #260 surface.

Reproduced on a real fixture: string cross-root path exits 2, the identical
payload with `[]` or `{}` exits 0.

Swept the class rather than the instance. Five more untyped read sites across
four other hooks, each one line from a type-strict or method-dependent call.
Census of what each can actually do:

  gsd-worktree-path-guard.js:173  BLOCKS  -> live bypass (the review's finding)
  gsd-prompt-guard.js:128         scanner -> silenced the injection scan, the
                                             same outcome as Major 2 by another
                                             route; verified empirically
  gsd-workflow-guard.js:206       advisory only (its exit-2 is the Bash
                                             force-add path, which reads
                                             `command`, not `file_path`)
  gsd-read-guard.js:141           advisory only
  gsd-read-injection-scanner.js:213  advisory only
  gsd-windsurf-pre-write.js:75    ALREADY TYPED — the shape adopted here

All six now read typed. The workflow-guard site keeps its truthiness fallback
(`(typeof x === 'string' && x) || ...`) because a bare type test would let an
empty `file_path` shortcut the `path` fallback.

Also declares one swept hit NOT fixed: `gsd-workflow-guard.js:175` reads
`command` untyped on a genuinely blocking path. Same shape, but not
exploitable — unlike file_path/path there is no second field carrying the
executable value, so a non-string command cannot smuggle a real `git add -f`
past the block. Left alone rather than widen this PR into the Bash path.

The regression gate is a SOURCE-level invariant, not a behavioural one, and
that is deliberate: the fixed read and the crashing read are black-box
identical — both end at exit 0, one via the catch and one via the early-out.
A test asserting exit 0 on a non-string payload passes against the unfixed
code, which is the same false-green the review flagged in the existing
`['non-string file_path (array)', []]` cases. Repeating it one level up would
be no better. tests/kimi-guard-typed-payload-reads.test.cjs fails if any hook
regresses to an untyped read (negative-controlled: it reports all five
pre-fix sites with correct file:line).

The behavioural cases requested — non-string file_path with NO `path` key —
are added to worktree-safety.test.cjs and labelled honestly as documenting the
explicit fail-open rather than detecting a revert.

Also states the relative-path premise (review Minor 5) at the early-out that
depends on it: "always safe" holds only while every runtime reaching there
resolves relative paths against the tool CWD. Claude Code satisfies it by
requiring absolute paths; kimi-cli's resolution behaviour is NOT verified here
and is recorded as an unverified premise rather than an asserted bypass.

Refs #2547

* docs(#2547): correct the changeset's closed-claim and fold the misattributed note

Review Major 3 also flagged the fragment: it said the non-string vector "threw
inside that guard's path.isAbsolute() ... `path` now wins outright", which
reads as closed when it was closed only conditionally. Rewritten to state what
is now true — closed unconditionally at all six read sites — and extended with
the Major 2 finding.

Review Minor 6 (the #2547 scope note living in a `pr: 2518` fragment) turns out
to understate the problem. Rendering the changelog and re-parsing it shows the
note is not merely misattributed — it is DROPPED. serializeChangelog emits each
fragment as a single `- ` bullet, and parseChangelog terminates a bullet at the
first non-continuation line, so everything after a blank line is lost on
re-parse. Audited all 44 fragments: exactly one was lossy —
2304-kimi-guard-tool-name.md, losing 656 of 1730 characters, i.e. precisely
that second paragraph. Folding it into this PR's fragment fixes the
attribution and the silent loss together; all 44 now round-trip losslessly.

That same mechanism is why the remaining nit — reformat this fragment's
~2,000-character paragraph for readability — is NOT applied. A paragraph break
or a bullet list would silently truncate the entry at the first blank line
(verified for both). The single-paragraph form is load-bearing under the
current serializer, not an authoring preference. Worth its own issue; noted in
the PR thread rather than worked around here.

Refs #2547

* chore(#2547): regenerate golden install parity after rebase onto next

Rebased onto `next` @ 9138271b (the PR had gone BEHIND by 20 commits; the
review's closing nit asked for it). The replay was CLEAN — no conflicts — and
that is exactly why this commit exists.

These fixtures are one key per installed file, so when the PR pins five hook
entries and the base rewrites others', the two edits land on different lines of
the same JSON. Git merges them silently and correctly AS TEXT while attesting
nothing about whether the merged hashes are still valid. Verified rather than
assumed: per-key equivalence against the old base showed the base had moved 22
of the 27 keys this PR pins in every runtime fixture, and
tests/golden-install-parity.test.cjs failed on 10 runtimes immediately after the
clean rebase. A push without this regen would have gone out red.

Regenerated with `npm run build && npm run gen:golden` under a throwaway
HOME/CLAUDE_CONFIG_DIR (the generators invoke the installer); live-profile
canary clean before and after.

Contamination check: every key differing from the base's committed fixture
resolves to a file this PR actually touches — the five guard hooks, under both
the `hooks/` and `.kimi/hooks/` install layouts, and nothing else. Derived from
the PR's changed-file set rather than a feature-name filter, which is what
would have mislabelled the registration surfaces.

Size baselines re-checked and NOT regenerated: this PR moves no workflow or
agent, and the base's own baselines are current (agent-size-budget,
workflow-size-budget, workflow-size, update-size-baseline all green).

Refs #2547

* chore(#2547): allowlist the field-shadowing security test in the injection scan

The new regression suite tripped the repo's own prompt-injection scan — a test
for the injection scanner setting off the injection scanner.

The fixture has to be a real injection phrase for the test to assert anything:
it is precisely the content gsd-prompt-guard must still scan once a
model-supplied `new_string` can no longer shadow the reconstructed
`edit[].new`. Weakening it to a benign string would make the suite vacuous.

Allowlisted rather than obfuscated, because that is this repo's established
convention for the class — tests/read-injection-scanner.security.test.cjs,
tests/security-prompt-injection.security.test.cjs,
tests/prompt-injection-scan.security.test.cjs and four others carry real
payloads as test DATA and are listed for exactly this reason. Splitting the
literal to dodge the grep would work but would make this one file inconsistent
with its five peers and leave the next reader wondering why.

Verified with the CI invocation itself (`scripts/prompt-injection-scan.sh
--diff upstream/next`): 26 files scanned, 0 findings. The .cjs codebase scan
does not cover tests/ and is unaffected (73 tests green).

Refs #2547

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-28 18:31:40 -04:00

661 lines
26 KiB
JavaScript

// allow-test-rule: source-text-is-the-product
// Workflow .md / agent .md / command .md / reference .md files — their text
// IS what the runtime loads. Testing text content tests the deployed contract.
// Per CONTRIBUTING.md exception matrix.
/**
* Tests for gsd-read-guard.js PreToolUse hook.
*
* The read guard intercepts Write/Edit tool calls on existing files and injects
* advisory guidance telling the model to Read the file first. This prevents
* infinite retry loops when non-Claude models (e.g. MiniMax M2.5 on OpenCode)
* attempt to edit files without reading them, hitting the runtime's
* "You must read file before overwriting it" error repeatedly.
*
* The hook is advisory-only (does not block) so Claude Code behavior is unaffected.
*/
process.env.GSD_TEST_MODE = '1';
const { test, describe, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const { createTempDir, cleanup } = require('./helpers.cjs');
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-read-guard.js');
/**
* Run the read guard hook with a given tool input payload.
* Returns { exitCode, stdout, stderr }.
*/
function runHook(payload, envOverrides = {}) {
const input = JSON.stringify(payload);
// Sanitize all Claude Code detection signals so positive-path tests work
// when the test runner itself is running inside Claude Code (#2344, #2520).
const env = {
...process.env,
CLAUDE_SESSION_ID: '',
CLAUDECODE: '',
CLAUDE_CODE_ENTRYPOINT: '',
CLAUDE_CODE_SSE_PORT: '',
CLAUDE_PROJECT_DIR: '',
...envOverrides,
};
try {
const stdout = execFileSync(process.execPath, [HOOK_PATH], {
input,
encoding: 'utf-8',
timeout: 5000,
stdio: ['pipe', 'pipe', 'pipe'],
env,
});
return { exitCode: 0, stdout: stdout.trim(), stderr: '' };
} catch (err) {
return {
exitCode: err.status ?? 1,
stdout: (err.stdout || '').toString().trim(),
stderr: (err.stderr || '').toString().trim(),
};
}
}
describe('gsd-read-guard hook', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempDir('gsd-read-guard-');
});
afterEach(() => {
cleanup(tmpDir);
});
// ─── Core: advisory on Write to existing file ───────────────────────────
test('injects read-first guidance when Write targets an existing file', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'console.log("hello");\n');
const result = runHook({
tool_name: 'Write',
tool_input: { file_path: filePath, content: 'console.log("world");\n' },
});
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'should produce output');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput, 'should have hookSpecificOutput');
assert.ok(output.hookSpecificOutput.additionalContext, 'should have additionalContext');
assert.ok(
output.hookSpecificOutput.additionalContext.includes('Read'),
'guidance should mention Read tool'
);
});
test('injects read-first guidance when Edit targets an existing file', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook({
tool_name: 'Edit',
tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' },
});
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'should produce output');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput.additionalContext.includes('Read'));
});
// ─── No-op cases: should NOT inject guidance ────────────────────────────
test('does nothing for Write to a new file (file does not exist)', () => {
const filePath = path.join(tmpDir, 'brand-new.js');
// File does NOT exist
const result = runHook({
tool_name: 'Write',
tool_input: { file_path: filePath, content: 'new content' },
});
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'should produce no output for new files');
});
test('does nothing for non-Write/Edit tools', () => {
const result = runHook({
tool_name: 'Bash',
tool_input: { command: 'echo hello' },
});
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '');
});
test('does nothing for Read tool', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'content');
const result = runHook({
tool_name: 'Read',
tool_input: { file_path: filePath },
});
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '');
});
// ─── Error resilience ──────────────────────────────────────────────────
test('exits cleanly on invalid JSON input', () => {
try {
const stdout = execFileSync(process.execPath, [HOOK_PATH], {
input: 'not json',
encoding: 'utf-8',
timeout: 5000,
stdio: ['pipe', 'pipe', 'pipe'],
});
// Should exit 0 silently
assert.equal(stdout.trim(), '');
} catch (err) {
assert.equal(err.status, 0, 'should exit 0 on parse error');
}
});
test('exits cleanly when tool_input is missing', () => {
const result = runHook({ tool_name: 'Write' });
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '');
});
// ─── Guidance content quality ──────────────────────────────────────────
test('guidance message includes the filename', () => {
const filePath = path.join(tmpDir, 'myfile.ts');
fs.writeFileSync(filePath, 'export const foo = 1;\n');
const result = runHook({
tool_name: 'Write',
tool_input: { file_path: filePath, content: 'export const foo = 2;\n' },
});
const output = JSON.parse(result.stdout);
assert.ok(
output.hookSpecificOutput.additionalContext.includes('myfile.ts'),
'guidance should include the filename being edited'
);
});
test('guidance message instructs to use Read tool before editing', () => {
const filePath = path.join(tmpDir, 'target.py');
fs.writeFileSync(filePath, 'x = 1\n');
const result = runHook({
tool_name: 'Edit',
tool_input: { file_path: filePath, old_string: 'x = 1', new_string: 'x = 2' },
});
const output = JSON.parse(result.stdout);
const ctx = output.hookSpecificOutput.additionalContext;
assert.ok(ctx.includes('Read'), 'must mention Read tool');
assert.ok(
ctx.includes('before') || ctx.includes('first'),
'must indicate Read should come before the edit'
);
});
// ─── Build / install integration ───────────────────────────────────────
test('hook is registered in build-hooks.js HOOKS_TO_COPY', () => {
const buildHooksPath = path.join(__dirname, '..', 'scripts', 'build-hooks.js');
const content = fs.readFileSync(buildHooksPath, 'utf8');
assert.ok(
content.includes('gsd-read-guard.js'),
'gsd-read-guard.js must be in HOOKS_TO_COPY so it ships in hooks/dist/'
);
});
test('hook is registered in install.js uninstall hook list', () => {
const installPath = path.join(__dirname, '..', 'bin', 'install.js');
const content = fs.readFileSync(installPath, 'utf8');
assert.ok(
content.includes("'gsd-read-guard.js'"),
'gsd-read-guard.js must be in the uninstall gsdHooks list'
);
});
test('exits cleanly when tool_input.file_path is non-string', () => {
const result = runHook({
tool_name: 'Write',
tool_input: { file_path: 12345, content: 'data' },
});
// file_path is a number — || '' yields '' — hook exits silently
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '');
});
// ─── Claude Code runtime skip (#1984) ─────────────────────────────────
test('skips advisory on Claude Code runtime (CLAUDE_SESSION_ID set)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDE_SESSION_ID: 'test-session-123' }
);
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'should produce no output on Claude Code');
});
});
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-2344-read-guard-claudecode-env.test.cjs — consolidation epic #1969 (B6 #1975)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-2344-read-guard-claudecode-env (consolidation epic #1969 B6 #1975)", () => {
/**
* Regression test for bug #2344
*
* gsd-read-guard.js checked process.env.CLAUDE_SESSION_ID to detect the
* Claude Code runtime and skip its advisory. However, Claude Code CLI exports
* CLAUDECODE=1, not CLAUDE_SESSION_ID. The skip never fired, so the
* READ-BEFORE-EDIT advisory injected on every Edit/Write call inside Claude
* Code — producing noise in long-running sessions.
*
* Fix: check CLAUDECODE (and CLAUDE_SESSION_ID for back-compat) before
* emitting the advisory.
*/
process.env.GSD_TEST_MODE = '1';
const { test, describe, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const { createTempDir, cleanup } = require('./helpers.cjs');
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-read-guard.js');
function runHook(payload, envOverrides = {}) {
const input = JSON.stringify(payload);
const env = {
...process.env,
CLAUDE_SESSION_ID: '',
CLAUDECODE: '',
CLAUDE_CODE_ENTRYPOINT: '',
CLAUDE_CODE_SSE_PORT: '',
CLAUDE_PROJECT_DIR: '',
...envOverrides,
};
try {
const stdout = execFileSync(process.execPath, [HOOK_PATH], {
input,
encoding: 'utf-8',
timeout: 5000,
stdio: ['pipe', 'pipe', 'pipe'],
env,
});
return { exitCode: 0, stdout: stdout.trim(), stderr: '' };
} catch (err) {
return {
exitCode: err.status ?? 1,
stdout: (err.stdout || '').toString().trim(),
stderr: (err.stderr || '').toString().trim(),
};
}
}
describe('bug #2344: read guard skips on CLAUDECODE env var', () => {
let tmpDir;
beforeEach(() => { tmpDir = createTempDir('gsd-read-guard-2344-'); });
afterEach(() => { cleanup(tmpDir); });
test('skips advisory when CLAUDECODE=1 is set (Claude Code CLI env)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDECODE: '1' }
);
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'advisory must not fire when CLAUDECODE=1');
});
test('skips advisory when CLAUDE_SESSION_ID is set (back-compat)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDE_SESSION_ID: 'test-session-123' }
);
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'advisory must not fire when CLAUDE_SESSION_ID is set');
});
test('still injects advisory when neither CLAUDECODE nor CLAUDE_SESSION_ID is set', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDECODE: '', CLAUDE_SESSION_ID: '' }
);
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'advisory should fire on non-Claude-Code runtimes');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput?.additionalContext?.includes('Read'));
});
});
});
}
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-2520-read-guard-hook-subprocess-env.test.cjs — consolidation epic #1969 (B6 #1975)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-2520-read-guard-hook-subprocess-env (consolidation epic #1969 B6 #1975)", () => {
/**
* Regression test for bug #2520
*
* The fix for #2344 added `|| process.env.CLAUDECODE` to the Claude Code
* skip check. That works in principle — CLAUDECODE=1 is propagated to Bash
* tool subprocesses — but it does NOT reach hook subprocesses on Claude Code
* v2.1.116. Claude Code applies a separate env filter when spawning
* PreToolUse hook commands; that filter drops bare CLAUDECODE and
* CLAUDE_SESSION_ID and keeps only CLAUDE_CODE_*-prefixed vars plus
* CLAUDE_PROJECT_DIR. `data.session_id` is, however, reliably delivered via
* the hook's stdin JSON payload (documented part of Claude Code's hook
* input schema).
*
* Fix: use `data.session_id` as the primary Claude Code signal, with
* CLAUDE_CODE_ENTRYPOINT / CLAUDE_CODE_SSE_PORT as env-var fallbacks, and
* keep legacy CLAUDECODE / CLAUDE_SESSION_ID for back-compat and
* future-proofing.
*/
process.env.GSD_TEST_MODE = '1';
const { test, describe, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const { createTempDir, cleanup } = require('./helpers.cjs');
const HOOK_PATH = path.join(__dirname, '..', 'hooks', 'gsd-read-guard.js');
/**
* Spawn the hook with an env that mirrors the actual Claude Code hook
* subprocess env: CLAUDECODE and CLAUDE_SESSION_ID are stripped, only
* CLAUDE_CODE_*-prefixed vars (plus CLAUDE_PROJECT_DIR) remain. Extra env
* overrides can be supplied via `envOverrides`.
*/
function runHookInClaudeCodeSubprocess(payload, envOverrides = {}) {
const input = JSON.stringify(payload);
const baseEnv = { ...process.env };
// Strip env vars Claude Code does NOT propagate to hook subprocesses.
delete baseEnv.CLAUDECODE;
delete baseEnv.CLAUDE_SESSION_ID;
const env = {
...baseEnv,
// Env vars Claude Code DOES propagate to hook subprocesses (observed on
// Claude Code CLI 2.1.116).
CLAUDE_CODE_ENTRYPOINT: 'cli',
CLAUDE_CODE_SSE_PORT: '51291',
CLAUDE_PROJECT_DIR: process.cwd(),
...envOverrides,
};
try {
const stdout = execFileSync(process.execPath, [HOOK_PATH], {
input,
encoding: 'utf-8',
timeout: 5000,
stdio: ['pipe', 'pipe', 'pipe'],
env,
});
return { exitCode: 0, stdout: stdout.trim(), stderr: '' };
} catch (err) {
return {
exitCode: err.status ?? 1,
stdout: (err.stdout || '').toString().trim(),
stderr: (err.stderr || '').toString().trim(),
};
}
}
describe('bug #2520: read guard detects Claude Code without relying on CLAUDECODE env', () => {
let tmpDir;
beforeEach(() => { tmpDir = createTempDir('gsd-read-guard-2520-'); });
afterEach(() => { cleanup(tmpDir); });
test('skips advisory when stdin payload includes session_id (Claude Code hook-subprocess env)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
// Isolate the stdin `session_id` signal by clearing the CLAUDE_CODE_*
// env fallbacks the helper normally provides. Without this the env
// fallback would rescue the skip even if session_id detection broke,
// hiding a regression of the primary signal.
const result = runHookInClaudeCodeSubprocess(
{
session_id: 'e7123e54-0977-45dd-848a-b9c8a45a5cd3',
tool_name: 'Edit',
tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' },
},
{ CLAUDE_CODE_ENTRYPOINT: '', CLAUDE_CODE_SSE_PORT: '', CLAUDE_PROJECT_DIR: '' },
);
assert.equal(result.exitCode, 0);
assert.equal(
result.stdout,
'',
'advisory must not fire when session_id is present on stdin (real Claude Code hook env)',
);
});
test('skips advisory when CLAUDE_CODE_ENTRYPOINT is set (env-var fallback, no session_id on stdin)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHookInClaudeCodeSubprocess(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDE_CODE_ENTRYPOINT: 'cli', CLAUDE_CODE_SSE_PORT: '' },
);
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'advisory must not fire when CLAUDE_CODE_ENTRYPOINT is set');
});
test('still injects advisory when no Claude Code signal is present (non-Claude host)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHookInClaudeCodeSubprocess(
{ tool_name: 'Edit', tool_input: { file_path: filePath, old_string: 'const x = 1;', new_string: 'const x = 2;' } },
{ CLAUDE_CODE_ENTRYPOINT: '', CLAUDE_CODE_SSE_PORT: '', CLAUDE_PROJECT_DIR: '' },
);
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'advisory should fire on non-Claude-Code hosts');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput?.additionalContext?.includes('Read'));
});
});
});
}
// ────────────────────────────────────────────────────────────────────────
// #2304 — Kimi tool vocabulary engages the read guard
// ────────────────────────────────────────────────────────────────────────
describe('#2304: Kimi tool vocabulary is normalized by the read guard', () => {
// Payload shapes mirror kimi-cli's actual tool schemas
// (src/kimi_cli/tools/file/{write,replace}.py): WriteFile takes
// `path`/`content`, StrReplaceFile takes `path` + `edit: Edit | list[Edit]`.
//
// SCOPE (#2547 finding 3): these cases omit `session_id`, so what they prove
// is that normalizeKimiPayload maps the Kimi tool VOCABULARY through to the
// Write/Edit branch — not that the advisory fires on a live Kimi turn. Every
// real kimi-cli payload carries a non-empty `session_id` (hooks/events.py
// `_base()` sets it unconditionally; kimisoul.py calls `set_session_id()` at
// the top of every turn), and the guard treats any non-empty `session_id` as
// "this is Claude Code, skip". Do NOT read a green here as evidence of
// production behaviour — the '#2547' describe below pins what actually
// happens against the production shape.
let tmpDir;
beforeEach(() => { tmpDir = createTempDir('gsd-read-guard-2304-'); });
afterEach(() => { cleanup(tmpDir); });
test('WriteFile on an existing file injects read-first guidance like Write', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'console.log("hello");\n');
const result = runHook({
tool_name: 'WriteFile',
tool_input: { path: filePath, content: 'console.log("world");\n' },
});
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'Kimi WriteFile should produce the advisory');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput?.additionalContext?.includes('Read'));
});
test('StrReplaceFile on an existing file injects guidance like Edit', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook({
tool_name: 'StrReplaceFile',
tool_input: { path: filePath, edit: { old: 'const x = 1;', new: 'const x = 2;' } },
});
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'Kimi StrReplaceFile should produce the advisory');
const output = JSON.parse(result.stdout);
assert.ok(output.hookSpecificOutput?.additionalContext?.includes('Read'));
});
test('module-qualified kimi_cli.tools.file:WriteFile is recognized', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'content\n');
const result = runHook({
tool_name: 'kimi_cli.tools.file:WriteFile',
tool_input: { path: filePath, content: 'replacement\n' },
});
assert.equal(result.exitCode, 0);
assert.ok(result.stdout.length > 0, 'module-qualified Kimi WriteFile should produce the advisory');
});
test('Kimi ReadFile stays out of scope (silent exit)', () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'content\n');
const result = runHook({
tool_name: 'kimi_cli.tools.file:ReadFile',
tool_input: { path: filePath },
});
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '', 'ReadFile is not a write tool — guard must stay silent');
});
});
// ────────────────────────────────────────────────────────────────────────
// #2547 finding 3 — the production Kimi payload shape (session_id present)
// ────────────────────────────────────────────────────────────────────────
describe('#2547: read guard against the production Kimi payload shape', () => {
// The #2304 cases above omit `session_id`. A live Kimi turn never does:
// - src/kimi_cli/hooks/events.py `_base()` returns
// {"hook_event_name", "session_id", "cwd"} — the field is unconditional;
// - src/kimi_cli/soul/kimisoul.py calls `set_session_id(session.id)` at the
// top of every turn, before tool dispatch, so the ContextVar holds a real
// UUID (its `default=""` only applies outside a turn).
//
// The guard's Claude Code check treats ANY non-empty `data.session_id` as
// "Claude Code already enforces read-before-edit, skip" (#2520), so on Kimi
// the advisory is DORMANT. These tests pin that real behaviour, which is what
// makes the #2304 cases above trustworthy as vocabulary-only coverage.
//
// This is a CHARACTERIZATION of a known gap, not an endorsement of it.
// Redesigning the runtime discrimination is explicitly OUT OF SCOPE for
// #2547. If a later change makes the advisory fire on Kimi, these tests are
// SUPPOSED to fail — update them then, rather than deleting the coverage.
let tmpDir;
beforeEach(() => { tmpDir = createTempDir('gsd-read-guard-2547-'); });
afterEach(() => { cleanup(tmpDir); });
const LIVE_SESSION_ID = 'e7123e54-0977-45dd-848a-b9c8a45a5cd3';
for (const [label, toolInput, toolName] of [
['StrReplaceFile', (p) => ({ path: p, edit: { old: 'const x = 1;', new: 'const x = 2;' } }), 'StrReplaceFile'],
['WriteFile', (p) => ({ path: p, content: 'replacement\n' }), 'WriteFile'],
['module-qualified WriteFile', (p) => ({ path: p, content: 'replacement\n' }), 'kimi_cli.tools.file:WriteFile'],
]) {
test(`${label} with a populated session_id is dormant (known gap, #2547)`, () => {
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const result = runHook({
session_id: LIVE_SESSION_ID,
tool_name: toolName,
tool_input: toolInput(filePath),
});
assert.equal(result.exitCode, 0);
assert.equal(result.stdout, '',
'The advisory is currently skipped on Kimi because the guard reads any ' +
'non-empty session_id as Claude Code. If this now emits, the runtime ' +
'discrimination changed — update this test and the #2304 block above.');
});
}
test('the ONLY difference is session_id — dropping it makes the same payload fire', () => {
// The false-green proof, asserted rather than described: one field flips the
// #2304 cases from firing to silent, and the live shape is the silent one.
const filePath = path.join(tmpDir, 'existing.js');
fs.writeFileSync(filePath, 'const x = 1;\n');
const toolInput = { path: filePath, edit: { old: 'const x = 1;', new: 'const x = 2;' } };
const withoutSession = runHook({ tool_name: 'StrReplaceFile', tool_input: toolInput });
const withSession = runHook({
session_id: LIVE_SESSION_ID,
tool_name: 'StrReplaceFile',
tool_input: toolInput,
});
assert.ok(withoutSession.stdout.length > 0,
'test-shape payload (no session_id) fires — this is what #2304 asserts');
assert.equal(withSession.stdout, '',
'production-shape payload (session_id present) is silent — so a green in ' +
'the #2304 block is evidence about vocabulary, not about production');
});
});