Files
msd-core/tests/kimi-guard-normalization-parity.test.cjs
0xdhx c61dd49d95 enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes

Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks
(decision: 'block', exit 2) a whole-file Write collapsing a curated
.planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md,
STATE.md) below 40% of its on-disk line count. Files under 40 lines
are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message)
bypasses for legitimate milestone resets.

Fix 3 of #973 — the only defense independent of per-agent tool config.
Registered on the Claude plugin surface (hooks.json), settings-json
runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi
spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and
INVENTORY regenerated; regression tests negative-controlled (16/16
RED with the hook absent, 16/16 GREEN with it present).

* chore(#2255): backfill changeset pr number to 2301

* enhance(#2255): address review — fail-closed reads, typed block output, registration, property test

Review fixes for trek-e's CHANGES_REQUESTED on PR #2301:

- Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES
  (no-shipping-drift test).
- Blocker 3: update the always-on hook enumerations in ADR-766 and
  CONTEXT.md from six to seven.
- Major 4: fail CLOSED on non-ENOENT read errors — only a missing file
  (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a
  typed readError field and the override still honored. Tested, with a
  negative control against the pre-fix hook.
- Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES
  budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the
  floor; sub-floor always exempt), boundary examples pinned.
- Major 6: block output now carries typed oldLines/newLines/
  overrideEnvVar fields; tests assert on those instead of regexing the
  free-form reason string.
- Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS
  bypass on macOS/Windows); limit+1 boundary tests added for both the
  floor and the ratio.

* enhance(#2255): engage the write guard on Kimi's native payload shape

The guard shipped with Claude-vocabulary checks (tool_name 'Write',
tool_input.file_path), which #2304 showed leaves a guard dormant on
Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli
forwards its native payload verbatim — tool_name 'WriteFile' (bare or
module-qualified) and tool_input.path per its tool schemas
(src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown
name, and exited 0.

Apply the same per-guard normalization PR #2326 gives the three
sibling guards (name + field mapping, inlined — hook scripts stage as
standalone files), and write the block reason to stderr as well as
stdout JSON: Kimi feeds stderr, not stdout, back to the model on
exit 2, so a stdout-only reason blocks without telling the model why
or naming the documented override.

Regression tests pipe Kimi-shaped payloads (engage, qualified-name,
stderr-reason) plus exemption pins (StrReplaceFile stays out of scope
by design; non-curated paths pass) — verified red against the pre-fix
guard, green after.

* enhance(#2255): rebase onto next; regenerate golden-parity fixtures

* enhance(#2255): wire the escape hatch into complete-milestone's reorganize step

Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP
reorganize — the tree's only legitimate milestone reset and the exact caller
GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the
rewrite through a shell write with the hatch set on the command (a hook
inherits the runtime env, so a bare Write cannot carry a per-step override),
and a binding test derives the env var name from the guard's typed output and
asserts (a) the workflow step sets it and (b) the guard passes the identical
catastrophic payload under it — so the next complete-milestone.md edit cannot
silently re-break the wiring.

* enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload

Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string
reconstruction were unreachable-by-effect — the guard exits 0 for any
tool_name !== 'Write', so nothing ever read the fields they set, leaving
guaranteed-surviving mutants against the Stryker bar. The map now carries
only WriteFile -> 'Write'; the StrReplaceFile exemption test message states
the fall-through it actually exercises.

* enhance(#2255): review minors — American spellings; writeSync before exit(2)

Minor 1: normalised/normalise -> American house style. Minor 2: the two
block paths wrote stdout+stderr via async pipe writes then exit(2) —
async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block
payload durable.

* enhance(#2255): assert stderr equals the typed reason, not raw prose

Minor 3: the last raw-text match in the suite pinned override-name prose on
stderr. The contract is "stderr carries the reason Kimi feeds back" — now
asserted as stderr non-empty and byte-equal to the parsed stdout.reason.

* enhance(#2255): bind the write-guard's Kimi normalization into the parity test

Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with
nothing binding it. This extends PR #2326's kimi-guard-normalization-parity
test (same path and helpers, authored as a superset so either merge order
resolves cleanly): sibling byte-parity is existence-gated zero-or-all —
trivially green until #2326 lands, full-strength after — and the write-guard
copy is bound semantically (map is the value-inverse of convertKimiToolName;
the Kimi name for Write must map, or the guard is dormant on Kimi; the
path -> file_path half must be present). Byte-parity is deliberately not
asserted for this copy: it legitimately omits the Edit-class mapping
(Major 1 — dead code in a Write-only guard).

* enhance(#2255): refresh golden-parity fixtures for revised guard + workflow

* chore(#2255): regenerate golden fixtures after rebase onto next

The committed fixture hashes were generated against a tree predating
next's latest 11 commits, which independently modified the same
install-parity surface. Rebased onto next and regenerated with
`npm run gen:golden`.

Verified: against upstream/next the regenerated fixtures differ by
exactly this PR's own entries -- hooks/gsd-write-guard.js (new),
hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and
gsd-core/workflows/complete-milestone.md. No unrelated drift.

* fix(#2255): regenerate workflow size baseline for complete-milestone

`complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2
review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step,
but tests/workflow-size-baseline.json was never regenerated. The
per-file workflow baseline test (issue #1074) failed on
ubuntu-latest/22 and both macOS shard 1/3 jobs.

The growth is justified: it is the escape-hatch binding requested in
review round 2 (the guard must not hard-block the tree's only
legitimate milestone reset), not incidental bloat.

Regenerated via `npm run size:baseline`; the diff is exactly the one
entry.

* chore(#2255): regenerate golden fixtures and size baseline after rebase onto next

* enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes

Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert
(no PreToolUse hook exists on Bash in this family; the write succeeded by
dodging the guard, not by the override firing) and the protection was prose.
The hatch is now a transport code consults: complete-milestone's reorganize
step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the
Write tool as the sanctioned path, and the guard — at the block point only —
verifies the sentinel is fresh (15 min) and names the pending target, then
CONSUMES it and allows that one write. Path-bound + single-use + freshness
keep it from becoming a standing unlock. The env var remains as the
interactive transport, where it can actually reach the hook.

Regression tests written first (negative control: 3 failed pre-fix): the
armed-sentinel Write passes and consumes; stale does not exempt; a token for
a different file neither exempts nor is consumed; the binding test now takes
the sentinel name from the guard's typed output (overrideSentinel), asserts
the step arms it, and asserts the step no longer routes the rewrite around
Write via a shell pipe.

Also in this commit, same file:
- m2: block emission is exception-safe — emitBlock() wraps both writeSync
  sites in their own try/catch that still exits 2, so an EPIPE can no longer
  convert fail-closed into the outer catch's fail-open.
- Header discloses the two reviewed design limits (cumulative sequential
  shrink; lexical match vs symlinked paths) per round-5 scoping.

* docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4)

USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the
runtime-hooks-surface registration comment, and the changeset now describe
both hatches — the single-use sentinel for workflow steps and the env var
for interactive use — instead of implying a per-step env can reach a hook.
The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)`
parenthetical the repo's fragments use (round-5 m4).

* chore(#2255): regenerate derived families on the rebased tree (full sweep)

Full generator sweep after rebasing onto next @ the body-parser-patched
lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every
regen delta verified to be either a PR-owned entry (gsd-write-guard.js,
complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's
committed value for entries our arbitrary-side conflict resolution had left
stale (all 18 runtime fixtures checked mechanically).

* test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync

The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY
retry budget; the sentinel disarm now rides it like every other teardown.

* chore(#2255): regenerate derived families after rebase onto next

Full sweep on the rebased tree (build -> gen-inventory-manifest ->
gen:golden -> size:baseline). Every delta is either a PR-owned entry
(hooks/gsd-write-guard.js, its registration surfaces
hooks/managed-hooks-registry.cjs and the two plugin buses,
gsd-core/workflows/complete-milestone.md) or exact convergence to
next's committed value across all 18 runtime fixtures.

* chore(#2255): regenerate derived families after rebase onto next @ a5180d96

Rebase onto current `next` (a5180d96) resolved 12 conflicting
golden-install-parity fixtures; all regenerated via the full generator
sweep (build, gen:golden, size:baseline) rather than a single generator.

`lint:generated-sync` reports every generated artifact in sync. All 45
differing fixture keys and the single workflow-size-baseline entry map
to files this PR actually touches; no foreign drift.

* fix(#2255): remove the stale unguarded reorganize_roadmap step (round-8 blocker)

complete-milestone.md carried a second ROADMAP-collapsing step,
`reorganize_roadmap`, distinct from the sentinel-armed
`reorganize_roadmap_and_delete_originals` this PR wired. It is a vestige
of the pre-archive-then-reorganize design: it sits BEFORE
archive_milestone, so executing it as written would collapse ROADMAP.md
before the archive snapshots the full phase detail — and its Write is
exactly the shape gsd-write-guard hard-blocks, with no hatch armed. The
file's own success criteria describe only one reorganize outcome
(Backlog-preserving, overwrite-in-place — the later step's properties),
and archive_milestone points forward to "the reorganize step".

Removed rather than wired, per the round-8 review's confirm-and-remove
option. A new binding test asserts the sentinel-armed step is the ONLY
reorganize step in the workflow, so an unguarded collapse step cannot be
silently reintroduced (negative-controlled: fails against the pre-fix
tree). Golden-parity fixtures and the size baseline regenerate for the
shrunk file; every changed fixture key is complete-milestone.md's own.

* test(#2255): document why the read-error injection is a path collision, not an fs monkeypatch

Round-8 nit: the non-ENOENT tests inject via a directory-at-target-path
collision instead of the repo's fs-method monkeypatch pattern. That is
deliberate, not drift — runHook exercises the hook as a spawnSync child
process, so an in-process fs.readFileSync patch (the pattern the cited
siblings use on require'd, in-process code) can never reach the code
under test. Record the reasoning at the injection site.

* chore(#2255): regenerate derived families after rebase onto next @ 0d08c320

Rebase onto current next (0d08c320) for the CONFLICTING/DIRTY state. All 32
conflicts were generated artifacts (19 golden-install-parity, 12 install-tree,
workflow-size-baseline); resolved arbitrarily and regenerated via a full
generator sweep (build, gen:golden, size:baseline, gen-inventory-manifest)
rather than hand-merged. No source conflicts.

Regen diff verified against the PR's changed-file set: 7 distinct differing
keys, all PR-owned (gsd-write-guard.js, managed-hooks-registry.cjs,
plugins/gsd-core.js, complete-milestone.md, and their .kimi mirrors).
lint:generated-sync clean.

* chore(#2255): regenerate derived families after rebase onto next @ 9138271b

Conflict set was 20 paths, every one a generated artifact, zero source
conflicts — resolved arbitrarily during the replay and regenerated here,
per the maintainer's round-9 recipe (never hand-merged).

Generator sweep (not just gen:golden): npm run build, gen:golden,
size:baseline, gen-inventory-manifest, gen:registry. INVENTORY-MANIFEST
came back byte-identical, so the merged value was already correct.

Regen diff verified == PR-touched entries: every differing leaf key
attributes to a file this PR changes (complete-milestone.md, the guard,
managed-hooks-registry.cjs, both plugin buses) — derived from the
changed-file set, not a feature-name filter.

* fix(#2255): port the #2595 Kimi-payload hardening the base landed for the sibling guards

Self-found via the round-10 base sweep: a8b40fa5 (#2547/PR #2595) hardened
gsd-prompt/read/worktree-path-guard's normalizeKimiPayload after this PR was
authored, and this guard carried all three pre-fix shapes:

- `path` now wins outright over a model-supplied `file_path` — kimi-cli's
  WriteFile schema carries no `file_path`, so a payload pairing a curated
  `path` with a spurious `file_path: ""` had the guard reading '' and
  exiting 0 while kimi-cli wrote to `path` (one-key bypass, no crash needed).
- null/primitive payloads pass through deliberately instead of throwing into
  the fail-open catch.
- the `file_path` read is typed (`typeof === 'string'`), closing the
  crash-to-allow on `[]`/`{}` — this is what the base-added
  kimi-guard-typed-payload-reads scan flags.

Negative-controlled: the shadowing payload exits 0 against the pre-fix guard
and 2 against this one.

* fix(#2255): round 9 Major 2 — stop printing the sentinel recipe in the agent-facing denial

The block message ended by telling the blocked agent exactly how to arm the
single-use sentinel — handing the #973 actor a mechanical self-authorization
in the moment it is blocked. The recipe is gone from the prose; the
corrective advice and the env override's name stay (the latter is a #2255
acceptance criterion, and a per-step env prefix cannot reach a hook anyway),
and the typed overrideSentinel field stays for the binding tests. The hatch
remains documented in USER-GUIDE.md and complete-milestone.md, where humans
and the workflow engine read.

* fix(#2255): round 9 Minors 1-2 — realpath-resolve the target before the curated match; disclose the /i Linux cost

Minor 1: a Write to a non-curated path that symlinks into a curated file was
not matched while writeFileSync followed the link — the target is now
realpath-resolved before the curated match (ENOENT keeps the lexical
resolution so new-file Writes still pass; any other realpath error falls
through to the read, which fails closed). Negative-controlled: the symlink
payload exits 0 against the pre-fix guard, 2 against this one. Test skips on
win32, where symlink creation needs privilege.

Minor 2: the header's design-limits block now names the unconditional /i
cost on case-sensitive Linux (a genuinely distinct .planning/roadmap.md is
also treated as curated) next to the stateless limit, and drops the closed
symlink limit.

* test(#2255): round 9 Minors 3-4 — CRLF counting pin + a passing Write leaves a fresh sentinel unburned

Minor 3: countLines' split('\n') is CRLF-safe for a count (the \r rides
along), confirmed by trace in the review — this pins it against this repo's
recurring CRLF regressions, on both sides of the compare and at the 40%
boundary.

Minor 4: consumeSentinelFor runs only after the ratio check would block, so
a within-tolerance Write never burns the workflow's token — true by
construction, previously un-asserted.

* fix(#2255): round 9 Major 3 — correct the stale env-var line in archive_milestone's summary

complete-milestone.md's "After archival" bullet still said the reorganize
happens "under GSD_ALLOW_PLANNING_SHRINK=1" — the wording from the round-2
design this PR's own history rejected in round 5 (a per-step env var cannot
reach a hook; setting it in a Bash step silently does nothing). It now points
at the sentinel mechanics the reorganize step actually documents, matching
that step and USER-GUIDE.md.

* docs(#2255): round 9 Major 1 — user-facing docs state the stateless per-Write limit

The changeset and USER-GUIDE described the guard as covering "catastrophically
shrinks" with no caveat, while the stateless design was disclosed only in the
hook header — an operator reading the shipped docs would conclude iterative
erosion is covered. Both surfaces now state the per-Write comparison and the
erosion non-goal explicitly, in line with what the guard does.

* fix(#2255): round 10 Blocker 1 — arm the hatch in the roadmapper's curated writes

gsd-roadmapper Step 7 Writes both .planning/ROADMAP.md and .planning/STATE.md
wholesale, and /gsd:new-milestone spawns it against the OUTGOING milestone's
files. new-milestone's `phases.clear` archives phase DIRECTORIES, never
ROADMAP.md, so nothing compacts it first, and no ordering rule forces
/gsd:complete-milestone to run before /gsd:new-milestone.

Measured against the shipped hook at the #973 file size (292 lines): a new
4-phase roadmap lands at 18.2% and an 8-phase one at 31.8% — both blocked; only
a 12-phase replacement (45.5%) clears. So the collapse is reachable on a
first-party flow, which is the failure mode #2255 para 4 says the guard must not
have.

Arms the single-use sentinel immediately before each Write, one per target
(the token is path-bound and single-use, so one arming cannot cover both), and
gates each arming on `[ -f ]` so the /gsd:new-project path — where the guard
already exempts via ENOENT — strands no unconsumed token.

Binding test derives the sentinel name from the guard's own typed output, so a
rename on either side fails the suite instead of silently unwiring the hatch.
Negative-controlled: 2 assertions fail against the pre-fix roadmapper.

* docs(#2255): round 10 Major 2 — scope the claim to the guarantee that ships

The sentinel is a plain file, so an agent that would reason past an advisory
can arm one with a single Bash call it is already permitted to make. The
"only defense independent of per-agent tool config" framing outran that, and
the changeset was on its way into CHANGELOG.md.

Retitles the claim on all three surfaces (changeset, guard header, USER-GUIDE)
to what the guard actually delivers: it blocks accidental and single-shot
collapse and is not a defense against a determined agent; what it converts is
"ignore a sentence" into "take one deliberate, path-bound, single-use,
auditable action".

Pinned by test on the DURABLE surfaces only — the guard header and USER-GUIDE.
The changeset fragment is deliberately not pinned: it is consumed at release,
so a test reading it would start failing the moment the release lands. The
bound-statement assertion normalizes comment markers and whitespace first, so
it pins the claim rather than the paragraph's line wrapping.

Negative-controlled: both assertions fail against the pre-fix surfaces.

* test(#2255): acknowledge the roadmapper growth from the round 10 Blocker 1 wiring

The emitted-attribution gate (#2719/#2767) flags gsd-roadmapper.md growing 1130
bytes without an acknowledgment. The growth is the Blocker 1 sentinel wiring
plus the rationale a future editor needs to keep it, so it gets an ack fragment
rather than a silencing regen — the gate's own message is explicit that there is
nothing left to regenerate.

Fragment is PR-scoped (2301-…) per the gate's naming instruction, and uses the
plain-string reason form the shipped fragments use.

Verified against the TRUE upstream tip, not the fork's origin/next: a stale
origin made this same gate report unrelated phantom drift (1 emitted path + 6
grown files + 5 stale acks) that vanishes when GSD_EMITTED_BASE is pinned.

* test(#2255): renumber the roadmapper PROSE_ALLOWLIST pin after the Step 7 wiring

CI red on shard 2/3, all four platforms. The #2751 gate keys PROSE_ALLOWLIST on
{file, line}; the Blocker 1 wiring added 18 lines above the allowlisted
parenthetical in agents/gsd-roadmapper.md, moving it 624 -> 642. Both halves of
the gate then fired: the moved line reads as a new offender, and the stale
entry no longer matches anything.

Line content at 642 is byte-identical to what the entry describes — a
descriptive "e.g." naming SDK queries a user could run — so this is a
renumber, not a re-classification.

Swept the defect class rather than the instance: agents/gsd-roadmapper.md is
the only line-pinned reference to any file this round changed.

Negative-controlled: both assertions fail against the un-renumbered allowlist.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:19:49 -04:00

217 lines
9.9 KiB
JavaScript

// allow-test-rule: source-text-is-the-product #2304 — this test's whole job is
// scanning hooks/*.js source text for the inlined KIMI_TOOL_NAMES copies; the
// text IS the artifact under test (the copies have no runtime binding).
/**
* Kimi guard-normalization parity test (#2304 / PR #2326 review Major 1;
* extended by PR #2301 review Major 2 for hooks/gsd-write-guard.js).
*
* The KIMI_TOOL_NAMES map + normalizeKimiPayload helper is deliberately
* inlined per hook script (a sibling require is a staging dependency that
* can fail silently — see the rationale comment in each guard), which
* leaves five hand-maintained copies plus their inverse in bin/install.js
* (claudeToKimiTools / convertKimiToolName). Nothing at runtime binds them.
*
* This test is that binding, with zero runtime coupling:
* 1. the five inlined copies are byte-identical;
* 2. every entry in each guard map is the value-inverse of what the
* installer's matcher vocabulary emits for that Claude tool;
* 3. every guard-relevant Claude tool the installer translates has a
* reverse entry — so a vocabulary extension or rename that updates
* convertKimiToolName without updating the guards fails HERE instead
* of leaving a guard silently dormant (the #2304 failure mode).
*
* gsd-write-guard.js is bound SEMANTICALLY, not byte-wise: its copy
* intentionally omits the Edit-class mapping (the guard exits 0 for any
* tool but Write, so StrReplaceFile/old_string handling there is dead code
* — #2301 review Major 1), so its map is checked against the installer
* inverse and for Write-dormancy, the only tool it inspects. It is still
* required to CARRY a block, so the copy cannot silently disappear.
*/
process.env.GSD_TEST_MODE = '1';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { convertKimiToolName } = require('../bin/install.js');
// Enumerated by scanning, never hardcoded: a sixth guard added later with its
// own copy of the block must be swept in automatically, or the copies diverge
// exactly the way this test exists to prevent (PR #2326 review M2). The
// dynamic scan is also what makes the no-shared-module decision safe.
const KIMI_MARKER = 'const KIMI_TOOL_NAMES';
const HOOKS_DIR = path.join(__dirname, '..', 'hooks');
const ALL_BLOCK_FILES = fs
.readdirSync(HOOKS_DIR)
.filter((f) => f.endsWith('.js'))
.filter((f) => fs.readFileSync(path.join(HOOKS_DIR, f), 'utf8').includes(KIMI_MARKER))
.map((f) => `hooks/${f}`)
.sort();
// Bound semantically (map-inverse + Write-dormancy), never byte-wise — see header.
const WRITE_GUARD_FILE = 'hooks/gsd-write-guard.js';
// The byte-identity cohort: every scanned guard except the deliberate subset.
const HOOK_FILES = ALL_BLOCK_FILES.filter((f) => f !== WRITE_GUARD_FILE);
// The five guards normalized for #2304. A scan that misses one of these is a
// broken scan, not a passing test — without this floor, an over-narrow filter
// would "pass" by finding nothing to check.
const KNOWN_NORMALIZED_GUARDS = [
'hooks/gsd-prompt-guard.js',
'hooks/gsd-read-guard.js',
'hooks/gsd-read-injection-scanner.js',
'hooks/gsd-workflow-guard.js',
'hooks/gsd-worktree-path-guard.js',
];
// Claude tool names whose PreToolUse/PostToolUse guards are registered with a
// translated matcher on Kimi (runtime-hooks-surface.cts buildKimiHooksTomlBlock):
// the write guards match WriteFile|StrReplaceFile, the injection scanner
// matches ReadFile, and gsd-workflow-guard.js matches Shell|WriteFile|StrReplaceFile.
const GUARD_RELEVANT_CLAUDE_TOOLS = ['Write', 'Edit', 'MultiEdit', 'Read', 'Bash'];
function extractBlock(file) {
const src = fs.readFileSync(path.join(__dirname, '..', file), 'utf8');
const start = src.indexOf('const KIMI_TOOL_NAMES');
assert.notEqual(start, -1, `${file}: KIMI_TOOL_NAMES block not found`);
const endMarker = ' return data;\n}';
const end = src.indexOf(endMarker, start);
assert.notEqual(end, -1, `${file}: normalizeKimiPayload end not found`);
return src.slice(start, end + endMarker.length);
}
function parseMap(block) {
// The guards declare `new Map([['KimiName', 'ClaudeName'], …])` (a Map so
// prototype keys resolve to undefined — review M1); parse the pair list.
const m = block.match(/const KIMI_TOOL_NAMES = new Map\(\[([\s\S]*?)\]\);/);
assert.ok(m, 'KIMI_TOOL_NAMES Map literal not parseable');
const entries = {};
for (const kv of m[1].matchAll(/\['(\w+)', '(\w+)'\]/g)) {
entries[kv[1]] = kv[2];
}
assert.ok(Object.keys(entries).length > 0, 'KIMI_TOOL_NAMES parsed empty');
return entries;
}
function assertMapIsInstallerInverse(map, file) {
for (const [kimiName, claudeName] of Object.entries(map)) {
const modulePath = convertKimiToolName(claudeName);
assert.ok(
typeof modulePath === 'string' && modulePath.endsWith(`:${kimiName}`),
`${file}: KIMI_TOOL_NAMES.${kimiName} -> '${claudeName}' is not the inverse of ` +
`convertKimiToolName('${claudeName}') = ${modulePath}`
);
}
}
describe('Kimi guard normalization parity', () => {
test('the scan finds every known normalized guard (floor — a scan that finds nothing must fail)', () => {
for (const known of KNOWN_NORMALIZED_GUARDS) {
assert.ok(
HOOK_FILES.includes(known),
`${known} carries no '${KIMI_MARKER}' block — either its normalization ` +
'was removed or the scan filter broke; both mean lost coverage'
);
}
});
test('the write guard carries a normalization block (semantically bound, but never absent)', () => {
assert.ok(
ALL_BLOCK_FILES.includes(WRITE_GUARD_FILE),
`${WRITE_GUARD_FILE} carries no '${KIMI_MARKER}' block — the shrink guard ` +
'is silently dormant on Kimi (#2304)'
);
});
test('all inlined copies of the normalization block are byte-identical', () => {
const blocks = HOOK_FILES.map(extractBlock);
for (let i = 1; i < blocks.length; i++) {
assert.equal(
blocks[i],
blocks[0],
`${HOOK_FILES[i]} normalization block diverges from ${HOOK_FILES[0]}`
);
}
});
test('every guard map is the value-inverse of the installer matcher vocabulary', () => {
assertMapIsInstallerInverse(parseMap(extractBlock(HOOK_FILES[0])), HOOK_FILES[0]);
assertMapIsInstallerInverse(parseMap(extractBlock(WRITE_GUARD_FILE)), WRITE_GUARD_FILE);
});
test('every guard-relevant Claude tool has a reverse entry (dormancy alarm)', () => {
const map = parseMap(extractBlock(HOOK_FILES[0]));
for (const claudeName of GUARD_RELEVANT_CLAUDE_TOOLS) {
const modulePath = convertKimiToolName(claudeName);
assert.ok(modulePath, `installer no longer maps ${claudeName} — update this test`);
const kimiName = modulePath.slice(modulePath.lastIndexOf(':') + 1);
assert.ok(
map[kimiName] !== undefined,
`Kimi name '${kimiName}' (from ${claudeName}) has no KIMI_TOOL_NAMES ` +
`reverse entry — the matching guard would be silently dormant on Kimi (#2304)`
);
}
});
test('gsd-write-guard.js maps the Kimi name for Write (its only inspected tool)', () => {
const map = parseMap(extractBlock(WRITE_GUARD_FILE));
const modulePath = convertKimiToolName('Write');
assert.ok(modulePath, "installer no longer maps 'Write' — update this test");
const kimiName = modulePath.slice(modulePath.lastIndexOf(':') + 1);
assert.equal(
map[kimiName],
'Write',
`${WRITE_GUARD_FILE}: Kimi name '${kimiName}' must map to 'Write' or the ` +
'shrink guard is silently dormant on Kimi (#2304)'
);
// The copy must also carry the payload-field half of the normalization —
// WriteFile delivers `path`, the guard reads `file_path`.
assert.ok(
extractBlock(WRITE_GUARD_FILE).includes('input.file_path'),
`${WRITE_GUARD_FILE}: normalizeKimiPayload no longer maps path -> file_path`
);
});
});
// The two shell guards (gsd-graphify-update.sh, gsd-phase-boundary.sh) carry
// the same #2304 normalization reimplemented in shell — a byte-identity
// assertion cannot span the JS↔shell boundary, so instead of faking one this
// block pins the two vocabulary facts each script depends on to the
// installer's live mapping. Behavior is covered by negative-controlled tests
// beside each hook's existing suite (graphify-auto-update.slow.test.cjs,
// hooks-opt-in.test.cjs); this block is only the vocabulary-drift alarm
// (a convertKimiToolName rename fails HERE).
describe('Kimi shell-guard vocabulary parity (#2304)', () => {
const readHook = (file) =>
fs.readFileSync(path.join(__dirname, '..', file), 'utf8');
test('gsd-graphify-update.sh maps the installer\'s Bash vocabulary back to Bash', () => {
const modulePath = convertKimiToolName('Bash');
assert.ok(modulePath, 'installer no longer maps Bash — update this test');
const kimiName = modulePath.slice(modulePath.lastIndexOf(':') + 1);
const src = readHook('hooks/gsd-graphify-update.sh');
assert.ok(
src.includes('TOOL_NAME="${TOOL_NAME##*:}"'),
'gsd-graphify-update.sh no longer strips the Kimi module-path prefix'
);
assert.ok(
src.includes(`[ "$TOOL_NAME" = "${kimiName}" ]`) && src.includes('TOOL_NAME="Bash"'),
`gsd-graphify-update.sh no longer maps Kimi '${kimiName}' to Bash — ` +
'the hook is silently dormant on Kimi (#2304)'
);
});
test('gsd-phase-boundary.sh prefers Kimi\'s authoritative tool_input.path, falls back to file_path (#2752)', () => {
const src = readHook('hooks/gsd-phase-boundary.sh');
assert.ok(
src.includes('(typeof i.path===\'string\'&&i.path)||(typeof i.file_path===\'string\'&&i.file_path)||\'\''),
'gsd-phase-boundary.sh no longer prefers tool_input.path over file_path — ' +
'path is the authoritative field (kimi-cli executes on it); a model-supplied ' +
'decoy file_path must not suppress or fabricate a reminder (#2752, mirrors #2595)'
);
});
});