9 Commits

Author SHA1 Message Date
Jakub Zych
1622b66907 test: remove Kimi tool-vocabulary hook tests
Some checks failed
Tests / PR mergeability (push) Successful in 17s
Tests / Base branch health (push) Successful in 9s
Tests / Detect test scope (push) Successful in 17s
Tests / lint-tests (push) Failing after 1m56s
Tests / plugin-validate (push) Successful in 1m2s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 21s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 18s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24) (push) Failing after 20s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 17s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 7s
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
Tests / Required tests (push) Has been cancelled
The Kimi payload normalization they exercised went away with the Kimi runtime.
The Claude-vocabulary force-add block test from the same suite is kept.
2026-10-06 20:09:18 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
91d5fdff6f chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces

Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse
advisory output alongside the existing additionalContext prose:

- gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName
- gsd-context-monitor.js: severity ('warning'|'critical')
- gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS
  + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523
  pattern)
- gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its
  findings array already existed from #3523)
- gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg,
  distinct from the existing force-add block leg's code

additionalContext stays byte-identical in every hook (verified per-hook
against the pristine HEAD version across a spread of payload shapes).

Migrates all 20 assertion sites named in the issue off
additionalContext.includes(...)/assert.match(...) substring-matching
onto the new typed fields, per CONTRIBUTING.md's prohibition on raw
text matching on test outputs.

Closes #3546

* test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper

Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called
gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so
each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's
own doc comment says is sized for plumbing reads (rev-parse/branch/log),
not write-heavy add/commit spawns. That file already documents the exact
same defect class from a prior incident (PR #3323) and exports
GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites -
commitN just wasn't using it. Observed failure: `git commit -m filler 9`
timed out under normal bench load, unrelated to any of this PR's own
diff (hooks/*.js + 5 other test files).

Not a flake: root-caused to the timeout bound being sized for the wrong
call class, per this repo's no-flakes rule.

* chore(#3546): backfill changeset PR number (#4167)

---------

Co-authored-by: sim <sim@local>
2026-09-01 22:32:54 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Tom Boucher
268ca7e32d fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites

* fix(#3504): harden hook injection patterns and force-add guard

* test(#3504): stage the scanner lib dependency in shared-hooks fixture

* chore(#3504): backfill changeset pr number

* test(#3504): build parity samples from fragments for the ci scan

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:19:35 -04:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
5fd5c81042 test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it

Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync,
each returning a typed discriminated union
{ outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }.
Every call is timeout-bounded; there is no unbounded path.

runGsdTools becomes an adapter over the seam. Its legacy
{ success, output, error, exitCode } shape and retry-once-on-kill behaviour
are preserved byte-identically, so none of its 136 caller files change.

Outcome discrimination was corrected against probed runtime behaviour rather
than assumption: a timeout and a maxBuffer overflow are identical on both
status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs
ENOBUFS). Overflow is therefore classified before timeout. This fixes a live
defect — the previous isKilled() treated an overflow as a kill, retried it for
a second full 60s run, and then reported "host OOM or scheduler contention"
for a child that had merely printed too much.

Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs,
which brought 31 previously unlinted shared helpers under the same rules their
sibling test files already obey, and fixes the 5 violations that surfaced —
including a bare npm invocation without shell:true in
tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now
routed through the existing portable runNpm helper.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): migrate every local spawn wrapper onto the process seam

Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions
with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name,
parameter list, return shape and post-processing (JSON parse, ANSI strip, env
sanitising, field extraction) — only the spawn mechanism changes, so no test
assertion moves.

The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather
than a fourth primitive; it is explicit rather than inferred from the file
extension, because guessing an interpreter from a path fails silently when a
script's name does not match its shebang.

Seven wrappers were previously unbounded and now carry an explicit timeout
sized to what each actually runs, not the seam default. Two of those seven
(gsd-write-guard, lint-docs-command-form) were absent from the issue's
inventory entirely and were found by scanning after the migration.

Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md
reference section covering the three primitives, the discriminated union, and
the two rules the seam enforces.

Scope disclosure recorded in the phase design notes: the issue scoped three
identifier names. A scan for local helpers that spawn AND return the spawn
result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded
direct git call sites. This change bounds 25 of those. The remaining surface
is the same defect class and is NOT closed by this PR.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify an externally-killed child as KILLED, not EXITED

Blocker found in this branch's own diff, independently confirmed by an
isolated reviewer.

A child killed by an external signal — a genuine bench OOM kill — makes
spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field.
The seam's "no error implies EXITED" rule therefore classified it as a clean
exit, and runGsdTools returned { success: false, exitCode: 1 } without
retrying. That silently defeated the #969 kill-discrimination for precisely
the case it was built for: the old isKilled() fired on `signal != null`,
retried once, then threw a labelled resource-starvation error. A real OOM
would have been reported as an ordinary assertion failure.

Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes
the adapter retry on TIMED_OUT or KILLED — reproducing the old
`killed || signal != null || code === 'ETIMEDOUT'` condition exactly.
SPAWN_FAILED still does not retry (matching the old behaviour, where signal
was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate
divergence: the old code retried it because signal was SIGTERM, burning a
second 60s run on a child that had merely printed too much.

All five outcomes verified against the live runtime rather than assumed:
SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT),
>1MB stdout -> buffer_overflow (ENOBUFS).

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): address standards-review findings on this branch

Three findings from the standards axis of the review, all in this branch's
own diff.

The CONTEXT.md glossary entry this branch introduced was already stale on the
branch's own last commit: it enumerated a 4-member OUTCOME while the code had
5, because the KILLED fix did not update it. That is precisely the drift the
"module changes update Domain-terms" gate exists to catch, so the entry now
lists all five and explains KILLED.

api-coverage-gate-e2e compared an outcome against the raw string 'exited'
rather than OUTCOME.EXITED, the only such outlier; the enum is now imported
and used. A sweep for the other four outcome literals found no further
comparison sites.

Three call sites hand the literal bash flag '-c' to the seam's first
parameter, which the JSDoc described as an absolute script path. Rather than
add a fourth primitive, the contract is corrected to match reality: the
parameter is renamed `target` and documented as the first argv element handed
to the interpreter — normally a script path, but for an interpreter invoked
with an inline program it may be that interpreter's own flag. No behaviour
change.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): assert the cross-platform timeout contract, not the macOS one

The remote runner failed on both Linux lanes (node 22 and node 24, identical)
while the same tests passed locally on macOS. Two assertions encoded a
platform-specific behaviour as a cross-platform guarantee.

When spawnSync times out, macOS preserves the child's partial stdout/stderr;
Linux discards it and returns empty strings. Verified on node v26.5.1 both
ways. The seam passes through whatever spawnSync hands it and cannot
manufacture output that was discarded, so the production code was correct —
the tests were wrong.

Both tests now assert the guarantee the seam actually makes on every
platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being
strings rather than undefined or a Buffer. The partial-content assertions are
retained behind an explicit process.platform === 'darwin' guard so the macOS
coverage is not lost, and the first test is renamed to say what it now
guarantees.

This is the failure mode the remote matrix exists to catch: local macOS
verification would have shipped it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout

Windows CI caught two defects the Linux matrix could not.

tests/context-predicates-query.test.cjs passes a 32K-char argv value. On
Windows that exceeds the argv limit and spawnSync fails with code
ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise,
status === null implies TIMED_OUT" — swallowed it, so the adapter retried a
spawn that can never succeed and then threw the resource-starvation error. The
old isKilled() returned false for that shape and returned an ordinary failure
result.

TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set.
Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG,
E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform
whose timeout errno differs classified correctly, so the greedy catch-all is no
longer needed.

The second defect is a contract regression I introduced and had claimed
otherwise. That same test asserts `typeof r.exitCode === 'number'`, and
toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the
assertion failed on type. The old code returned `err.status ?? 1` on every
non-retried failure path. The adapter now returns 1 again for both, and the
comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is
retracted: the seam keeps the richer truth (exitCode null plus a distinct
outcome), the legacy adapter keeps the old numeric contract its callers
actually depend on.

Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT ->
SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL ->
KILLED; clean exit -> EXITED.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 01:06:39 -04:00
0xdhx
a8b40fa53f fix(#2547): fail closed on crashing and path-shadowing Kimi payloads (#2595)
* fix(#2547): fail closed on a malformed Kimi edit list in normalizeKimiPayload

`normalizeKimiPayload` rebuilt old_string/new_string with
`String(e.old ?? '')`. `??` guards the value, not the dereference, so a
nullish entry in a Kimi `edit` list threw a TypeError at the top of the
handler, before any tool dispatch. Each guard's outer
`catch { process.exit(0) }` swallowed that crash and emitted the same exit
code as "nothing to report" — turning a should-BLOCK call into a silent
allow.

Two hard blocks were bypassable:

  * gsd-worktree-path-guard's cross-git-root write block (#260) — a
    StrReplaceFile write whose path resolves to a different git root is
    correctly blocked with a well-formed edit list, and silently allowed
    with `edit: [null]`.
  * gsd-workflow-guard's force-add block on agent-* branches — a Shell
    payload carrying a spurious `edit: [null]` field walks past it. The
    Bash path never reads `edit`; the field only has to be present to
    trigger the crash.

Fixed with `e?.old` / `e?.new`, landed identically in all five copies so
tests/kimi-guard-normalization-parity.test.cjs's byte-identity assertion
still holds.

The crash boundary is nullish specifically, not "non-object": `('x').old`
and `(7).old` are legal reads yielding undefined, so string/number entries
never threw. Both are kept as controls proving the fix did not change
their behaviour.

Regression coverage is folded into the owning suites per CONTRIBUTING.md
(no new bug-* files). Negative-controlled: the nullish cases exit 0
against pre-fix guards and exit 2 after, with positive controls (the
equivalent well-formed payload blocks) and negative controls (in-worktree
writes and benign commands still pass) alongside.

Refs #2547

* test(#2547): exercise the production Kimi payload shape in read-guard tests

The `#2304: Kimi tool vocabulary engages the read guard` cases send
payloads with no `session_id`, and runHook injects none. A live Kimi turn
always carries one — kimi-cli's hooks/events.py `_base()` sets it
unconditionally, and soul/kimisoul.py calls `set_session_id()` at the top
of every turn before tool dispatch, so the ContextVar's `default=""` never
reaches a tool call.

gsd-read-guard treats any non-empty `data.session_id` as "Claude Code
already enforces read-before-edit, skip" (#2520). So the advisory those
tests assert fires only for a shape production never sends: the tests were
green, and the guard was dormant on Kimi. A sibling #2520 case in the same
file asserts the skip when `session_id` IS present — both passed, and the
production shape hits the skip.

Two changes, test-validity only:

  * Retitle the #2304 block to say what it proves — the tool VOCABULARY is
    normalized through to the Write/Edit branch — with a comment warning
    not to read it as production evidence.
  * Add a #2547 block asserting behaviour against the production shape
    (session_id populated), including a case that pins the delta directly:
    the same payload fires without session_id and is silent with it.

The #2547 block characterizes a known gap; it does not endorse it.
Redesigning how the guard discriminates runtimes is explicitly out of
scope for #2547. If a later change makes the advisory fire on Kimi these
tests are supposed to fail — update them then rather than dropping the
coverage.

Refs #2547

* docs(#2547): scope the Kimi guard-engagement claim to what Kimi enforces

#2518 engaged the guards' Kimi matchers and the release notes describe the
result as "All seven guard hooks now engage on Kimi", singling out the
prompt-injection read scanner as "the security-relevant guard" taken "from
silently dormant to engaged". That is not achievable for the scanner at
the emit layer.

gsd-read-injection-scanner.js is a PostToolUse hook, and kimi-cli's
dispatch never inspects PostToolUse hook results: src/kimi_cli/soul/
toolset.py awaits PreToolUse and honours `result.action == "block"`, but
fires PostToolUse via asyncio.create_task() and returns the ToolResult
without awaiting it — the done_callback only retrieves the task's own
exception. So no output shape the scanner emits can block or flag a Kimi
tool call, and `security.injection_blocking` cannot take effect there.
Reshaping the scanner's output would not change this; the enforcement gap
is in kimi-cli's PostToolUse handling, which is out of scope here.

This corrects the claim rather than the code — there is no gsd-core emit
fix that would make it true:

  * .changeset/2304-kimi-guard-tool-name.md — the fragment is unreleased,
    so it would otherwise ship this as a CHANGELOG security claim.
    Headline narrowed to "normalize Kimi's payload shape" and a scope
    paragraph added naming what actually blocks on Kimi (the two
    PreToolUse blocks) versus what cannot.
  * docs/migration/kimi-to-kimi-code.md — the scanner was listed under
    "Every GSD `PreToolUse` guard"; it is PostToolUse. Corrected, and the
    "What about the dormant guards?" section now splits enforceable from
    not-enforceable instead of saying Phase 0 "fixed all seven".
  * hooks/gsd-read-injection-scanner.js — the same scope note in the
    file's own Kimi rationale comment, where the next contributor to touch
    the normalization will actually read it. Comment only; the shared
    normalization block is untouched and byte-identity still holds.

Refs #2547

* chore(#2547): regenerate golden install-parity fixtures for the guard fix

The golden install-parity fixtures record a content hash per installed
file, so changing the five guard hooks changes their hashes across every
runtime's fixture. Regenerated with the full sweep (build, gen:golden,
size:baseline) rather than a single generator — running gen:golden alone
leaves tests/workflow-size-baseline.json stale and loses CI jobs to a
regeneration that looked complete.

The size baselines came out unchanged (no workflow or agent bodies
touched) and the hash delta is confined to exactly the five guards:
gsd-prompt-guard, gsd-read-guard, gsd-read-injection-scanner,
gsd-workflow-guard, gsd-worktree-path-guard.

Refs #2547

* fix(#2547): guard the String() coercion in normalizeKimiPayload too

Found by adversarial review of the first commit, then reproduced against
pristine next: `e?.old` closes the nullish dereference but leaves a second
route to the same crash-to-allow.

`{"toString": null}` is valid JSON, and coercing it throws
`TypeError: Cannot convert object to primitive value` — so an edit entry
that IS a well-formed object still crashes normalization, still lands in
the outer `catch { process.exit(0) }`, and still downgrades a should-BLOCK
call to a silent allow. Confirmed on both hard blocks:

  {"tool_name":"Shell","tool_input":{
     "command":"git add -f secret.env",
     "edit":[{"old":{"toString":null},"new":"x"}]}}      -> exit 0 (was)

  {"tool_name":"StrReplaceFile","tool_input":{
     "path":"<main-repo>/src/index.ts",
     "edit":[{"old":{"toString":null},"new":"x"}]}}      -> exit 0 (was)

Both exit 2 now.

The coercion is wrapped rather than replaced with a `typeof === 'string'`
test on purpose. Degrading only the non-coercible entry keeps
stringification identical for every value that CAN coerce — numbers,
arrays, plain objects — which matters because gsd-prompt-guard scans
new_string for injection patterns, and a `typeof` test would silently stop
scanning content that reaches that scan today (e.g. `new: ["ignore all
previous instructions"]` currently stringifies and is scanned). Verified:
zero behaviour change across string, number, bool, null, array-of-strings,
nested array, plain object and `__proto__`-keyed input; only the throwing
case changes, from crash to ''.

Regression cases are negative-controlled against the previous commit: the
four new coercion-trap tests fail with only the `e?.old` fix in place and
pass with this one.

Refs #2547

* chore(#2547): cover the String() coercion vector in the changeset

The release note described only the nullish-dereference route. Both routes
reach the same fail-open, so both belong in the changelog entry, along with
why the coercion is wrapped rather than type-tested.

Refs #2547

* chore(#2547): point the changeset fragment at the real PR number

The fragment has to exist before `gh pr create` runs, so it carried the
issue number as a placeholder. Corrected to 2595 now that the PR is open.

Refs #2547

* fix(#2547): make Kimi's `path` authoritative over a model-supplied `file_path`

normalizeKimiPayload copied Kimi's `path` into `file_path` only when
`file_path === undefined`, so any `file_path` the model chose to include won
outright. Every guard reads `file_path`; kimi-cli executes on `path`. The guard
therefore inspected one file while the write landed on another.

This bypass needs no crash. A payload pairing a cross-root `path` with a
spurious `file_path: ""` left gsd-worktree-path-guard reading an empty string
and exiting 0, while the identical write without the extra key blocked — the
same cross-root write the #260 block exists to catch. The shadowing also
preserved a non-string `file_path` (`[]`), which threw inside that guard's
path.isAbsolute() and reached its outer `catch { process.exit(0) }`: the same
crash-to-allow the rest of #2547 closes, reached through the guard's own read
rather than through normalization.

Reachability is not speculative. kimi-cli's soul/toolset.py json-parses the
model's raw tool arguments and passes that dict verbatim as tool_input to
PreToolUse, performing typed validation only later inside tool.call() — after
the hook has already decided. So the model controls extra keys in tool_input at
the moment the guard runs. kimi-cli's file tools carry no `file_path` field at
all (src/kimi_cli/tools/file/write.py, replace.py), so a `file_path` in a Kimi
payload is always model-supplied.

`path` now wins outright. Overwriting can only ever narrow what a guard inspects
to the path that will actually be written, so it cannot under-block.
Normalization returns early for non-Kimi tool names, so the native Claude Code
contract (file_path governs) is untouched.

Landed identically across all five inlined copies; the byte-identity assertion
in tests/kimi-guard-normalization-parity.test.cjs enforces that.

* test(#2547): cover the file_path-shadowing bypass in the #260 guard suite

Four cases, each exiting 0 (bypass) against the pre-fix guards: a spurious
empty-string file_path, an in-worktree decoy file_path, and non-string
file_path values (array and object) that additionally crashed
path.isAbsolute() into the outer catch.

Two controls that are not bypass cases and matter as much:

  - an in-worktree write carrying a cross-root DECOY file_path must still exit
    0. Pre-fix this blocked, because the decoy won; the guard now follows the
    path kimi-cli executes on in both directions, so the fix narrows what is
    inspected without over-blocking.

  - a native Claude Edit (no `path` field) must still block on file_path alone.
    normalizeKimiPayload returns early for non-Kimi tool names, and this pins
    that the non-Kimi contract did not move. It passes both pre- and post-fix
    by design.

Negative-controlled: run against the pre-fix hooks, the four bypass cases and
the decoy control fail, and the native-Claude control passes.

* test(#2547): back the totality claim with property tests over fc.anything()

This PR claims the fix "makes normalization total over the inputs JSON can
express" — a for-all guarantee — while the tests backing it are example-based,
each shape added reactively after a crash was found by hand (the String()
coercion trap was itself found by adversarial review after the first commit
shipped). Example-based tests cannot substantiate a for-all claim; they record
the counterexamples someone happened to think of.

Four properties over fc.anything(), which is exactly the JSON-expressible
domain the claim names:

  (a) totality over any tool_input
  (b) totality over any edit list — the crash surface both #2547 fixes targeted
  (c) `path` always wins over any model-supplied `file_path` (the review blocker
      invariant: a guard reading file_path can never be aimed at a file other
      than the one kimi-cli writes)
  (d) a non-Kimi tool_name passes through untouched — the native Claude contract

normalizeKimiPayload is inlined per hook with no runtime binding, so there is
nothing to require. The block is extracted from hook source and evaluated via
the SAME extraction contract kimi-guard-normalization-parity.test.cjs uses, so
a source edit that breaks one breaks both instead of silently testing a stale
block. An extraction floor test fails loudly if the extraction yields a no-op.

Non-vacuous, and checked rather than assumed: against pristine pre-#2547 `next`,
(a), (b) and (c) all FAIL and (d) passes. (a) needed the fix that makes it
meaningful — a bare fc.anything() for tool_input passed even against the live
defect, because arbitrary generation essentially never invents the `edit` key
the crash lives behind, so the generator is biased onto the keys normalization
actually reads and unioned back with unbiased input.

* chore(#2547): cover the shadowing vector in the changeset and regen goldens

Golden install-parity churn is hash-only, on exactly the five hook files this
round changed. gsd-phase-boundary.sh is deliberately unchanged.

* test(#2547): make the property test able to kill the coercion mutant

Review Major 1: the generative test added to stop the NEXT counterexample
could not kill the one it was written for. Reproduced the reviewer's matrix
independently — against the shipped generator, a mutant reverting `editText`
to the unguarded `String(v ?? '')` passed all four properties.

Cause, confirmed by measurement: the edit-array ENTRIES were bare
`fc.anything()`, which essentially never invents an `old`/`new` key, so
`e?.old` was always undefined and `String(undefined ?? '')` never coerced
anything. That is the same vacuity the file's own comment describes one level
up, reproduced one level down.

The review's prescribed fix — bias the entry onto `{old, new}` — is necessary
but NOT sufficient, and this is the part worth recording: measured over 20,000
draws, bare `fc.anything()` yields a non-coercible value 3 times (0.015%). At
`numRuns: 200` an `old` key holding a hostile value essentially never
co-occurs, and the mutant survives the entry bias too. Both levels need
biasing — the entry onto the keys normalization reads, and the VALUE onto the
shape that actually throws.

`{"toString": <non-function>}` is that shape and stays inside the
"JSON-expressible" domain the claim names (JSON.parse produces it verbatim);
`fc.anything({withNullPrototype: true})` would also kill the mutant but widens
the domain past what the PR asserts, so it is not used.

Verified: M1 now dies at every seed tried (1/7/42/99/4242/31337, failing
within 3-31 cases) while HEAD stays green at all of them.

Also closes three coverage gaps the review listed as nits — properties (e)
totality over any JSON value as the WHOLE payload, (f) the tool_output →
tool_response mapping (including that an existing tool_response is not
clobbered), and (g) an empty edit list reconstructing nothing.

Property (e) required a one-line fix in the normalizer itself: `JSON.parse
('null')` is null, and null/primitive payloads threw on the `data.tool_name`
read — falsifying the "total over the inputs JSON can express" claim. Harmless
in practice (the throw landed in the same fail-open catch as the exit 0 it now
takes deliberately), but the claim should be true as stated. Landed
byte-identically across all five copies.

Refs #2547

* fix(#2547): make old_string/new_string authoritative, not merely defaulted

Review Major 2: the shadowing class the prior round's BLOCKER closed for
`file_path` survived one field over. `old_string`/`new_string` were still
reconstructed only `if (input.<field> === undefined)`, so a model-supplied
value won.

The argument for making `path` authoritative applies verbatim here.
kimi-cli's StrReplaceFile schema is `path` + `edit` only
(src/kimi_cli/tools/file/replace.py @ 4a550ef) and carries no
`old_string`/`new_string` at all, so either key appearing in a Kimi payload is
always model-supplied — exactly like `file_path`.

Verified end-to-end against the reviewer's payload: a cross-root write
carrying `new_string: ""` alongside an injected `edit[].new` left
gsd-prompt-guard reading '' and returning at its `if (!content)` guard, so the
injection advisory never fired and the reconstructed content was never
scanned. `new_string: null` behaved identically. Negative-controlled: both
produce empty output against pre-fix source and fire the advisory after.

Chose unconditional reconstruction over the offered `typeof` alternative
deliberately. A type test closes `""`/`null` but leaves the interesting case
open — a benign NON-EMPTY decoy (`new_string: "chore: tidy"`) shadows just as
effectively and passes any type test. The new suite includes that case
specifically; it is what discriminates between the two candidate fixes.

Also pins the kimi-cli SHA in the authoritative-path comment, as requested —
it cited file names with no version while the issue pins 4a550ef.

Landed byte-identically across all five inlined copies; the parity test's
byte-identity assertion holds.

Refs #2547

* fix(#2547): close the non-string file_path crash-to-allow at every read site

Review Major 3: the crash-to-allow was closed only as a side effect of `path`
masking the bad value, while the changeset read as though it were closed
outright. Confirmed both of the review's reachability claims: `[]`/`{}`/`42`
are truthy, survive the `!rawFilePath` early-out, and throw inside
path.isAbsolute() into the outer `catch { process.exit(0) }`; and normalization
returns early for native Claude Code payloads (KIMI_TOOL_NAMES has no 'Edit'
entry), so `{"tool_name":"Edit","tool_input":{"file_path":[]}}` reached it
untouched — this guard's original #260 surface.

Reproduced on a real fixture: string cross-root path exits 2, the identical
payload with `[]` or `{}` exits 0.

Swept the class rather than the instance. Five more untyped read sites across
four other hooks, each one line from a type-strict or method-dependent call.
Census of what each can actually do:

  gsd-worktree-path-guard.js:173  BLOCKS  -> live bypass (the review's finding)
  gsd-prompt-guard.js:128         scanner -> silenced the injection scan, the
                                             same outcome as Major 2 by another
                                             route; verified empirically
  gsd-workflow-guard.js:206       advisory only (its exit-2 is the Bash
                                             force-add path, which reads
                                             `command`, not `file_path`)
  gsd-read-guard.js:141           advisory only
  gsd-read-injection-scanner.js:213  advisory only
  gsd-windsurf-pre-write.js:75    ALREADY TYPED — the shape adopted here

All six now read typed. The workflow-guard site keeps its truthiness fallback
(`(typeof x === 'string' && x) || ...`) because a bare type test would let an
empty `file_path` shortcut the `path` fallback.

Also declares one swept hit NOT fixed: `gsd-workflow-guard.js:175` reads
`command` untyped on a genuinely blocking path. Same shape, but not
exploitable — unlike file_path/path there is no second field carrying the
executable value, so a non-string command cannot smuggle a real `git add -f`
past the block. Left alone rather than widen this PR into the Bash path.

The regression gate is a SOURCE-level invariant, not a behavioural one, and
that is deliberate: the fixed read and the crashing read are black-box
identical — both end at exit 0, one via the catch and one via the early-out.
A test asserting exit 0 on a non-string payload passes against the unfixed
code, which is the same false-green the review flagged in the existing
`['non-string file_path (array)', []]` cases. Repeating it one level up would
be no better. tests/kimi-guard-typed-payload-reads.test.cjs fails if any hook
regresses to an untyped read (negative-controlled: it reports all five
pre-fix sites with correct file:line).

The behavioural cases requested — non-string file_path with NO `path` key —
are added to worktree-safety.test.cjs and labelled honestly as documenting the
explicit fail-open rather than detecting a revert.

Also states the relative-path premise (review Minor 5) at the early-out that
depends on it: "always safe" holds only while every runtime reaching there
resolves relative paths against the tool CWD. Claude Code satisfies it by
requiring absolute paths; kimi-cli's resolution behaviour is NOT verified here
and is recorded as an unverified premise rather than an asserted bypass.

Refs #2547

* docs(#2547): correct the changeset's closed-claim and fold the misattributed note

Review Major 3 also flagged the fragment: it said the non-string vector "threw
inside that guard's path.isAbsolute() ... `path` now wins outright", which
reads as closed when it was closed only conditionally. Rewritten to state what
is now true — closed unconditionally at all six read sites — and extended with
the Major 2 finding.

Review Minor 6 (the #2547 scope note living in a `pr: 2518` fragment) turns out
to understate the problem. Rendering the changelog and re-parsing it shows the
note is not merely misattributed — it is DROPPED. serializeChangelog emits each
fragment as a single `- ` bullet, and parseChangelog terminates a bullet at the
first non-continuation line, so everything after a blank line is lost on
re-parse. Audited all 44 fragments: exactly one was lossy —
2304-kimi-guard-tool-name.md, losing 656 of 1730 characters, i.e. precisely
that second paragraph. Folding it into this PR's fragment fixes the
attribution and the silent loss together; all 44 now round-trip losslessly.

That same mechanism is why the remaining nit — reformat this fragment's
~2,000-character paragraph for readability — is NOT applied. A paragraph break
or a bullet list would silently truncate the entry at the first blank line
(verified for both). The single-paragraph form is load-bearing under the
current serializer, not an authoring preference. Worth its own issue; noted in
the PR thread rather than worked around here.

Refs #2547

* chore(#2547): regenerate golden install parity after rebase onto next

Rebased onto `next` @ 9138271b (the PR had gone BEHIND by 20 commits; the
review's closing nit asked for it). The replay was CLEAN — no conflicts — and
that is exactly why this commit exists.

These fixtures are one key per installed file, so when the PR pins five hook
entries and the base rewrites others', the two edits land on different lines of
the same JSON. Git merges them silently and correctly AS TEXT while attesting
nothing about whether the merged hashes are still valid. Verified rather than
assumed: per-key equivalence against the old base showed the base had moved 22
of the 27 keys this PR pins in every runtime fixture, and
tests/golden-install-parity.test.cjs failed on 10 runtimes immediately after the
clean rebase. A push without this regen would have gone out red.

Regenerated with `npm run build && npm run gen:golden` under a throwaway
HOME/CLAUDE_CONFIG_DIR (the generators invoke the installer); live-profile
canary clean before and after.

Contamination check: every key differing from the base's committed fixture
resolves to a file this PR actually touches — the five guard hooks, under both
the `hooks/` and `.kimi/hooks/` install layouts, and nothing else. Derived from
the PR's changed-file set rather than a feature-name filter, which is what
would have mislabelled the registration surfaces.

Size baselines re-checked and NOT regenerated: this PR moves no workflow or
agent, and the base's own baselines are current (agent-size-budget,
workflow-size-budget, workflow-size, update-size-baseline all green).

Refs #2547

* chore(#2547): allowlist the field-shadowing security test in the injection scan

The new regression suite tripped the repo's own prompt-injection scan — a test
for the injection scanner setting off the injection scanner.

The fixture has to be a real injection phrase for the test to assert anything:
it is precisely the content gsd-prompt-guard must still scan once a
model-supplied `new_string` can no longer shadow the reconstructed
`edit[].new`. Weakening it to a benign string would make the suite vacuous.

Allowlisted rather than obfuscated, because that is this repo's established
convention for the class — tests/read-injection-scanner.security.test.cjs,
tests/security-prompt-injection.security.test.cjs,
tests/prompt-injection-scan.security.test.cjs and four others carry real
payloads as test DATA and are listed for exactly this reason. Splitting the
literal to dodge the grep would work but would make this one file inconsistent
with its five peers and leave the next reader wondering why.

Verified with the CI invocation itself (`scripts/prompt-injection-scan.sh
--diff upstream/next`): 26 files scanned, 0 findings. The .cjs codebase scan
does not cover tests/ and is unaffected (73 tests green).

Refs #2547

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-28 18:31:40 -04:00
Tom Boucher
7e905aa137 feat(#2505): Phase 0 — Kimi PreToolUse guard vocabulary normalization (precondition; carries PR #2326 forward) (#2518)
* fix(#2304): normalize Kimi tool vocabulary in PreToolUse guard payload checks

The Kimi [[hooks]] registrations translate the matcher to Kimi's tool
vocabulary (WriteFile|StrReplaceFile) but the guard scripts early-exit
unless the payload's tool_name is a Claude name (Write/Edit/MultiEdit),
so every guard was dormant on Kimi: the matcher fired, the script saw
WriteFile, and exit(0)'d.

Normalize the payload's tool_name at the top of each guard
(WriteFile -> Write, StrReplaceFile -> Edit; bare or module-qualified
kimi_cli.tools.file:* forms) before the check. Inlined per guard rather
than a hooks/lib/ helper because hook scripts are staged as standalone
files on every hook surface, and a sibling require is a staging
dependency that can fail silently.

Regression tests pipe Kimi-vocabulary payloads at each guard and assert
it engages (typed fields: exit status, decision, hookSpecificOutput) —
verified red against the pre-fix scripts, green after.

* fix(#2304): normalize Kimi tool_input fields and route block reasons to stderr

Cross-AI review of the initial fix, verified against kimi-cli source,
found the tool_name normalization alone leaves the guards dormant on a
real Kimi runtime: kimi-cli forwards tool_input verbatim
(src/kimi_cli/hooks/events.py), and its tool schemas
(src/kimi_cli/tools/file/{write,replace}.py) use path/content and
edit.old/edit.new (single Edit or list) — not Claude's
file_path/old_string/new_string. The guards read file_path, got '',
and exited 0 past the now-open tool_name gate.

Extend the per-guard normalization to the payload fields
(path -> file_path, edit -> old_string/new_string with list flattening),
and write the worktree guard's block reason to stderr as well as the
stdout JSON — Kimi feeds stderr, not stdout, back to the model on
exit 2 (docs/en/customization/hooks.md exit-code table).

Regression tests rewritten to Kimi's actual payload shapes (plus an
edit-list case and a stderr-reason assertion) — verified red against
the name-only fix, green after.

* fix(#2304): join all edit[] entries into old_string, matching new_string

Review nit on #2326: old_string took only edits[0].old while new_string
joined the whole list. Symmetric join removes the latent trap for any
future consumer sizing before/after content (e.g. the #2255 write guard).

* fix(#2304): normalize Kimi ReadFile vocabulary in read-injection scanner

Review Major 2 on #2326: gsd-read-injection-scanner.js had the identical
dormancy — its Kimi matcher fires on 'ReadFile' but the SCANNED_TOOLS
check only knew 'Read', so injected content in read files was never
flagged on Kimi installs.

Folds the same inlined normalization block into the scanner and extends
the shared KIMI_TOOL_NAMES map with ReadFile:'Read' in all four copies so
they stay byte-identical. Harmless in the three write guards: a
normalized 'Read' falls out of their Write/Edit allowlist exactly as the
unmapped name did. Field mapping verified against kimi-cli upstream
(src/kimi_cli/tools/file/read.py Params.path); the existing
path->file_path copy covers the scanner's file_path read.

* test(#2304): parity test binding the four inlined Kimi normalization copies

Review Major 1 on #2326: KIMI_TOOL_NAMES + normalizeKimiPayload is
deliberately inlined in four hook scripts (staging-dependency rationale,
unchanged), with the inverse table in bin/install.js — five
hand-maintained surfaces and nothing binding them.

Static binding, zero runtime coupling:
- the four inlined blocks must be byte-identical;
- each guard-map entry must be the value-inverse of
  convertKimiToolName() for its Claude name;
- every guard-relevant Claude tool (Write/Edit/MultiEdit/Read) must have
  a reverse entry — a vocabulary rename or extension that updates the
  installer without updating the guards now fails in CI instead of
  leaving a guard silently dormant (the #2304 recurrence door).

Negative-controlled: diverging one copy or dropping a map entry fails
the suite against the fixed code.

* test(#2304): regenerate golden parity fixtures for guard hook changes

CI red on #2326: all 10 golden-parity failures were the staged guard
hooks drifting from their fixtures. Regenerated with npm run gen:golden
(after npm run build) under throwaway HOME/CLAUDE_CONFIG_DIR; diff
verified to change exactly the four PR-touched guard entries per
surface, nothing else.

* test(#2304): regression tests for Kimi ReadFile engaging the scanner

Mirrors the per-guard Kimi vocabulary tests the PR added for the three
write guards: bare and module-qualified ReadFile produce the advisory,
path exclusions still apply post-normalization, unknown Kimi names stay
fail-open. Negative-controlled against the pre-fold scanner (the two
positive cases fail there; exclusion/fall-through correctly pass on
both sides).

* fix(#2304): normalize Kimi Shell vocabulary in workflow guard

Withdraws the disclosed out-of-scope split: verification showed the
Bash->Shell case needs NO different mapping — kimi-cli's Shell.Params
names its field `command` (src/kimi_cli/tools/shell/__init__.py), same
as Claude's Bash — and the guard's write branch (Write/Edit/MultiEdit
allowlist) was ALSO dormant on Kimi under its Shell|WriteFile|
StrReplaceFile matcher. Same defect class as the other four hooks.

Folds the identical inlined block into gsd-workflow-guard.js and
extends the shared map with Shell:'Bash' in all five copies (harmless
outside the workflow guard: a normalized Bash falls out of the other
guards' checks as before). Parity test now binds five copies and adds
Bash to the dormancy alarm. New workflow-guard test file exercises the
observable block (force-add on a worktree-agent branch): Shell bare and
module-qualified block with WORKTREE_AGENT_FORCE_ADD_FORBIDDEN, benign
Shell passes, Claude Bash unchanged — negative-controlled against the
pre-fold guard (the two Kimi cases fail there). Golden parity fixtures
regenerated; diff verified to change exactly the five guard entries per
surface.

* fix(#2304): map Kimi tool_output and route workflow-guard block to stderr

Third-party review (cross-AI verifier) caught two gaps in the revision:

1. Kimi PostToolUse events carry `tool_output`, not `tool_response`
   (kimi-cli src/kimi_cli/hooks/events.py post_tool_use()), so the
   read-injection scanner — which reads data.tool_response — was STILL
   dormant on real Kimi payloads; the earlier tests passed because they
   sent Claude-shaped payloads. The shared normalization block now maps
   tool_output -> tool_response (inert in PreToolUse guards, where the
   field is absent), and the scanner's Kimi tests send the real shape.

2. The workflow guard's force-add block wrote its reason to stdout only.
   Kimi's exit-2 protocol feeds stderr back to the model — the exact
   fix this PR already applied to the other blocking guard — so the
   newly-awakened block would have been a silent denial. Reason now
   also routed to stderr, asserted in the test.

Also: the scanner's "unknown name" test now uses a genuinely unmapped
name (FetchURL) — Shell stopped qualifying when it entered the map —
and the workflow guard's write branch (WriteFile advisory,
StrReplaceFile .planning pass) gains behavioral coverage. All five
copies stay byte-identical (parity test green); golden fixtures
regenerated, diff verified to the five guard entries per surface.
Negative-controlled: 3 new assertions fail against the pre-fix hooks.

* docs(#2304): update changeset to cover the full five-guard fix

Review round 2 (2026-07-18) flagged the changeset as stale: it was
written for the first commit and still described only the three guards
named in the issue. The shipped diff grew to five guards plus two
payload dimensions the original body never mentioned. The body now
names gsd-read-injection-scanner and gsd-workflow-guard, the ReadFile
and Shell vocabulary entries, the tool_output -> tool_response mapping,
and the workflow guard's stderr block-reason routing.

* test(#2304): regenerate kilo golden fixture after #2305 landed on next

The branch's fixture sweep predates 50efae13 (fix(#2305), PR #2327),
which made Kilo ship the five shared guard hooks. Rebased onto next and
re-ran the full generator sweep (gen:golden, size:baseline, and the
four registry/contract generators); the only delta across all of them
is kilo.json's five guard-hook hashes, matching this PR's hook edits.

* fix(#2304): fold Kimi normalization into the two shell hooks

The 2026-07-19 review found the last two guards with the #2304 dormancy:

- hooks/gsd-graphify-update.sh gated on tool_name == "Bash" but is
  registered on Kimi with matcher 'Shell' — Gate 1 never matched and the
  auto-rebuild was silently dormant. kimi-cli's Shell.Params names its
  field `command` (src/kimi_cli/tools/shell/__init__.py), same as Claude
  Bash, so only the name needs mapping: strip the module-path prefix,
  map Shell -> Bash.
- hooks/gsd-phase-boundary.sh read only tool_input.file_path, but Kimi's
  file tools name the field `path` (src/kimi_cli/tools/file/write.py +
  replace.py) — the hook read '' and .planning/ writes went undetected.
  Falls back to tool_input.path when file_path is absent, mirroring
  normalizeKimiPayload's precedence in the JS guards.

The normalization is reimplemented in shell — a byte-identity assertion
cannot span the JS<->shell boundary, so the parity test gains a
shell-guard vocabulary block that pins both scripts' mapping facts to
convertKimiToolName's live vocabulary instead of faking a byte binding.
Behavior is covered by negative-controlled tests beside each hook's
existing suite (verified red against the pre-fix scripts): Kimi Shell
dispatch (bare + module-qualified) with a WriteFile negative control in
graphify-auto-update.slow.test.cjs, and Kimi path detection, file_path
precedence, and a non-.planning negative control in hooks-opt-in.test.cjs.

Changeset updated to name all seven guards; golden install-parity
fixtures regenerated (diff is exactly the two hook entries per runtime;
size baselines unchanged).

* fix(#2304): use a Map for KIMI_TOOL_NAMES so prototype keys cannot pass the guard fall-through

A bare bracket lookup on an object literal resolves 'constructor',
'__proto__', 'toString', 'valueOf' and 'hasOwnProperty' through
Object.prototype to truthy functions/objects, so `if (!mapped)` failed
to short-circuit and data.tool_name was assigned a non-string. Map.get
returns undefined for those keys — the same shape the repo already uses
in canonicalizeRuntimeName (src/runtime-name-policy.cts). Applied
identically to all five inlined copies (review M1, PR #2326).

No new bypass class: unrecognized strings already fail open by design;
this fixes the lookup being wrong, not the posture.

* test(#2304): enumerate normalized guards by scanning hooks/, not a hardcoded list

The parity test's file list was a literal five-entry array — a sixth guard
with its own copy-pasted normalization block would be silently uncovered,
the exact divergence mode the test exists to prevent (review M2). Now the
list is a scan of hooks/*.js for the KIMI_TOOL_NAMES marker, with a floor
assertion so a scan that finds nothing fails instead of passing vacuously.
Also parses the Map declaration introduced by the M1 fix, and carries the
allow-test-rule annotation documenting the source-text scanning (review m4).

* test(#2304): parse hook JSON output instead of substring-matching raw stdout

workflow-guard.test.cjs asserted on unparsed stdout while read-guard.test.cjs
in the same PR parses the JSON envelope first — match the better pattern at
all four assertion sites (review m5).

* test(#2304): regenerate golden parity fixtures after Map conversion in the five guards

* docs(#2304): reset changeset pr:0 placeholder for Phase 0 PR (#2507)

The closed PR #2326's changeset carried pr:2326. Phase 0 of epic #2505
re-lands this fix on a fresh branch; the pr: field will be backfilled
to the real Phase 0 PR number immediately after gh pr create returns.

* docs(changeset): backfill PR #2518 for Phase 0 (#2507)

---------

Co-authored-by: 0xdhx <darkhawkx@gmail.com>
2026-07-21 23:42:02 -04:00