Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.
harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.
Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.
Closes#2627
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Routes standalone agent-discipline prompt modules to the capability
ecosystem instead of the core skill layer. Records three grounds: skills/
is a generated 1:1 projection of commands/gsd (gen-plugin-skills.cjs, gated
by lint:generated-sync), ADR-857 D4 contribution hooks exist precisely so
prompt-woven behavior can leave core, and the self-rated-confidence
mechanism these asks center on is measured weak in
references/honest-verifier.md:25-29.
Scopes the denial with an explicit does-NOT-cover section so a future
triage keyword match cannot misapply it to loop-step-attached prompt fixes
or to externally-measured calibration, and makes the revisit condition
measurable against the honest-verifier baseline.
Closes#2614
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver
Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.
worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.
resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.
Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: rebuild tracked state-transition.cjs to match #2400 source
The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit 2bcfaa2e2) added the progress.total_plans frontmatter sync to source but the tracked bin/lib/state-transition.cjs was never rebuilt, so the fix was not shipping to consumers of the compiled artifact. The mandatory build:lib step for Phase 2 surfaced the drift; recompiling makes the already-merged, already-changelogged #2400 fix effective. Artifact-only resync (no source/test change); drift class tracked by #2591.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs
The gemini and claude blocks in gsd-core/workflows/review.md were the only
two of the ten prompt-fed reviewer legs with both `2>/dev/null` and no
empty-output guard. Any failure that wrote no stdout — CLI missing,
unauthenticated, rate-limited, crashed — left a zero-byte review file with
the only diagnostic evidence already discarded.
write_reviews substitutes each file's raw content into a `## <Reviewer>
Review` section with no marker separating "empty/failed" from "ran cleanly,
nothing to report", so a failed lane silently degraded the advertised
N-reviewer consensus to N-1 while present_results reported success. Both
/gsd:review and /gsd:plan-review-convergence share the invoke_reviewers
step, so both were affected.
Both legs now redirect stderr to a `.err` sidecar and write a diagnostic
stub with the captured stderr appended when the review file comes back
empty — the same shape the codex and cursor legs already use.
Scope limit, noted in the block comment: the guard only runs if the block
itself completes. A host Bash-tool timeout that kills the whole block skips
it, the same hard bound the OpenCode block already documents. The existing
timeout guidance (#2194) covers that case and is unchanged.
Regression test extracts the two dispatch blocks verbatim from the workflow
and runs them under bash against a failing CLI stub, asserting the review
file is non-empty and carries a diagnosable message plus the captured
stderr. It fails against pre-fix review.md (4 of 5 cases).
Golden install-parity fixtures and the workflow size baseline are
regenerated: one review.md hash per fixture, one size entry.
Fixes#2494
* fix(#2494): stamp the changeset fragment with the filed PR number
The fragment was authored with the documented `pr: 0` placeholder because
the PR number does not exist until `gh pr create` returns. Now that this
PR is #2592, stamp it — scripts/changeset/parse.cjs requires pr > 0, so
the placeholder would fail the Changeset Required check.
* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards
PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md).
Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single
canonical owner — as a pure, additive extension. The 17 locked exports and
PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical,
so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs)
stays green.
New pure round-trippable model (ADR Decision 4):
- PhaseId { project, milestone, phase, subphase?, plan? }.
- parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token
`GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens
(`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new
parser — normalizePhaseName and every legacy reader keep accepting those
tokens unchanged (conservative default; no existing path gains a throw).
- renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug`
with a slug guard that sanitizes path-traversal input.
- getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone
from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning
the `vN.0` form (parity with READING-A). The optional parameter keeps the helper
pure (no config read) and byte-compatible — every existing single-arg caller
resolves to the unchanged READING-A body (ADR Decision 6).
- extractPhaseToken(dirName, convention?): bracket dir branch GATED on
convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is
string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal
family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no
string-only discriminator is complete — an ungated auto-detect silently
reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit
convention signal keeps every existing convention-less call site byte-identical
(pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs).
- comparator: no new code — comparePhaseNum already orders the dot-decimal
`PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a
PR-2 resolution concern (bracketQualifiedKey), not core grammar.
- SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999}
non-milestone guard; the bracket-prefix reading is gated the same way (an
ungated read called `P0.0-foundation` a sentinel), legacy leading-int form
unchanged.
- BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately
more permissive than parsePhaseId — a read-tolerance source for PR-2, not the
emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt
owner so PR-2 builds every bracket read regex from the canonical source and
check:phase-id-drift stays green stack-wide.
The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*`
grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the
config permits parses. parsePhaseId has no live callers in PR-1, so this grammar
choice is forward-facing for PR-2 with zero PR-1 behavior impact.
Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips,
full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases),
extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens,
sentinel + slug guards, bare-token rejection, exported-source behavioral
assertions, and two generative fast-check properties: render∘parse identity over
well-formed displays, and the toDir/disk↔display bijection. Plus a #2043
numeric-tail characterization (single- AND multi-digit rows) in
tests/phase-id.test.cjs pinning the convention-less reading byte-identical.
The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish)
and rebuilt by CI, so it is intentionally not committed.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3)
PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4).
B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers,
over-padded numbers, multi-space separators, stray whitespace), so
render(parse(x)) === x did not hold for every well-formed x as ADR-612
Decision 4 requires. Both branches now enforce canonicality by construction:
parse permissively, rebuild the canonical string via the same emit path
(renderPhaseId for display, a hand-rebuilt token for dir/token), and throw
"parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's
entry is removed — the match anchors now reject leading/trailing whitespace
outright, folding into the existing "not a bracket phase id" rejection.
M1 (major): toDir only ever guarded the slug; project/milestone/phase/
subphase were interpolated unsanitized, so a hand-built PhaseId (a
structural, not nominal, type) could smuggle a path-traversal segment onto
disk. Every field is now validated against the exact shape parsePhaseId
itself would produce before use.
M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling
trailing hyphen in the emitted dir name. toDir now throws in that case.
M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable
from the dir-branch's plan tail, so it silently broke the disk<->identity
bijection on read-back. toDir now rejects all-digit slugs.
Nits: toDir now rejects a non-string slug instead of coercing it to the
literal token 'undefined'/'null'; sentinel boundary tests added for
milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999},
not an inclusive range — these were already correct, now locked by test).
Test-first: every new assertion (concrete examples + fast-check mutation
property for B1; concrete cases for M1-M3 and the nits) was written and
confirmed red before the implementation changes, per repo TDD convention.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(#2249): reformat changeset body to house convention (review Mi2)
The fragment added in ab26190a was a plain paragraph — no bold headline,
no trailing issue reference. Reformat to the repo's
`**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g.
.changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md).
Uses (#2249), the issue every commit on this branch references, not the
PR number already carried in frontmatter (`pr: 2258`) — the changelog
serializer appends `(#{pr})` unconditionally, so a body also ending in
`(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered
bullet directly via parseFragment + serializeChangelog: it now reads
`... (#2249) (#2258)`, matching the dominant convention across the other
fragments (frontmatter pr = merged PR, body reference = originating issue).
Also moved the docs-exempt marker back before the paragraph -> after it
(matching the file's original order): the marker sits on its own line and
is stripped before the body is used, but placing it first left a leading
blank line in front of the bold headline once reformatted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2)
PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
two property-generator coverage gaps the reviewer flagged; no source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged).
Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no
property fed a 3+-digit milestone/phase/subphase/plan through parse/render/
toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen
numArb to 1–999 so the round-trip and disk↔display bijection properties both
span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no
leading zero, so canonicality still holds). Add a concrete regression pinning
the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and
toDirs to 'GSD.100-05-feature' without truncation.
Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered
milestone/phase pad + whitespace mutations but never a subphase pad. Add
unpad-subphase / overpad-subphase to the mutation set and a generated
`includeSub` boolean that decides whether the canonical carries a `.SS`
(forced in for the subphase mutations so there is always a `.SS` to mutate);
non-subphase mutations keep their original no-subphase coverage.
Non-vacuity verified against the compiled lib by temporarily probing each
widened/new property and confirming it fails: round-trip counterexample
["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit
tokens are genuinely generated and reach the body; a no-op unpad-subphase
mutation trips the mutated===canonical guard (counterexample
["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached
with a subphase present. Probes reverted; numRuns unchanged.
Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail;
`npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major)
BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that
re-derived the grammar as an unbounded `\d+` literal instead of consuming
PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket
path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug
whose first word is a year) over-collected the token as `01-14-2026` instead of
`01-14`.
Interpolating the cap verbatim at every position was rejected on evidence: the
bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent.
The exactly-2 cap at the others would under-collect ids toDir itself emits —
`PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit
sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and
`[GSD.100] 05` is a pinned regression. Those positions are delimiter-
disambiguated (a required field separator; a dot a slug can never contain),
not heuristically recognized, so they have no year collision to defend against.
Upstream draws the same line for the same reason: core-utils/phase cap the
paired PLAN component while the leading phase component stays unbounded.
So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and
each position takes the width its delimiter affords: leading unbounded, dash-1
and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner
seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of
the token grammar.
Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE
instead of re-spelling it as a literal, so the emit-side gate and the read-side
token source are one rule — the same single-owner discipline this fix is about.
Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant).
Refs #2249
* test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major)
The comment block alone cannot hold the divergence: src/phase-id.cts is exempt
from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not
catch the bracket token source drifting from the seam. Per the Generative Fix
Divergence rule, the divergence is pinned behaviorally instead.
Surface 6 joins the existing #2232 parity gate rather than starting a rival one:
the review named the bracket token source "a sixth continuation-recognition
site", and continuation-grammar-parity.test.cjs is already the invariant-named
home where the five #2043 sites agree with the owner on a shared width corpus.
Surface 6 asserts the same contract at the bracket run's slug-adjacent position
(`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so
the bracket path now fails the same gate the other five do.
A second block pins the DELIBERATE half — the wider canonical width at the
delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out
of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would
look like a cleanup rather than a regression.
The generative property ties the READ side to the EMIT side metamorphically: for
every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that
id's numeric run — no more, no less. It needed a new arbitrary: the existing
slugArb generates one [a-z0-9] word and so can never produce the number-leading
slug the collision requires.
Probe-falsified, both directions (probes reverted):
- reverting the source to the old unbounded `\d+` fails 8: the parity gate
reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the
review's scenario verbatim — and the property shrinks to
["A",1,1,undefined,"100-a"].
- interpolating the seam at EVERY position (the rejected verbatim option) leaves
the repro and parity green but fails the divergence gate `'02' !== '02-105'`
and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the
evidence that a verbatim cap under-collects ids toDir emits.
Width 2 stays green under both probes — the corpus agrees with the owner exactly
where the old and new rules coincide, so the gate discriminates rather than
merely mirroring the regex.
Refs #2249
* docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major)
* test(#2249): pin deterministic grammar boundary cases (re-review m1)
PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes
the m1 proof gap — the grammar's bounds were exercised only incidentally
through the fast-check domain (1-999, [a-z0-9] slugs). No source change
(src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged).
Adds a deterministic boundary block (7 describe groups, +22 tests) pinning
the compiled lib's CURRENT behavior — a proof gap, not a behavior gap:
- m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse
(display + dir) -> render/toDir round-trip byte-equality. The plan
position is identity-symmetric (parse/render accept 99/100/101) but toDir
drops it (filename-surface dimension only).
- m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs
99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps
the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of
the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry.
- m1.3 leading-zero 007 -> not-canonical rejection at every position/form.
- m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing
slug (never stored, never mis-read as a plan) and rejects a line
terminator; toDir's allow-list sanitizer collapses each to a safe
[a-z0-9-] token or rejects sanitize-to-empty.
- m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test);
an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE;
an abs-path string is not a bracket id; an abs-path dir slug is dropped to
a clean tuple.
- m1.6 whitespace-only -> not-a-bracket-phase-id.
- m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral):
garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes.
No accept-not-reject case is a src bug: parse never STORES an abusive slug
(dropped from the identity tuple) and toDir independently re-sanitizes on
emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by
parse is the documented positional design (toDir drops the plan; the
read-token caps it) — divergence pinned, not papered over.
Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both
its parse-side and emit-side), ran -> 8 distinct named failures, reverted ->
66/66 green. Confirms every new group executes and can fail.
Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar +
continuation-grammar-parity + collision-characterization + phase-id family
175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green
except one pre-existing, unrelated env failure (npm-integrity-gate: a live
npm-audit advisory in the production dep tree — reproduces with this change
stashed; no package.json/lock change here).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase
These two commands never parse --validate (silent no-op); the flag is
real only for /gsd-quick. Remove the false flag-table rows and CLI
examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/
ko-KR/pt-BR mirrors), and correct the manager.flags.execute example
from --validate to --cross-ai (a flag execute-phase actually parses).
/gsd-quick's real --validate docs are left untouched.
Ref #2197
* docs(#2197): add changeset for --validate docs removal
---------
Co-authored-by: CI Rebase Check <ci@gsd-redux>
* test(#2576): add failing regression for padded resolves_phase compare
* fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos
* chore(#2576): backfill changeset pr to 2597
* test(#2576): drop fast-check property test (cross-platform-fragile on Windows CI)
lint-fix-has-regression-test.cjs: new gate that fails if a fix(#NNNN)
or feat(#NNNN) commit has zero behavioral test files (*.test.cjs,
excluding auto-generated fixtures/baselines) in its diff. Wired into
lint:ci so it runs before PR creation.
Missing regression tests added:
- #2429: codex local scope does not set $HOME/.agents skills home;
global scope does (tests/runtime-artifact-layout.test.cjs)
- #2279: map-codebase instructions say to overwrite existing dates,
not just replace [YYYY-MM-DD] placeholders (tests/commands.test.cjs)
* test(#2474): update dispatch gate test for dual-gate behavior
The #2772 test asserted the gate reads USE_WORKTREES_FOR_PLAN only.
Update to accept the dual-gate (USE_WORKTREES + USE_WORKTREES_FOR_PLAN).
* fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too
The per-plan dispatch condition checked only USE_WORKTREES_FOR_PLAN
(submodule-derived), ignoring the project-level USE_WORKTREES flag.
Add USE_WORKTREES to the gate. Net-negative edit: compress two
nearby prose lines to offset the added shell condition (93353 bytes,
down from 93368).
Closes#2474
* docs(#2474): backfill changeset PR number (2561)
* fix: merge coverage gate into single-process check (#2474)
The test:coverage:unit script chained two c8 invocations with &&:
the first ran tests and wrote coverage data to .nyc_output/, the
second read that data for per-file branch checks. On fast CI runners
(ubuntu/24), the second process started before the filesystem flushed
the first process's writes — a classic TOCTOU race that caused
intermittent coverage gate failures.
Replace the two-process chain with a single c8 invocation that
generates both text and json-summary reports, followed by a Node
script (scripts/check-coverage-gate.cjs) that reads the JSON summary
once and checks both overall and per-file thresholds. No filesystem
race is possible because the JSON report is fully written before the
check script reads it.
* fix(#2429): scope Codex skills home override to --global only
The skills-kind home override (redirecting skills to $HOME/.agents) was
applied regardless of scope. Gate it behind scope === 'global' so
--local installs keep skills project-local under the config directory.
Closes#2429
* docs(#2429): backfill changeset PR number (2553)
* fix(#2400): warn on planned-phase no-op + sync progress.total_plans
Bug A: When STATE.md Current Position has no recognized labels (narrative
prose), emit a warning field so the workflow detects the no-op instead
of continuing with stale state.
Bug B: Sync progress.total_plans in the YAML frontmatter when a plan
count is provided, preventing contradictory state between frontmatter
(0) and body (actual count). This writes the explicitly-provided count,
not a re-derivation from disk (#500 safe).
Closes#2400
* docs(#2400): backfill changeset PR number (2552)
* test(#2366): regression tests for parseCoverageMatrix scoping bugs
Bug 1: summary table outside matrix not parsed as data
Bug 2: multi-section matrix with repeated headers parses correctly
Bug 3: markdown emphasis on decision cell is stripped
* fix(#2366): scope parseCoverageMatrix to recognized coverage tables
Replace latching sawHeader with contextual inMatrix tracking that
resets on non-pipe lines, preventing summary tables from being parsed
as data (bug 1). Allow multiple headers for multi-section matrices
(bug 2). Strip markdown emphasis from decision cells before validation
(bug 3).
Closes#2366
* fix(#2366): update representative-corpus test to expect correct behavior
The test previously documented the known-buggy parseCoverageMatrix behavior.
Now that the fix is in place, test against the expected correct output
(expectedBlock, expectedCounts, expectedErrorCount) instead of the
currentBuggyOutput snapshot.
* docs(#2366): backfill changeset PR number (2551)
* fix(#2279): reword date stamping to overwrite existing dates on Update runs
The map-codebase agent and workflow instructions only said to replace
[YYYY-MM-DD] placeholders, but Update-path files already contain concrete
dates from the prior run. Reword to SET the date stamps unconditionally,
overwriting whatever date is already there.
Closes#2279
* docs(#2279): backfill changeset PR number (2550)
* test(#2269): regression test for query commit --files scoping
Verify secure-phase.md, validate-phase.md, and next.md all pass --files
to their query commit calls.
* fix(#2269): add --files to three unscoped query commit call sites
secure-phase.md, validate-phase.md, and next.md were the only 3 of 65
query commit call sites that omitted --files, causing blanket staging
of .planning/ and committing unrelated files. Add --files with the
specific artifact path to each.
Closes#2269
* docs(#2269): backfill changeset PR number (2549)
* test(#1995): regression test for agent-<id> branch namespace
Add failing-first tests proving that normalizeCleanupManifestEntry and
planWorktreeRecordAgent reject Claude Code's current agent-<id> isolation
branches (only worktree-agent-<id> is accepted). Boundary tests cover both
namespaces plus rejection cases.
* fix(#1995): widen worktree branch regex to accept agent-<id> namespace
Claude Code's isolation="worktree" branch naming changed from
worktree-agent-<id> to agent-<id>. Widen the regex in all 7 locations
from ^worktree-agent-[A-Za-z0-9._/-]+$ to ^(worktree-)?agent-[A-Za-z0-9._/-]+$
so both namespaces are accepted. Introduce a shared WORKTREE_AGENT_BRANCH_RE
constant in src/worktree-safety.cts to prevent future drift.
Closes#1995
* fix(#1995): update workflow guards, test assertions, and baselines
Widen the branch-check regex in execute-phase.md and execute-plan.md.
Update all test assertions that checked for ^worktree-agent- to expect
the widened ^(worktree-)?agent- pattern. Regenerate golden-install-parity
fixtures, agent-size-baseline, and workflow-size-baseline.
Closes#1995
* fix(#1995): update extractCwdGuardBash sanity check for widened regex
The e2e test's sanity check verified the extracted bash block contained
'worktree-agent-'. After widening to '(worktree-)?agent-', update the
check to match the new pattern.
* fix(#1995): widen missed workflow-guard branch check + changeset + lint fixes
- hooks/gsd-workflow-guard.js: widen startsWith('worktree-agent-') to
/^(worktree-)?agent-/ regex — same defect class, was missed in prior commit
- tests/worktree.test.cjs: fix indentation regression from prior edit
- Add .changeset/1995-worktree-agent-branch-namespace.md (pr:0 placeholder)
Found by orthogonal code review (Step 4).
* fix(#1995): regenerate golden + size baselines for workflow-guard change
* docs(#1995): backfill changeset PR number (2548)
The kimi-variant-disambiguation test set only HOME in the spawnSync env,
but os.homedir() on Windows resolves USERPROFILE — so the installer never
found the probe config files and the 'variant mismatch' warning never
fired. Every other installer test in the repo sets both HOME and
USERPROFILE (agent-skills, augment-upgrades, antigravity-upgrades, etc.);
this one was newly written for Phase 5 and missed the pattern.
Also: both new test files carried allow-test-rule exemptions without the
'see #NNN' issue reference required by ADR-456, failing lint-allow-test-rule-refs.
* fix(#2460): pi before_provider_request fail-opens without explicit override
pi/gsd.cjs's buildBeforeProviderRequestHandler unconditionally rewrote
payload.model to the built-in pi/sonnet tier default (claude-sonnet-5)
via resolveTierEntry's catalog fallback. For any pi user on a non-Anthropic
provider (kimi-coding, zai, openrouter, openai-codex, minimax, ...), this
silently broke every request: pi's chosen model was replaced with one the
active provider did not know.
The fix inspects model_profile_overrides.pi[tier] explicitly BEFORE calling
resolveTierEntry (which falls back to the built-in catalog and would mask
the 'user did not opt in' signal). When the user has not set an override
(or set it to null), the handler returns undefined — fail-open — and pi's
chosen model flows through untouched. Only an explicit opt-in via
model_profile_overrides.pi[tier] steers.
Tests:
- the ACTUALLY-REGISTERED handler fail-opens when no override configured
(was: 'steers to default-tier model-catalog pi id' — encoded the bug).
- new test reproducing the reporter's exact repro (model: 'k3' → undefined).
- override path: explicit model_profile_overrides.pi.sonnet config steers
to the user-configured model id.
- defensive: explicit null override also fail-opens.
Per the reporter's suggested fix#1 of #2460.
* test(#2460): regen pi golden parity + install tree fixtures
pi/gsd.cjs changed → pi install hash changed → regenerate the parity
+ install-tree fixtures via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
* fix(#2460): treat empty-string override as fail-open (M1 review)
Per code-review M1 + security M1: an explicit empty-string override
(`{ pi: { sonnet: "" } }`) silently bypassed the fail-open guard because
the check was `=== undefined || === null` only. resolveTierEntry's falsy
`if (userRaw)` then fell back to the built-in catalog and rewrote
payload.model to claude-sonnet-5 — re-introducing the exact bug this PR
fixes, via a degenerate config shape.
Fix: widen the guard to also reject `''`. The test now exercises both
null and '' in a loop, asserting fail-open for both.
* fix(#2460): clear hono/@hono/node-server moderate advisories via npm override
GHSA-v422-hmwv-36x6-class advisories (3 moderate) appeared during this
PR's session:
- @hono/node-server <2.0.5 (path traversal on Windows via encoded paths)
- hono 4.3.3 - 4.12.26 (API Gateway v1 adapter drops distinct repeated
request header values)
- both transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol
/sdk@1.29.0
The npm audit fix re-resolved hono to 4.12.31 (within the existing ^4.11.4
range declared by MCP SDK), clearing the hono advisory without an override.
The @hono/node-server advisory cannot be re-resolved the same way: MCP SDK
pins @hono/node-server@^1.19.9, and the fix requires 2.0.5+. There is no
MCP SDK release that allows @hono/node-server@2.x (latest 1.29.0 is the
most recent), and bumping @anthropic-ai/claude-agent-sdk to 0.3.x does
not help (its peerDependency is still @modelcontextprotocol/sdk@^1.29.0).
The override (sibling to the existing 'qs' and 'body-parser' entries) is
therefore the only available tool — distinct from the body-parser case in
re-resolution.
Verified: npm audit --omit=dev reports 0/0/0/0 advisories.
Also: regenerated the pi golden-install-parity fixture (the pi/gsd.cjs
change in this PR altered the pi install hash).
* docs(changeset): add Fixed fragment for #2460 PR
The two-otters-jog.md changeset was created earlier but lost during the
cherry-pick detour to fix#2454 PR 1's npm advisory cascade. Recreating
here with PR number 2499 backfilled (no placeholder cycle needed).
* feat(#2454): PR 2 — cmdAgentSkills fallback reads installed agent prompt
When no agent_skills config entry exists for a given agent type (the common
case on AGENTS-native runtimes), cmdAgentSkills previously returned empty
output. Workflows that inject ${AGENT_SKILLS_*} into subagent dispatch
prompts then carried nothing — the persona was lost.
The fallback: resolve the runtime's agents directory via checkAgentsInstalled
and read <agentsDir>/<agentType>.md. The installed agent prompt content
(now present for kimi-code via the flat-skills install layout) flows into
the dispatch prompt so the persona survives even without explicit config
opt-in. This is the reporter's suggested fix#2 from #2454.
The fallback triggers for ALL runtimes (not just kimi-code) when no config
entry exists — it is strictly additive (returns content the previous empty
path could not). If the agent file is not found on disk, the block stays
empty (same as before).
* docs(changeset): Phase 3 agent-skills fallback Added (#2510)
* docs(changeset): backfill PR #2521 for Phase 3 (#2510)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout
PR 1 registered the kimi-code EoS descriptor with empty artifactLayout
(SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test).
PR 2 fills in the install surface:
- src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill
function. Today it delegates to convertClaudeCommandToKimiSkill (Python
kimi-cli) because Kimi Code uses the same Agent Skills format + /skill:
invocation per official docs. The distinct function name lets a future
divergence land cleanly if Kimi Code's skill format evolves independently.
- gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS.
- capabilities/kimi-code/capability.json: artifactLayout.global now declares
the skills kind with converter='convertClaudeCommandToKimiCodeSkill' +
home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code
docs: merge_all_available_skills = true default).
- tests/installer-migration-install.integration.test.cjs: REMOVE the
SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface.
- Regenerated capability-registry + capability-matrix + golden install
parity + install tree fixtures for kimi-code.
* fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump
- src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to
SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path
can dispatch off the descriptor's converter string.
- tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count
26 → 27 (added convertClaudeCommandToKimiCodeSkill).
* fix(#2454): remove home override from kimi-code skills (inherit configDir)
The home:'.kimi-code' override made the install plan resolve skills dest
to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's
temp configDir to miss the install. Removing it lets skills inherit
configDir like most runtimes.
* fix(#2454): kimi-code install contract surface is flat-skills (no agents)
Kimi Code has NO custom named subagents (per official docs: 3 built-in
coder/explore/plan only). The kimi-skills-agents surface expects agents/
gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed
to flat-skills which only checks for skills/gsd-* dirs.
* docs(changeset): Phase 2 kimi-code install layout Added (#2509)
* docs(changeset): backfill PR #2520 for Phase 2 (#2509)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI)
PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting
GSD's kimi support into two distinct products per the user's directive:
- kimi (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python)
- kimi-code (new): Moonshot's Node Kimi Code CLI (~/.kimi-code,
runtime: node, KIMI_CODE_HOME env)
Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json
descriptors, not hardcoded branches in install.js. The new descriptor uses
the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml
hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs).
Critical Kimi Code constraint reflected in the descriptor:
hostIntegration.dispatch.namedDispatch: false
hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan']
hostBehaviors.namedSubagentsSupported: false
Kimi Code's official docs confirm only 3 built-in subagents with NO custom-
subagent registration (the [subagent] table only has timeout_ms). The
kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in
kimi-code's artifactLayout.
Schema adjustments:
- subagentToolkit set to 'undocumented' (the existing escape hatch); the
schema enum (full/read-only) lacks a 'limited'/'built-in-only' value.
A follow-up PR can extend the schema enum to add 'built-in-only' as a
first-class axis value reflecting Kimi Code's documented model.
Registration:
- capabilities/kimi-code/capability.json (new descriptor, modeled on codex)
- bin/install.js: allRuntimes array + --all list + --kimi-code flag
- gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases
(kimi-code, kimicode, kimi_code)
- src/runtime-name-policy.cts: FALLBACK_ALIASES map
- gsd-core/bin/lib/capability-registry.cjs: regenerated via
scripts/gen-capability-registry.cjs --write
Tests:
- tests/multi-runtime-select.test.cjs updated for the new runtime count (18)
+ new --kimi-code flag test + 'All' shortcut renumbered 18 → 19.
Out of scope for PR 1 (follow-up PRs in the sequence):
- Install-time decision logic (kimi vs kimi-code detection / prompt)
- agent-install-check semantics for kimi-code (verify Agent Skills presence)
- cmdAgentSkills fallback returning subagent prompt content
- Workflow template mapping (named agents → built-in coder/explore/plan)
- Migration guidance for users currently on 'kimi' who are actually on Kimi Code
- Schema enum extension for subagentToolkit: 'built-in-only'
Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS).
* fix(#2454): complete drift-guard registrations for kimi-code runtime
The drift guards caught every surface that pins runtime enumeration. Each
update is mechanical, driven by the guard's named failure mode:
- src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code
- src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the
isKimiCode predicate generator
- bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream
entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19
- gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry
(null/null/null — same as kimi, no model tier defaults until configured)
- docs/reference/capability-matrix.md: regenerated via
scripts/gen-capability-matrix.cjs --write (kimi-code row added)
- tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP:
kimi-code → '.kimi-code'
- tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden
(the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated)
The capability-registry is already regenerated from the prior commit.
* test(#2454): update drift-guard tests for kimi-code runtime registration
Multiple drift guards pin runtime enumeration counts and option numbering.
Each update is mechanical, driven by the guard's named failure mode:
- tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17);
'all 16 flags' → 'all 17 flags' in test names + messages.
- tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering
cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15,
trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice
test for kimi-code (option 11). Prompt test updated for new numbering.
- tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains
kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs);
EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per
docs, same as Python kimi/opencode).
- tests/global-config-home-fragment.test.cjs: table-count test renamed
13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier).
* fix(#2454): empty artifactLayout for kimi-code (PR 1 scope)
The skills kind requires a converter (existing converters are per-runtime
like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only
registers the descriptor; the actual Agent Skills converter (and a new
'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside
the install-time decision logic. Empty artifactLayout.global is valid and
means 'nothing to install yet via the layout seam'.
Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs
(localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as
option 11 in install.js's buildRuntimePromptText (renumbered downstream
options 11..17 → 12..18, All 18 → 19).
* fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode)
The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved)
for the new kimi-code runtime id. Property names with hyphens are awkward for
consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new
runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing
the conventional PascalCase flag name. The 16 prior single-word runtime ids
are unaffected (the regex finds no hyphens).
* fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures
- tests/runtime-flags.test.cjs drift guard: use proper kebab-case
conversion (isKimiCode → kimi-code, not 'kimicode') so the registry
comparison doesn't false-positive on hyphenated runtime ids.
- tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae
single-choice tests for the renumbered options (kilo 11→12, opencode
12→13, pi 13→14, qwen 14→15, trae 15→16).
- tests/install.test.cjs: Kilo integration option 11→12, prompt test
regex updated.
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
The kimi-code install produces the standard GSD install layout (skills,
contexts, references, etc.) — 436 paths, same shape as other runtimes
that have no custom converter yet.
* fix(#2454): add kimi-code install contract + global config home fragment
- src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code
→ '.kimi-code' so getGlobalConfigHomeFragment returns the correct path
instead of falling through to the default '.claude'.
- tests/installer-migration-install.integration.test.cjs
RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for
PR 1; PR 2 will specialize once the Agent Skills converter lands).
- tests/multi-runtime-select.test.cjs: fix space-separated-choices test
for the renumbered kilo option (11 → 12).
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
kimi-code.json: regenerated after rebasing onto current next (new
planner-reversibility.md from #2471 etc. now included).
* test(#2454): skip kimi-code install contract until PR 2 ships install layout
The end-to-end install test (tests/installer-migration-install.integration
.test.cjs) asserts every allRuntimes entry installs a runtime-specific
artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the
capability descriptor + flags + labels, but the install LAYOUT (Agent
Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands
in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and
self-removing — PR 2 removes the entry alongside adding the install
surface, restoring the contract loop to full coverage.
* fix(#2454): restore compact model-catalog.json format (M1 review)
Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations'
commit used python json.dump(indent=2) which inflated the file from 165→607
lines (every nested entry got expanded) and lost the trailing newline. The
semantic change was just a 3-line kimi-code entry. Restored the original
hybrid format (top-level indent=2 + inner entries' one-line style) and
added kimi-code in matching form.
Regenerated golden install parity + install tree fixtures since the
model-catalog.json hash changed.
* fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code)
CI lint-tests job failed on the glossary drift guard
(scripts/check-glossary-refs.cjs --check):
✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but
bin/install.js's allRuntimes array has 18.
✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js
(missing from CONTEXT.md's list: kimi-code).
Missed in the prior commits because gsd-test does not run the glossary
check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims
to 18 values + kimi-code in the member list.
* chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511)
* docs(changeset): Phase 1 kimi-code runtime Added (#2511)
* test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands
* docs(changeset): backfill PR #2519 for Phase 1 (#2511)
* fix(#2304): normalize Kimi tool vocabulary in PreToolUse guard payload checks
The Kimi [[hooks]] registrations translate the matcher to Kimi's tool
vocabulary (WriteFile|StrReplaceFile) but the guard scripts early-exit
unless the payload's tool_name is a Claude name (Write/Edit/MultiEdit),
so every guard was dormant on Kimi: the matcher fired, the script saw
WriteFile, and exit(0)'d.
Normalize the payload's tool_name at the top of each guard
(WriteFile -> Write, StrReplaceFile -> Edit; bare or module-qualified
kimi_cli.tools.file:* forms) before the check. Inlined per guard rather
than a hooks/lib/ helper because hook scripts are staged as standalone
files on every hook surface, and a sibling require is a staging
dependency that can fail silently.
Regression tests pipe Kimi-vocabulary payloads at each guard and assert
it engages (typed fields: exit status, decision, hookSpecificOutput) —
verified red against the pre-fix scripts, green after.
* fix(#2304): normalize Kimi tool_input fields and route block reasons to stderr
Cross-AI review of the initial fix, verified against kimi-cli source,
found the tool_name normalization alone leaves the guards dormant on a
real Kimi runtime: kimi-cli forwards tool_input verbatim
(src/kimi_cli/hooks/events.py), and its tool schemas
(src/kimi_cli/tools/file/{write,replace}.py) use path/content and
edit.old/edit.new (single Edit or list) — not Claude's
file_path/old_string/new_string. The guards read file_path, got '',
and exited 0 past the now-open tool_name gate.
Extend the per-guard normalization to the payload fields
(path -> file_path, edit -> old_string/new_string with list flattening),
and write the worktree guard's block reason to stderr as well as the
stdout JSON — Kimi feeds stderr, not stdout, back to the model on
exit 2 (docs/en/customization/hooks.md exit-code table).
Regression tests rewritten to Kimi's actual payload shapes (plus an
edit-list case and a stderr-reason assertion) — verified red against
the name-only fix, green after.
* fix(#2304): join all edit[] entries into old_string, matching new_string
Review nit on #2326: old_string took only edits[0].old while new_string
joined the whole list. Symmetric join removes the latent trap for any
future consumer sizing before/after content (e.g. the #2255 write guard).
* fix(#2304): normalize Kimi ReadFile vocabulary in read-injection scanner
Review Major 2 on #2326: gsd-read-injection-scanner.js had the identical
dormancy — its Kimi matcher fires on 'ReadFile' but the SCANNED_TOOLS
check only knew 'Read', so injected content in read files was never
flagged on Kimi installs.
Folds the same inlined normalization block into the scanner and extends
the shared KIMI_TOOL_NAMES map with ReadFile:'Read' in all four copies so
they stay byte-identical. Harmless in the three write guards: a
normalized 'Read' falls out of their Write/Edit allowlist exactly as the
unmapped name did. Field mapping verified against kimi-cli upstream
(src/kimi_cli/tools/file/read.py Params.path); the existing
path->file_path copy covers the scanner's file_path read.
* test(#2304): parity test binding the four inlined Kimi normalization copies
Review Major 1 on #2326: KIMI_TOOL_NAMES + normalizeKimiPayload is
deliberately inlined in four hook scripts (staging-dependency rationale,
unchanged), with the inverse table in bin/install.js — five
hand-maintained surfaces and nothing binding them.
Static binding, zero runtime coupling:
- the four inlined blocks must be byte-identical;
- each guard-map entry must be the value-inverse of
convertKimiToolName() for its Claude name;
- every guard-relevant Claude tool (Write/Edit/MultiEdit/Read) must have
a reverse entry — a vocabulary rename or extension that updates the
installer without updating the guards now fails in CI instead of
leaving a guard silently dormant (the #2304 recurrence door).
Negative-controlled: diverging one copy or dropping a map entry fails
the suite against the fixed code.
* test(#2304): regenerate golden parity fixtures for guard hook changes
CI red on #2326: all 10 golden-parity failures were the staged guard
hooks drifting from their fixtures. Regenerated with npm run gen:golden
(after npm run build) under throwaway HOME/CLAUDE_CONFIG_DIR; diff
verified to change exactly the four PR-touched guard entries per
surface, nothing else.
* test(#2304): regression tests for Kimi ReadFile engaging the scanner
Mirrors the per-guard Kimi vocabulary tests the PR added for the three
write guards: bare and module-qualified ReadFile produce the advisory,
path exclusions still apply post-normalization, unknown Kimi names stay
fail-open. Negative-controlled against the pre-fold scanner (the two
positive cases fail there; exclusion/fall-through correctly pass on
both sides).
* fix(#2304): normalize Kimi Shell vocabulary in workflow guard
Withdraws the disclosed out-of-scope split: verification showed the
Bash->Shell case needs NO different mapping — kimi-cli's Shell.Params
names its field `command` (src/kimi_cli/tools/shell/__init__.py), same
as Claude's Bash — and the guard's write branch (Write/Edit/MultiEdit
allowlist) was ALSO dormant on Kimi under its Shell|WriteFile|
StrReplaceFile matcher. Same defect class as the other four hooks.
Folds the identical inlined block into gsd-workflow-guard.js and
extends the shared map with Shell:'Bash' in all five copies (harmless
outside the workflow guard: a normalized Bash falls out of the other
guards' checks as before). Parity test now binds five copies and adds
Bash to the dormancy alarm. New workflow-guard test file exercises the
observable block (force-add on a worktree-agent branch): Shell bare and
module-qualified block with WORKTREE_AGENT_FORCE_ADD_FORBIDDEN, benign
Shell passes, Claude Bash unchanged — negative-controlled against the
pre-fold guard (the two Kimi cases fail there). Golden parity fixtures
regenerated; diff verified to change exactly the five guard entries per
surface.
* fix(#2304): map Kimi tool_output and route workflow-guard block to stderr
Third-party review (cross-AI verifier) caught two gaps in the revision:
1. Kimi PostToolUse events carry `tool_output`, not `tool_response`
(kimi-cli src/kimi_cli/hooks/events.py post_tool_use()), so the
read-injection scanner — which reads data.tool_response — was STILL
dormant on real Kimi payloads; the earlier tests passed because they
sent Claude-shaped payloads. The shared normalization block now maps
tool_output -> tool_response (inert in PreToolUse guards, where the
field is absent), and the scanner's Kimi tests send the real shape.
2. The workflow guard's force-add block wrote its reason to stdout only.
Kimi's exit-2 protocol feeds stderr back to the model — the exact
fix this PR already applied to the other blocking guard — so the
newly-awakened block would have been a silent denial. Reason now
also routed to stderr, asserted in the test.
Also: the scanner's "unknown name" test now uses a genuinely unmapped
name (FetchURL) — Shell stopped qualifying when it entered the map —
and the workflow guard's write branch (WriteFile advisory,
StrReplaceFile .planning pass) gains behavioral coverage. All five
copies stay byte-identical (parity test green); golden fixtures
regenerated, diff verified to the five guard entries per surface.
Negative-controlled: 3 new assertions fail against the pre-fix hooks.
* docs(#2304): update changeset to cover the full five-guard fix
Review round 2 (2026-07-18) flagged the changeset as stale: it was
written for the first commit and still described only the three guards
named in the issue. The shipped diff grew to five guards plus two
payload dimensions the original body never mentioned. The body now
names gsd-read-injection-scanner and gsd-workflow-guard, the ReadFile
and Shell vocabulary entries, the tool_output -> tool_response mapping,
and the workflow guard's stderr block-reason routing.
* test(#2304): regenerate kilo golden fixture after #2305 landed on next
The branch's fixture sweep predates 50efae13 (fix(#2305), PR #2327),
which made Kilo ship the five shared guard hooks. Rebased onto next and
re-ran the full generator sweep (gen:golden, size:baseline, and the
four registry/contract generators); the only delta across all of them
is kilo.json's five guard-hook hashes, matching this PR's hook edits.
* fix(#2304): fold Kimi normalization into the two shell hooks
The 2026-07-19 review found the last two guards with the #2304 dormancy:
- hooks/gsd-graphify-update.sh gated on tool_name == "Bash" but is
registered on Kimi with matcher 'Shell' — Gate 1 never matched and the
auto-rebuild was silently dormant. kimi-cli's Shell.Params names its
field `command` (src/kimi_cli/tools/shell/__init__.py), same as Claude
Bash, so only the name needs mapping: strip the module-path prefix,
map Shell -> Bash.
- hooks/gsd-phase-boundary.sh read only tool_input.file_path, but Kimi's
file tools name the field `path` (src/kimi_cli/tools/file/write.py +
replace.py) — the hook read '' and .planning/ writes went undetected.
Falls back to tool_input.path when file_path is absent, mirroring
normalizeKimiPayload's precedence in the JS guards.
The normalization is reimplemented in shell — a byte-identity assertion
cannot span the JS<->shell boundary, so the parity test gains a
shell-guard vocabulary block that pins both scripts' mapping facts to
convertKimiToolName's live vocabulary instead of faking a byte binding.
Behavior is covered by negative-controlled tests beside each hook's
existing suite (verified red against the pre-fix scripts): Kimi Shell
dispatch (bare + module-qualified) with a WriteFile negative control in
graphify-auto-update.slow.test.cjs, and Kimi path detection, file_path
precedence, and a non-.planning negative control in hooks-opt-in.test.cjs.
Changeset updated to name all seven guards; golden install-parity
fixtures regenerated (diff is exactly the two hook entries per runtime;
size baselines unchanged).
* fix(#2304): use a Map for KIMI_TOOL_NAMES so prototype keys cannot pass the guard fall-through
A bare bracket lookup on an object literal resolves 'constructor',
'__proto__', 'toString', 'valueOf' and 'hasOwnProperty' through
Object.prototype to truthy functions/objects, so `if (!mapped)` failed
to short-circuit and data.tool_name was assigned a non-string. Map.get
returns undefined for those keys — the same shape the repo already uses
in canonicalizeRuntimeName (src/runtime-name-policy.cts). Applied
identically to all five inlined copies (review M1, PR #2326).
No new bypass class: unrecognized strings already fail open by design;
this fixes the lookup being wrong, not the posture.
* test(#2304): enumerate normalized guards by scanning hooks/, not a hardcoded list
The parity test's file list was a literal five-entry array — a sixth guard
with its own copy-pasted normalization block would be silently uncovered,
the exact divergence mode the test exists to prevent (review M2). Now the
list is a scan of hooks/*.js for the KIMI_TOOL_NAMES marker, with a floor
assertion so a scan that finds nothing fails instead of passing vacuously.
Also parses the Map declaration introduced by the M1 fix, and carries the
allow-test-rule annotation documenting the source-text scanning (review m4).
* test(#2304): parse hook JSON output instead of substring-matching raw stdout
workflow-guard.test.cjs asserted on unparsed stdout while read-guard.test.cjs
in the same PR parses the JSON envelope first — match the better pattern at
all four assertion sites (review m5).
* test(#2304): regenerate golden parity fixtures after Map conversion in the five guards
* docs(#2304): reset changeset pr:0 placeholder for Phase 0 PR (#2507)
The closed PR #2326's changeset carried pr:2326. Phase 0 of epic #2505
re-lands this fix on a fresh branch; the pr: field will be backfilled
to the real Phase 0 PR number immediately after gh pr create returns.
* docs(changeset): backfill PR #2518 for Phase 0 (#2507)
---------
Co-authored-by: 0xdhx <darkhawkx@gmail.com>
The finalize job created the release/hotfix -> main merge-back PR but
never merged it, so every release and hotfix needed a manual merge click
on main. The opposite direction (main -> next) is already admin-merged by
auto-backmerge.yml; this direction was the asymmetric manual step where
the recurring divergence got hand-reconciled.
Adds a finalize step (after "Verify publish", so main only absorbs a
confirmed-published release) that finds the open merge-back PR, polls
until GitHub settles its mergeability, and admin-merges it with a merge
commit ONLY when MERGEABLE. A CONFLICTING PR is left open for manual
resolution rather than force-merged. Non-fatal (continue-on-error): the
tag + npm publish already happened, so a merge-back that can't complete
(org PR policy, token) must not fail the release.
With the main-is-ancestor-of-next invariant restored (#2504), this merge
is clean every release -- verified: a simulated 1.8.1 hotfix merges to
main producing [1.8.1]->[1.8.0]->[1.7.0]->[1.6.1] and version 1.8.1 with
no conflict, and main's hardened auto-backmerge.yml survives the merge.
Guarded by three assertions in release-backmerge-invariants.test.cjs
(step present + admin-merge, gates on MERGEABLE, continue-on-error);
verified they fail when --admin or the MERGEABLE guard or the
continue-on-error is removed.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Delivers the #2506 blast-radius fix to main NOW instead of deferring to
the next release. Deferring specifically fails for a hotfix: 1.8.x hotfix
branches are cut from the immutable v1.8.0 tag (which lacks this
hardening), so a hotfix would carry the un-hardened workflow to main and
never converge it. Converging now makes main's backmerge robust
regardless of whether the next release is a minor or a hotfix.
The only functional change is `continue-on-error: true` on the two
version-sync steps (verified byte-diff vs next). main and next are now
identical on auto-backmerge.yml.
The main->next auto-backmerge fails after nearly every release, leaving
main not an ancestor of next, so the following release->main merge-back
conflicts. Root cause is a copy-shuffling loop: auto-backmerge.yml must be
identical on main and next, but `-s ours` (main->next) and the release-tree
merge-back (release->main) each overwrite one copy wholesale, so a fix
applied to one copy is repeatedly overwritten by the copy that lacks it.
The build:lib step proves it: added to main (329233fc8), overwritten by
the 1.7.0 merge-back, re-added to next (#2281), never on old-main -> the
1.7.0 backmerge ran on main's broken copy and failed at "Sync next's
version" (npm version -> gen-capability-registry needs the gitignored
capability-ledger from build:lib -> absent -> step fails -> "Open PR"
skipped -> no PR -> main never becomes an ancestor of next).
Two-part durable fix:
1. Blast-radius containment: mark the version-sync steps continue-on-error.
The job's load-bearing purpose is opening + admin-merging the back-merge
PR (the ancestry that keeps release->main clean). A version-sync failure
(missing build:lib after a copy regression, or any npm-version lifecycle
hiccup) can no longer abort that PR. A sync failure now costs only a
stale next version, trivially re-synced -- never a broken back-merge.
2. Required-steps gate: tests/release-backmerge-invariants.test.cjs parses
the workflow YAML and asserts build:lib runs before the version-sync,
both steps are continue-on-error, the ancestry steps exist, and the
finalize timeout is >= 30 (sibling #2281 regression). It runs on every
branch, so a PR shipping a fix-less copy fails at PR time instead of at
release time -- which is exactly what #1855/#1928/#1990 did undetected.
Verified the test fails on both regression modes (continue-on-error removed;
build:lib step removed) and passes on the fixed workflow.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Reconciles main to the v1.8.0 tree while preserving the [1.7.0] CHANGELOG
section that release/1.8.0 omitted (root-caused in #2502: the main->next
auto-backmerge failed after 1.7.0, so next never received 1.7.0's promoted
CHANGELOG). Both parents kept so the v1.7.0 and v1.6.1 tags remain in
main's ancestry. This also delivers the fixed auto-backmerge.yml (build:lib
step, #2281) to main so the main->next backmerge stops failing.
* fix(#2488): strip leading terminators so changeset bullets survive re-parse
A fragment body beginning with a line terminator rendered as an empty
`- ` bullet followed by an orphaned paragraph. `parseChangelog` treats a
non-indented line as terminating a bullet, so `github-release-notes.cjs`
silently dropped the entry when re-parsing CHANGELOG.md to build the
GitHub Release body.
Two independent causes, both in scripts/changeset/parse.cjs:
1. `extractDocsExempt` stripped trailing terminators but not leading
ones. `DOCS_EXEMPT_RE` is `^...$` under /m, so removing a first-line
`<!-- docs-exempt -->` marker left the `\n` that `$` does not consume.
2. `parseFragment` preserved the post-frontmatter body verbatim, so a
blank line between the closing `---` and the first content line
produced the same leading `\n` with no marker involved.
8 of 256 pending fragments were affected, split 4/4 across the two
causes — including the OpenCode MCP binding, the pi extension, and the
EoS adapters, all of which would have vanished from the v1.8.0 release
notes.
Regression tests cover both causes in LF and CRLF form, plus an
end-to-end serializeChangelog -> parseChangelog round-trip that pins the
user-visible defect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2488): regenerate golden install fixtures for parse.cjs
scripts/ ships in the npm package and the installer, so the golden
install-parity fixtures record a content hash for every shipped file.
Editing scripts/changeset/parse.cjs drifts that hash and fails all 18
per-runtime parity tests.
Regenerated via `npm run gen:golden`. The diff is exactly one line per
fixture — the scripts/changeset/parse.cjs hash — with no unrelated drift.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
lint-allow-test-rule-refs (ADR-456) requires a #NNN issue ref on the SAME line
as allow-test-rule:. CI-only lint (lint:ci), so it passed the local lint gate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>