45 Commits

Author SHA1 Message Date
Jakub Zych
bb6adb3200 chore: release 2.0.0
Some checks failed
Tests / conformance test (macos-latest, 24) (pull_request) Blocked by required conditions
Tests / conformance test (windows-latest, 24, shard 1/3) (pull_request) Blocked by required conditions
Tests / conformance test (windows-latest, 24, shard 2/3) (pull_request) Blocked by required conditions
Tests / conformance test (windows-latest, 24, shard 3/3) (pull_request) Blocked by required conditions
Tests / Required tests (pull_request) Blocked by required conditions
Changeset Required / PR mergeability (pull_request) Successful in 17s
Default Flip Documentation / PR mergeability (pull_request) Successful in 10s
Dependabot Auto-Merge / auto-merge (pull_request) Has been skipped
Docs Required / PR mergeability (pull_request) Successful in 10s
Mutation Testing / PR mergeability (pull_request) Successful in 10s
Security Scan / PR mergeability (pull_request) Successful in 12s
Tests / PR mergeability (pull_request) Successful in 9s
Tests / Base branch health (pull_request) Successful in 10s
Tests / Detect test scope (pull_request) Successful in 17s
PR Target Validator / validate-target (pull_request_target) Successful in 14s
Branch Cleanup / Delete merged PR branch (pull_request) Failing after 10s
Branch Cleanup / Weekly orphaned branch sweep (pull_request) Has been skipped
Changeset Required / changeset-lint (pull_request) Successful in 29s
Default Flip Documentation / default-flip-documentation (pull_request) Successful in 30s
Docs Required / docs-lint (pull_request) Successful in 5m1s
Mutation Testing / Detect changed covered modules (pull_request) Successful in 19s
Security Scan / security (pull_request) Successful in 5m5s
Tests / lint-tests (pull_request) Failing after 1m48s
Tests / plugin-validate (pull_request) Successful in 1m4s
Tests / test (ubuntu-latest, 24, shard 1/3) (pull_request) Failing after 19s
Tests / test (ubuntu-latest, 24, shard 2/3) (pull_request) Failing after 20s
Tests / test (ubuntu-latest, 24, shard 3/3) (pull_request) Failing after 19s
Tests / test (ubuntu-latest, 24) (pull_request) Failing after 20s
Tests / test (inert CI) (pull_request) Has been skipped
Tests / QA loop walk (smell ratchet) (pull_request) Failing after 20s
Mutation Testing / Stryker (${{ matrix.name }}) (pull_request) Has been skipped
Tests / Coverage gate (merged shards) (pull_request) Has been skipped
Tests / Publish emitted-baseline artifact (pull_request) Has been skipped
Mutation Testing / Stryker mutation score (changed files only) (pull_request) Has been skipped
2026-10-09 10:54:05 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
github-actions[bot]
3697b94655 chore: sync next package version to 1.14.0 2026-09-14 00:01:49 +00:00
github-actions[bot]
b3906c66f6 chore: sync next package version to 1.13.0 2026-09-06 02:10:29 +00:00
Behruz Nassre Esfahani
5ad9a36f35 fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by
spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id
was a hardcoded literal, so `--host` chose only the argv RENDERING while the
LEVEL always came from the installed plan-checker's frontmatter — `low` under
every shipped model profile. Every prompt-fed lane therefore ran at a fast
structural verifier's effort, and because the rendered argument is a CLI config
override it silently beat the effort the operator had configured for that CLI.
At `low` a large source-grounded prompt makes a model end its turn with no final
message, so the lane came back empty and its stub read as a crash.

Effort is a property of the review, so the lane declares it. Two new fields on
ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` —
carried through each capability manifest and the generated registry, set on the
three lanes with an argv effort channel and null on the other nine. A new pure
`resolveLaneEffort()` resolves config key -> lane default -> nothing, where
"nothing" emits no effort argument at all and the reviewer CLI's own
configuration decides; `inherit` selects that path explicitly and an
unrecognized level falls back to the lane default rather than being forwarded to
a CLI that would reject it. The host's negotiated effortSurface still gates the
rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving
in-process also removes up to twelve subprocess spawns per review.

The empty-output stub now names the effort the lane ran at and distinguishes a
clean exit from a timeout kill, a non-zero exit, and a process that never ran —
`status` is null for both a timeout and a signal, so those were indistinguishable
before. The hint is hedged: a clean empty exit is most often a model stopping
short, but it is also consistent with a CLI writing its output elsewhere.

Also: the capability validator now knows both fields, rejects a malformed key or
an out-of-vocabulary default, and rejects a default declared without a config
key (a level the operator could never override). An existing end-to-end row in
tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures
the lane's own key and pins the decoupling in the same real spawn, with the
agent execution tier set to a level that must not appear.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long.

Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive.

Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:25:44 -04:00
Tom Boucher
c0fd2e3f4c feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route

Extends tests/host-integration.test.cjs with negotiateHostCapabilities
maxConcurrency negotiation coverage (test matrix rows 1-15, including a
fast-check property test) and a new #3673 dispatch-capacity CLI route
describe block spawning the real gsd-tools.cjs (rows 16-25). Extends
tests/host-integration-validator-parity.test.cjs with an all-19-descriptor
maxConcurrency presence/validity sweep (row 26) and adds a hostile-input
validator test to host-integration.test.cjs (row 27).

The dispatch.maxConcurrency field does not exist yet, so these tests fail.

* feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query

Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration
Interface (ADR-1239 Phase 1), following the existing dispatch.isolation
sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES
floors, and a negotiateHostCapabilities branch that passes through a positive
safe integer and fails closed to 1 otherwise (no engine-side reduction, per
the design doc's explicit rejection of a min(host,engine) rule).

capability-validator.cjs gains parity validation for the new field (optional,
positive safe integer or the "undocumented" sentinel — mirroring isolation's
"added after existing descriptors" treatment).

gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling
of `dispatch-isolation` with no side effects: live env
(GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence.

All 19 capabilities/*/capability.json descriptors now declare
dispatch.maxConcurrency: claude carries the one cited value (20, per
code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented"
(not yet researched for this axis).

* docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix

Updates docs/reference/host-integration-interface.md's dispatch struct entry
(also backfilling the previously-undocumented isolation/backgroundDispatch
sub-fields found stale in the same table) and adds fail-closed/live-transport
precedence prose for the new maxConcurrency field.

Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in
docs/reference/host-integration-capability-matrix.md — required for
tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check,
which asserts every declared dispatch sub-axis is documented there.

Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist
and example descriptor block to mention maxConcurrency (and, likewise
backfilling a stale gap, isolation/backgroundDispatch).

* fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check

Exports isPositiveSafeInteger from src/host-integration.cts as the single
source of truth for the dispatch.maxConcurrency positive-safe-integer
contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity
now both call it instead of independently reimplementing the same predicate.

Also removes the __proto__/constructor/prototype reserved-name branch from
capability-validator.cjs's maxConcurrency check — copy-pasted from the
string-enum fields above it, but unreachable for a numeric field (the
generic positive-safe-integer branch already rejects any string) and absent
from maxDepth, the field the code's own comment claims to mirror.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:12 -04:00
Tom Boucher
62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00
github-actions[bot]
4a705cf19f chore: sync next package version to 1.12.0 2026-08-30 03:07:03 +00:00
Tom Boucher
aaf47c5fc2 fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget

No prompt cap can reach any CLI reviewer lane, by any configuration. Two
independent defects compound: all nine `transport: spawn` lanes declare
`promptBudgetKey: null`, so `budgetFor` returns on its first line; and the
documented global `review.max_prompt_tokens` is advertised in the schema
manifest but declared nowhere, so the resolver never materializes it and
`budgetFor`'s fallback is dead code.

Adds to tests/reviewer-config-federation.test.cjs, which already owns the
per-reviewer budget config-set/config-get idiom:

- a CLI lane inherits the global cap (RED: reports null)
- an http lane with the -1 sentinel inherits the global cap (RED: reports null)
- the resolved review surface carries max_prompt_tokens at all (RED: absent)
- per-lane overrides the global on a CLI lane
- the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as
  unset, 1 is the smallest real budget — the regression budgetFor's own comment
  warns about
- anti-tightening pins that must stay green: an empty config leaves every lane
  null, the three existing budgeted lanes are unchanged, and config-set still
  rejects a per-reviewer key naming something that is not a declared lane
- a fast-check property over the resolution contract itself, with -1, 0 and
  non-finite inputs generated explicitly rather than left to chance

Every row was reproduced by hand against the real CLI before being written, so
the RED/GREEN split is observed rather than predicted.

Refs #3691

* fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve

No prompt cap could reach any CLI reviewer lane, by any configuration. Two
independent defects compounded.

The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor,
gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so
`budgetFor` returned on its first line and `review-lane plan` reported
`promptBudget: null` no matter what was configured. Each now declares
`review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset
sentinel the three local-server lanes already use.

Separately, the central `review.max_prompt_tokens` was listed in the schema
manifest's validKeys and documented as a supported setting, but declared
nowhere — the resolved surface is built from capability declarations plus the
defaults manifest, and neither carried it. `configGet` returned undefined and
`budgetFor`'s documented fallback was dead code. It is now declared with a
`null` default, exactly as docs/CONFIGURATION.md already specified, so the
default behavior is unchanged: nothing configured means nothing trims.

Two things the diagnosis had not predicted, found and fixed while implementing:

- `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded
  registration site that `mergeReviewerLanes` prefers over the capability
  registry on a slug collision. Editing only the capability files left every
  CLI lane still null. Both sites now agree.
- The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked
  the capability edits; regenerated with `npm run gen:capability-registry`
  rather than hand-edited.

docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one —
today ollama, lm_studio and llama_cpp". That is false as of this change and is
corrected rather than left to rot.

The trim-versus-refuse question the issue raises is deliberately not taken up
here: the refusal path already exists for the case that matters — a reviewer
whose minimum set exceeds its budget is skipped rather than sent a misleading
prompt — and trimming above that floor is the documented, shipped design of the
feature. Changing it would alter behavior for the three lanes that already
work, which is not what the issue asks for.

Fixes #3691

* fix(#3691): document the new global and narrow an invariant this change obsoleted

The full suite surfaced two consequences of giving every CLI lane a budget key.

`review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in
the planning-config reference, which config-field-docs guards. Documented,
including the sentinel semantics a reader needs: a per-lane value overrides the
global, `-1` means unset and inherits it, and `0` means "do not trim that lane"
and is not unset.

The #2797 federation guard asserted that "a lane with no model flag and no host
owns no config keys". That held only because budget keys existed solely on the
three local-server lanes, all of which have hosts. A lane can now legitimately
own a config key for a third reason, so qwen tripped it.

The assertion is narrowed rather than weakened: such a lane must still own no
model key and no host key, and may own at most its own
`review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is
strictly more specific in the dimensions that still matter. Proven to still
bite: hypothetically giving qwen a `review.models.qwen` key fails it with
`model/host: review.models.qwen`. The name and comment cite #3691 for why the
premise changed, so a reader sees a deliberate narrowing, not erosion.

Checked the sibling assertions in that describe block; the other three do not
rest on the obsolete premise and are untouched.

Refs #3691

* fix(#3685): port the write-flag content-change contract to its three sibling sites

#3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which
reported `fs.existsSync(path)` rather than whether the transaction wrote
anything. Three sibling sites carried the identical defect and are ported here.

- `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded.
  `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and
  the flag reports it. #2640/#2974 already fixed `state_updated` at this same
  call site and left this one behind, so the correct shape was adjacent.
- `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` —
  byte-identical to #3685's bug in a different command.
- `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never
  consulting the MILESTONES.md write.

`gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display
and never branches on it, so the flip from always-true to content-based changes
no workflow behavior. Verified by reading the step, not assumed.

One trap found while implementing: the obvious in-memory
`finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s
shipped shape verbatim — gives a FALSE POSITIVE for milestone completion.
`platformWriteSync` normalizes Markdown at write time, and the milestone-closure
transform regenerates `## Current Position` fresh on every call, so its
pre-normalize output always differs from the already-normalized file on disk
even when the persisted bytes are identical. The comparison is therefore made
against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are
left untouched — their repeat-no-op tests pass, so they are not exposed to this
artifact.

`milestones_updated` has no reachable no-op: the MILESTONES.md write
unconditionally appends an entry every call. Only the true direction is pinned,
documented inline rather than faked with a passing test.

Refs #3685

* fix(#3685): compare write-flag content through the writer's own normalizer

An independent reviewer disproved a claim made while porting #3685's contract
to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to
the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`.

`platformWriteSync` normalizes on write — CRLF stripped, blank-line runs
collapsed, a blank line inserted after a heading, a single trailing newline
enforced. Every flag that compares the PRE-normalization in-memory string
against the on-disk pre-image can therefore report a change when the persisted
bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading
the file after the write; the other sites compared raw strings.

All of them now go through one exported seam,
`contentChangedAfterNormalize(filePath, before, after)`, which normalizes both
sides exactly as the writer does. That removes the extra disk read the milestone
workaround needed, and makes the sites agree by construction rather than by
four independent implementations of one rule — the divergence the repo names as
an anti-pattern.

Reachability, stated precisely rather than uniformly: the seam is load-bearing
at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and
`stateUpdated`, where section-rewrite logic genuinely regenerates content into a
different-but-normalization-equivalent shape. At
`updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch
never reassigns `content`, so the raw comparison was already correct there. The
first analysis claimed the reverse; this is the corrected finding.

Also fixes an unsound test premise the remote suite caught. The byte-identity
precondition in `roadmap_updated is false when ROADMAP.md comes out
byte-identical` asserted against a hand-authored, un-normalized fixture — so the
very first write reformatted it and the file could not come back identical. The
fixture is now written already-normalized, so the assertion compares a
normalized pre-image against a normalized post-image and still fails if the flag
regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS
too, and the earlier local check simply never exercised it.

The sibling true-direction and milestone tests were checked for the same premise
and do not share it — they assert `notEqual`, or compare two post-write states
produced through the same normalizing seam.

Refs #3685

* chore(changeset): backfill PR number for #3691 fragment

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:39:51 -04:00
github-actions[bot]
97ca69a53e chore: sync next package version to 1.11.0 2026-08-19 13:51:49 +00:00
Tom Boucher
3ab0007164 enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19)

preserveUserArtifacts held user files only in an in-memory Map across the
wipe, so any process death between preserve and restore lost them outright.

Seven call sites, not the four the issue records. Three of them never called
the helper at all - they open-coded the same read/wipe/write - so searching
for callers under-counted by construction; the extra sites were found by
sweeping for the pattern instead.

The worst is the mainline install path, where the crash window spans the
entire gsd-core tree copy rather than a single rmSync.

Adds src/user-artifact-staging.cts: durable on-disk staging with a record
written after the copies land as the commit point, plus recovery of orphaned
batches on the next run - without recovery the staged bytes survive but the
user's file is still gone, which would pass its own test while delivering
nothing.

Routes copyPreservingSymlink through installFs() so staging cannot bypass the
install fs seam, and reunites its symlink-safety docblock with the function it
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-3574 with four claims disproved by implementation

Implementing Phase 6 disproved four statements the ADR rests on. The central
decision - no single materializer - is unaffected and stands.

Corrected: decision 3 was already satisfied, so nothing was extracted; the
agents-bypass runtime set omitted claude, kilo and opencode, and closing it
needed three new pieces of descriptor contract rather than proceeding on its
own terms; three of the four blockers the layout comment names were already
stale; and F19 is seven call sites, not four.

Records the generalizable lesson: the defect is the pattern of holding user
data in memory across a wipe, not the helper, so searching for callers of the
helper under-counts by construction.

Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and
notes that copyPreservingSymlink needed routing through the install fs seam
before it could be reused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close dangling-symlink blind spot and harden staging recovery

An adversarial review found the F19 staging work shipped red and unsafe.

Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween
missed dangling symlinks in both its root check and its per-segment walk,
because it probed with existsSync, which is false for a link whose target does
not exist. Fixing only the new module would have reused a guard that was
itself blind. This guard protects the whole install tree.

Recovery no longer throws: it degrades per entry and per file, so one bad
batch cannot block the others. Previously an unrecoverable entry propagated
out of the first statement of install and uninstall, before the cleanup that
would have removed it - wedging the installer permanently.

Partial fs adapters now throw on any omitted method instead of silently
reaching the real filesystem, closing the trap that let a test poison list
pass while real IO happened.

Staged names must be flat, recovery refuses a dangling destination symlink,
and a batch whose recovery genuinely failed is no longer swept - it was
discarding the only durable copy of the file it had just failed to restore.

Replaces three tests that could not fail, including the one labelled negative
proof.

Known limitation, documented not closed: concurrent installs sharing a staging
key can still lose a batch. A real fix needs a cross-process lock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enh(#2875): make the descriptor authoritative for the agents kind

Deletes the inline agent-staging loop in bin/install.js and the
_DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from
its capability descriptor instead of an inline hostBehaviors dispatch.

Closing it needed three pieces of contract the descriptor pipeline never had,
all reducible to one missing input - per-agent resolution context: a
frontmatter-extensions step for claude's effort and disallowedTools, per-agent
model-override resolution for kilo and opencode, and a named branding
converter for hermes, whose rewrite data was already declared.

Seven runtimes were on the loop, not the six the design recorded - kimi-code
was found by a golden fixture, not by analysis. claude-local and kimi-code
both silently lost their agents mid-change; the fixtures caught both and the
cause was fixed rather than the fixtures regenerated.

A parity harness gates the migration: both pipelines over identical inputs,
byte-identical output including filenames, per runtime. It is demonstrated
red before being trusted. Surface and install paths converge for all seven,
which also fixes surface previously writing no agents for these runtimes.

Codex's config.toml strip stays put - it mutates host config, which no
descriptor kind models.

Also routes install-model-override-resolver and install-effort-resolver
through the install fs seam. Both leaked real filesystem IO from the install
call tree; the stricter adapter is what exposed them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): record the agents-descriptor migration and correct the ADR count

The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host
integration guide told readers to join a set that is gone. Replaces that with
what is now true - declare an agents entry and it installs, on the surface
path as well as install - and points anyone needing a per-agent transform at
the three extension points rather than at a new inline branch.

Corrects the ADR amendment: seven runtimes were on the inline loop, not six.
kimi-code was found by a golden fixture going red, not by reading. That is the
third short count this phase, all from enumerating by symbol or set membership
when the thing that matters is a behavior.

Adds the Changed changeset for the surface-path convergence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-2866 - claude global always wrote agents on disk

The claude row's global=[skills] described what capability.json declared, not
what the installer wrote. bin/install.js's inline agent-staging loop was never
scope-gated and never consulted the descriptor, so a claude --global install
has always written agents/gsd-*.md.

Phase 6 closes the gap by deleting that loop and declaring agents on claude's
descriptor at global scope. On-disk bytes are unchanged - the golden fixtures
did not move, which is the evidence that the descriptor, not the installer,
was incomplete.

#2218 is unaffected: agents are not trigger-bearing, so the wider row does not
introduce a new shadowing case.

Records the warning that an incomplete descriptor is invisible while a second
code path silently does its work, and only surfaces when the two are forced
into agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close review findings across staging, agents and the parity harness

Two independent reviews of this branch found defects the local gates missed.

Security: a dangling symlink at a migration destination allowed writing
outside configDir - the same class this change claimed to close, missed at the
terminal write of the flow being added. The staging-root resolver threw as the
first statement of install and uninstall, so a hostile symlink bricked both,
and symlinked-configDir users lost uninstall as well as install; it now
degrades instead of aborting. Recovery gained a source-side symlink check and
now refuses a relative destDir, which resolved against cwd. Converter dispatch
gained a runtime allowlist - lint-time validation stopped mattering once this
branch promoted that dispatch from the surface path to real installs.

Correctness: claude --local --minimal exited 1 because the minimal profile
legitimately yields zero agents and the new path treated that as a failure.
cline --local silently lost its agents - its descriptor declared none while
the deleted loop wrote them unconditionally. The agents prune was widened to
any gsd-* entry and destroyed user files it never owned.

The parity harness, on which the migration's safety argument rested, drove a
synthetic registry and never byte-compared the shipped descriptors; two of its
trap rows could not fail. It now drives the real registry across 13
runtime-scope rows including kimi-code and cline-local, and its red-proof is
demonstrated by corrupting a live capability.json. Three goldens that had
encoded the cline regression as expected behavior were corrected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close findings from both mandated review engines

/security-review found the staging source-side walk honouring
GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write
destination. A symlinked files/ component dereferenced because
copyPreservingSymlink lstats the leaf only, so an intermediate link is
followed. The source walk no longer honours the opt-in; the destination check
still does.

/code-review spec axis found this branch had reintroduced its own bug:
migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after
the legacy dir was wiped and before the staged batch was restored, so a
planted symlink bricked uninstall permanently and orphaned the batch. Refusal
kept, abort removed.

kimi-code local silently lost its agents, the same class as the cline bug, and
the parity harness recorded that exclusion as intentional - the third test in
this branch to pin a regression as correct.

--minimal now creates an empty agents/ dir that never existed. Behaviour
restored rather than softening the changeset, so its byte-identical claim
stays true.

Standards axis: try/finally removed from twelve test bodies, fast-check
properties added for parseOwnerPid, boundary coverage at the grace window and
the ancestor-probe depth, a parity assertion for the staging-root helper
duplicated across two files, and the 8-deep config walk deduplicated.

Records 60-review.json with every finding and disposition from five passes,
including the smells left unfixed and why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): prune stale agents unconditionally in minimal mode

The previous round stopped an empty agents/ directory being created when the
resolved profile yields no agents. That was implemented by skipping the agents
kind entirely, which also skipped its stale-agent prune - so a full to minimal
downgrade left stale gsd-* agents behind.

The deleted inline loop pruned unconditionally and only skipped writing. Those
are three separate conditions, not one: prune always, write only when there is
something to write, create the directory only when writing.

Both call sites now run _removeGsdEntries before the empty-staged early exit.
The symlink-escape guard moved with it, since the prune also touches dest.
Codex .toml agents and the config.toml stanzas are cleaned again, and
user-owned agents are still preserved.

The agents/ directory is left in place after a prune empties it, matching
every sibling kind - none of them remove the destination directory itself.

Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh
install, so fixture generation is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): document interrupted-install recovery for user-owned files

The durable-staging fix is invisible to the user it protects. Someone whose
install died mid-flight has no way to know USER-PROFILE.md was staged before
the delete, that the next run restores it, or that recovery happens at the
start of that run rather than in the background.

Written as the task the user has - finish the interrupted command - rather
than as a description of the mechanism, and states what it will not do:
overwrite a file already present, or touch staging belonging to another
install still running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2875): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2875): assert the J8 model override without building a regex

CodeQL flagged incomplete string escaping: the assertion interpolated the
override value into a RegExp while escaping only forward slashes, which is
meaningless in a constructor, leaving real metacharacters unescaped.

The failure direction was the dangerous one - a metacharacter would have made
the match more permissive, so the row would pass when it should fail. That
matters here because J8 exists precisely because an earlier revision was a
tautology; the rewrite reintroduced a different way for the same assertion to
stop discriminating.

Replaced with a line-wise exact match, so no regex is constructed at all.
Swept the other test files this branch adds; no sibling instances.

lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full
metachar-escape copy, so a single slash replace slipped under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 17:25:53 -04:00
0xdhx
0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00
Tom Boucher
cf6de5e1c0 feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution

23 tests over the 50-test-matrix rows. RED by construction:
resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet,
and the validator silently ignores triggerPrecedence today.

Written in the per-runtime describe idiom the other four
runtime-artifact-layout suites use, not a table.

The rows that carry the weight: windsurf must NOT report a shadow it
does not have, since its global scope emits only agents and agents are
not trigger-bearing; agents and kimi-agents must be absent from the
output for every runtime; and reordering a runtime's triggerPrecedence
must flip the winner, which is the only assertion that proves the axis
is read rather than decorative.

Stems are injected, never scanned, so the surface is assertable with no
filesystem.

* feat(#2871): resolve triggers and host precedence, not just placement

resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger
a runtime emits, with the scope and kind that produced it, whether the
host registers it directly or only through a router, and which artifact
shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers
need placement only and the issue requires them unchanged.

AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The
host-integration matrix models command and dispatch as separate interface
points: an agent is invoked through the Agent tool's subagent_type, not
by typing a slash trigger, and _copyStaged never applies the kind prefix
to an agents entry. So agents and kimi-agents are excluded from the
surface entirely, and this commit amends ADR-2866 with a dated
correction. #2218's conclusion is unchanged -- the collision is strictly
commands-vs-skills, and claude's local /gsd-* trigger surface is still
fully shadowed -- but the ADR implied the local agents surface was lost
too, and it is not.

That correction is what makes windsurf come out right. Its global scope
emits only agents, so it has no global trigger and its local commands
are unshadowed. Model agents as trigger-bearing and windsurf falsely
reports a full shadow.

The triggerPrecedence axis lands on all 19 descriptors as an ordered
kind list, one value with one owner, rather than a numeric rank spread
across N kind entries with nothing keeping them consistent. Validation
uses a required-with-default shape that has no precedent in this
validator -- every existing axis is hard-required -- so a third-party
capability.json omitting the field still validates, which is what
ADR-894's additive-only contract promises.

Winner resolution reads Phase 1's scope rank first, then the kind
ordering. A test reorders the axis and asserts the winner flips, since
an axis that is added, validated and never consulted would pass every
other assertion.

shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this
issue's out-of-scope note.

Verified via the remote runner.

* fix(#2871): single-source namespacedByDir and close two test gaps

Four findings from the isolated adversarial review.

The namespacedByDir rule had reached three copies -- install-engine,
surface, and the new trigger resolver -- one of which carried a
hand-written keep-in-sync comment and no assertion. That is this repo's
generative-fix-divergence class. Extracted to one exported predicate all
three now call. Verified by diverging one copy deliberately: the existing
#816 parity test failed, and passes again on revert.

The omission test was vacuous. Row 16 asserted that a descriptor without
triggerPrecedence still validates, but built its fixture from claude's
shipped descriptor -- which this PR had just added the axis to. It now
clones and deletes the key, following the shippedDescriptorWithout
pattern, and asserts both that validation passes and that the resolver
still picks the right winner from the default. The second half is what
makes it prove anything.

resolveTriggerSurface silently dropped an unrecognized scope while every
sibling in this epic throws. Two phases of one epic should not disagree
about whether an invalid scope is an error, so it now rejects through the
same shared validator; an empty scope list still returns empty rather
than throwing.

The ADR amendment had been spliced into the middle of the References
list, orphaning its last bullet. Moved to the top, after the header
block, which is where ADR-3660 and ADR-1016 both put dated amendments.
No lint checks markdown structure, so this was green while malformed.

* fix(#2871): single-source the command filename composition too

The earlier fix shared the namespacedByDir boolean but left the
filename composition around it written twice -- once in _copyStaged as
what actually gets written, once in resolveTriggerSurface as what gets
predicted. The predictor could go stale silently.

One exported helper now composes it for both. The entry.name asymmetry
that looked like it would block extraction does not: entry.name is
filtered to end in .md and stem is entry.name minus those three
characters, so the two branches are the same string by construction.

Divergence proven to fail: injecting a marker into the helper broke the
trigger-surface suite; reverting restored 25/25. The four sibling layout
suites hold at 227 unchanged.

* docs(#2871): correct the ADR timing notes that this phase makes stale

The Amended by back-links on ADR-3660 and ADR-1016 were written in
Phase 0, when the widenings they describe had not shipped. Each carried
a forward-looking clause -- "the module changes at Phase 2, not before,
until then this module resolves placement only" -- which becomes false
the moment this PR merges. ADR-2866's own Amends header and its
reciprocal-notes section carried the same tense.

All four now describe what shipped. This is a tense and status
correction on Accepted ADRs, not a change to any decision.

Worth stating because it is the failure mode this epic keeps meeting:
gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI
would have caught either the missing back-link in Phase 0 or these stale
clauses now. They stay correct only because someone checks.

* chore(#2871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:25:42 -04:00
Tom Boucher
653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
github-actions[bot]
854c93533c chore: sync next package version to 1.9.1 2026-07-31 13:12:09 +00:00
github-actions[bot]
4232a79396 chore: sync next package version to 1.9.0 2026-07-31 03:14:04 +00:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
0408276791 chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema

Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the
ADR's swap amendment: a federated config slice lives inside a
capabilities/<id>/capability.json, and three of the five key families had
no capability directory until 5a created them.

Four key families move to the lanes that use them; the central-schema
removal and the federated addition land in this one commit because the
exclusivity invariant fails the build on a key present in both.
review.max_prompt_tokens, review.default_reviewers and
review.reviewer_instances describe policy ACROSS lanes and stay central.

Two things the issue did not name, both found while building it:

1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys
   against manifest.validKeys only, and two of the four families
   (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>)
   were pattern-backed. That is not cosmetic: isCentralConfigKey consults
   those patterns and mergeFederatedConfig skips every key for which it
   returns true, so declaring a slice while the pattern survived would
   have shipped an INERT slice behind a green gate — the exact
   half-migrated shape the invariant exists to prevent. The gate now
   loads the patterns from the same manifest the runtime reads.

2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a
   federated key always resolves to its declared default. The three
   fallback guards in review.md checked only empty-or-"null", so a user
   who set the GLOBAL review.max_prompt_tokens would have silently lost
   trimming on the HTTP lanes. The guards now treat 0 as unset.

D9 says review.models.<slug> is owned by "the lane whose slug it names".
That is false for one lane: the shipped key is review.models.agy while
the slug is antigravity. Ownership follows the lane; the key name is
preserved, because renaming would break every config that sets it.

Existing tests updated rather than left asserting the old world:
config-get on a cleared federated key yields empty instead of
not-found (what #2046 actually protects — never persisting the literal
"null" — is unchanged and still asserted); the config-schema dynamic
pattern representative moves to reviewer_instances; the
prototype-pollution guard case moves to a surviving dynamic prefix so
alert #26 keeps its coverage, with a new case asserting the old key is
now rejected earlier; and Phase 2's harvest-widening inertness assertion
becomes an ownership assertion, since Phase 4 is what consumes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives

A federated config key always resolves to its declared default, so an
unset per-lane prompt budget needed a value the workflow could treat as
'not configured'. The first cut used 0 — which is wrong: 0 is already a
LEGITIMATE per-lane budget meaning 'do not trim this lane' (the
early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it
as unset would have silently switched a user who deliberately disabled
trimming for one lane onto the global budget.

The sentinel is now -1, which is not a valid token budget, so all three
states stay distinguishable: unset falls back to global, an explicit 0
disables trimming for that lane, and an explicit N is used. Locked by
three CLI round-trip tests.

Surfaced by the isolated security reviewer before it crashed mid-run;
verified independently against the shipped trim guard rather than taken
on trust.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): update central-registration assertions and stay under the review.md cap

The remote runner caught both; my local sweep missed the files.

1. tests/plan-review-convergence.test.cjs asserted the three local-server
   host keys are in VALID_CONFIG_KEYS. They are federated to their lane
   capabilities now, and the exclusivity invariant forbids a key living
   in both places. What #2306-local actually protects is that config-set
   ACCEPTS them, so that is what is asserted — via isValidConfigKey, the
   predicate config-set itself uses, which spans central and federated.
   A second assertion pins federated ownership, so a silent reversion
   back to the central schema fails too.

2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap
   is a red line, not a budget to raise. The three per-lane budget guard
   comments were near-identical; condensed to one terse line each. 61371
   bytes, 69 to spare. Real extraction to workflows/review/modes/ is
   Phase 5b/6 work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs

Isolated security review findings.

MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse
error and returned [], while its sibling loadCentralConfigKeys, reading
the SAME file, writes to stderr and throws ExitError(1) on that identical
failure class. Fail-open here defeats the gate this function exists to
feed: with zero patterns, validateCrossCapability's pattern-collision
check silently passes and an inert federated slice ships green. It was
masked in the one production call site only because loadCentralConfigKeys
runs first against the same path — a coincidence of ordering, not a
guarantee, and this function is exported and called standalone. The two
now share a contract: ENOENT is the legitimate absent case, anything else
throws loudly. A single unparseable PATTERN is still skipped, which
degrades to "checked less" rather than blocking every build. The branch
had zero coverage; it now has two tests (malformed JSON, EISDIR).

MINOR — docs/CONFIGURATION.md still listed review.models.qwen and
review.models.cursor as settable, ~770 lines below this PR's own new
Ownership section. Those lanes take no model flag, so they declare no
model key and config-set now rejects them. Rows removed; the missing
review.models.agy row added; the per-reviewer budget row corrected to
name only the lanes that own a budget key, and to document that a
per-lane 0 disables trimming for that lane.

Also fixes a shadowed "raw" binding introduced by the fail-closed change,
which made the generator unrequirable — caught immediately by its own
--check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2797): backfill changeset pr number to 2841

* fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env

tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next
stays green, and it is currently blocking at least three unrelated PRs
(#2841, #2832, #2827) with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

The existing loop comment attributes this to `git add .` rehashing O(n^2)
blobs "before the object write had landed" and works around it by staging
one path per iteration. That is not the cause: sequential execFileSync
calls cannot race each other's object writes, and the failure persisted
after that change — it simply moved to a lower commit index.

The cause is that the git() helper inherited the runner's environment. A
leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's
index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the
blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole
operation. In every case `git commit` then cannot resolve a blob it just
staged, which is precisely the error above.

Verified by negative control: with GIT_DIR exported, this test fails on
the unfixed helper (the git commands operate on the wrong repository
entirely); with the helper stripping GIT_* it passes. The single-path
staging is kept — it is genuinely less work — but it is no longer load
bearing.

Found while shipping #2797. Fixed in place rather than deferred: it is a
defect surfaced during the work, and it is blocking other contributors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2452): build the base-advance commits empty, removing the lost-object class

The base-ref guard has been failing in CI with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

It is currently red on at least three unrelated PRs (#2841, #2832,
#2827) while next stays green.

Two theories have now been tried and neither held. #1881 blamed `git
add .` rehashing O(n^2) blobs and switched to staging one path per
iteration; the failure moved from commit 32 to commit 25 and carried on.
The preceding commit here made the git helper hermetic against leaked
GIT_* environment — that IS a real vulnerability (with GIT_DIR exported
the helper operates on the wrong repository entirely, proven by negative
control) but it produces a different error than CI reports, so it is not
demonstrably the cause either.

Neither trigger reproduces off-CI, so this stops guessing at the trigger
and removes the failure CLASS instead. The loop needs base-branch DEPTH
and nothing else: no assertion reads these commits' contents, and
base-side files cannot appear in `origin/base...HEAD` regardless.
`--allow-empty` writes no blob and no tree, so there is no object for the
index to reference and lose. It is also far less work than 60
write+hash+index cycles.

The guard still proves its mechanism: the test asserts that a --depth=1
base fetch FAILS and a full fetch resolves, so a broken topology would
surface immediately rather than passing vacuously.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Test <test@example.com>
2026-07-30 08:52:56 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
Tom Boucher
d39ca0c004 Merge pull request #2490 from open-gsd/docs/2481-adr-1239-effort-axis
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
2026-07-21 20:14:27 -04:00
github-actions[bot]
8cf747724f chore: sync next package version to 1.8.0 2026-07-22 00:06:08 +00:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
6b196ef638 fix(#2423): finalize job syncs next package.json after final release (#2437)
* fix(release): finalize job calls sync-next-version.cjs to bump next after final release (#2423)

The release pipeline's 'finalize' job shipped X.Y.0 to npm 'latest' and
merged to main, but never bumped 'next' to match. 'scripts/sync-next-version.cjs'
exists exactly for this — its docstring promises to run 'for every release
type (rc / hotfix / final)' — but it was wired only into the 'rc' job
(release.yml:479), not 'finalize'. As a result, after 1.7.0 shipped on
2026-07-15, 'next' stayed at 1.7.0-rc.6 and every npm script banner on
'next' (and feature branches cut from it) reported the stale rc version.

Regression of #1104 — closed incomplete (covered rc only, not final).

This patch:
  - adds a 'Sync next branch to the published release' step to the
    'finalize' job, mirroring the rc job's pattern at line 479 (uses
    VERSION from inputs.version rather than PRE_VERSION from steps.prerelease,
    since finalize does not run the prerelease step);
  - gates it on !inputs.dry_run + continue-on-error:true (matches rc);
  - adds tests/release-finalize-syncs-next-version.test.cjs — a structural
    YAML assertion that fails against the pre-fix workflow and passes after.
    Failing-first demonstrated during development; the test parses job blocks
    by indentation rather than grep so it stays valid as the file grows.

Root-cause diagnosis: scripts/sync-next-version.cjs:14 docstring admits
'used by release.yml's rc job, which has no next-targeting PR of its own'.
release.yml:506-685 (finalize job) had no sync-next-version step before
this patch. Every prior rc.N release has a matching 'chore: sync next
package version to 1.7.0-rc.N' commit; there is no such commit for 1.7.0.

* chore(release): sync next package version to 1.7.0 (#2423)

Replays the canonical 'chore: sync next package version to <v>' commit
that the release pipeline's rc job auto-produces via scripts/sync-next-version.cjs,
for the 1.7.0 final release that shipped on 2026-07-15 (commit dd4c90f82
'chore: finalize v1.7.0' on main). Without this, 'next' (and every feature
branch cut from it) carried 1.7.0-rc.6 indefinitely and reported it in
every npm script banner (e.g. 'lint:ci').

Bumps 43 synchronized manifests via the npm 'version' lifecycle hook
(scripts/sync-manifest-versions.cjs --stage + scripts/gen-capability-registry.cjs
--write), matching the file set of commit 27f69cc48 ('chore: sync next
package version to 1.7.0-rc.6') and commit dd4c90f82 ('chore: finalize
v1.7.0').

This is the immediate Layer-1 repair for #2423. Layer-2 (prevent recurrence)
is the workflow patch in the previous commit; Layer-3 (regression test)
ships with it. Future X.Y.0 final releases will produce this commit
automatically once the workflow fix lands.

* test(release): tighten #2423 dry-run gate assertion to the sync step

Code review of fix/2423 found that test #3 ('gates sync-next-version on
!inputs.dry_run') asserted too loosely: it scanned the entire finalize
block for any '!inputs.dry_run' line, so it would still pass if the
gate were stripped from the sync-next-version step specifically — the
exact regression the test name promises to catch. The finalize job has
multiple steps with their own !inputs.dry_run gates (e.g. Verify
publish), so the loose version masked the very bug it claimed to detect.

Tighten by extracting the specific YAML step block containing
'scripts/sync-next-version.cjs' and asserting the gate appears within
THAT step's lines, not anywhere in the job. Verified the tightened test:

  - PASSES against the post-fix workflow (sync step has its own gate)
  - FAILS when the sync step's gate is stripped (even when other steps
    in finalize retain their own !inputs.dry_run gates) — the exact
    regression that previously slipped through

Adds extractStepBlockContaining(jobBlock, marker) helper alongside the
existing extractJobBlock(text, jobName). Reuses the same indentation-
based parsing, so it stays valid as the file grows.

* chore(changeset): backfill pr:2437 in .changeset/sturdy-ibex-jump.md

CLAUDE.md changeset convention: 'Use placeholder pr:0 during initial commit.
Backfill immediately after gh api POST /pulls returns the real number.'

PR #2437 created from branch fix/2423-release-finalize-sync-next-version.

* fix(#2423): add see #2423 to allow-test-rule exemption per ADR-456

CI lint-allow-test-rule-refs failed on PR #2437: ADR-456 requires new
allow-test-rule exemptions added after the ADR's acceptance to include a
tracking issue number in the comment, in the form
  // allow-test-rule: <reason> (see #NNN)
The exemption added in commit 976c8b0a2 lacked this ref. Fixed.

Verified locally:
  node scripts/lint-allow-test-rule-refs.cjs
    → ok lint-allow-test-rule-refs: 173 grandfathered exemption(s) tracked, no novel untracked offenders
2026-07-19 14:33:19 -04:00
github-actions[bot]
27f69cc489 chore: sync next package version to 1.7.0-rc.6 2026-07-12 05:40:45 +00:00
Tom Boucher
c6ce110efa feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter,
branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace
normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into
descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json,
read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches
in install-engine.cts folded to descriptor flags (claude+hermes descriptors
updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity
for qwen/hermes/claude(global+local).

UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents
artifact-layout kind + convertClaudeAgentToQwenAgent converter (name +
description + tools YAML block list; color/model dropped). qwen routed onto
the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES).
UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the
descriptor-gated hook-writer loop (activates only for qwen).

Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors +
no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file
validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix
+ how-to updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 15:44:19 -04:00
github-actions[bot]
451e58e6d5 chore: sync next package version to 1.7.0-rc.5 2026-07-10 17:09:38 +00:00
Tom Boucher
fdd5e401eb feat(architecture): [EoS/claude] drive claude through the imperative adapter + descriptor-driven hostBehaviors (#2086)
Fold Claude Code's install/uninstall onto the Embeddable Orchestration System
(ADR-1239 Phase D). claude is GSD's tier-1 reference host, but its install path
was still driven by 13 hardcoded `runtime === 'claude'` string-equality branches
scattered across bin/install.js rather than the public Host-Integration Interface.

- Route install()/uninstall() through `createImperativeAdapter({runtime})` — the
  adapter delegates to the SAME installRuntimeArtifacts/uninstallRuntimeArtifacts
  engine calls, so output is byte-identical (proven pre/post, both scopes).
- Replace all 13 `runtime === 'claude'` / `runtime !== 'claude'` branches with
  descriptor-driven `runtime.hostBehaviors` lookups on capabilities/claude/
  capability.json (attributionSource, authorsCanonicalWorkflow, localInstallStyle,
  permissionsSchema, settingsFileByScope, sourceMarkerFile, agentFrontmatterExtensions,
  ownsClaudePaths, nativeModelAliases, skillsGlobalOnboarding). Behavior is
  identical; the brittle string-equality coupling (the add-a-host tax) is gone.
- Single-source the scattered literal 'claude' defaults/rosters behind DEFAULT_RUNTIME.
- Extend golden-install-parity to assert the claude LOCAL legacy layout is
  byte-identical too (AC1 "both scopes"); exclude the platform-varying
  settings.local.json (same reason settings.json is excluded).
- New tests/claude-imperative-reference.test.cjs: adapter kind, programmatic-cli
  profile, fail-closed negotiation on a corrupted/partial descriptor, and an AC2
  source guard that no `runtime === 'claude'` branch remains.

No user-visible install-output change (internal architecture only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:31:53 -04:00
github-actions[bot]
a6263bba8a chore: sync next package version to 1.7.0-rc.4 2026-07-07 06:08:27 +00:00
github-actions[bot]
a828d8207c chore: sync next package version to 1.7.0-rc.3 2026-07-06 02:19:50 +00:00
github-actions[bot]
8f791817f4 chore: sync next package version to 1.7.0-rc.2 2026-07-03 04:15:36 +00:00
github-actions[bot]
feb7036399 chore: sync next package version to 1.7.0-rc.1 2026-06-30 13:28:49 +00:00
Tom Boucher
a3d3c2a445 refactor(#1756): derive getDirName from a documented runtime.localConfigDir descriptor axis (#1757)
ADR-1239 Phase B (parent #1679). getDirName was a hand-maintained 15-branch
if-chain mapping each runtime to its local content-rewrite dot-dir. Relocate
those values into a documented runtime.localConfigDir descriptor field; derive
getDirName from registry.runtimes[id].runtime.localConfigDir (fallback .claude).

- 16 capability.json gain runtime.localConfigDir (byte-identical values)
- capability-validator.cjs requires it (non-empty dot-dir); registry regenerated
- docs/reference/capability-manifest.md documents the field + the three
  divergent values (copilot=.github, antigravity=.agents, kimi=.kimi-code)
- drift-guard test: golden value map + key-set equality both ways

Byte-identical install output for all 16 runtimes (golden-parity harness #1730).

Closes #1756

Co-authored-by: review-bot <review-bot@gsd>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 12:44:09 -04:00
Tom Boucher
cf2e66b39e feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten

Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1708): address review findings on typed dispatch-flatten

Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): backfill backgroundDispatch in role:runtime test fixtures

Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate

fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): add changeset for typed dispatch-flatten

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1708): remove stray temp PR-body file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): add issue ref to bug-853 allow-test-rule annotations

ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:37:33 -04:00
Tom Boucher
30d4b85de5 feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module

ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1684): validate and document host-integration axes (16 runtimes)

Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add host-integration capability matrix and adr amendment

New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1684): harden dispatch negotiation edge cases

Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1684): register host-integration.cjs in lint-ignore and inventory

New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add changeset fragment for host-integration interface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add how-to for sourcing a host's integration axes

Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 11:33:06 -04:00
github-actions[bot]
548a986ffc chore: sync next package version to 1.6.0 2026-06-24 23:09:44 +00:00
github-actions[bot]
2e406e8e05 chore: sync next package version to 1.6.0-rc.3 2026-06-24 03:52:20 +00:00
github-actions[bot]
f20dc691c1 chore: sync next package version to 1.6.0-rc.2 2026-06-22 05:22:23 +00:00
github-actions[bot]
4441b20b4a chore: sync next package version to 1.6.0-rc.1 2026-06-20 19:58:40 +00:00
Tom Boucher
33ccf5f89d fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout

Claude Code project-local installs now write command files as flat
gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md
(subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form)
matching hooks, statusline, and all cross-command references.

- capabilities/claude/capability.json: local destSubpath commands/gsd → commands
- bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites
- bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup
- bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys
- legacy migration: preserves dev-preferences.md across reinstall and uninstall
- gsd-core/bin/lib/capability-registry.cjs: regenerated
- 6 new regression tests (L0–L5) in bug-1367-*.test.cjs
- Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests
- scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1367): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:27 -04:00
Tom Boucher
2421cf1b4a feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1)

Make the capability manifest versioned — the data substrate the Capability
Ecosystem (ADR-1244) keys off:

- capability.json gains a REQUIRED semver `version` plus the optional
  ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`,
  `provenance`); the build-time conformance validator enforces them via a new
  `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay).
- All 32 native capabilities stamped with `version` (= package version,
  lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep
  that keeps them in sync, and the issue-844 regression guard is extended.
- Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in
  version strings; range/integrity fields are shape-validated (satisfaction
  and the load-time gate are deferred to Phase 2/4).
- Capability rel-paths emitted forward-slash for cross-platform git correctness.

Closes #1430

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1430): add changeset for versioned capability manifest

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:45:53 -04:00
Tom Boucher
94872662e9 feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor

postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in
runtime-hooks-surface.cts) now select the event-name dialect from
registry.runtimes[id].runtime.hookEvents instead of the hardcoded
(runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini'
→ AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the
applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents
'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude
dialect (matches the old else).

The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact;
runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED —
hookEvents (2-value) is too coarse to drive them (the event set differs within
hookEvents='claude'); per-event-set drive tracked in #1076.

Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND
pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for
gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread).

Closes #1077

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI

hooks/dist is gitignored and absent in scoped/windows CI jobs that do not
pre-run build:hooks. Without it, install() finds no hook files and all
AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty,
failing every hook-presence assertion. Added an idempotent ensureHooksDist()
called in a top-level before() — mirrors the pattern from
bug-376-claude-js-hook-gsd-rewriter.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors

Purely additive: four new fields on all 16 runtime capability.json descriptors,
validator extended with three new closed-vocab sets, registry regenerated.
Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new
required fields so all 255 capability-registry tests continue to pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): drive per-event hook guards from extendedHookEvents descriptor

Replace hardcoded runtime-name checks (isQwen||runtime==='claude',
runtime==='claude', isGemini) in applySettingsJsonHooks with a single
descriptor-driven extendedEvents array derived from the new opts field.
Remove isQwen and isGemini derivations (no remaining uses after the three
guard blocks are migrated). Wire extendedHookEvents from the capability
registry in bin/install.js call site. Add behavioral regression test
(enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is
purely descriptor-based and runtime-name-agnostic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY

- Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs
  and read installSurface / writesSharedSettings / permissionWriter from
  runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2).
- ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface.
- Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read
  installSurface directly from capMap descriptor bodies, breaking the require cycle
  (adapter now requires the generated registry; gen-script must not require the adapter).
- Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests)
  pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract.
- Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip

- Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in
  applySettingsJsonHooks (src/runtime-hooks-surface.cts).
- Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with
  hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations
  (ADR-857 phase 5g drive 3).
- Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks
  call site in bin/install.js using the established
  _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom.
- Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites
  proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name;
  (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously
  hardcoded to skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016

Add Decision 7a documenting the four axes added in the 5f-completion pass,
update axis counts from "six" to "twelve", note 5f-completion drives as done
in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done
and 5g (InstallPlan capstone) remains the only open phase.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level)

The test fixture makeRuntimeCapMap did not include installSurface in the runtime object,
so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw.
Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct
installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests.
The gate implementation already reads r.installSurface correctly from the descriptor level.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests

GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a
runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none,
codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors.

GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events
(BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family
events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'.

Added 10 rejection tests in suite 27 covering each gate + each new field validator.
All 16 real runtimes satisfy both gates (verified before coding).
Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES,
VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup

1. bin/install.js: add explicit literal fallback for hooksSurface when the committed
   capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json').
   The descriptor is always the source of truth in normal operation.

2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty)
   to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty.
   ensureHooksDist() in before() guarantees hook files exist.

3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime
   already fully covered by Test 1's deepStrictEqual over all four fields.

4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the
   parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface
   absent → typeof r.installSurface !== 'string' → soft-skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete)

ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in
runtime-config-adapter-registry collects install-level descriptor axes into
the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope
section updated: 5g capstone is complete, phase 5 fully materialized.
ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves).
ADR-58 Implementation note added (2026-06-11): realized in
runtime-config-adapter-registry (co-located with adapter-selection).
CONTEXT.md Runtime Config Adapter Registry entry extended to document
resolveInstallPlan and both-halves realization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1082): update install drift guard to the resolveInstallPlan seam (5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 22:36:47 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00