Files
msd-core/docs/adr/2782-reviewer-lane-capability-surface.md
0xdhx 0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00

62 KiB
Raw Blame History

ADR-2782: Reviewer Lane — the cross-AI reviewer handoff becomes a declared capability surface

  • Status: Accepted
  • Date: 2026-07-28
  • Amended: 2026-07-29 by Phase 1 (#2794) — D1's flag becomes flags[]; D2's promptChannel gains none, outputChannel gains file-arg with a companion outputArg; D8's uniqueness invariant restated over the flattened flag set. All four are additive widenings of closed enums, each forced by a shipped lane the original survey did not cover. See Amendments at the end.
  • Issue: #2782 (epic); Phase 0 tracked by #2793
  • Amends: ADR-857 (extension points as data — extends D7/D8 in the same "amend, not reverse" sense ADR-1244 D8 established) · ADR-894 (adds a role-typed body and a third role) · ADR-1016 (the runtime body is no longer the only body a role: "runtime" capability may carry; its closed-vocabulary principle is upheld, not relaxed — see D6) · ADR-1244 (D5 gains a fourth executable-surface disclosure class; D9's matrix gains a lane column)
  • Unchanged and explicitly out of scope: ADR-0011 (reviewer selection precedence) · ADR-1517 (the REVIEWS.md contract and reviewer instances)
  • Subsumes: #2690 (core single-sourcing — lands as Phase 1 under this ADR rather than as its own design)

Context

A cross-AI reviewer lane — one external CLI or model endpoint that /gsd:review hands a plan to for independent review — is declared today in three unrelated places, none of which is the capability system, and none of which a third party can extend.

1. The roster is half registry-derived, half hardcoded. src/review-reviewer-selection.cts derives slugs from runtime.hostBehaviors.reviewerCli === true (:40-49), then concatenates a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail (:32-38) for five reviewers that have no capabilities/<id>/ directory at all. The module's own comment says exactly this. Six capabilities carry the flag; five reviewers have no descriptor of any kind.

2. The invocation contract is prose. gsd-core/workflows/review.md is 1070 lines; invoke_reviewers spans roughly 60% of it as hand-authored per-CLI bash. Each leg re-implements probe, argv shape, model lookup, effort channel, timeout, stderr capture, and empty-output policy.

3. The output contract is prose. write_reviews hardcodes per-reviewer section headings, including two literal instance names.

Three structural consequences follow, and they are why this is a capability question rather than only a refactor question.

(a) reviewerCli is a bare boolean in an undocumented, unvalidated bag. hostBehaviors appears zero times in docs/reference/capability-manifest.md — not in the envelope table, not in the runtime-body axis table — and scripts/gen-capability-registry.cjs does not validate its keys. The one field that decides reviewer membership is unspecified, unvalidated, and carries no invocation data. A capability author can discover it only by reading src/review-reviewer-selection.cts:47.

(b) Reviewer-ness is welded to runtime-ness, and the runtime body structurally cannot hold a lane contract. capability-manifest.md:141 states the runtime body is "a closed 8-axis (plus 4 install-surface) vocabulary; no feature-only fields (skills, agents, steps, contributions, gates, hooks) are permitted," and gen-capability-registry.cjs:505 enforces the consequence — a role: "runtime" capability is stored whole into runtimes[] and its config/steps/ contributions/gates are never harvested. A reviewer lane therefore cannot own its own federated config keys. That is why review.models.*, review.ollama_host, review.lm_studio_host, review.llama_cpp_host, and review.max_prompt_tokens_per_reviewer.* all live in the central schema instead of with the lane that uses them — the exact half-migrated shape the config-key exclusivity invariant exists to prevent.

(c) A reviewer that is not a GSD install target has nowhere to live. gemini, coderabbit, ollama, lm_studio, and llama_cpp are review or model CLIs GSD never installs into. There is no capabilities/<id>/ for them, so they are a hardcoded tail by necessity, not by choice.

Net: adding a reviewer lane is a core patch. It means editing the roster module or a runtime descriptor, hand-authoring a bash leg, hand-adding a write_reviews heading, adding central config keys, and updating five prose surfaces. #2718 was that patch in flight (PR #2776, closed in favor of this design); #2781 is the documentation drift it produced. Cross-cutting fixes land per-leg: #2494 and #2605 were the same empty-output defect filed twice; #2475 (effort channel), #2589 (model lookup), #2295 (resolved-model recording), and #2272 (flag parity) are the same shape.

What a survey of the twelve lanes actually shows

The design was drafted assuming one lane shape. Reading all twelve legs disproved that, and the correction is the most consequential decision in this ADR (D2).

Family Lanes Shape
Spawned CLI gemini, claude, codex, coderabbit, opencode, qwen, cursor, antigravity, kimi-code Binary + argv; prompt via stdin or argv; stdout captured, stderr to a .err sidecar
OpenAI-compatible HTTP ollama, lm_studio, llama_cpp No binary. curl to /v1/chat/completions on a user-configured host; model discovered via GET /v1/models piped through jq

Three of twelve lanes are not spawned binaries at all. Timeout floors genuinely diverge — a measured ~570 s for Codex at xhigh effort and ~525 s for headless Claude drive a 900 000 ms floor with 1 200 000 ms for those two, while the Antigravity leg runs a 600 s external cap over a 540 s native --print-timeout, and the HTTP lanes use 120 s. Five lanes require jq on PATH. The Antigravity leg carries a deliberate three-layer fallback for an upstream stdout bug.

Divergence between lanes is real and frequently correct. The value of a descriptor is therefore one place where divergence is declared, not one behavior imposed on every lane.

Decisions

D1 — A reviewer body on the capability manifest, admissible on two roles

A reviewer lane is declared as data in a reviewer body:

{
  "id": "acme-reviewer",
  "role": "reviewer",
  "version": "1.0.0",
  "title": "Acme Review CLI",
  "description": "Cross-AI plan review lane backed by the Acme CLI.",
  "tier": "full",
  "requires": [],
  "engines": { "gsd": ">=1.9.0" },

  "reviewer": {
    "slug": "acme",
    "flags": ["--acme"],
    "transport": "spawn",
    "probe": { "kind": "command-exists", "binary": "acme" },
    "invoke": {
      "binary": "acme",
      "args": ["review", "--format", "text"],
      "promptChannel": "stdin",
      "outputChannel": "stdout",
      "modelArg": "--model",
      "effortChannel": "argv"
    },
    "timeoutFloorMs": 900000,
    "emptyOutput": "stub-with-stderr",
    "reviewsSection": "Acme Review",
    "evidenceClass": "source-grounded",
    "requiresBinaries": [],
    "promptBudgetKey": null,
    "handler": null
  },

  "config": {
    "review.models.acme": {
      "type": "string",
      "default": "",
      "description": "Model passed to the Acme reviewer lane."
    }
  }
}

The body is admissible on role: "runtime" — so the six capabilities that are both install targets and reviewers (claude, codex, cursor, opencode, qwen, antigravity) keep exactly one manifest — and on a new role: "reviewer" (D3) for lanes that are not install targets.

This is the amendment to ADR-1016: a role: "runtime" capability may now carry a reviewer body alongside its runtime body. The runtime body itself remains closed and unchanged; no feature-only field becomes permissible on it. A lane body is a third thing, not a relaxation of the second.

Because a lane may own a federated config slice, gen-capability-registry.cjs must harvest config from a lane-bearing capability of either role — the specific limitation at :505 that context (b) describes.

D2 — transport is a closed discriminator, and it selects the invoke sub-shape

reviewer.transport is a closed enum: spawn | openai-http.

spawn openai-http
invoke.binary required forbidden
invoke.args required (array) forbidden
invoke.promptChannel stdin | argv | argv-file-ref | none forbidden
invoke.outputChannel stdout | file-arg forbidden
invoke.outputArg required iff outputChannel: "file-arg", else forbidden forbidden
invoke.hostConfigKey forbidden required (dotted config key holding the base URL)
invoke.path forbidden required (e.g. /v1/chat/completions)
invoke.modelDiscovery forbidden closed enum: none | first-from-models-endpoint
invoke.modelArg optional forbidden (model travels in the JSON body)
invoke.effortChannel closed enum: none | argv | env none
invoke.env optional: object of environment name/value pairs, string values only forbidden (no child process to carry an environment)

A manifest declaring fields from both sub-shapes, or neither, fails validation. The discriminator is explicit rather than inferred from field presence: inference leaves a manifest with both — or with neither — carrying undefined meaning, which is precisely what a closed vocabulary exists to prevent.

promptChannel: "argv-file-ref" exists because two lanes (cursor, kimi-code) take the prompt as an argv argument, and passing a full plan set inline would approach the 32 767-character Windows execFileSync ceiling. The file-reference form passes a short instruction naming a prompt file in the run directory. That instruction must also carry the absolute repository root, because an argv-fed CLI does not reliably inherit the review's working directory — the existing cursor and kimi-code legs already do this by hand (review.md:447-448, :550-552).

outputChannel is a required, named, closed-enum field rather than an implicit assumption, because the alternative — a lane that writes its review to a file and prints nothing — is a shape a real CLI can take, and an unnamed assumption is the thing a later contributor silently violates.

Amended 2026-07-29 (#2794): this ADR originally recorded outputChannel as having "exactly one member (stdout) today" and described the file-writing lane as a shape a real CLI could take. It already does. codex captures its review through its own -o/--output-last-message <FILE> and discards stdout, because on Windows it writes process-teardown output to stdout after the final message, and a stdout redirect would append that noise to a non-empty file — slipping past the empty-output guard as a silently polluted review (#1698). The enum therefore ships with two members, and file-arg carries a companion outputArg naming the argument that takes the path: knowing the review lands in a file is useless without it.

promptChannel likewise gains none. coderabbit is fed no prompt at all — it reviews the working-tree diff and accepts neither a prompt nor a model flag (review.md:367). The original three-member enum had no way to say "this lane receives nothing", which would have forced Phase 2 to either invent a sentinel or mis-declare the lane.

Both were found by building Phase 1's descriptor table against all eleven shipped legs. That is the same evidence path that produced openai-http in the first place, and it is the process working: the vocabulary widens on a lane that exists, never on speculation.

Three further declared fields carry per-lane divergence that would otherwise live only in prose:

Field Values Why it exists
evidenceClass source-grounded | diff-only CodeRabbit reviews a diff, not the source tree, and its findings are deliberately down-weighted in synthesis (review.md:367). Today that caveat is a prose annotation a reader may miss; declaring it lets write_reviews render the caveat from data
requiresBinaries string[] External tools the lane needs on PATH — jq for five lanes. A missing prerequisite reports the lane unavailable with an install hint rather than running it into an empty review
promptBudgetKey dotted config key | null Per-lane prompt trimming (prepare_trimmed_prompt_for_reviewer, review.md:646-704) is keyed per slug today; the key becomes the lane's own federated config (D9)

Naming note: the field is requiresBinaries, not requires. The envelope already carries a requires (capability-id dependencies, ADR-1244). Two fields named requires at different nesting depths with unrelated semantics is a defect waiting to happen; the collision was caught in review of this ADR and renamed here rather than left for a downstream phase to trip over.

D3 — A third role, role: "reviewer", for lanes that are not install targets

gemini, coderabbit, ollama, lm_studio, and llama_cpp become first-party capabilities with a reviewer body, no runtime body, and no install surface — which is the honest description of what they are. runtimeCompat is not required for this role (it declares which host runtimes a feature surfaces through; a lane surfaces through none).

tier remains required, because it is the source of truth for install-profile membership. A role: "reviewer" capability therefore receives profile membership from deriveProfileMembership (gen-capability-registry.cjs:201-213) like any other. That membership is inert: the capability contributes no artifacts, so there is nothing to install. This is stated explicitly because a reader encountering a lane in an install profile would otherwise reasonably assume it installs something.

Rejected: one role for every lane, splitting codex into codex + codex-reviewer. It is the cleaner discriminator and was rejected for churn — six manifests would each fragment into two capabilities and two ids, complicating roster derivation for no gain.

D4 — The reviewer body is optional and absent-safe at every layer

This is a normative MUST, and it governs every downstream phase.

  1. A capability with no reviewer body is simply not a lane. This is never a validation error. Most runtime capabilities are install targets only; a validator that errors on an absent body would break the majority of the registry.
  2. An overlay declaring a role or a field this GSD version does not know is skipped with a warning via the existing engines.gsd hard gate (ADR-1244 D6) — never a crash. This is the forward half: a capability built for a newer GSD degrades to discovered-but-inactive.
  3. An unknown field inside a reviewer body is ignored with a warning rather than failing validation.
  4. A lane naming an unknown handler fails closed — the lane is unavailable; the registry does not crash.
  5. A capability with no reviewer body must not perturb its disclosure signature (D5). An absent body that changed the signature would force spurious re-consent across every installed capability.

The asymmetry is deliberate and is Postel's Law applied with a boundary: liberal in what a manifest may omit, strict in what it asserts. Permissiveness about absence is forward compatibility; permissiveness about assertions would be an untyped escape hatch.

Absent-safe governs discovery, never explicit selection

Rules 1–5 describe what happens when the system is looking for lanes. They do not apply once a user has named one. If a user runs /gsd:review --acme and the acme lane is unavailable — because its capability was skipped under rule 2, because its handler failed closed under rule 4, because a prerequisite binary is missing, or because its egress destination changed (D5) — that is an error, surfaced and non-silent. It is not an informational note, and the run does not quietly proceed with a thinner reviewer set.

This is called out because the current implementation does the opposite: an unavailable explicitly-requested reviewer is recorded as an info (review-reviewer-selection.cts:246-248). The workflow's own guidance already names why that is wrong — "a cross-AI review that silently drops a lane is blind in one eye" (review.md:304) — and a design whose whole premise is more lanes from less trusted sources must not inherit a silently-degrading selector. Correcting this is Phase 1's responsibility, because Phase 1 is where the selector is single-sourced.

In one line: not finding a lane nobody asked for is normal; failing to run a lane somebody asked for is an error.

Where warnings surface

"Skipped with a warning" means nothing unless a human sees it. Warnings arising at build time (registry generation over first-party capabilities) are emitted by gen-capability-registry.cjs on stderr, and fail the build only where D8's uniqueness invariants are breached. Warnings arising at load time — an overlay skipped by engines.gsd, an unknown field, a handler that failed closed — surface on the /gsd:review run that would have used the lane, and in gsd capability list, which is where a user goes to ask why a capability is inactive. A warning written only to a build log nobody reads is not a warning.

D5 — A fourth executable-surface disclosure class: the reviewer lane

ADR-1244 D5 rule 2 requires that executable surfaces be disclosed and consented at install, and names three classes: hooks, command modules, and mcpServers. A reviewer lane is a fourth, and it is materially different from the other three: it receives data. A lane is piped the plan text, the requirements, the research findings, and the CONTEXT.md decisions, and its output is read back into REVIEWS.md. That is an egress channel for the most sensitive artifacts GSD produces.

Making lanes pluggable without a disclosure class would open a data-exfiltration path behind a manifest field. The trust work is therefore the gating requirement of this design, not polish.

discloseExecutableSurfaces gains a reviewer-lane surface that discloses, by transport:

  • spawn — the binary and its full declared args, in both rendered and raw form, exactly as MCP servers already disclose argv/rawArgs (capability-trust.cts:688-690).
  • openai-http — the destination host URL resolved from hostConfigKey, plus the hostConfigKey itself. Disclosing curl would be technically true and practically meaningless; the destination is the disclosure that matters. A localhost destination is still disclosed, and is distinguished from a remote one.

Both forms additionally disclose the egress payload classes — plan text, requirements, research findings, CONTEXT.md decisions — rather than an unhelpful "sends data to the tool".

Disclosing the binary without its args is insufficient, and this is not hypothetical. A lane declaring binary: "python3" with innocuous args could, in a later version, change args to ["-c", "<arbitrary program>"] without the binary changing at all. That is precisely the bug class #1459 already fixed for MCP servers, and a binary-only disclosure would reopen it. args is therefore disclosed and signature-bound.

The lane folds into disclosureSignature / signatureForManifest as stable sorted JSON, exactly as env/cwd do for MCP servers (#1459). executableSetChanged treats any of the following as an executable-set change for the auto-update re-consent trigger (ADR-1244 D5 rule 4): adding or removing a lane, changing its slug, transport, binary, args, hostConfigKey, promptChannel or handler, or changing any other field of its declared invoke object — the residual added by #2483, which is what stops this list going stale again. See the 2026-08-05 amendment: an enumeration of "the fields that matter" had already fallen eight fields behind by the time env arrived, so the signature no longer relies on one.

The egress destination is re-verified at invocation, not only at install

hostConfigKey is the one consent-bound value that does not live in the SHA-pinned bundle. It names a key in .planning/config.json, which is user- and CI-editable at any time with no re-install and no integrity check — unlike every existing consent-bound field (command, args, env, cwd, url), all of which come from the manifest itself (capability-trust.cts:74-125).

Left unaddressed, this is a real hole: a lane consented against http://localhost:8080 could be silently redirected to a remote host by a later config edit — including one arriving through an ordinary pull request touching .planning/config.json — and every subsequent review run would egress plans, requirements, research, and decisions to the new destination with no re-prompt.

Therefore, normatively:

  1. The consent record binds the resolved host, not merely the config key.
  2. Before invoking an openai-http lane the runtime re-resolves hostConfigKey and compares the result against the consented host.
  3. On mismatch the lane is blocked, not silently redirected; the user is told the destination changed and must re-consent. A blocked lane reports like any other unavailable lane — it never degrades to running against the new host.
  4. This check lives on the invocation path (Phase 5b), not only in the install path.

A host change is a change of who receives the user's plans. It is the most security-relevant mutation in this design, and it must not be reachable by editing a JSON file.

Implementation note added by Phase 3 (#2796) — how rule 1 is actually satisfied.

The resolved host is deliberately excluded from the disclosure signature, and a reader comparing rule 1 to capability-trust.cts must not mistake that for the rule being unimplemented.

signatureForManifest(manifest, stagedDir?) is the single consent key that both the loader and the lifecycle compute, explicitly so the two "can never drift". The loader has no config resolver — hostConfigKey names a key in .planning/config.json, which is outside the SHA-pinned bundle. Folding the resolved host into the signature would therefore make the loader and the lifecycle compute different signatures for the same manifest, producing a permanent false-mismatch loop that re-prompts forever.

So the binding is split, and rule 1 still holds end to end:

  • the signature binds the manifest-derived lane fields — slug, transport, binary, args, hostConfigKey, promptChannel, handler, plus every other declared invoke field via the #2483 residual (2026-08-05: the enumeration alone was eight fields short, including defaultHost, which is itself an egress destination);
  • the consent record additionally stores the resolved host, which is what rule 1 requires;
  • Phase 5b re-resolves and compares at invocation and blocks on mismatch, which is where rule 4 already places the check.

reviewsSection and timeoutFloorMs are excluded for a different reason: a cosmetic change must not force re-consent, because a prompt carrying no security information is how users learn to click through — the same failure this decision cites when rejecting a per-run egress prompt.

(Recorded here rather than in the PR that made the decision. A squash-merged PR body is not a durable record: it is invisible to anyone reading the ADR later, which is exactly who needs this.)

Stated honestly, and consistent with ADR-1244 D5's own acknowledgment that there is no sandbox: even with the above, consent-at-install remains a weaker gate for a standing egress channel than for a hook. A user consents once; the lane thereafter receives every plan on every review run. Disclosure plus destination re-verification makes the channel visible, pinned, and revocable — it does not make it safe. A per-run egress prompt was considered and rejected as consent fatigue that trains users to approve blindly.

D6 — handler is a closed enum of first-party names; third-party lanes are data-only

Lane divergence is real (context above), so the descriptor must not promise uniformity. Where a lane needs genuinely imperative behavior — the Antigravity three-layer fallback is the canary — reviewer.handler names an imperative module by closed first-party name, rather than growing conditionals inside data.

The enum ships with exactly these members. null is the default and covers eight of the twelve lanes:

handler Lanes What it owns that data cannot express
null the other eight Nothing — the declared vocabulary suffices
"antigravity" antigravity The three-layer fallback for an upstream stdout bug; the two-level timeout (a 600 s external timeout/gtimeout cap wrapping a 540 s native --print-timeout, review.md:560); and the stale-response watermark guard that rejects a cached conversation from a prior run
"openai-compatible" ollama, lm_studio, llama_cpp Model discovery against /v1/models, the JSON request/response shape, and the served-model mismatch warning raised when the responding model differs from the one requested (review.md:794-797)

Enumerating the members here is deliberate. A "closed enum" whose membership is left to the implementing phase is not closed, and three separate phases would each have invented a different list.

On timeoutFloorMs and the Antigravity two-level timeout. The descriptor carries one scalar, timeoutFloorMs — the outer wall-clock bound every lane gets. A lane whose tool has its own internal timeout (Antigravity's --print-timeout is the only current case) expresses that inner bound in its handler, not in the descriptor. Adding a second declarative timeout field to serve one first-party lane would be speculative generality; the handler seam exists for exactly this. The delegation is stated here so a reader does not wonder where the measured 540 s went.

This upholds rather than relaxes ADR-1016. That ADR's core principle is that "a runtime that needs a shape no existing primitive expresses is supported by adding a first-party primitive … never by embedding arbitrary code or an open escape hatch in the descriptor," and its §Alternatives #2 explicitly rejected an open escape hatch. handler is the same construction as ADR-1016 Decision 3's closed ConverterName: the descriptor references a first-party function by name and never embeds it.

The consequence must be stated plainly, because it caps this epic's headline claim. Third-party lanes are data-only. A third-party CLI needing a shape the closed vocabulary lacks is blocked on a first-party PR. The honest claim is most lanes, declaratively — not any lane.

The escalation path is the ADR-1016 model, and it is documented rather than implied: file an issue naming the primitive the vocabulary lacks; it is reviewed and added first-party. D2 is the worked example of that path already functioning — the openai-http transport exists precisely because a survey produced evidence that three real lanes did not fit, and the vocabulary widened on evidence rather than on speculation.

Two real CLIs that would NOT fit today, named so the boundary is a known quantity rather than a surprise for the first third-party author who hits it:

  • Aider mutates the repository by default — it edits files and commits. The vocabulary has no way to declare "this tool must be invoked read-only", and the existing lanes achieve that only by asking politely inside the prompt text (review.md:448: "Do not edit any files"), which a coding-agent CLI is under no obligation to honor. A repo-mutating reviewer is a materially different safety posture from a read-only one, and the descriptor does not currently express it.
  • Plandex requires a stateful session (plandex new) before a review turn. D2's single binary + args + prompt-channel shape describes one invocation; it cannot express a two-phase setup-then-invoke sequence.

Neither is a reason to reject this design — both are exactly the "file an issue naming the primitive" case, and both would likely be served by two future primitives: a declared invocation-safety posture, and a setup phase on the descriptor. They are recorded because an ADR claiming a closed vocabulary is sufficient, without naming what it excludes, is claiming more than it verified.

Revisiting this to permit a third-party handler module confined to the capability install root (the ADR-1244 D7 model, which does allow third-party command modules) would genuinely deliver "any plugin can ship a lane." It is rejected here because it reverses an ADR-1016 rejection rather than amending it, and because D7 itself calls third-party code execution the highest-risk surface and sequences it last. It should be revisited only with its own ADR and its own evidence.

D7 — probe.kind is a closed enum wider than existence, and every probe is bounded

probe.kind is a closed enum:

Kind Fields Semantics
command-exists binary command -v <binary>
command-capability binary, needle, timeoutMs <binary> --help bounded, matched against needle
http-reachable hostConfigKey, path, timeoutMs Bounded GET; reachable ⇒ available

command-exists alone is structurally insufficient, and the evidence is concrete: kimi is claimed by both Kimi Code CLI (Node) and the legacy Python kimi-cli — which is a separate, first-party, non-reviewer runtime capability in this repo. An existence-only probe registers the wrong tool. This was found in review of PR #2776 and is the reason the vocabulary ships wider than one member.

Every probe that starts a process or a connection MUST be bounded. This repo carries a named Unbounded Subprocesses defect class, and the original Kimi probe was a live instance of it: an unbounded kimi --help | grep that ran on every /gsd:review invocation regardless of which flags were passed, so a user whose Kimi binary waited on a first-run consent or auth prompt would hang every future review — including reviews that never asked for that lane.

command-capability bounds via external timeout, falling back to gtimeout (the precedent already set by the Antigravity block at review.md:560). Stock macOS ships neither; where no bounding mechanism is available the probe is skipped and the lane reported unavailable, which degrades a lane rather than hanging a command.

D8 — Uniqueness is a build-time conformance invariant

Across the merged first-party ∪ overlay set, reviewer.slug, reviewer.flags, and reviewer.reviewsSection are each unique — for flags, over the flattened set of every lane's flags, since one lane may declare several. A collision fails the build gate.

Amended 2026-07-29 (#2794): D1 originally declared a singular flag, and this invariant was stated over it. antigravity is selected by both --antigravity and --agy (review.md, docs/COMMANDS.md), which a single-valued field cannot express, so the field is flags: string[] and uniqueness flattens across lanes. reviewsSection uniqueness is not cosmetic: two lanes sharing a heading would silently merge their output in REVIEWS.md, producing a review that appears to have consensus it does not have.

An overlay lane colliding with a first-party lane is rejected, first-party winning — the existing id-uniqueness precedent (capability-manifest.md:167).

Reviewer instances are not lanes. review.reviewer_instances.<name> = {cli, model?, agent?} (ADR-1517) lets one model-capable adapter run as several reviewer identities. Instances resolve through a lane and continue to; they do not participate in the roster, the flag set, or this uniqueness check.

D9 — Reviewer config keys become federated, and the roster derives from declared lanes

review.models.*, review.<host>_host, and review.max_prompt_tokens_per_reviewer.* move from the central schema to federated config slices owned by their lane capabilities. Key names and existing .planning/config.json files are unchanged; only validation provenance moves, so no user migration is required. Per the config-key exclusivity invariant (capability-manifest.md:173), the central-schema removal and the federated addition must land in the same commit or the build gate fails on a key present in both.

Ownership is per-key and per-lane, so that no key is owned twice — Phase 4 implements this table rather than re-deriving it:

Key Owner Notes
review.models.<slug> the lane whose slug it names One key per lane; a lane with no model override declares none
review.ollama_host ollama The hostConfigKey its own descriptor points at (D2)
review.lm_studio_host lm_studio ditto
review.llama_cpp_host llama_cpp ditto
review.max_prompt_tokens_per_reviewer.<slug> the lane whose slug it names The lane's promptBudgetKey (D2) resolves to this
review.max_prompt_tokens stays central A global default across all lanes; owned by no single lane, so federating it would be wrong
review.default_reviewers stays central Selection policy over lanes (ADR-0011), not a property of any lane
review.reviewer_instances stays central Instance→lane mapping (ADR-1517); an instance is not a lane (D8)

The last three rows matter as much as the first five: a key that describes policy across lanes must not be federated into one, and the exclusivity invariant would not catch that error — it only catches a key owned twice, not a key owned by the wrong side.

KNOWN_REVIEWER_SLUGS derives from declared reviewer bodies. hostBehaviors.reviewerCli survives as a derived legacy alias for one release and is then removed. Where both a body and the alias are present, the body wins. The field is undocumented, so external users are unlikely — but "undocumented" is not "unused", which is why it gets a deprecation window and a changeset note rather than a silent removal. The removal is owned by a named phase (#2801), not left implicit.

Consequences

Positive.

  • Adding a reviewer becomes one manifest installed through gsd capability install <url> — no core patch, no workflow edit, no release cycle — for any lane the vocabulary expresses.
  • A cross-cutting fix (empty output, effort channel, model lookup) becomes a single-site change covering every lane, retiring the #2494 → #2605 cadence.
  • The roster gets one generated source, which makes the DEFECT.GENERATIVE-FIX parity assertion for #2781 mechanical rather than per-lane.
  • Third-party lanes arrive behind the existing trust gate — disclosure, consent, SHA pin, engines.gsd, reserved namespaces — instead of as an unreviewable prose block.
  • A lane owns its own configuration, closing a half-migrated config surface.

Negative, and accepted.

  • The closed vocabulary must grow, under review, when a genuinely new lane shape appears. This is intentional friction and it is the trust boundary. D2 shows the cost is real: the first survey already forced one widening.
  • Third-party lanes are data-only (D6). "Any plugin can ship a reviewer lane" overstates what this delivers; the ADR and the epic should both say most.
  • Consent-at-install is a weaker gate for a standing egress channel than for a hook (D5). There is no sandbox.
  • discloseExecutableSurfaces is already cyclomatic 51 / cognitive 99 with five dependents. Adding a fourth class lands in an existing hotspot; the implementing phase should extract per-class helpers rather than grow the switch, and should expect the mutation gate to bite.
  • Two declaration mechanisms coexist for one release (D9).
  • Normalizing empty-output handling is observable on lanes that previously returned nothing silently. That is a bug fix that breaks a workaround, and it needs a changeset note rather than a silent correction.

Explicitly unchanged: reviewer selection precedence (ADR-0011), the REVIEWS.md contract (ADR-1517), and every existing lane's observable command shape.

Implementation phases (dependency-ordered)

Verified with /adr-phase-coverage: every deliverable is claimed by exactly one phase, every hand-off lands, and every user-facing capability has a phase that wires its entry point.

Every decision is mapped to the phase that delivers it, so no decision is left to be "handled somewhere".

Phase Issue Delivers Deliverable
0 #2793 — This ADR
1 #2794 D4 (explicit-selection carve-out only) Core single-sourced invocation descriptor + DEFECT.GENERATIVE-FIX parity assertion; corrects the silently-degrading selector — closes #2690
2 #2795 D1, D2, D3, D4, D7, D8 Manifest reviewer body and the third role; transport and probe.kind closed enums; registry harvest, validation, uniqueness; the absent-safe invariant and its warning channel
3 #2796 D5 The fourth trust-disclosure class: binary + args / host + hostConfigKey, egress payload classes, signature binding
4 #2797 D9 (config half) Federated config migration per the ownership table, same-commit
5a #2798 D9 (roster half) The 11 existing lanes declare reviewer bodies; roster derives; hardcoded tail deleted
5b #2799 D6, D5 (invocation-time host re-verification) invoke_reviewers / write_reviews iterate lanes; the antigravity and openai-compatible handler modules ported from the existing bash legs; the kimi-code lane — closes #2718
6 #2800 — Docs, hostBehaviors documentation gap, capability matrix, locale parity gate — closes #2781
7 #2801 D9 (alias removal) Remove the hostBehaviors.reviewerCli alias, the release after 5a

Two mappings are worth calling out because a reader would otherwise assume the wrong phase. D6's handler modules are code, not data — porting Antigravity's ~100-line three-layer fallback and the three OpenAI-compatible lanes into named first-party modules is Phase 5b's work, delivered alongside the iteration that calls them. And D5 splits across two phases: the disclosure itself is Phase 3, but the invocation-time destination re-verification necessarily lands in Phase 5b, because that is where the invocation path is built.

Why kimi-code lands in 5b and not 5a. 5a makes the roster derive from declared bodies, but 5b is what makes invoke_reviewers iterate them. The eleven existing lanes already have hand-authored legs, so declaring them in 5a changes nothing observable. kimi-code is net-new with no leg — declaring it in 5a would make it selectable but not invocable: present in --all, selected, and producing an empty section for the whole 5a → 5b window. Landing it with the iteration keeps the Phase 1 parity assertion green across the entire migration.

Alternatives considered

  1. A single unified invoke shape. The design this ADR started from. Rejected on evidence: a read of all twelve legs found three that are HTTP endpoints with no binary (see Context). Had it shipped, Phase 2 would have bolted on an implicit second shape or stranded three lanes in the hardcoded tail this epic exists to delete.
  2. Transport inferred from field presence (binary ⇒ spawn, hostConfigKey ⇒ http). Fewer fields; rejected because a manifest with both or neither has undefined meaning.
  3. A spawn-only body, leaving the three HTTP lanes in core. Smaller and sooner; rejected because it preserves a hardcoded tail and permanently bars a third party from shipping a local-model lane — the epic's own problem statement in miniature.
  4. Core descriptor table only (#2690 as filed). Single-sources invocation inside review-reviewer-selection.cts and collapses the eleven blocks. Cheaper and lands sooner, and it does fix the cross-cutting-defect cadence — but it does not make lanes installable: still a core patch, still no trust gate, still no federated config. Not discarded — adopted as Phase 1, so the descriptor shape is designed once under this ADR rather than twice.
  5. Keep hostBehaviors.reviewerCli, just document and validate it. Cheapest, and it does close the documentation gap. Rejected because it leaves problems (b) and (c) intact: a lane still cannot own its config, and the five non-installable reviewers still have nowhere to live.
  6. A third-party handler module confined to the install root. See D6 — the only option that genuinely delivers "any plugin"; rejected here as reversing rather than amending ADR-1016, and as the surface ADR-1244 D7 sequences last. Revisit with its own ADR.
  7. Route lanes through MCP. Rejected: reviewers are batch, single-shot, ten-to-twenty-minute invocations. An MCP server lifecycle adds nothing, and mcpServers disclosure already covers the cases that genuinely are servers.
  8. One role: "reviewer" for every lane, splitting the six dual-purpose runtimes. Cleaner discriminator; rejected for churn (D3).

Amendments

2026-07-29 — vocabulary widened by Phase 1 (#2794)

Phase 1 built the core descriptor table against all eleven shipped legs, which is the first time every lane's contract was written down in one place. That surfaced four cases the original survey did not cover. All four are additive widenings of closed enums, none reverses a decision, and each is forced by a lane that exists today rather than by a hypothetical:

# Decision Was Is Forced by
1 D2 promptChannel: stdin | argv | argv-file-ref adds none coderabbit is fed no prompt — it reviews the working-tree diff
2 D2 outputChannel: stdout ("exactly one member today") adds file-arg codex already writes via -o/--output-last-message and discards stdout (#1698)
3 D2 — adds outputArg, required iff file-arg knowing the review lands in a file is useless without the argument naming it
4 D1, D8 flag: string flags: string[], uniqueness flattened antigravity is selected by both --antigravity and --agy

Why this is the process working, not a design failure. D2 already records that the original draft assumed one lane shape and that reading all twelve legs disproved it — openai-http exists because a survey produced evidence, not because anyone predicted it. These four are the same mechanism at the next level of detail: the vocabulary widens when a real lane does not fit, under review, and never on speculation. D6's escalation path ("file an issue naming the primitive the vocabulary lacks") is for third parties; a first-party phase that finds the gap while implementing amends the ADR directly, which is what happened here.

What this does not change. No decision is reversed. transport remains a closed two-member discriminator selecting the invoke sub-shape; probe.kind and handler are untouched; the absent-safe invariant (D4), the disclosure class (D5), and the config-ownership table (D9) are unaffected. Phase 2 (#2795) implements the manifest validator against the vocabulary as amended here, which is the point of amending rather than leaving it for Phase 2 to rediscover.

2026-07-29 — three factual corrections from Phase 2 (#2795)

Implementing the validator required reading the code each claim rests on. Three statements above did not survive that reading. None changes a decision; each would have misdirected a later phase, which is precisely why they are corrected here rather than worked around in code.

1. The cause of the stranded config keys was misattributed (Context (b), Scope of changes, D9).

The ADR attributes reviewer config keys living centrally to the runtime body forbidding feature-only fields. That is not the mechanism. FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME is ['skills','agents','steps','contributions','gates','hooks','activationKey'] — config is not in it, and a role: "runtime" capability carrying a config slice passes validation today. The real cause is two harvest sites that never read it:

  • gen-capability-registry.cjs nested its config-harvest loop inside the role === 'feature' branch, so a non-feature capability's config was silently dropped from configKeys/configSchema.
  • validateCrossCapability opened its config-key ownership loop with if (cap.role !== 'feature') … continue, so a non-feature capability was exempt from both single-ownership and the central-schema collision check.

Phase 2 fixes both by filtering on the presence of a config slice rather than on the role. This matters for Phase 4 (#2797), which would otherwise have been designed against a constraint that does not exist — and it means the exclusivity invariant was, until now, unenforced for every non-feature capability rather than merely unused.

2. D3's profile-membership claim is inverted.

D3 states that a role: "reviewer" capability "receives profile membership from deriveProfileMembership (gen-capability-registry.cjs:201-213) like any other" and that the membership is inert. It receives no membership: that function skips any capability without a non-empty skills array, and a lane-only capability has none. The intended outcome — a reviewer capability installs nothing — holds exactly as D3 wanted, and tier remains required as the source of truth for the requires-closure. Only the stated mechanism was wrong, and a Phase 5a author following D3 would have gone looking for membership that is not there.

3. The specified capability folder names for two lanes would fail the build.

The Scope-of-changes section and #2798 both name capabilities/lm_studio/ and capabilities/llama_cpp/. Both would be rejected: id must equal the folder name and match KEBAB_RE (/^[a-z][a-z0-9-]*$/), which does not admit _. Three namespaces are in play for one lane and they are deliberately not the same string:

value casing fixed by
capability id / folder lm-studio kebab the id conformance invariant
reviewer.slug lm_studio snake the shipped roster and review.lm_studio_host, which D9 leaves unchanged
reviewer.flags --lm-studio kebab the shipped flag

Phase 2 therefore validates reviewer.slug against its own pattern rather than reusing KEBAB_RE, which would have rejected two shipped lanes. Phase 5a must create capabilities/lm-studio/ and capabilities/llama-cpp/, each declaring the snake-case slug.

(Corrected 2026-07-29: this paragraph first recorded the pattern as /^[a-z][a-z0-9_-]*$/. Phase 2's own security review caught that as a divergence from Phase 1's exported LANE_SLUG_RE, which permits a leading digit — a model-named lane such as 4o-mini would have been accepted by the core descriptor and rejected by the manifest validator, reintroducing exactly the translation layer this epic deletes. The shipped pattern is /^[a-z0-9][a-z0-9_-]*$/, and a parity assertion now fails if the two ever drift again.)

2026-07-29 — Phase 4 and Phase 5a are swapped (ordering correction from Phase 5a)

The phase table above runs Phase 4 (federated config) before Phase 5a (lane declarations), and #2798 states "Depends on Phases 2 and 4". That ordering is inverted, and it makes Phase 4 unsatisfiable.

D9's ownership table assigns review.ollama_host to the ollama lane, review.lm_studio_host to lm_studio, and review.llama_cpp_host to llama_cpp. A federated config slice lives inside a capabilities/<id>/capability.json — and those capability directories do not exist until Phase 5a creates them. Verified before the swap: capabilities/{ollama,lm_studio,llama_cpp,gemini,coderabbit} were all absent, and only the six reviewerCli-flagged runtime capabilities existed.

So in the stated order Phase 4 has nowhere to put three of its five key families, and its own "Done when" — "review.<host>_host owned by lane capabilities" — cannot be met. Shipping it unmet would be a failed deployment under CI.GATE.acceptance-criteria-required.

5a's stated dependency on Phase 4 is likewise unfounded: declaring a reviewer body requires only the manifest vocabulary from Phase 2. The real dependency graph is Phase 2 → 5a → 4, with 5a → 5b → 6 unchanged. Nothing about either phase's content changes — only their order.

A second correction, to #2798's acceptance list. It requires "docs/INVENTORY.md updated + node scripts/gen-inventory-manifest.cjs --write run after build:lib". That rests on a false premise: the inventory catalogs bin/lib/*.cjs modules, not capability directories — antigravity, opencode and qwen appear zero times in it, and INVENTORY-MANIFEST.json's six families contain no capabilities/ entry at all. gen-inventory-manifest.cjs --check passes with the five new capability directories added and no inventory edit. The item is vacuous for this phase, and inventing an edit to satisfy it would introduce drift rather than prevent it.

2026-07-30 — vocabulary widened by Phase 5b (#2799)

Phase 5b is the cutover: it deletes the ~640 lines of hand-authored bash and runs every lane from the declaration. Building the resolver against all twelve legs — the first time each leg's runtime contract, not just its shape, had to be reproduced — surfaced five gaps. All five are additive, each is forced by a lane that ships today, and none reverses a decision. This is the same mechanism D2 and the Phase 1 amendment record, at the next level of detail.

# Decision Was Is Forced by
1 D6 handler: null | antigravity | openai-compatible adds opencode opencode's review is RECONSTRUCTED from assistant text parts of a --format json stream; a plain stdout copy writes the raw JSON envelope into REVIEWS.md (#1936). Admitted under the enum's own second arm — a documented upstream defect data cannot express — exactly as antigravity was
2 D1 model key implicit as review.models.<slug> adds reviewer.modelConfigKey antigravity's slug is antigravity but its shipped key is review.models.agy. Resolving by slug misses it and silently ignores a configured model, disabling the pinned-model escape hatch #2073 added
3 D2 invoke.args a fixed array an argv template over a closed four-member placeholder set ({{model}}, {{effort}}, {{output}}, {{prompt}}) The injected pieces do not all go in the same place: codex injects the model after its exec subcommand and the output file later still, while five lanes end with a bare - that must stay last. Positional splicing produced codex --model M -o F exec --ephemeral …, which is not a valid invocation
4 D2 openai-http invoke had no default adds invoke.defaultHost and invoke.fallbackModel Phase 4 federated every review.*_host with a default of "", so the real fallback (http://localhost:11434, llama3, …) existed only inside the bash leg. A data-driven lane would POST to a garbage URL
5 D7 — kimi-code lands with a command-capability probe Net-new lane, per the phase table. kimi is claimed by both Kimi Code CLI and the legacy Python kimi-cli (analysis from closed PR #2776, credit @drungrin)

modelConfigKey is OPTIONAL, and that is D4 rule 2 rather than a convenience. It did not exist before this phase, so requiring it would fail validation on every reviewer manifest authored against an earlier GSD. Absent reads as null.

D5 rule 1 was recorded as delivered by Phase 3 and was not implemented. The implementation note added to D5 on 2026-07-29 states that "the consent record additionally stores the resolved host". It did not: ConsentRecord carried no host field, recordProjectConsent accepted none, and nothing in the tree bound one — so this phase's rule-4 comparison had no baseline to compare against. Phase 5b implements it, as an optional reviewerHost that isValidConsentRecord does not require, so no record already on disk is invalidated and no re-consent storm fires (D4 rule 5). It stays out of disclosureSignature for the reason that note gives. Recorded here because the ADR asserting a rule was delivered is precisely what would stop a later phase from checking.

Three runtime dependencies leave the review path, and two of them were platform holes rather than mere overhead: jq (absent on stock Windows/Git-Bash, #2589 — it gated five lanes), curl, and the external timeout/gtimeout the Antigravity leg probed for. Stock macOS ships neither killer, so D7's "where no bounding mechanism is available the probe is skipped" carve-out was, in practice, that lane running unbounded on every stock Mac. spawnSync's native timeout is always available, so the bound is now unconditional and that carve-out is obsolete.

The DEFECT.GENERATIVE-FIX parity gate is re-pointed. Phase 1's assertion required a literal <!-- reviewer-lane: <slug> --> per lane inside invoke_reviewers and a literal ## <Section> Review per lane inside write_reviews — the exact text this phase deletes. Those two families could not be kept without keeping the hand-maintained per-lane blocks the epic exists to remove, so they are replaced by descriptor ↔ registry parity in both directions (the registry is what the runtime iterates once lanes are data) plus an anti-parity assertion that fires if a bespoke leg is ever re-added. That also gives #2781/Phase 6 the mechanical single source its docs and locale gate needs, which per-leg text could never provide.

2026-08-03 — D2 spawn invoke vocabulary widened by #2483 (invoke.env)

The claude lane was the only reviewer additionally inheriting the invoking user's global CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory — a context asymmetry against the workflow's own independent-review premise, since gemini sees only the assembled prompt and codex runs --ephemeral. Closing it needs two environment variables set for that one spawn. Additive, and forced by a lane that ships today.

# Decision Was Is Forced by
1 D2 spawn invoke carried no way to shape the child's environment adds invoke.env — optional, an object of environment name/value pairs with string values only; forbidden on openai-http The claude lane must spawn with CLAUDE_CODE_DISABLE_CLAUDE_MDS=1 CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 (#2483). The pairs are static per lane, so this is declared data — not a handler, which D6 reserves for behavior data cannot express

Declared data rather than a handler, and D6 is the wrong authority for it. An earlier revision of this change cited D6 in the source comment. D6 governs the closed handler enum — imperative behavior admitted first-party — and says nothing about the invoke field vocabulary, which is D2's territory. The citation did not cover the widening, which is why this entry exists rather than a code comment pointing at the wrong decision.

env is OPTIONAL, per D4 rule 2, exactly as modelConfigKey was in the Phase 5b entry above: it did not exist before this change, so requiring it would fail validation on every reviewer manifest authored against an earlier GSD. Absent means the lane inherits the environment unchanged.

Forbidden on openai-http, and registered in the discriminator to make that enforceable. An openai-http lane spawns no child, so an environment pair there has no referent. The first implementation validated env's shape but left it out of SPAWN_ONLY_INVOKE_FIELDS — the list the openai-http arm rejects against — so it was accepted on that transport in silence, alone among the spawn fields. Both registrations are required; neither implies the other.

What this does not change. No decision is reversed. transport remains a closed two-member discriminator; effortChannel stays in neither field list because D2 defines it for both transports, so it is shared rather than spawn-only.

env IS added to the D5 disclosure signature, and to the human consent prompt. An installed overlay reviewer body reaches resolveLanePlan and is executed: routeReviewLane builds its lane map from mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true})) (D8, #2927 / #3062), and that merge is field-identical per D1 — it admits the overlay body without deep-validating invoke, precisely because the invocation seam is where a lane is re-validated before it runs. So a third-party manifest can declare env on a reviewer lane and have those pairs applied to the spawned child. Undisclosed, that is arbitrary code execution behind a consent prompt that never mentioned it (NODE_OPTIONS=--require ./evil.js; LD_PRELOAD on POSIX). D5 already folds env into MCP-server disclosure and names that exact shape as the reason; reviewer lanes now carry the identical treatment.

2026-08-05: the earlier reading — that env needed no disclosure because manifest invoke fields never reach resolveLanePlan — is withdrawn. It was true when written and #3062 retired it. See git history for the superseded text.

The enumeration was the defect, not the missing name. env was the ninth invoke field that reaches resolveLanePlan without being bound by the D5 lane signature; the other eight were defaultHost, path, outputChannel, outputArg, modelArg, effortChannel, modelDiscovery and fallbackModel. Two of those are egress-relevant on their own — defaultHost is the destination the manifest itself declares, used whenever hostConfigKey resolves to nothing (configured ?? declaredDefault), and path completes the URL — so a lane with an unresolved config key disclosed "(unresolved …)" while shipping the D5 egress payload classes to an address of the manifest's choosing. Adding a ninth name would have left a tenth open, so the lane element instead carries a residual of every other declared invoke key, mirroring the rawConfig completeness backstop the MCP surface has carried since #1459 finding 5. env and defaultHost are additionally named explicitly, mirroring that same line's deliberate explicit-then-backstop overlap.

D4.5 is preserved one level down, and the cost it guards against does not arise here anyway. The residual element is appended to the lane tuple only when the lane declares something beyond the eight already-bound fields, so a lane declaring none keeps a byte-identical signature. State the scope of that property honestly, because it is easy to oversell in both directions:

  • It is vacuous for any VALID lane, and that is the honest statement. Once the residual covers the lane body's outer fields too, a lane that produces no residual is one declaring no flags, no probe, no emptyOutput, no evidenceClass, no requiresBinaries and no promptBudgetKey — i.e. a body the validator rejects. Measured across the twelve first-party reviewer capabilities: zero are in the byte-identical class. The conditional append is still correct — it keeps the signature minimal and means the residual element carries information when present — but it is a property of the encoding, not a claim that anyone's signature is unchanged.
  • And no capability is re-prompted regardless. A code change to disclosureSignature cannot invalidate an existing consent: hasProjectConsent matches on the recomputed bundle contentHash — the signature is explicitly "no longer the security binding" (#1459 CB-1/CB-2) — and the upgrade path's executableSetChanged(old, new) compares two disclosures both computed by the current code, so widening the signature shifts both sides equally.
  • What the widening actually buys is therefore forward-looking and is the whole point: an upgrade whose manifest edits env, defaultHost, or any other declared invoke field now registers as an executable-surface change and re-consents, where previously it could change what the lane runs in silence. First-party capabilities never enter this path at all — the install flow blocks a first-party id before trust evaluation.

Validation is defence in depth; consent is the boundary. invoke.env is validated for object shape, POSIX name grammar and string values, and a denylist refuses execution-primitive names outright on a reviewer lane — PATH, NODE_OPTIONS, LD_PRELOAD, DYLD_INSERT_LIBRARIES, BASH_ENV, PYTHONPATH, PERL5OPT, RUBYOPT, GIT_SSH_COMMAND, JAVA_TOOL_OPTIONS and siblings, matched case-insensitively because Windows environment lookup is. PATH is included deliberately: it is the most complete primitive of the set, and a lane needing a specific executable declares an absolute invoke.binary rather than reshaping the child's PATH.

State the limit plainly, because the list invites being mistaken for the control: it cannot be complete against an arbitrary third-party child, and disclosure runs before validation — on manifests validation would reject. So the boundary remains install-time consent, which shows every declared pair and binds it to the signature; execution-primitive names additionally carry a warning line in the prompt. A name missing from both lists costs a quieter line on a value the user is still shown.

One inconsistency this entry closes, and it was real. Consequences above states that adding a reviewer is "one manifest … no core patch", and CONTEXT.md, gsd-core/workflows/review.md and resolveLanePlan's own header all describe overlay manifests reaching the resolver — while the runtime, until #3062, built laneBySlug solely from the first-party table and rejected every slug absent from it. Four documents on one side, the runtime on the other. #3062 resolved it in the documents' favour, which is what makes the disclosure above mandatory rather than defensive.