Files
msd-core/docs/how-to/ship-a-reviewer-lane.md
Tom Boucher 653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00

16 KiB
Raw Blame History

How to ship a reviewer lane in your capability

Goal: Declare a reviewer lane in a capability manifest so /gsd-review discovers your external review CLI or model endpoint, offers a flag for it, invokes it, and renders its output into REVIEWS.md — without patching GSD core.

Prerequisites: You already have a capability (capability.json), or you are creating one. The reviewer tool is installed and works from your shell. GSD 1.9.0 or later.

Before 1.9.0 a reviewer was a core patch: a hardcoded roster entry, a hand-authored bash leg in review.md, a hardcoded output heading, and central config keys. From 1.9.0 a lane is manifest data, so shipping a reviewer is shipping a capability. See ADR-2782 for the decision record.


Decide which shape your lane takes

A reviewer body is admissible on two roles. Pick by whether GSD installs into your tool.

Your situation Use Why
Your capability is already a runtime GSD installs into (it has a runtime body) and that same CLI can also review Keep role: "runtime", add a reviewer body One manifest stays one manifest — this is how codex, cursor, and antigravity ship
Your tool only reviews — GSD never installs commands, agents, or skills into it role: "reviewer" The honest description: a lane with no install surface, like gemini, coderabbit, and ollama
Your capability adds planning steps, gates, or contributions role: "feature" — and a separate lane capability A feature manifest may not carry a reviewer body; the validator rejects it

A role: "reviewer" capability must carry a reviewer body, must not carry a runtime body, and must not carry any feature-only field — skills, agents, steps, contributions, gates, hooks, or activationKey. A lane owns no artifacts and wires no loop extension point.


Declare a spawned-CLI lane

Most lanes are transport: "spawn" — GSD runs a binary and reads its output. Add a reviewer block to your manifest:

{
  "id": "acme-review",
  "role": "reviewer",
  "version": "1.0.0",
  "title": "Acme Review CLI",
  "description": "Acme CLI — cross-AI /gsd-review reviewer lane only; not a GSD install target.",
  "tier": "full",
  "requires": [],
  "engines": { "gsd": ">=1.9.0" },

  "reviewer": {
    "slug": "acme",
    "flags": ["--acme"],
    "transport": "spawn",
    "probe": { "kind": "command-exists", "binary": "acme" },
    "invoke": {
      "binary": "acme",
      "args": ["review", "{{model}}", "-p", "-"],
      "promptChannel": "stdin",
      "outputChannel": "stdout",
      "modelArg": "--model",
      "effortChannel": "none"
    },
    "timeoutFloorMs": 900000,
    "emptyOutput": "stub-with-stderr",
    "reviewsSection": "Acme",
    "evidenceClass": "source-grounded",
    "requiresBinaries": [],
    "promptBudgetKey": null,
    "modelConfigKey": "review.models.acme",
    "handler": null
  },

  "config": {
    "review.models.acme": {
      "type": "string",
      "default": "",
      "description": "Model passed to the Acme reviewer lane."
    }
  }
}

Four fields decide whether the lane works at all, so get these right first:

  • invoke.args carries the {{model}}, {{prompt}}, {{effort}}, and {{output}} placeholders. GSD substitutes them; anything else is passed through literally.
  • promptChannel says how the plan text reaches the tool — stdin when it reads a pipe, argv or argv-file-ref when it takes the prompt path as an argument, none when the tool reads the working tree itself (as CodeRabbit does).
  • outputChannel is stdout, or file-arg when the tool writes to a path you name (then also declare invoke.outputArg, as codex does with -o).
  • timeoutFloorMs is the measured wall-clock floor for your tool. Lane divergence here is expected and correct — the descriptor exists to declare divergence in one place, not to impose one number on every lane.

reviewsSection is the heading your findings render under in REVIEWS.md. It must be unique across every installed lane; two lanes sharing a heading would silently merge their output into apparent consensus that never happened.

For the full field table — every type, enum member, and default — see Capability manifest § Reviewer body.


Declare an OpenAI-compatible HTTP lane instead

If your reviewer is a served model endpoint rather than a CLI, use transport: "openai-http". The invoke block takes a different shape — a destination, not a binary:

"reviewer": {
  "slug": "acme_local",
  "flags": ["--acme-local"],
  "transport": "openai-http",
  "probe": {
    "kind": "http-reachable",
    "hostConfigKey": "review.acme_host",
    "path": "/v1/models",
    "timeoutMs": 2000
  },
  "invoke": {
    "hostConfigKey": "review.acme_host",
    "defaultHost": "http://localhost:8080",
    "path": "/v1/chat/completions",
    "modelDiscovery": "first-from-models-endpoint",
    "fallbackModel": "acme-7b",
    "effortChannel": "none"
  },
  "timeoutFloorMs": 120000,
  "emptyOutput": "stub-with-stderr",
  "reviewsSection": "Acme Local",
  "evidenceClass": "source-grounded",
  "requiresBinaries": [],
  "promptBudgetKey": "review.max_prompt_tokens_per_reviewer.acme_local",
  "modelConfigKey": "review.models.acme_local",
  "handler": "openai-compatible"
}

handler: "openai-compatible" is what gives you model discovery against /v1/models, the request and response shape, and the served-model mismatch warning. Declare the matching review.acme_host key in your config block alongside the model key.

Every probe is bounded. An unbounded --help | grep probe is a named defect in this repo. http-reachable requires timeoutMs; so does command-capability, the probe kind you use when a bare binary name is ambiguous.


Own your lane's config keys

Declare the lane's keys in your own manifest config block, never in the central schema. A key present in both is a build failure, not a warning — federated ownership is exclusive.

Name modelConfigKey and promptBudgetKey to match keys you actually declare. A lane pointing at a key nobody owns resolves to nothing, which reads to the user as "my model override is being ignored."

Users then set them the ordinary way, in .planning/config.json:

{ "review": { "models": { "acme": "acme-large" } } }

Build and install

If your capability lives in the GSD repo, regenerate the committed registry and check for drift:

npm run gen:capability-registry
npm run lint:generated-sync

If you are shipping out-of-tree, package and install it like any other capability:

gsd capability install <url>

Uniqueness is checked across the merged first-party ∪ overlay set — a duplicate slug, a duplicate entry in flags, or a duplicate reviewsSection collides. Expect that collision to be quiet. gsd capability install does not run the cross-capability check; it runs at load time, and a colliding overlay is dropped from the active set with a warning rather than failing the install. First-party always wins.

That failure mode is worth internalizing before you debug it: the install command reports success, and your lane simply never appears. If a lane you just installed is missing from gsd-tools review-lane sections, suspect a name collision before you suspect the probe. Malformed values inside the body — a slug outside ^[a-z0-9][a-z0-9_-]*$, an enum member that does not exist, an outputArg without outputChannel: "file-arg" — are ordinary validation errors and are reported directly.

Two naming rules are easy to conflate, so keep them apart. Your slug may not be __proto__, constructor, or prototype — a prototype-pollution guard, not a namespace policy; any other grammatical slug is yours, including one starting gsd-. Your capability id, separately, may not begin with gsd-, gsd-core-, or anthropic-; those prefixes are reserved so nothing can impersonate a first-party capability.

An unknown field inside your reviewer body behaves differently: it is a non-fatal warning on stderr, never a build failure. A manifest built against a newer GSD degrades visibly instead of crashing.

Publish it so people can find it

From GSD 1.9.1, a lane has its own discoverability catalog: the Reviewer Lane Registry. Listing is a documentation PR — append one entry to docs/registries/reviewers.json, regenerate, open a PR. Register once; your GitHub Releases are the update channel from then on.

Follow List your reviewer lane in the registry for the task flow.

The Reviewer Lane Registry is for lanes that are not install targets in their own right. If your reviewer body rides on a role: "runtime" capability, list that capability under whichever catalog matches its primary install shape instead — one entry, not two.


Verify the lane resolves

Check that GSD sees your lane before you run a real review:

gsd-tools review-lane sections
gsd-tools review-lane flags

sections lists every slug with the heading it renders under; flags lists every selector flag. Your lane appears in both, or it is not installed.

Then dry-check the invocation plan for your lane alone:

gsd-tools review-lane plan --selected acme

A resolvable lane returns "ok": true with its section, transport, and prompt path. Once that is green, run it for real against a planned phase:

/gsd-review --phase 3 --acme

If the lane is absent from --all, the probe is the usual culprit: command-exists fails silently when the binary is not on PATH in the environment GSD runs in.


Know what your users are consenting to

A reviewer lane is a fourth executable-surface disclosure class, alongside hooks, command modules, and MCP servers — and it is the only one that receives data. Your lane is piped plan text, requirements, research findings, and CONTEXT.md decisions. Install-time disclosure says so plainly:

  reviewer lane (1): an external reviewer receives plan/review data on every run
    - acme -> acme review --model acme-large -p -
        sends: plan text, requirements, research findings, CONTEXT.md decisions

Be clear-eyed about what that buys, because your users are trusting your judgment as much as the mechanism. ADR-2782 D5 says it plainly: disclosure and host pinning make the channel "visible, pinned, and revocable — it does not make it safe", and consent-at-install is a weaker gate for a standing egress channel than for a hook. A user consents once; your lane thereafter receives every plan on every review run. A per-run prompt was considered and rejected as consent fatigue. Design your lane as if that single consent is the only one you will ever get, because it is.

Three consequences you should design for:

  • Your args are signature-bound, not just your binary. Changing binary, args, hostConfigKey, promptChannel, or handler in a new version re-triggers consent on update. Changing reviewsSection or timeoutFloorMs does not — a cosmetic prompt is how users learn to click through.
  • An openai-http lane binds the resolved host, not just the config key. GSD re-resolves hostConfigKey before every invocation and blocks the lane if the destination changed, rather than silently sending plans somewhere new. Users see the lane refuse and must re-consent. Point defaultHost at the address you actually mean.
  • What the user saw is only tamper-evident because the bundle is pinned. The disclosure is trustworthy because the downloaded capability is integrity-checked and its hash recorded at install; a later change to your args or hostConfigKey shows up as a changed signature rather than sliding in quietly. Keep engines.gsd honest for the same reason — it is a hard gate, so a lane declaring a range it does not actually work on is blocked at install and skipped at load rather than failing confusingly at review time.

For the reasoning behind consent-plus-integrity rather than a sandbox, see The capability trust model.


Migrate off the removed reviewerCli flag

Before 1.9.0, a runtime capability declared itself a reviewer with a boolean in the open host-behaviors bag:

"runtime": { "hostBehaviors": { "reviewerCli": true } }

That flag carried no invocation data — it only added the capability id to the roster, leaving the probe, argv shape, timeout, and output policy hardcoded in GSD core. It was superseded by the reviewer body in 1.9.0, kept working for one release as a derived alias, and was removed in the release after that.

Symptom. Your capability installs and validates exactly as before, but /gsd-review no longer offers your flag and your lane never runs. On a registry build or a capability install you will see:

⚠ capability "your-cap" runtime.hostBehaviors.reviewerCli was removed (ADR-2782 D9)
  — ignored, and it contributes no reviewer lane. Declare a `reviewer` body instead;
  see docs/how-to/ship-a-reviewer-lane.md

Fix. Delete the flag and declare a reviewer body, following Declare a spawned-CLI lane above. Your reviewer.slug should be whatever the flag used to contribute — your capability id — so existing review.default_reviewers entries and --<slug> flags keep working:

{
  "id": "your-cap",
  "role": "runtime",
  "runtime": { "hostBehaviors": { } },
  "reviewer": {
    "slug": "your-cap",
    "flags": ["--your-cap"]
  }
}

Then rebuild and re-verify with the steps in Build and install and Verify the lane resolves.

Two things the migration buys you beyond restoring the lane: your invocation shape becomes declared data rather than something GSD core has to know about, and your lane can own its own config keys (see Own your lane's config keys).

Nothing else about your capability changes. A manifest still carrying the removed key parses, validates, and installs exactly as before — it simply contributes no lane, and says so. The key is inert, not fatal.


Conditionals: when the vocabulary does not fit your tool

Third-party lanes are data-only. handler is a closed enum of first-party names (antigravity, openai-compatible, opencode, or null) — you may reference an existing member, but you cannot ship your own handler module.

Your tool What to do
Runs one command, reads a prompt, writes a review Declare it — the vocabulary covers this, which is eight of the twelve shipped lanes
Is an OpenAI-compatible endpoint transport: "openai-http" with handler: "openai-compatible"
Needs a stateful setup turn before it can review (Plandex-style new then review) Not expressible today — the descriptor describes one invocation. File an issue naming the primitive
Edits files or commits by default (Aider-style) Not expressible today — there is no way to declare a read-only invocation posture, and the prompt asking politely is not a guarantee. File an issue naming the primitive
Needs genuinely imperative behavior for an upstream bug File an issue. Named handlers are added first-party after review, the same path that widened the vocabulary to add openai-http

Filing the issue is the supported route, not a workaround. The openai-http transport exists because three real lanes did not fit and the vocabulary widened on that evidence.