* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal Inverts the Phase 5a rows that assert the derived legacy alias still contributes a reviewer slug, and adds the removal-warning coverage the alias's exit needs (ADR-2782 D9). RED against unmodified production code, by design: the six shipped manifests still declare the key and collectReviewerWarnings emits nothing for hostBehaviors. Refs #2801 * chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias ADR-2782 D9, Phase 7 — the final phase of epic #2782. The derived legacy alias survived one release (Phase 5a shipped in 1.9.0; 1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer body is now the only route onto the reviewer roster. - deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli - the key is stripped from the six manifests that carried it; each already declares a reviewer body whose slug equals its capability id, so the derived roster is unchanged at the same twelve slugs - collectReviewerWarnings emits a presence-based, non-fatal removal notice for any manifest still declaring the key, reaching both the build-time registry generation and the third-party overlay load path. The check runs before the reviewer-body early-return, because the manifest it exists for is the alias-only one that has no body. - hostBehaviors stays an open, unvalidated bag for its other 59 keys; this adds one keyed removal notice, not general validation Refs #2801 * refactor(#2801): give the reviewer-warning channel a typed IR Review finding: the new tests asserted with String#includes() on the warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on Test Outputs' bans in favor of a typed intermediate representation. Adds the IR beside the renderer rather than replacing it, which is the shape that section prescribes and bin/verify-reapply-patches.cjs already models: - REVIEWER_WARNING, a frozen code enum - REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one symbol instead of duplicating a literal - collectReviewerWarningRecords(cap), returning typed records collectReviewerWarnings(cap) keeps its exact string[] contract as a thin map over the records, so both production consumers are untouched. Every section-K row now asserts on record.code/field/capId and none on the rendered message. Locks the code surface, asserts the renderer stays one-to-one with the records, and migrates the pre-existing Phase 2 test on the same channel off prose matching. Refs #2801 * test(#2801): invert the section F alias fall-through regression row Caught by the remote runner: 2 unique failures on both Node lanes out of 31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's isolated-security-review regressions — asserted that a blank reviewer.slug falls through to the hostBehaviors.reviewerCli alias rather than dropping the lane. That is the direct inverse of this phase's contract. The original rationale held only while the alias existed. With it gone there is nothing to fall through to: a blank body is not a declaration, and a declaration is the only route onto the roster. Inverted rather than deleted — the row carries the adversarial-review provenance for the slug trim, and removing a security regression guard to make a change pass is backwards. The duplicate row added earlier in section C is dropped instead; section F is its canonical home. Also corrects two count strings Phase 5b left at eleven while asserting twelve, which would misreport on failure. Refs #2801 * docs(#2801): give the removed reviewerCli flag a migration path The Reference edit alone satisfied CI — a file under docs/ moved, so lint-docs-required.cjs was green — while the task-oriented quadrant said nothing about the removal. A maintainer whose lane had just gone silent would have found the field documented as removed and no page telling them what to do about it. Adds a migration section to the how-to: the symptom, the verbatim warning they will see, the before/after manifest, and the note to keep the reviewer slug equal to the capability id so existing review.default_reviewers entries and --<slug> flags survive. Refs #2801 * chore(#2801): backfill changeset pr number to 3272 * feat(#2801): close the runtime.hostBehaviors vocabulary ADR-1016 closes twelve descriptor axes and rejects an open escape hatch in the descriptor. It never mentioned runtime.hostBehaviors, and that silence was read as permission: 59 keys across 18 manifests, 39 of them set by a single capability, validated by nothing. The reference docs went further and attributed the open seam to ADR-1016, which does not mention the field at all. KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the alias removal notice, reaching both build-time generation and overlay install. Warning, never error, for the reason this phase exists: an error would hard-break an out-of-tree descriptor carrying a bespoke key with no deprecation window, which is what reviewerCli was given a release to avoid. Escalation is a separate decision. reviewerCli is excluded from the unknown-key sweep so it keeps its own notice with the migration pointer rather than drawing two records. A parity test binds the vocabulary to the shipped manifests in both directions, and a second asserts no shipped capability draws a notice, so the closure is provably inert in-tree. Records the decision and the miscitation as an ADR-1016 amendment. Refs #2801 * fix(#2801): bound and sanitize the unknown-key diagnostics Two findings from an isolated adversarial review of the closure commit, both proven by execution rather than asserted. MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep had no ceiling. An installed third-party manifest is bounded only by MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced 800,000 records and ~139MB of message text, retained for the registry's lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a summary carrying omittedCount, mirroring capability-loader's existing slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars. MINOR, newly reachable: manifest-supplied key names were interpolated raw. Unlike cap.id, which validateCapability gates on KEBAB_RE before these diagnostics run, hostBehaviors keys have no grammar check anywhere, so ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New describeKey replaces C0/C1 controls and clips at 80 chars. The file already had describeValue for this and applied it only to values. Both fixes land on the pre-existing reviewer.* sweep too — it carried the identical pair, and fixing only the new copy would leave the same defect one screen from its own fix. Refs #2801 --------- Co-authored-by: sim <sim@local>
16 KiB
How to ship a reviewer lane in your capability
Goal: Declare a reviewer lane in a capability manifest so /gsd-review discovers your external review CLI or model endpoint, offers a flag for it, invokes it, and renders its output into REVIEWS.md — without patching GSD core.
Prerequisites: You already have a capability (capability.json), or you are creating one. The reviewer tool is installed and works from your shell. GSD 1.9.0 or later.
Before 1.9.0 a reviewer was a core patch: a hardcoded roster entry, a hand-authored bash leg in review.md, a hardcoded output heading, and central config keys. From 1.9.0 a lane is manifest data, so shipping a reviewer is shipping a capability. See ADR-2782 for the decision record.
Decide which shape your lane takes
A reviewer body is admissible on two roles. Pick by whether GSD installs into your tool.
| Your situation | Use | Why |
|---|---|---|
Your capability is already a runtime GSD installs into (it has a runtime body) and that same CLI can also review |
Keep role: "runtime", add a reviewer body |
One manifest stays one manifest — this is how codex, cursor, and antigravity ship |
| Your tool only reviews — GSD never installs commands, agents, or skills into it | role: "reviewer" |
The honest description: a lane with no install surface, like gemini, coderabbit, and ollama |
| Your capability adds planning steps, gates, or contributions | role: "feature" — and a separate lane capability |
A feature manifest may not carry a reviewer body; the validator rejects it |
A role: "reviewer" capability must carry a reviewer body, must not carry a runtime body, and must not carry any feature-only field — skills, agents, steps, contributions, gates, hooks, or activationKey. A lane owns no artifacts and wires no loop extension point.
Declare a spawned-CLI lane
Most lanes are transport: "spawn" — GSD runs a binary and reads its output. Add a reviewer block to your manifest:
{
"id": "acme-review",
"role": "reviewer",
"version": "1.0.0",
"title": "Acme Review CLI",
"description": "Acme CLI — cross-AI /gsd-review reviewer lane only; not a GSD install target.",
"tier": "full",
"requires": [],
"engines": { "gsd": ">=1.9.0" },
"reviewer": {
"slug": "acme",
"flags": ["--acme"],
"transport": "spawn",
"probe": { "kind": "command-exists", "binary": "acme" },
"invoke": {
"binary": "acme",
"args": ["review", "{{model}}", "-p", "-"],
"promptChannel": "stdin",
"outputChannel": "stdout",
"modelArg": "--model",
"effortChannel": "none"
},
"timeoutFloorMs": 900000,
"emptyOutput": "stub-with-stderr",
"reviewsSection": "Acme",
"evidenceClass": "source-grounded",
"requiresBinaries": [],
"promptBudgetKey": null,
"modelConfigKey": "review.models.acme",
"handler": null
},
"config": {
"review.models.acme": {
"type": "string",
"default": "",
"description": "Model passed to the Acme reviewer lane."
}
}
}
Four fields decide whether the lane works at all, so get these right first:
invoke.argscarries the{{model}},{{prompt}},{{effort}}, and{{output}}placeholders. GSD substitutes them; anything else is passed through literally.promptChannelsays how the plan text reaches the tool —stdinwhen it reads a pipe,argvorargv-file-refwhen it takes the prompt path as an argument,nonewhen the tool reads the working tree itself (as CodeRabbit does).outputChannelisstdout, orfile-argwhen the tool writes to a path you name (then also declareinvoke.outputArg, ascodexdoes with-o).timeoutFloorMsis the measured wall-clock floor for your tool. Lane divergence here is expected and correct — the descriptor exists to declare divergence in one place, not to impose one number on every lane.
reviewsSection is the heading your findings render under in REVIEWS.md. It must be unique across every installed lane; two lanes sharing a heading would silently merge their output into apparent consensus that never happened.
For the full field table — every type, enum member, and default — see Capability manifest § Reviewer body.
Declare an OpenAI-compatible HTTP lane instead
If your reviewer is a served model endpoint rather than a CLI, use transport: "openai-http". The invoke block takes a different shape — a destination, not a binary:
"reviewer": {
"slug": "acme_local",
"flags": ["--acme-local"],
"transport": "openai-http",
"probe": {
"kind": "http-reachable",
"hostConfigKey": "review.acme_host",
"path": "/v1/models",
"timeoutMs": 2000
},
"invoke": {
"hostConfigKey": "review.acme_host",
"defaultHost": "http://localhost:8080",
"path": "/v1/chat/completions",
"modelDiscovery": "first-from-models-endpoint",
"fallbackModel": "acme-7b",
"effortChannel": "none"
},
"timeoutFloorMs": 120000,
"emptyOutput": "stub-with-stderr",
"reviewsSection": "Acme Local",
"evidenceClass": "source-grounded",
"requiresBinaries": [],
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.acme_local",
"modelConfigKey": "review.models.acme_local",
"handler": "openai-compatible"
}
handler: "openai-compatible" is what gives you model discovery against /v1/models, the request and response shape, and the served-model mismatch warning. Declare the matching review.acme_host key in your config block alongside the model key.
Every probe is bounded. An unbounded --help | grep probe is a named defect in this repo. http-reachable requires timeoutMs; so does command-capability, the probe kind you use when a bare binary name is ambiguous.
Own your lane's config keys
Declare the lane's keys in your own manifest config block, never in the central schema. A key present in both is a build failure, not a warning — federated ownership is exclusive.
Name modelConfigKey and promptBudgetKey to match keys you actually declare. A lane pointing at a key nobody owns resolves to nothing, which reads to the user as "my model override is being ignored."
Users then set them the ordinary way, in .planning/config.json:
{ "review": { "models": { "acme": "acme-large" } } }
Build and install
If your capability lives in the GSD repo, regenerate the committed registry and check for drift:
npm run gen:capability-registry
npm run lint:generated-sync
If you are shipping out-of-tree, package and install it like any other capability:
gsd capability install <url>
Uniqueness is checked across the merged first-party ∪ overlay set — a duplicate slug, a duplicate entry in flags, or a duplicate reviewsSection collides. Expect that collision to be quiet. gsd capability install does not run the cross-capability check; it runs at load time, and a colliding overlay is dropped from the active set with a warning rather than failing the install. First-party always wins.
That failure mode is worth internalizing before you debug it: the install command reports success, and your lane simply never appears. If a lane you just installed is missing from gsd-tools review-lane sections, suspect a name collision before you suspect the probe. Malformed values inside the body — a slug outside ^[a-z0-9][a-z0-9_-]*$, an enum member that does not exist, an outputArg without outputChannel: "file-arg" — are ordinary validation errors and are reported directly.
Two naming rules are easy to conflate, so keep them apart. Your slug may not be __proto__, constructor, or prototype — a prototype-pollution guard, not a namespace policy; any other grammatical slug is yours, including one starting gsd-. Your capability id, separately, may not begin with gsd-, gsd-core-, or anthropic-; those prefixes are reserved so nothing can impersonate a first-party capability.
An unknown field inside your reviewer body behaves differently: it is a non-fatal warning on stderr, never a build failure. A manifest built against a newer GSD degrades visibly instead of crashing.
Publish it so people can find it
From GSD 1.9.1, a lane has its own discoverability catalog: the Reviewer Lane Registry. Listing is a documentation PR — append one entry to docs/registries/reviewers.json, regenerate, open a PR. Register once; your GitHub Releases are the update channel from then on.
Follow List your reviewer lane in the registry for the task flow.
The Reviewer Lane Registry is for lanes that are not install targets in their own right. If your reviewer body rides on a role: "runtime" capability, list that capability under whichever catalog matches its primary install shape instead — one entry, not two.
Verify the lane resolves
Check that GSD sees your lane before you run a real review:
gsd-tools review-lane sections
gsd-tools review-lane flags
sections lists every slug with the heading it renders under; flags lists every selector flag. Your lane appears in both, or it is not installed.
Then dry-check the invocation plan for your lane alone:
gsd-tools review-lane plan --selected acme
A resolvable lane returns "ok": true with its section, transport, and prompt path. Once that is green, run it for real against a planned phase:
/gsd-review --phase 3 --acme
If the lane is absent from --all, the probe is the usual culprit: command-exists fails silently when the binary is not on PATH in the environment GSD runs in.
Know what your users are consenting to
A reviewer lane is a fourth executable-surface disclosure class, alongside hooks, command modules, and MCP servers — and it is the only one that receives data. Your lane is piped plan text, requirements, research findings, and CONTEXT.md decisions. Install-time disclosure says so plainly:
reviewer lane (1): an external reviewer receives plan/review data on every run
- acme -> acme review --model acme-large -p -
sends: plan text, requirements, research findings, CONTEXT.md decisions
Be clear-eyed about what that buys, because your users are trusting your judgment as much as the mechanism. ADR-2782 D5 says it plainly: disclosure and host pinning make the channel "visible, pinned, and revocable — it does not make it safe", and consent-at-install is a weaker gate for a standing egress channel than for a hook. A user consents once; your lane thereafter receives every plan on every review run. A per-run prompt was considered and rejected as consent fatigue. Design your lane as if that single consent is the only one you will ever get, because it is.
Three consequences you should design for:
- Your
argsare signature-bound, not just your binary. Changingbinary,args,hostConfigKey,promptChannel, orhandlerin a new version re-triggers consent on update. ChangingreviewsSectionortimeoutFloorMsdoes not — a cosmetic prompt is how users learn to click through. - An
openai-httplane binds the resolved host, not just the config key. GSD re-resolveshostConfigKeybefore every invocation and blocks the lane if the destination changed, rather than silently sending plans somewhere new. Users see the lane refuse and must re-consent. PointdefaultHostat the address you actually mean. - What the user saw is only tamper-evident because the bundle is pinned. The disclosure is trustworthy because the downloaded capability is integrity-checked and its hash recorded at install; a later change to your
argsorhostConfigKeyshows up as a changed signature rather than sliding in quietly. Keepengines.gsdhonest for the same reason — it is a hard gate, so a lane declaring a range it does not actually work on is blocked at install and skipped at load rather than failing confusingly at review time.
For the reasoning behind consent-plus-integrity rather than a sandbox, see The capability trust model.
Migrate off the removed reviewerCli flag
Before 1.9.0, a runtime capability declared itself a reviewer with a boolean in the open host-behaviors bag:
"runtime": { "hostBehaviors": { "reviewerCli": true } }
That flag carried no invocation data — it only added the capability id to the roster, leaving the probe, argv shape, timeout, and output policy hardcoded in GSD core. It was superseded by the reviewer body in 1.9.0, kept working for one release as a derived alias, and was removed in the release after that.
Symptom. Your capability installs and validates exactly as before, but /gsd-review no longer offers your flag and your lane never runs. On a registry build or a capability install you will see:
⚠ capability "your-cap" runtime.hostBehaviors.reviewerCli was removed (ADR-2782 D9)
— ignored, and it contributes no reviewer lane. Declare a `reviewer` body instead;
see docs/how-to/ship-a-reviewer-lane.md
Fix. Delete the flag and declare a reviewer body, following Declare a spawned-CLI lane above. Your reviewer.slug should be whatever the flag used to contribute — your capability id — so existing review.default_reviewers entries and --<slug> flags keep working:
{
"id": "your-cap",
"role": "runtime",
"runtime": { "hostBehaviors": { } },
"reviewer": {
"slug": "your-cap",
"flags": ["--your-cap"]
}
}
Then rebuild and re-verify with the steps in Build and install and Verify the lane resolves.
Two things the migration buys you beyond restoring the lane: your invocation shape becomes declared data rather than something GSD core has to know about, and your lane can own its own config keys (see Own your lane's config keys).
Nothing else about your capability changes. A manifest still carrying the removed key parses, validates, and installs exactly as before — it simply contributes no lane, and says so. The key is inert, not fatal.
Conditionals: when the vocabulary does not fit your tool
Third-party lanes are data-only. handler is a closed enum of first-party names (antigravity, openai-compatible, opencode, or null) — you may reference an existing member, but you cannot ship your own handler module.
| Your tool | What to do |
|---|---|
| Runs one command, reads a prompt, writes a review | Declare it — the vocabulary covers this, which is eight of the twelve shipped lanes |
| Is an OpenAI-compatible endpoint | transport: "openai-http" with handler: "openai-compatible" |
Needs a stateful setup turn before it can review (Plandex-style new then review) |
Not expressible today — the descriptor describes one invocation. File an issue naming the primitive |
| Edits files or commits by default (Aider-style) | Not expressible today — there is no way to declare a read-only invocation posture, and the prompt asking politely is not a guarantee. File an issue naming the primitive |
| Needs genuinely imperative behavior for an upstream bug | File an issue. Named handlers are added first-party after review, the same path that widened the vocabulary to add openai-http |
Filing the issue is the supported route, not a workaround. The openai-http transport exists because three real lanes did not fit and the vocabulary widened on that evidence.
Related
- Capability manifest — the full
reviewerbody field table and validation rules - Capability manifest →
hostBehaviors— why the bag is unvalidated, and the one removal notice inside it - Set up cross-AI review — the user-facing side: choosing, configuring, and running reviewers
- Develop a Capability for GSD 1.5+ — manifests, registry generation, and federated config
- Publish a capability — versioning,
engines.gsd, and distribution - The capability trust model — disclosure, consent, and integrity
- ADR-2782 — why the lane became a capability surface, and what it deliberately excludes