Files
msd-core/docs/reference/capability-matrix.md
Tom Boucher 6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00

187 lines
11 KiB
Markdown

# Capability matrix reference
> **Generated file — do not edit by hand.**
> This matrix is generated from the capability registry by
> `scripts/gen-capability-matrix.cjs` and kept honest by a drift guard
> (`tests/capability-matrix-sync.test.cjs` runs `--check`). Any manual edit is
> overwritten on the next generation run. To change a capability's declared
> metadata, edit the corresponding `capabilities/<id>/capability.json` and run
> `node scripts/gen-capability-matrix.cjs --write`.
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
[Capability manifest fields](#manifest-field-reference) —
[The capability trust model](../explanation/capability-trust-model.md)
---
## Column definitions
| Column | Description |
|---|---|
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: `gsd-`, `gsd-core-`, `anthropic-`. |
| **role** | `feature` — extends what the loop does; `runtime` — adapts GSD to a specific AI runtime/IDE. |
| **tier** | `core` — always active; `standard` — active when the runtime supports it; `full` — opt-in or runtime-specific. |
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. `—` means the capability declares no range. |
| **extension points** | The loop points this capability registers hooks into (from the registry's `byLoopPoint` index). `—` means it registers none (typical for runtime capabilities, whose job is surface emission). |
| **hook kinds** | Which of `step`, `contribution`, `gate` the capability's hooks use. `—` means none. |
| **source** | `first-party` — ships with GSD Core; `third-party` — installed from an external source via `gsd capability install`. |
> **On versions.** This matrix intentionally omits a per-capability `version`
> column. First-party capabilities are versioned **in lockstep** with the GSD
> Core package (their `capability.json` `version` always equals the GSD release
> version), so a per-row version would simply repeat the package version and
> churn the committed file on every release. The stable host-compatibility
> signal — `engines.gsd` — is shown instead. A third-party capability's exact
> version is recorded in the per-runtime ledger (`.gsd-capabilities.json`) at
> install time.
---
## Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD
Core package and are stamped with the package version at release (per
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
to third-party capabilities.
### Feature capabilities (role: feature) — 20
Feature capabilities extend what the loop does — contributing research,
planning, execution, verification, or ship artefacts at the loop extension
points.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `ai-integration` | feature | full | `>=1.6.0` | `plan:pre`, `verify:pre` | step, contribution, gate | first-party |
| `assumption-delta` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
| `audit` | feature | full | `>=1.6.0` | — | — | first-party |
| `broken-windows` | feature | full | `>=1.7.0` | `ship:pre` | gate | first-party |
| `claude-orchestration` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:pre` | contribution | first-party |
| `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party |
| `drift` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post` | gate | first-party |
| `external-job` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party |
| `gap-analysis` | feature | standard | `>=1.6.0` | `plan:post` | gate | first-party |
| `graphify` | feature | full | `>=1.6.0` | — | — | first-party |
| `intel` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `mempalace` | feature | full | `>=1.6.0` | `discuss:pre`, `discuss:post`, `plan:pre`, `plan:post`, `execute:wave:post`, `verify:post`, `ship:post` | step, contribution | first-party |
| `nyquist` | feature | full | `>=1.6.0` | `verify:post` | step | first-party |
| `pattern-mapper` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `profile-pipeline` | feature | full | `>=1.6.0` | — | — | first-party |
| `research` | feature | standard | `>=1.6.0` | `plan:pre` | step | first-party |
| `schema-gate` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
| `security` | feature | full | `>=1.6.0` | `plan:pre`, `verify:post`, `ship:pre` | step, contribution, gate | first-party |
| `tdd` | feature | full | `>=1.6.0` | `plan:pre`, `execute:post` | contribution, gate | first-party |
| `ui` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post`, `verify:post` | step, gate | first-party |
### Runtime capabilities (role: runtime) — 19
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are `—`.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `antigravity` | runtime | core | `>=1.6.0` | — | — | first-party |
| `augment` | runtime | core | `>=1.6.0` | — | — | first-party |
| `claude` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cline` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codebuddy` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codex` | runtime | core | `>=1.6.0` | — | — | first-party |
| `copilot` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cursor` | runtime | core | `>=1.6.0` | — | — | first-party |
| `hermes` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kilo` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kimi` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kimi-code` | runtime | core | `>=1.7.0` | — | — | first-party |
| `opencode` | runtime | core | `>=1.6.0` | — | — | first-party |
| `pi` | runtime | core | `>=1.7.0` | — | — | first-party |
| `qwen` | runtime | core | `>=1.6.0` | — | — | first-party |
| `trae` | runtime | core | `>=1.6.0` | — | — | first-party |
| `vscode` | runtime | core | `>=1.7.0` | — | — | first-party |
| `windsurf` | runtime | core | `>=1.6.0` | — | — | first-party |
| `zcode` | runtime | core | `>=1.6.0` | — | — | first-party |
### Reviewer capabilities (role: reviewer) — 5
Reviewer capabilities declare a cross-AI **reviewer lane** — one external CLI or
model endpoint `/gsd:review` hands a plan to (ADR-2782 D3). They are not install
targets: they emit no skills, agents, hooks or surface files, so their
extension-point and hook-kind cells are `—`. A host that is *also* a reviewer
(Claude, Codex, Cursor, OpenCode, Qwen, Antigravity) keeps one manifest and
appears under **runtime** above, carrying its lane alongside its runtime body;
only lanes that GSD never installs into appear here.
Because a lane receives the plan text, requirements, research findings and
`CONTEXT.md` decisions, it is a disclosed executable surface and is consent-gated
at install like any other — see
[the trust model](../explanation/capability-trust-model.md).
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `coderabbit` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `gemini` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `llama-cpp` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `lm-studio` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `ollama` | reviewer | full | `>=1.8.0` | — | — | first-party |
---
## Third-party capabilities
This matrix is the **first-party catalogue**: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via `gsd capability install <spec>` it enters the **runtime
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is `gsd capability list` (see the
[`gsd capability` command reference](gsd-capability-command.md)), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with `source` = `third-party`.
### Column values for third-party rows
| Column | Value |
|---|---|
| **id** | As declared in `capability.json`. Must not use reserved prefixes (`gsd-`, `gsd-core-`, `anthropic-`). |
| **role** | `feature` or `runtime`, as declared. |
| **tier** | `core`, `standard`, or `full`, as declared. |
| **engines.gsd** | Range from `capability.json`; verified at install and at each load. |
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
| **hook kinds** | `step`, `contribution`, and/or `gate` as declared. Disclosed in the consent summary at install. |
| **source** | `third-party` |
### Community registry
Whether GSD operates or advertises a central community registry of third-party
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
ship regardless of that decision; URL/git/npm/tarball import does not depend on
a central registry.
---
## Manifest field reference
The fields below are defined in `capability.json` and govern how a capability
appears in this matrix. For the full schema, see
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
and the [capability manifest reference](capability-manifest.md).
| Field | Required | Type | Purpose |
|---|---|---|---|
| `version` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
| `engines.gsd` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
| `compatVersions` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
| `integrity` | No | `sha512-<base64>` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
| `provenance` | No | `{ sourceRepo, commit }` | Source provenance; populated in CI for first-party/curated capabilities. |
---
## Related documents
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
- [Capability manifest reference](capability-manifest.md) — the full `capability.json` schema
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)