Files
msd-core/docs/reference/capability-matrix.md
Tom Boucher 7372d99a26 enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales

The reviewer lane roster was hand-enumerated across five documentation
surfaces and three workflow files that had drifted apart: --kimi-code was
missing from all four translated COMMANDS.md mirrors, --coderabbit from
every workflow forwarding list, and --antigravity from FEATURES.md.

Adds checkReviewerDocsParity, a second pure gate deliberately separate from
checkReviewerLaneParity so a stale doc cannot make the runtime checker look
red. Workflows now derive their flag lists from a new review-lane flags
query instead of hand-enumerating them, which also retires the unanchored
grep that matched --agy inside --antigravity.

Documents the previously absent reviewer body and hostBehaviors field in
the capability manifest reference.

Closes #2800
Closes #2781
Closes #2272

* fix(#2800): key the docs parity table arm on first-cell position

Review found the flag arm was file-scoped, so the forwarding row that lists
every flag in its third cell satisfied it on its own. Deleting a lane's own
reviewer-table row -- the #2781 regression this gate exists to prevent --
therefore passed undetected.

Arm 4 keys on the FIRST table cell, which separates a lane row from the
forwarding row structurally and in every locale. Regression test included.

* fix(#2800): shape-filter the flags subcommand output

All three consumers read review-lane flags through an unquoted command
substitution so the output word-splits into loop items. Phase 2 admits
third-party overlay lanes, so an overlay flag containing whitespace would
inject a second loop item and one containing a glob would expand against
the cwd. Emit only well-formed flags so neither reaches the shell.

* fix(#2800): remove the regex length ceiling and count only prose mentions

Review found two real defects in the docs parity gate.

The never-throws contract was false: building a RegExp from a declared flag
or section title throws SyntaxError past ~100k chars, and Phase 2 admits
overlay lanes whose declared strings are untrusted in length. Every one of
these matches is literal, so String.includes replaces the regex outright,
which also deletes escapeLiteral and the llama.cpp escaping it existed for.

Arm 1 was context-blind: a flag mentioned only inside a fenced example or a
commented-out row counted as documented. Both are stripped before matching.

Also advertises all 13 lane flags in the argument-hint and corrects a stale
eleven-lane count in the slug grammar note.

* test(#2800): repoint the convergence suite off deleted workflow text

The derived flag loop deleted the literal per-flag grep lines four tests
matched on. Two of those failed loudly. The behavioral and property tests
failed SILENTLY instead: their end marker no longer resolved, so the parse
block extracted empty and both passed vacuously, and the property test's
gsd_run stub had a no-op default that hid it.

All now share one extractor and execute the real deployed block through a
gsd_run shim backed by the actual binary. The whitelist assertions become an
anti-parity check: re-adding a hand-written flag list must fail.

Also repairs two vacuous cases in the docs parity suite. The unreadable-doc
test called its own mock rather than the reader, and the integration test
bounded nothing, so a doc losing its marker would have been silently skipped
and still passed green.

* fix(#2800): run the derived flag loop after the launcher preamble

The remote matrix caught a real runtime bug, not a test artifact. In
autonomous.md and plan-review-convergence.md the launcher preamble that
defines gsd_run lives in a separate, LATER bash fence than the derived loop.
Each fence is its own shell, so gsd_run was undefined where the loop ran:
the command substitution yielded nothing and zero reviewer flags would have
been forwarded. Worse than the drift this epic fixes, and silent.

The whole CONVERGENCE_ARGS construction moves as one unit, because the
--max-cycles append sits between the loop and the preamble and would
otherwise have run against an uninitialized variable and then been dropped
by the relocated initializer.

Also documents all 13 lane flags in help/modes/full.md, which the repo gates
bidirectionally against each command's argument-hint.

* test(#2800): repoint the two converge suites off deleted flag literals

Both asserted workflow.includes('--codex') against the hand-enumerated list
the derived loop removed. They now assert the derivation itself, keep --all
and --text (convergence controls, still literal), and add an anti-parity
guard so re-adding a hardcoded list fails.

The lost pass-through proof is replaced with a real one: every flag the
tests used to hardcode is asserted present in the actual roster emitted by
the binary, which is the property the old assertion was protecting.

* test(#2800): acknowledge the workflow byte growth from the derived flag loop

* chore(#2800): backfill changeset pr number to 2882

* fix(#2800): strip HTML comments to a fixed point in the parity gate

CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the
single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick
construction smuggles a commented-out row past the gate and it counts as
documented. Not an injection risk here since nothing is rendered, but it is
the exact false pass this helper exists to prevent.

Strips to a fixed point, then treats any surviving opener as unterminated so
the multi-line branch closes it on a later line. Terminates because every
pass strictly shortens the string.

* test(#2800): pin the comment-smuggling regression with a real reproducer

The obvious fixture for this class does not reproduce it: <!--<!---->-->
leaves a dangling --> rather than a live <!--, and is caught either way, so
it would have passed with and without the fix. The join-trick construction
(<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely
regresses on the single-pass strip and is what the test now uses.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:14:13 -04:00

187 lines
11 KiB
Markdown

# Capability matrix reference
> **Generated file — do not edit by hand.**
> This matrix is generated from the capability registry by
> `scripts/gen-capability-matrix.cjs` and kept honest by a drift guard
> (`tests/capability-matrix-sync.test.cjs` runs `--check`). Any manual edit is
> overwritten on the next generation run. To change a capability's declared
> metadata, edit the corresponding `capabilities/<id>/capability.json` and run
> `node scripts/gen-capability-matrix.cjs --write`.
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
[Capability manifest fields](#manifest-field-reference) —
[The capability trust model](../explanation/capability-trust-model.md)
---
## Column definitions
| Column | Description |
|---|---|
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: `gsd-`, `gsd-core-`, `anthropic-`. |
| **role** | `feature` — extends what the loop does; `runtime` — adapts GSD to a specific AI runtime/IDE; `reviewer` — declares a cross-AI reviewer lane (ADR-2782). A capability may be both a runtime and a reviewer. |
| **tier** | `core` — always active; `standard` — active when the runtime supports it; `full` — opt-in or runtime-specific. |
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. `—` means the capability declares no range. |
| **extension points** | The loop points this capability registers hooks into (from the registry's `byLoopPoint` index). `—` means it registers none (typical for runtime capabilities, whose job is surface emission). |
| **hook kinds** | Which of `step`, `contribution`, `gate` the capability's hooks use. `—` means none. |
| **source** | `first-party` — ships with GSD Core; `third-party` — installed from an external source via `gsd capability install`. |
> **On versions.** This matrix intentionally omits a per-capability `version`
> column. First-party capabilities are versioned **in lockstep** with the GSD
> Core package (their `capability.json` `version` always equals the GSD release
> version), so a per-row version would simply repeat the package version and
> churn the committed file on every release. The stable host-compatibility
> signal — `engines.gsd` — is shown instead. A third-party capability's exact
> version is recorded in the per-runtime ledger (`.gsd-capabilities.json`) at
> install time.
---
## Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD
Core package and are stamped with the package version at release (per
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
to third-party capabilities.
### Feature capabilities (role: feature) — 20
Feature capabilities extend what the loop does — contributing research,
planning, execution, verification, or ship artefacts at the loop extension
points.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `ai-integration` | feature | full | `>=1.6.0` | `plan:pre`, `verify:pre` | step, contribution, gate | first-party |
| `assumption-delta` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
| `audit` | feature | full | `>=1.6.0` | — | — | first-party |
| `broken-windows` | feature | full | `>=1.7.0` | `ship:pre` | gate | first-party |
| `claude-orchestration` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:pre` | contribution | first-party |
| `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party |
| `drift` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post` | gate | first-party |
| `external-job` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party |
| `gap-analysis` | feature | standard | `>=1.6.0` | `plan:post` | gate | first-party |
| `graphify` | feature | full | `>=1.6.0` | — | — | first-party |
| `intel` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `mempalace` | feature | full | `>=1.6.0` | `discuss:pre`, `discuss:post`, `plan:pre`, `plan:post`, `execute:wave:post`, `verify:post`, `ship:post` | step, contribution | first-party |
| `nyquist` | feature | full | `>=1.6.0` | `verify:post` | step | first-party |
| `pattern-mapper` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `profile-pipeline` | feature | full | `>=1.6.0` | — | — | first-party |
| `research` | feature | standard | `>=1.6.0` | `plan:pre` | step | first-party |
| `schema-gate` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
| `security` | feature | full | `>=1.6.0` | `plan:pre`, `verify:post`, `ship:pre` | step, contribution, gate | first-party |
| `tdd` | feature | full | `>=1.6.0` | `plan:pre`, `execute:post` | contribution, gate | first-party |
| `ui` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post`, `verify:post` | step, gate | first-party |
### Runtime capabilities (role: runtime) — 19
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are `—`.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `antigravity` | runtime | core | `>=1.6.0` | — | — | first-party |
| `augment` | runtime | core | `>=1.6.0` | — | — | first-party |
| `claude` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cline` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codebuddy` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codex` | runtime | core | `>=1.6.0` | — | — | first-party |
| `copilot` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cursor` | runtime | core | `>=1.6.0` | — | — | first-party |
| `hermes` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kilo` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kimi` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kimi-code` | runtime | core | `>=1.7.0` | — | — | first-party |
| `opencode` | runtime | core | `>=1.6.0` | — | — | first-party |
| `pi` | runtime | core | `>=1.7.0` | — | — | first-party |
| `qwen` | runtime | core | `>=1.6.0` | — | — | first-party |
| `trae` | runtime | core | `>=1.6.0` | — | — | first-party |
| `vscode` | runtime | core | `>=1.7.0` | — | — | first-party |
| `windsurf` | runtime | core | `>=1.6.0` | — | — | first-party |
| `zcode` | runtime | core | `>=1.6.0` | — | — | first-party |
### Reviewer capabilities (role: reviewer) — 5
Reviewer capabilities declare a cross-AI **reviewer lane** — one external CLI or
model endpoint `/gsd:review` hands a plan to (ADR-2782 D3). They are not install
targets: they emit no skills, agents, hooks or surface files, so their
extension-point and hook-kind cells are `—`. A host that is *also* a reviewer
(Claude, Codex, Cursor, OpenCode, Qwen, Antigravity) keeps one manifest and
appears under **runtime** above, carrying its lane alongside its runtime body;
only lanes that GSD never installs into appear here.
Because a lane receives the plan text, requirements, research findings and
`CONTEXT.md` decisions, it is a disclosed executable surface and is consent-gated
at install like any other — see
[the trust model](../explanation/capability-trust-model.md).
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `coderabbit` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `gemini` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `llama-cpp` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `lm-studio` | reviewer | full | `>=1.8.0` | — | — | first-party |
| `ollama` | reviewer | full | `>=1.8.0` | — | — | first-party |
---
## Third-party capabilities
This matrix is the **first-party catalogue**: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via `gsd capability install <spec>` it enters the **runtime
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is `gsd capability list` (see the
[`gsd capability` command reference](gsd-capability-command.md)), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with `source` = `third-party`.
### Column values for third-party rows
| Column | Value |
|---|---|
| **id** | As declared in `capability.json`. Must not use reserved prefixes (`gsd-`, `gsd-core-`, `anthropic-`). |
| **role** | `feature`, `runtime`, or `reviewer`, as declared. |
| **tier** | `core`, `standard`, or `full`, as declared. |
| **engines.gsd** | Range from `capability.json`; verified at install and at each load. |
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
| **hook kinds** | `step`, `contribution`, and/or `gate` as declared. Disclosed in the consent summary at install. |
| **source** | `third-party` |
### Community registry
Whether GSD operates or advertises a central community registry of third-party
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
ship regardless of that decision; URL/git/npm/tarball import does not depend on
a central registry.
---
## Manifest field reference
The fields below are defined in `capability.json` and govern how a capability
appears in this matrix. For the full schema, see
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
and the [capability manifest reference](capability-manifest.md).
| Field | Required | Type | Purpose |
|---|---|---|---|
| `version` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
| `engines.gsd` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
| `compatVersions` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
| `integrity` | No | `sha512-<base64>` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
| `provenance` | No | `{ sourceRepo, commit }` | Source provenance; populated in CI for first-party/curated capabilities. |
---
## Related documents
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
- [Capability manifest reference](capability-manifest.md) — the full `capability.json` schema
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)