* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com>
11 KiB
Capability matrix reference
Generated file — do not edit by hand. This matrix is generated from the capability registry by
scripts/gen-capability-matrix.cjsand kept honest by a drift guard (tests/capability-matrix-sync.test.cjsruns--check). Any manual edit is overwritten on the next generation run. To change a capability's declared metadata, edit the correspondingcapabilities/<id>/capability.jsonand runnode scripts/gen-capability-matrix.cjs --write.
See also: ADR-1244 — Capability manifest fields — The capability trust model
Column definitions
| Column | Description |
|---|---|
| id | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: gsd-, gsd-core-, anthropic-. |
| role | feature — extends what the loop does; runtime — adapts GSD to a specific AI runtime/IDE; reviewer — declares a cross-AI reviewer lane (ADR-2782). A capability may be both a runtime and a reviewer. |
| tier | core — always active; standard — active when the runtime supports it; full — opt-in or runtime-specific. |
| engines.gsd | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. — means the capability declares no range. |
| extension points | The loop points this capability registers hooks into (from the registry's byLoopPoint index). — means it registers none (typical for runtime capabilities, whose job is surface emission). |
| hook kinds | Which of step, contribution, gate the capability's hooks use. — means none. |
| source | first-party — ships with GSD Core; third-party — installed from an external source via gsd capability install. |
On versions. This matrix intentionally omits a per-capability
versioncolumn. First-party capabilities are versioned in lockstep with the GSD Core package (theircapability.jsonversionalways equals the GSD release version), so a per-row version would simply repeat the package version and churn the committed file on every release. The stable host-compatibility signal —engines.gsd— is shown instead. A third-party capability's exact version is recorded in the per-runtime ledger (.gsd-capabilities.json) at install time.
Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD Core package and are stamped with the package version at release (per ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied to third-party capabilities.
Feature capabilities (role: feature) — 20
Feature capabilities extend what the loop does — contributing research, planning, execution, verification, or ship artefacts at the loop extension points.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
ai-integration |
feature | full | >=1.6.0 |
plan:pre, verify:pre |
step, contribution, gate | first-party |
assumption-delta |
feature | full | >=1.6.0 |
plan:pre |
contribution | first-party |
audit |
feature | full | >=1.6.0 |
— | — | first-party |
broken-windows |
feature | full | >=1.7.0 |
ship:pre |
gate | first-party |
claude-orchestration |
feature | full | >=1.7.0 |
plan:post, execute:wave:pre |
contribution | first-party |
code-review |
feature | full | >=1.6.0 |
execute:post |
step | first-party |
drift |
feature | full | >=1.6.0 |
plan:pre, execute:wave:post |
gate | first-party |
external-job |
feature | full | >=1.7.0 |
plan:post, execute:wave:post |
contribution | first-party |
gap-analysis |
feature | standard | >=1.6.0 |
plan:post |
gate | first-party |
graphify |
feature | full | >=1.6.0 |
— | — | first-party |
intel |
feature | full | >=1.6.0 |
plan:pre |
step | first-party |
mempalace |
feature | full | >=1.6.0 |
discuss:pre, discuss:post, plan:pre, plan:post, execute:wave:post, verify:post, ship:post |
step, contribution | first-party |
nyquist |
feature | full | >=1.6.0 |
verify:post |
step | first-party |
pattern-mapper |
feature | full | >=1.6.0 |
plan:pre |
step | first-party |
profile-pipeline |
feature | full | >=1.6.0 |
— | — | first-party |
research |
feature | standard | >=1.6.0 |
plan:pre |
step | first-party |
schema-gate |
feature | full | >=1.6.0 |
plan:pre |
contribution | first-party |
security |
feature | full | >=1.6.0 |
plan:pre, verify:post, ship:pre |
step, contribution, gate | first-party |
tdd |
feature | full | >=1.6.0 |
plan:pre, execute:post |
contribution, gate | first-party |
ui |
feature | full | >=1.6.0 |
plan:pre, execute:wave:post, verify:post |
step, gate | first-party |
Runtime capabilities (role: runtime) — 19
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are —.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
antigravity |
runtime | core | >=1.6.0 |
— | — | first-party |
augment |
runtime | core | >=1.6.0 |
— | — | first-party |
claude |
runtime | core | >=1.6.0 |
— | — | first-party |
cline |
runtime | core | >=1.6.0 |
— | — | first-party |
codebuddy |
runtime | core | >=1.6.0 |
— | — | first-party |
codex |
runtime | core | >=1.6.0 |
— | — | first-party |
copilot |
runtime | core | >=1.6.0 |
— | — | first-party |
cursor |
runtime | core | >=1.6.0 |
— | — | first-party |
hermes |
runtime | core | >=1.6.0 |
— | — | first-party |
kilo |
runtime | core | >=1.6.0 |
— | — | first-party |
kimi |
runtime | core | >=1.6.0 |
— | — | first-party |
kimi-code |
runtime | core | >=1.7.0 |
— | — | first-party |
opencode |
runtime | core | >=1.6.0 |
— | — | first-party |
pi |
runtime | core | >=1.7.0 |
— | — | first-party |
qwen |
runtime | core | >=1.6.0 |
— | — | first-party |
trae |
runtime | core | >=1.6.0 |
— | — | first-party |
vscode |
runtime | core | >=1.7.0 |
— | — | first-party |
windsurf |
runtime | core | >=1.6.0 |
— | — | first-party |
zcode |
runtime | core | >=1.6.0 |
— | — | first-party |
Reviewer capabilities (role: reviewer) — 5
Reviewer capabilities declare a cross-AI reviewer lane — one external CLI or
model endpoint /gsd:review hands a plan to (ADR-2782 D3). They are not install
targets: they emit no skills, agents, hooks or surface files, so their
extension-point and hook-kind cells are —. A host that is also a reviewer
(Claude, Codex, Cursor, OpenCode, Qwen, Antigravity) keeps one manifest and
appears under runtime above, carrying its lane alongside its runtime body;
only lanes that GSD never installs into appear here.
Because a lane receives the plan text, requirements, research findings and
CONTEXT.md decisions, it is a disclosed executable surface and is consent-gated
at install like any other — see
the trust model.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
coderabbit |
reviewer | full | >=1.8.0 |
— | — | first-party |
gemini |
reviewer | full | >=1.8.0 |
— | — | first-party |
llama-cpp |
reviewer | full | >=1.8.0 |
— | — | first-party |
lm-studio |
reviewer | full | >=1.8.0 |
— | — | first-party |
ollama |
reviewer | full | >=1.8.0 |
— | — | first-party |
Third-party capabilities
This matrix is the first-party catalogue: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via gsd capability install <spec> it enters the runtime
registry overlay (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is gsd capability list (see the
gsd capability command reference), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with source = third-party.
Column values for third-party rows
| Column | Value |
|---|---|
| id | As declared in capability.json. Must not use reserved prefixes (gsd-, gsd-core-, anthropic-). |
| role | feature, runtime, or reviewer, as declared. |
| tier | core, standard, or full, as declared. |
| engines.gsd | Range from capability.json; verified at install and at each load. |
| extension points | The loop points the capability registers into, validated against the known 12 identifiers. |
| hook kinds | step, contribution, and/or gate as declared. Disclosed in the consent summary at install. |
| source | third-party |
Community registry
Whether GSD operates or advertises a central community registry of third-party capabilities is TBD/TBA (PRD). The matrix mechanic and all manifest fields ship regardless of that decision; URL/git/npm/tarball import does not depend on a central registry.
Manifest field reference
The fields below are defined in capability.json and govern how a capability
appears in this matrix. For the full schema, see
ADR-1244 D1
and the capability manifest reference.
| Field | Required | Type | Purpose |
|---|---|---|---|
version |
Yes | semver string | Capability version. The registry rejects manifests without it. |
engines.gsd |
Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
compatVersions |
No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
integrity |
No | sha512-<base64> |
SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
provenance |
No | { sourceRepo, commit } |
Source provenance; populated in CI for first-party/curated capabilities. |
Related documents
- ADR-1244 — Capability Ecosystem
- The capability trust model — why the trust rules are structured as they are
- The phase loop — the 12 loop extension points in context
- Capability manifest reference — the full
capability.jsonschema - ADR-857 — the original capability architecture (D7/D8 extended by ADR-1244)