feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)

* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-06-19 12:45:26 -04:00
committed by GitHub
parent 34bc096ec2
commit 0d56f544d2
16 changed files with 507 additions and 231 deletions

View File

@@ -0,0 +1,5 @@
---
type: Added
pr: 1458
---
**Capability matrix reference** — a generated catalogue (`docs/reference/capability-matrix.md`) of every first-party capability's role, tier, extension points, hook kinds, and `engines.gsd`, generated from the committed registry and kept honest by a CI drift guard so it can never fall out of sync with the actual capability set. (#1458)

View File

@@ -188,13 +188,13 @@ ADR-1244 D3 fetch-and-stage seam (`gsd-core/bin/lib/capability-source.cjs`). Pri
ADR-1244 D4 per-runtime install manifest (`gsd-core/bin/lib/capability-ledger.cjs`). Leaf module (only `node:fs`/`node:path` plus `shell-command-projection`'s `platformWriteSync`). Records `{ id, version, source, integrity, files[], sharedEdits[{file,marker}] }` per installed capability in `.gsd-capabilities.json` at the runtime config dir root. Exports: `readLedger` (structural-validated, never throws), `writeLedger` (atomic via `platformWriteSync`), `recordInstall` (idempotent, prototype-pollution-guarded), `removeEntry`, and `reconcile` (reports orphans whose `files[]` are missing on disk; hardened against non-string/`..` members; never mutates). Serves as the atomic commit point for Phase-4 upgrade/remove and the reconciliation basis for detecting stale entries after out-of-band deletions.
### Capability Trust Gate
ADR-1244 Phase 4 (D5) PURE policy module (`gsd-core/bin/lib/capability-trust.cjs`). Computes *what* a capability would do and *whether* policy permits it; performs no mutation and no I/O beyond existence-checking declared artifacts. Exports: `discloseExecutableSurfaces(manifest, stagedDir?)` (enumerates the three executable surfaces — `hooks`, command modules, `mcpServers` — and flags `hasExecutable`); `evaluateInstallTrust(args)` (composes source policy + reserved-namespace + engines gate + disclosure into `{ allowed, requiresConsent, disclosure, engines, blockReasons }`); `evaluateSourceAllowed(parsed, strictKnownRegistries)` enforcing `capabilities.strict_known_registries` (unset/null → permissive-with-consent; `[]` → block all external; non-empty → host-based allowlist, never substring); `checkEngines(manifest, hostVersion)` (engines.gsd hard gate via `semverSatisfies` + `compatVersions` graceful-downgrade picking the newest working version); `executableSetChanged(old, new)` (auto-update re-consent trigger); `checkReservedNamespace` (`gsd-`/`gsd-core-`/`anthropic-`). The barrier is consent + integrity + reversibility, NOT a sandbox — see `docs/explanation/the-capability-trust-model.md`.
ADR-1244 Phase 4 (D5) PURE policy module (`gsd-core/bin/lib/capability-trust.cjs`). Computes *what* a capability would do and *whether* policy permits it; performs no mutation and no I/O beyond existence-checking declared artifacts. Exports: `discloseExecutableSurfaces(manifest, stagedDir?)` (enumerates the three executable surfaces — `hooks`, command modules, `mcpServers` — and flags `hasExecutable`); `evaluateInstallTrust(args)` (composes source policy + reserved-namespace + engines gate + disclosure into `{ allowed, requiresConsent, disclosure, engines, blockReasons }`); `evaluateSourceAllowed(parsed, strictKnownRegistries)` enforcing `capabilities.strict_known_registries` (unset/null → permissive-with-consent; `[]` → block all external; non-empty → host-based allowlist, never substring); `checkEngines(manifest, hostVersion)` (engines.gsd hard gate via `semverSatisfies` + `compatVersions` graceful-downgrade picking the newest working version); `executableSetChanged(old, new)` (auto-update re-consent trigger); `checkReservedNamespace` (`gsd-`/`gsd-core-`/`anthropic-`). The barrier is consent + integrity + reversibility, NOT a sandbox — see `docs/explanation/capability-trust-model.md`.
### Capability Lifecycle
ADR-1244 Phase 4 (D5+D6) orchestration seam (`gsd-core/bin/lib/capability-lifecycle.cjs`) composing the source resolver, ledger, and trust gate into the mutating operations. Exports: `installCapability` (pre-fetch source gate → resolve copy-only with `promote:false` → trust verdict → promote + apply marker-stamped shared edits → **ledger commit**; nothing written on block/abort), `upgradeCapability` (atomic stage-then-swap: old set aside, new swapped in, shared edits re-derived, **ledger committed**, backup dropped; re-prompts when the executable set changed), `removeCapability` (strip only `_gsdCapability`-marked shared-config entries — user hand-edits preserved — delete exactly the ledger-recorded files, then drop the entry; `CAPABILITY_DATA` preserved unless `removeData`), `reconcileCapabilities` (crash recovery driven by the ledger's `_pending {kind,backupName,sharedFiles}` INTENT — not a version comparison: roll an uncommitted upgrade back by restoring the backup, an uncommitted fresh install away entirely, and re-sync shared config from the winning bundle, guaranteeing no half-state), plus `applyCapabilitySharedEdits`/`stripCapabilitySharedEdits` (marker-isolated JSON edits, prototype-pollution-guarded). All four mutating ops + reconcile take a cross-process lock (`.gsd/capabilities/.lock`, atomic stale-steal) so a concurrent reconcile can't clear a live intent. Capability code never executes during any operation. The source resolver's `promote:false`/`skipEnginesGate` options are the seams that let this module own the swap/commit ordering and the engines gate (with `compatVersions` downgrade hint).
### Capability Command Dispatch
ADR-1244 Phase 5 (D7) registry-driven dispatch of capability command families. First-party families (`graphify`/`intel`/`audit`, shipped in `bin/lib/`) dispatch via `dispatchCapabilityCommand` (`gsd-core/bin/gsd-tools.cjs`) against the FROZEN `capability-registry.cjs` `commandFamilies` (confined to `bin/lib/`) — unchanged. Third-party (installed overlay) families dispatch via `dispatchOverlayCapabilityCommand`: after the first-party path returns false, it calls `loadRegistry({ includeInstalled, cwd })` and dispatches a family iff its `capId` is in `_overlay.commandRoots` — which `capability-loader.cjs` populates ONLY for accepted overlay capabilities that (a) declare `commands` AND (b) have a **committed** ledger entry (present, non-`_pending`); a bundle dropped on disk with no ledger entry is NOT command-dispatchable (consent gate). The router module is `require()`'d FROM the capability's install root via `defaultRequireFromInstallRoot` (bare-`.cjs` basename + `realpath` containment, rejecting `..` traversal and symlink escape); same own-property/function/sync-only guards as the first-party path. Wired into the `runCommand` default arm before "Unknown command". Project-scope ledgers live in the repo tree and are only as trustworthy as the repo — see `docs/explanation/the-capability-trust-model.md` "project-scope trust boundary".
ADR-1244 Phase 5 (D7) registry-driven dispatch of capability command families. First-party families (`graphify`/`intel`/`audit`, shipped in `bin/lib/`) dispatch via `dispatchCapabilityCommand` (`gsd-core/bin/gsd-tools.cjs`) against the FROZEN `capability-registry.cjs` `commandFamilies` (confined to `bin/lib/`) — unchanged. Third-party (installed overlay) families dispatch via `dispatchOverlayCapabilityCommand`: after the first-party path returns false, it calls `loadRegistry({ includeInstalled, cwd })` and dispatches a family iff its `capId` is in `_overlay.commandRoots` — which `capability-loader.cjs` populates ONLY for accepted overlay capabilities that (a) declare `commands` AND (b) have a **committed** ledger entry (present, non-`_pending`); a bundle dropped on disk with no ledger entry is NOT command-dispatchable (consent gate). The router module is `require()`'d FROM the capability's install root via `defaultRequireFromInstallRoot` (bare-`.cjs` basename + `realpath` containment, rejecting `..` traversal and symlink escape); same own-property/function/sync-only guards as the first-party path. Wired into the `runCommand` default arm before "Unknown command". Project-scope ledgers live in the repo tree and are only as trustworthy as the repo — see `docs/explanation/capability-trust-model.md` "project-scope trust boundary".
### Loop Extension Point
A named, stable site on a host loop step (per-step `pre`/`post` plus per-wave in Execute; 12 total) where Capabilities register hooks. Three hook kinds: `step` (runs as its own sequenced unit), `contribution` (injects into the core step's prompt/context), and `gate` (checks and optionally blocks via a declared `blocking` flag). Each hook declares the artifacts it produces and consumes; hook order is derived by topological sort of that produces/consumes graph (capability-id tiebreak), which also defines data flow — file-artifact based, surviving `/clear` and fresh executor contexts. Hooks are surfaced by runtime resolution with concrete projection: the workflow calls a query that resolves the active hooks and returns fully-rendered, ordered markdown for the executor. Failure is default-resilient — a non-gate hook that errors is skipped with a warning; a hook may opt into `onError: halt`. Part of the Capability system. ADR-857 phase 3c ships the registry-consuming query layer: `gsd-core/bin/lib/loop-resolver.cjs` exposes `resolveLoopHooks({ point, registry, config })` (pure, no I/O), `renderLoopHooks(resolved)` (pure markdown renderer), and `cmdLoopRenderHooks(cwd, point, raw, opts)` (I/O entry point); activated via `gsd-tools loop render-hooks <point>` which emits `{ point, activeHooks[], rendered }`. Activation is driven by `when` (dotted config key resolved against `loadConfig`), with inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard. The first phase-6 cutovers wiring workflows to this query have landed — ui-phase at `plan:pre` and ui-review at `verify:post` (in `plan-phase.md`/`autonomous.md`); further per-feature cutovers are ongoing.

View File

@@ -311,7 +311,7 @@ CJS command family routers dispatch through `CommandRoutingHub`. The hub owns th
Command families declared by capabilities (`commands: [{ family, module, router }]`) are dispatched from the registry rather than a hardcoded switch. The `runCommand` default arm tries, in order:
1. **First-party** — `dispatchCapabilityCommand` against the frozen `capability-registry.cjs` `commandFamilies`, loading the router from `bin/lib/`. The in-tree families (`graphify`, `intel`, `audit`) reach their routers this way (the legacy hardcoded switch is retired).
2. **Third-party (installed overlay)** — `dispatchOverlayCapabilityCommand` calls `loadRegistry({ includeInstalled })` and dispatches a family only when its `capId` appears in `_overlay.commandRoots`. The loader lists a command root **only** for an accepted overlay capability with a **committed** ledger entry (consent gate), and the router module is `require()`'d **from that capability's install root**, confined by basename validation + `realpath` containment (rejecting `..` traversal and symlink escape). This is the one point where third-party capability code executes; see [the capability trust model](explanation/the-capability-trust-model.md) for the consent + confinement + project-scope trust boundary.
2. **Third-party (installed overlay)** — `dispatchOverlayCapabilityCommand` calls `loadRegistry({ includeInstalled })` and dispatches a family only when its `capId` appears in `_overlay.commandRoots`. The loader lists a command root **only** for an accepted overlay capability with a **committed** ledger entry (consent gate), and the router module is `require()`'d **from that capability's install root**, confined by basename validation + `realpath` containment (rejecting `..` traversal and symlink escape). This is the one point where third-party capability code executes; see [the capability trust model](explanation/capability-trust-model.md) for the consent + confinement + project-scope trust boundary.
Both paths share the same guards: prototype-pollution-safe command keys, an own-property router check, and synchronous-only routers (an async router is a fail-fast error).

View File

@@ -1672,7 +1672,7 @@ The check is also run as part of `npm test` via `tests/enh-2789-description-budg
A capability can ship its own command family by declaring `commands: [{ family, module, router }]` in its `capability.json` (ADR-1244 D7). Once the capability is **installed and consented** (a committed entry exists in the per-runtime `.gsd-capabilities.json` ledger), running `gsd-tools <family> …` (equivalently the `gsd <family>` wrapper) dispatches to the capability's router. The first-party families `graphify`, `intel`, and `audit-uat`/`audit-open` use exactly this registry-driven seam.
Dispatch is gated for safety: the router module is loaded **only from the capability's own install root** (a bare `.cjs` basename, traversal- and symlink-confined), and a capability that is merely present on disk **without** a committed ledger entry is **not** command-dispatchable (its declarative skills/agents/config still load). A project-scoped capability's commands are only as trustworthy as the repository they ship in — see [The capability trust model](explanation/the-capability-trust-model.md).
Dispatch is gated for safety: the router module is loaded **only from the capability's own install root** (a bare `.cjs` basename, traversal- and symlink-confined), and a capability that is merely present on disk **without** a committed ledger entry is **not** command-dispatchable (its declarative skills/agents/config still load). A project-scoped capability's commands are only as trustworthy as the repository they ship in — see [The capability trust model](explanation/capability-trust-model.md).
---

View File

@@ -552,7 +552,7 @@ Setting the parent object (`agent_skills_security`) directly is not supported; u
## Capability Trust (`capabilities.*`)
Policy for installing and updating third-party capabilities (ADR-1244). These keys govern the trust gate; they have no effect if you only ever use the native first-party capabilities shipped with GSD. They are **policy inputs** read by the `gsd capability` command flow, which passes the resulting decision into the capability lifecycle — `strict_known_registries` gates whether a source may be installed at all; `auto_update` is consulted by the `update`/`outdated` flow (which always re-prompts when a new version's executable surface set changes). The full rationale — including why there is no sandbox — is in [The capability trust model](explanation/the-capability-trust-model.md).
Policy for installing and updating third-party capabilities (ADR-1244). These keys govern the trust gate; they have no effect if you only ever use the native first-party capabilities shipped with GSD. They are **policy inputs** read by the `gsd capability` command flow, which passes the resulting decision into the capability lifecycle — `strict_known_registries` gates whether a source may be installed at all; `auto_update` is consulted by the `update`/`outdated` flow (which always re-prompts when a new version's executable surface set changes). The full rationale — including why there is no sandbox — is in [The capability trust model](explanation/capability-trust-model.md).
| Setting | Type | Default | Description |
|---------|------|---------|-------------|

View File

@@ -396,7 +396,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `capability-registry.cjs` | Generated central Capability Registry — role-partitioned index of all co-located capability declarations (`capabilities/<id>/capability.json`); emitted by `scripts/gen-capability-registry.cjs --write` (ADR-894 §5) |
| `capability-source.cjs` | Capability source resolver (ADR-1244 D3) — `resolveCapabilitySource(spec, opts)` fetches and stages a capability from local path, git (https/ssh/git transports only), npm pack (no lifecycle scripts), tarball (sha512 integrity verify before extraction), or registry (stub); tar-slip/symlink rejection; atomic staging to `$GSD_HOME/.gsd/capabilities/<id>/`; no capability code executes during install |
| `capability-state.cjs` | Unified capability-state resolver (ADR-857 phase 4b/6) — composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; exports pure `resolveCapabilityState`, reusable `resolveCapabilityRuntimeState`, and I/O handler `cmdCapabilityState`; command surface: `gsd-tools capability state [--config-dir <path>]` emitting `{ runtimeConfigDir, capabilities[] }` |
| `capability-trust.cjs` | Capability trust gate (ADR-1244 Phase 4, D5 + compatibility half of D6) — PURE policy module: `discloseExecutableSurfaces` (hooks/command modules/mcpServers), `evaluateInstallTrust` (compose source policy + reserved-namespace + engines gate + disclosure → allowed/requiresConsent/blockReasons), `evaluateSourceAllowed` (`strict_known_registries`: permissive/lockdown/host-allowlist), `checkEngines` (engines.gsd hard gate + `compatVersions` graceful-downgrade), `executableSetChanged` (auto-update re-consent trigger); no sandbox — see `docs/explanation/the-capability-trust-model.md` |
| `capability-trust.cjs` | Capability trust gate (ADR-1244 Phase 4, D5 + compatibility half of D6) — PURE policy module: `discloseExecutableSurfaces` (hooks/command modules/mcpServers), `evaluateInstallTrust` (compose source policy + reserved-namespace + engines gate + disclosure → allowed/requiresConsent/blockReasons), `evaluateSourceAllowed` (`strict_known_registries`: permissive/lockdown/host-allowlist), `checkEngines` (engines.gsd hard gate + `compatVersions` graceful-downgrade), `executableSetChanged` (auto-update re-consent trigger); no sandbox — see `docs/explanation/capability-trust-model.md` |
| `capability-validator.cjs` | Shared runtime-callable capability validator (ADR-1244 D2) — extracted from `scripts/gen-capability-registry.cjs` so the build-time generator and the runtime overlay loader share ONE validation implementation (generative-parity guarded); exports `validateCapability`/`validateCrossCapability`/`validateVersionEnvelope`/`validateConsumesGlobal`/… plus the closed-vocabulary sets and `SEMVER_RE` |
| `capability-writer.cjs` | Capability State Writer (ADR-1213) — write-side inverse of the resolver; projects desired per-capability enabled/gates onto surface + config substrates, then re-resolves (assert-and-report); exports `setCapabilityState` and I/O handler `cmdCapabilitySet`; command surface: `gsd-tools capability set <id> [--on\|--off] [--gate <key>=<true\|false>]` |
| `check-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools check` |

View File

@@ -56,6 +56,9 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [PLAN.md schema](reference/plan-md.md) — field-by-field reference for `.planning/phases/<N>/PLAN.md`
- [Planning artifacts](reference/planning-artifacts.md) — all `.planning/` files and their roles
- [Review and verification capabilities](reference/review-verification-capabilities.md) — code review, security, and Nyquist capability ownership and hook contracts
- [Capability matrix](reference/capability-matrix.md) — generated catalogue of every capability's role, tier, extension points, hook kinds, and `engines.gsd`
- [Capability manifest](reference/capability-manifest.md) — the full `capability.json` schema and validation rules
- [`gsd capability` command](reference/gsd-capability-command.md) — install / update / remove / list reference for third-party capabilities
---
@@ -65,7 +68,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [The phase loop](explanation/the-phase-loop.md) — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- [Multi-agent orchestration](explanation/multi-agent-orchestration.md) — how subagents are spawned, scoped, and coordinated
- [Security model](explanation/security-model.md) — trust boundaries, permissions, and safe automation
- [The capability trust model](explanation/the-capability-trust-model.md) — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- [The capability trust model](explanation/capability-trust-model.md) — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- [Architecture](ARCHITECTURE.md) — system architecture, agent model, and data flow
- [Discuss modes](workflow-discuss-mode.md) — assumptions mode vs interview mode for `/gsd-discuss-phase`
- [Context monitoring](context-monitor.md) — context window monitoring hook architecture

View File

@@ -3,8 +3,8 @@
> **Explanation** — This document describes *why* GSD draws its trust
> boundaries where it does, and *what the trade-offs are*. It is not a
> step-by-step guide to installing capabilities; for that, see the how-to
> guides for [importing a capability](../how-to/) and
> [version management](../how-to/). For the decision record, see
> guides for [importing a capability](../how-to/import-a-capability-from-a-url.md) and
> [version management](../how-to/version-a-capability.md). For the decision record, see
> [ADR-1244 D5](../adr/1244-capability-ecosystem.md#d5--trust-model-artifact-parity-is-full-trust-posture-is-tiered).
> For the capability field reference, see the
> [capability matrix](../reference/capability-matrix.md).
@@ -172,6 +172,15 @@ What it does not defend against: a malicious capability where the author
themselves publishes a bad bundle. The SHA is honest about what you are
installing; it says nothing about whether what you are installing is safe.
It also pins **only the top-level bundle**, not an `npm`-sourced capability's
resolved dependency graph. `--ignore-scripts` and copy-only staging stop
install-time execution, but when a command module is later `require()`'d, Node
resolves and runs its transitive dependencies — which the bundle SHA does not
cover (the Wiz / VS Code lesson). For the `npm` source kind, a green integrity
check means "the package tarball is the one you pinned," not "every line of code
that will run is the code you reviewed." Authors who want a stronger guarantee
should vendor their dependencies or ship a lockfile.
### Auto-update off by default, re-consent on executable-set change
When auto-update is enabled for a third-party capability, each update is
@@ -200,13 +209,54 @@ rejected at the conformance gate. This prevents impersonation: a malicious
actor cannot publish a capability called `gsd-security` and exploit a user's
implicit trust in the GSD namespace.
### `strictKnownRegistries` for managed environments
### `capabilities.strict_known_registries` for managed environments
Teams or enterprises that want to constrain which capability sources are
permissible can set `strictKnownRegistries` in managed or project config to an
explicit allowlist of URLs or registry names. Setting it to `[]` blocks all
external installs. This gives an administrator a policy lever that operates
before the user even sees a consent prompt.
permissible set `capabilities.strict_known_registries` in config. Its semantics:
- **unset / `null`** *(default)* — permissive: external installs (git / npm /
tarball) are allowed, each still passing the consent + integrity gate. Local
filesystem installs are always allowed.
- **`[]`** *(explicit empty array)* — lockdown: **all external installs are
blocked**; local-only.
- **non-empty list** — a **host-based** allowlist: only sources whose host
matches an entry (exact host or a subdomain of it — `github.com` matches
`api.github.com` but never `evilgithub.com`; the literal token `npm` permits
the npm source kind). A malformed (non-array) value **fails closed**.
This gives an administrator a policy lever that operates before the user even
sees a consent prompt. The default is permissive-with-consent (not Obsidian-style
restricted-by-default), because the epic deliberately chose decentralised import
with the consent prompt as the default barrier and lockdown one config key away.
### Command dispatch: where third-party code runs (1.6.0)
A capability may declare a **command family** (`commands: [{ family, module,
router }]`); `gsd-tools <family>` dispatches it by `require()`-ing the router.
This is the one place a third-party capability's own code executes, so it is
gated twice. **Consent:** a third-party family is dispatchable only if its
capability has a **committed ledger entry** — i.e. you installed it through the
lifecycle and consented. A bundle merely present on disk with no install record
still contributes its declarative surfaces but is **not** command-dispatchable.
**Confinement:** the router module loads only from the capability's own install
root (bare-`.cjs` basename, `realpath`-confined, rejecting `..` traversal and
symlink escape); a first-party command can never be shadowed by a third-party one.
#### The project-scope trust boundary
Capabilities install **globally** (`$GSD_HOME/.gsd/capabilities/`) or
**project-scoped** (`<projectRoot>/.gsd/capabilities/`). The consent record is
the ledger — and a project-scope ledger lives **inside the repository**. A repo
you check out can therefore ship a capability bundle *and* a ledger that marks
it committed; running `gsd-tools <its-family>` in that repo executes its code.
This is the same boundary that already applies to a repo's build scripts or a
project-scoped capability's hooks, and it is **narrower** — a command never fires
on its own; you have to type it. GSD cannot cryptographically distinguish a
genuine project-local install from a forged project ledger, so: a **global**
install's consent record lives outside any repo and is trustworthy; a
**project-scoped** capability is only as trustworthy as the repository it ships
in. Review repos before running `gsd` commands in them, and prefer global
installs for capabilities you want to trust across projects.
---

View File

@@ -1,126 +0,0 @@
# The capability trust model
> **Status:** Implemented in ADR-1244 Phase 4 (`gsd-core/bin/lib/capability-trust.cjs`,
> `capability-lifecycle.cjs`). This document is the *honest* explanation of why GSD trusts
> third-party capabilities the way it does — including what it does **not** protect against.
GSD capabilities can ship the same executable artifacts GSD itself ships — hook scripts,
command modules, and (for third parties) MCP server declarations. That is full first-party
parity, chosen deliberately (epic #1244, alternative 2). Parity means a third-party
capability is, in the limit, **arbitrary code you have invited into your agent runtime**.
This page explains the barrier GSD puts in front of that invitation, and why the barrier is
*consent + integrity + reversibility* rather than a sandbox.
## Why there is no sandbox (Ask #1)
ADR-857 D8 originally earned "no sandbox" on a *declarative-only* premise: capabilities were
data, not code. This feature reverses that premise, so the justification must be re-derived on
grounds that survive third-party **code execution**. It is:
**GSD has no process boundary to interpose.** A capability's executable surfaces all run inside
the host agent runtime's own trust context:
- **hook scripts** run as the runtime's configured hook commands (your shell, your
permissions);
- **command modules** are `require()`'d directly into the GSD CLI's own Node process;
- **MCP servers** are spawned by the host runtime, not by GSD.
GSD never forks a sandboxed child to run any of this. It *could* not meaningfully do so: the
runtime loads hooks and MCP servers itself, and a command module is in-process by definition.
A `sandboxTier` enum on capabilities that GSD does not actually enforce would be security
theater — worse than nothing, because it implies a boundary that isn't there.
The runtime descriptor's existing `sandboxTier` axis (`none` | `codex-agent-sandbox`,
`runtime-config-adapter-registry.cts`) is the **runtime's own** sandbox over everything it
executes — capability code included. That sandbox keeps applying. But it belongs to the
runtime (e.g. Codex's `workspace-write`), not to GSD, and GSD does not extend or simulate it.
So the barrier GSD owns is the same posture the Obsidian and VS Code marketplaces settled on
after their own supply-chain incidents:
1. **No execution at install.** Staging is copy-only. `npm` sources use `--ignore-scripts`;
there is no `postinstall`. Installing a capability never runs its code.
2. **Integrity before extraction.** Tarball sources verify a `sha512` digest over the fetched
bytes *before* anything is written to the install root; a mismatch aborts.
3. **Disclosure + explicit consent.** Every executable surface a capability declares is
disclosed at install, and installing one that ships *any* executable surface requires
explicit consent. Declining aborts cleanly, writing nothing.
4. **Namespace reservation.** Third parties cannot claim the `gsd-`, `gsd-core-`, or
`anthropic-` id prefixes, so a hostile capability cannot impersonate a first-party one.
5. **Reversibility.** The ledger records exactly what each install wrote; `remove` deletes
exactly those files and surgically strips exactly the shared-config entries the capability
added, touching nothing the user owns.
Consent is meaningful precisely because (1)–(2) guarantee nothing has run *yet* when you are
asked, and (4)–(5) guarantee you can fully undo it. That is the trade GSD makes instead of a
sandbox it cannot honestly provide.
## `strictKnownRegistries` default (Ask #4a)
`capabilities.strict_known_registries` gates **install itself**, not merely auto-update.
| Value | Meaning |
| --- | --- |
| unset / `null` *(default)* | External installs (git / npm / tarball / registry) are allowed, each still passing the consent + integrity gate. Local filesystem installs always allowed. |
| `[]` *(explicit empty array)* | **Blocks all external installs** — managed/enterprise lockdown. Local installs only. |
| `["github.com", "registry.example.com"]` | Allowlist — only sources whose host matches an entry (exact host or a subdomain of it) are permitted. |
The default is **permissive-with-consent**, not Obsidian-style restricted-by-default. The epic
deliberately chose decentralized URL/git distribution over a curated-registry gate (alternative
3), so the consent prompt is the default barrier and full lockdown is one config key away. The
allowlist match is **host-based**, not substring: `github.com` does not match
`evilgithub.com`.
## The npm transitive-dependency boundary (Ask #4b)
`integrity` (sha512) pins **only the top-level fetched artifact** — the tarball, or the output
of `npm pack`. It does **not** pin the resolved dependency graph.
`--ignore-scripts` and copy-only staging stop *install-time* execution (no `postinstall`). But
when a capability's command module is later `require()`'d, Node resolves and executes that
module's **transitive dependency tree**, and sha512 says nothing about those packages — they
can be mutable or compromised independently of the pinned top artifact. This is exactly the
Wiz / VS Code supply-chain lesson.
For the `npm` source kind, therefore: a green integrity check means "the package tarball is the
one you pinned," **not** "every line of code that will run is the code you reviewed." Authors
who want a stronger guarantee should vendor their dependencies into the capability or ship a
lockfile; the consent prompt always discloses when a capability ships command modules, which is
the surface through which transitive code reaches your process.
## Where third-party code runs: command dispatch (Phase 5 / D7)
A capability may declare a **command family** (`commands: [{ family, module, router }]`). When you
run `gsd-tools <family> …`, GSD dispatches it by `require()`-ing the named module's router function.
This is the one place a third-party capability's own code executes in the GSD CLI process, so it is
gated twice:
1. **Consent.** A third-party family is dispatchable only if its capability has a **committed
ledger entry** — i.e. you installed it through the lifecycle (and, for executable surfaces,
consented). The loader records a dispatchable "command root" only for committed (non-`_pending`)
capabilities; a bundle merely *present* on disk with no ledger entry still contributes its
declarative surfaces (skills/agents/config) but is **not** command-dispatchable.
2. **Confinement.** The router module is loaded **from the capability's own install root** — a bare
`.cjs` basename, `realpath`-confined to that directory, rejecting `..` traversal and symlink
escape. A capability can never reach code outside its own bundle, and a first-party command
(`graphify`/`intel`/`audit`, which ship in `bin/lib/`) can never be shadowed by a third-party one.
### The project-scope trust boundary (be honest about it)
Capabilities can be installed **globally** (under your home, `$GSD_HOME/.gsd/capabilities/`) or
**project-scoped** (under a repository, `<projectRoot>/.gsd/capabilities/`). The consent signal is
the ledger, and the project-scope ledger (`<projectRoot>/.gsd-capabilities.json`) lives **inside the
repository**. A repository you check out can therefore ship both a capability bundle *and* a
ledger that marks it "committed." Running `gsd-tools <that-family>` inside such a repo will execute
its code.
This is the same trust boundary that already applies to a repository's build scripts, npm
`postinstall`, or a project-scoped capability's hooks (which GSD has loaded since Phase 2) — and it
is **narrower** than those, because a command never fires on its own: you have to type it. GSD does
not (and with plain files cannot) cryptographically distinguish a genuine project-local install from
a forged project ledger. The honest guidance:
- A **global** install's consent record lives outside any repo and is trustworthy.
- A **project-scoped** capability is only as trustworthy as the repository it ships in — treat
running its commands like running that repo's other code. Review repos before running `gsd`
commands in them, and prefer global installs for capabilities you want to trust across projects.

View File

@@ -108,4 +108,4 @@ gsd capability remove <id> --scope project
- [How to version and upgrade a capability](version-a-capability.md)
- [Develop a Capability for GSD 1.5+](develop-a-capability.md)
- [Turn a capability off (and keep it off)](turn-a-capability-off.md)
- [The capability trust model](../explanation/the-capability-trust-model.md) — why removal is surgical and reversible
- [The capability trust model](../explanation/capability-trust-model.md) — why removal is surgical and reversible

View File

@@ -2,14 +2,15 @@
> **Generated file — do not edit by hand.**
> This matrix is generated from the capability registry by
> `scripts/gen-capability-matrix.cjs` (introduced in 1.6.0) and kept honest
> by a drift guard in CI. Any manual edit will be overwritten on the next
> generation run. To change a capability's declared metadata, edit the
> corresponding `capabilities/<id>/capability.json` and rebuild.
> `scripts/gen-capability-matrix.cjs` and kept honest by a drift guard
> (`tests/capability-matrix-sync.test.cjs` runs `--check`). Any manual edit is
> overwritten on the next generation run. To change a capability's declared
> metadata, edit the corresponding `capabilities/<id>/capability.json` and run
> `node scripts/gen-capability-matrix.cjs --write`.
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
[Capability manifest fields](#manifest-field-reference) —
[Trust model explanation](../explanation/capability-trust-model.md)
[The capability trust model](../explanation/capability-trust-model.md)
---
@@ -17,96 +18,96 @@ See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
| Column | Description |
|---|---|
| **id** | Canonical capability identifier; must be unique across first- and third-party capabilities. Reserved prefixes: `gsd-`, `gsd-core-`, `anthropic-`. |
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: `gsd-`, `gsd-core-`, `anthropic-`. |
| **role** | `feature` — extends what the loop does; `runtime` — adapts GSD to a specific AI runtime/IDE. |
| **tier** | `core` — always active; `standard` — active when the runtime supports it; `full` — opt-in or runtime-specific. |
| **version** | Semver version of the capability. Values shown are placeholders; the generator stamps exact per-capability versions from `capability.json` at release. |
| **engines.gsd** | Semver range expressing host-version compatibility. A hard gate at install and at load. |
| **extension points** | Loop extension points this capability registers into. See [the phase loop](../explanation/the-phase-loop.md) for the full ordered list. `see capability.json` means the generator would emit the precise set; only well-known registrations are listed here. |
| **hook kinds** | Subset of `step`, `contribution`, `gate` that the capability's hooks use. |
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. `—` means the capability declares no range. |
| **extension points** | The loop points this capability registers hooks into (from the registry's `byLoopPoint` index). `—` means it registers none (typical for runtime capabilities, whose job is surface emission). |
| **hook kinds** | Which of `step`, `contribution`, `gate` the capability's hooks use. `—` means none. |
| **source** | `first-party` — ships with GSD Core; `third-party` — installed from an external source via `gsd capability install`. |
> **On versions.** This matrix intentionally omits a per-capability `version`
> column. First-party capabilities are versioned **in lockstep** with the GSD
> Core package (their `capability.json` `version` always equals the GSD release
> version), so a per-row version would simply repeat the package version and
> churn the committed file on every release. The stable host-compatibility
> signal — `engines.gsd` — is shown instead. A third-party capability's exact
> version is recorded in the per-runtime ledger (`.gsd-capabilities.json`) at
> install time.
---
## Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD
Core package and are stamped with the package version at release (per
ADR-1244 D6). They are not subject to the consent or integrity-pin flow
applied to third-party capabilities.
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
to third-party capabilities.
### Feature capabilities (role: feature)
### Feature capabilities (role: feature) — 16
Feature capabilities extend what the five-step loop does — contributing
research, planning, execution, verification, or ship artefacts.
Feature capabilities extend what the loop does — contributing research,
planning, execution, verification, or ship artefacts at the loop extension
points.
| id | role | tier | version | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|---|
| `research` | feature | standard | 1.6.0 | `>=1.6.0` | `discuss:pre`, `plan:pre` | step, contribution | first-party |
| `ui` | feature | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `ai-integration` | feature | standard | 1.6.0 | `>=1.6.0` | see capability.json | step, gate | first-party |
| `security` | feature | full | 1.6.0 | `>=1.6.0` | `execute:pre`, `verify:pre` | gate | first-party |
| `code-review` | feature | standard | 1.6.0 | `>=1.6.0` | `verify:pre`, `verify:post` | step, gate | first-party |
| `schema-gate` | feature | standard | 1.6.0 | `>=1.6.0` | `execute:pre` | gate | first-party |
| `pattern-mapper` | feature | standard | 1.6.0 | `>=1.6.0` | see capability.json | contribution | first-party |
| `nyquist` | feature | full | 1.6.0 | `>=1.6.0` | see capability.json | step, gate | first-party |
| `validation` | feature | standard | 1.6.0 | `>=1.6.0` | `verify:pre`, `verify:post` | step, gate | first-party |
| `graphify` | feature | full | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `intel` | feature | standard | 1.6.0 | `>=1.6.0` | see capability.json | step, contribution | first-party |
| `audit` | feature | standard | 1.6.0 | `>=1.6.0` | see capability.json | step, gate | first-party |
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `ai-integration` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `audit` | feature | full | `>=1.6.0` | — | — | first-party |
| `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party |
| `drift` | feature | full | `>=1.6.0` | `execute:wave:post` | gate | first-party |
| `gap-analysis` | feature | standard | `>=1.6.0` | `plan:post` | gate | first-party |
| `graphify` | feature | full | `>=1.6.0` | — | — | first-party |
| `intel` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `mempalace` | feature | full | `>=1.6.0` | `discuss:pre`, `discuss:post`, `plan:pre`, `plan:post`, `execute:wave:post`, `verify:post`, `ship:post` | step, contribution | first-party |
| `nyquist` | feature | full | `>=1.6.0` | `verify:post` | step | first-party |
| `pattern-mapper` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
| `profile-pipeline` | feature | full | `>=1.6.0` | — | — | first-party |
| `research` | feature | standard | `>=1.6.0` | `plan:pre` | step | first-party |
| `schema-gate` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
| `security` | feature | full | `>=1.6.0` | `plan:pre`, `verify:post`, `ship:pre` | step, contribution, gate | first-party |
| `tdd` | feature | full | `>=1.6.0` | `plan:pre`, `execute:post` | contribution, gate | first-party |
| `ui` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post`, `verify:post` | step, gate | first-party |
> **Note:** version `1.6.0` is the placeholder the generator replaces with the
> actual per-capability `version` field from each `capability.json`. The 12
> loop extension points available to feature capabilities are, in order:
> `discuss:pre`, `discuss:post`, `plan:pre`, `plan:post`, `execute:pre`,
> `execute:wave:pre`, `execute:wave:post`, `execute:post`, `verify:pre`,
> `verify:post`, `ship:pre`, `ship:post`. A capability registers into the
> subset it needs; registration of all 12 is unusual.
### Runtime capabilities (role: runtime)
### Runtime capabilities (role: runtime) — 16
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files appropriate for that
host environment.
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are `—`.
| id | role | tier | version | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|---|
| `claude` | runtime | core | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `codex` | runtime | core | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `gemini` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `antigravity` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `cline` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `cursor` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `opencode` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `kilo` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `copilot` | runtime | full | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `augment` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `trae` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
| `qwen` | runtime | standard | 1.6.0 | `>=1.6.0` | see capability.json | step | first-party |
> **Note:** runtime capabilities typically do not register into the 12 loop
> extension points in the same way feature capabilities do — their primary
> responsibility is surface emission (skills, agents, config). Exact hook
> registrations, where they exist, are emitted by the generator into the
> `extension points` cell.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
| `antigravity` | runtime | core | `>=1.6.0` | — | — | first-party |
| `augment` | runtime | core | `>=1.6.0` | — | — | first-party |
| `claude` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cline` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codebuddy` | runtime | core | `>=1.6.0` | — | — | first-party |
| `codex` | runtime | core | `>=1.6.0` | — | — | first-party |
| `copilot` | runtime | core | `>=1.6.0` | — | — | first-party |
| `cursor` | runtime | core | `>=1.6.0` | — | — | first-party |
| `gemini` | runtime | core | `>=1.6.0` | — | — | first-party |
| `hermes` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kilo` | runtime | core | `>=1.6.0` | — | — | first-party |
| `kimi` | runtime | core | `>=1.6.0` | — | — | first-party |
| `opencode` | runtime | core | `>=1.6.0` | — | — | first-party |
| `qwen` | runtime | core | `>=1.6.0` | — | — | first-party |
| `trae` | runtime | core | `>=1.6.0` | — | — | first-party |
| `windsurf` | runtime | core | `>=1.6.0` | — | — | first-party |
---
## Third-party capabilities
Once a user installs a third-party capability via `gsd capability install
<spec>`, it enters the **runtime registry overlay** (ADR-1244 D2) and appears
in this matrix on their machine alongside native capabilities. Third-party
rows use the same column schema as first-party rows.
### How a third-party row is produced
The generator reads the capability's `capability.json` from the per-scope
install root (`~/.gsd/capabilities/<id>/` for global installs;
`.gsd/capabilities/<id>/` for project-scoped installs), validates it against
the same conformance rules applied to native manifests, and emits a row
identical in shape to the native rows above. The only difference is the
`source` column, which shows `third-party`.
This matrix is the **first-party catalogue**: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via `gsd capability install <spec>` it enters the **runtime
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is `gsd capability list` (see the
[`gsd capability` command reference](gsd-capability-command.md)), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with `source` = `third-party`.
### Column values for third-party rows
@@ -115,41 +116,41 @@ identical in shape to the native rows above. The only difference is the
| **id** | As declared in `capability.json`. Must not use reserved prefixes (`gsd-`, `gsd-core-`, `anthropic-`). |
| **role** | `feature` or `runtime`, as declared. |
| **tier** | `core`, `standard`, or `full`, as declared. |
| **version** | Semver from `capability.json`; the value recorded in the ledger at install time. |
| **engines.gsd** | Range from `capability.json`; verified at install and at each load. |
| **extension points** | As declared in `capability.json`. Validated against the known 12 extension-point identifiers. |
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
| **hook kinds** | `step`, `contribution`, and/or `gate` as declared. Disclosed in the consent summary at install. |
| **source** | `third-party` |
### Community registry
Whether GSD operates or advertises a central community registry of third-party
capabilities is **TBD/TBA** (see [PRD-1244 §8](../prd/1244-capability-ecosystem.md#8-open-questions--decisions-deferred)).
The matrix mechanic and all manifest fields ship in 1.6.0 regardless of that
decision. URL/git/npm/tarball import does not depend on a central registry.
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
ship regardless of that decision; URL/git/npm/tarball import does not depend on
a central registry.
---
## Manifest field reference
The fields below are defined in `capability.json` and govern how a capability
appears in this matrix. For the full schema, see [ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest).
appears in this matrix. For the full schema, see
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
and the [capability manifest reference](capability-manifest.md).
| Field | Required | Type | Purpose |
|---|---|---|---|
| `version` | **Yes** | semver string | Capability version. The registry rejects manifests without this field. |
| `version` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
| `engines.gsd` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
| `compatVersions` | No | object: cap-version → min-gsd-version | Graceful downgrade table for sources that enumerate versions (git tags, registry, npm). |
| `integrity` | No | `sha512-<base64>` | SHA-512 digest of the capability bundle. Verified before extraction when present; mismatch aborts. |
| `provenance` | No | `{ sourceRepo, commit }` | Source provenance. SHOULD be present for first-party and curated capabilities; populated in CI. |
| `compatVersions` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
| `integrity` | No | `sha512-<base64>` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
| `provenance` | No | `{ sourceRepo, commit }` | Source provenance; populated in CI for first-party/curated capabilities. |
---
## Related documents
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
- [Capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured the way they are
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture; D7 and D8 extended by ADR-1244
- [ADR-894](../adr/894-capability-declaration-format.md) — capability declaration format
- [ADR-1016](../adr/1016-runtime-capability-descriptor.md) — runtime capability descriptor
- [Capability manifest reference](capability-manifest.md) — the full `capability.json` schema
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)

View File

@@ -0,0 +1,284 @@
#!/usr/bin/env node
'use strict';
/**
* gen-capability-matrix.cjs — ADR-1244 Phase 6 (Decision D9).
*
* Generates docs/reference/capability-matrix.md FROM the committed capability
* registry (gsd-core/bin/lib/capability-registry.cjs), so the matrix can never
* drift from the actual capability set. Kept honest by a drift guard
* (tests/capability-matrix-sync.test.cjs runs `--check`).
*
* The matrix is RELEASE-STABLE by design: it does NOT embed each capability's
* exact `version` (which tracks the GSD package version in lockstep and would
* churn the committed file — and trip the drift guard — on every release). It
* shows `engines.gsd` (the stable host-compatibility RANGE) instead, and notes
* the version-lockstep rule in prose. The committed matrix therefore changes
* only on intentional capability edits (add/remove a capability, change its
* tier/role/engines/extension-points/hook-kinds) — never on a version bump.
*
* Usage:
* node scripts/gen-capability-matrix.cjs # print to stdout
* node scripts/gen-capability-matrix.cjs --write # write the committed file
* node scripts/gen-capability-matrix.cjs --check # exit 1 if the committed file is stale
*/
const fs = require('fs');
const path = require('path');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
const ROOT = path.resolve(__dirname, '..');
const REGISTRY_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs');
const MATRIX_PATH = path.join(ROOT, 'docs', 'reference', 'capability-matrix.md');
/** Canonical loop extension points, in order (mirrors the phase loop). */
const LOOP_POINTS = [
'discuss:pre', 'discuss:post',
'plan:pre', 'plan:post',
'execute:pre', 'execute:wave:pre', 'execute:wave:post', 'execute:post',
'verify:pre', 'verify:post',
'ship:pre', 'ship:post',
];
const POINT_ORDER = new Map(LOOP_POINTS.map((p, i) => [p, i]));
/**
* Build a capId → { points:Set, kinds:Set } map from the registry's byLoopPoint
* index — the authoritative record of which loop points each capability registers
* into and with which hook kind (step / contribution / gate).
*/
function extensionsByCapability(registry) {
const out = new Map();
const byPoint = registry.byLoopPoint || {};
const KIND = { steps: 'step', contributions: 'contribution', gates: 'gate' };
for (const point of Object.keys(byPoint)) {
const reg = byPoint[point] || {};
for (const arrKey of ['steps', 'contributions', 'gates']) {
for (const hook of reg[arrKey] || []) {
const capId = hook && hook.capId;
if (typeof capId !== 'string') continue;
let e = out.get(capId);
if (!e) { e = { points: new Set(), kinds: new Set() }; out.set(capId, e); }
e.points.add(point);
e.kinds.add(KIND[arrKey]);
}
}
}
return out;
}
function fmtPoints(set) {
if (!set || set.size === 0) return '—';
for (const p of set) {
// Surface a typo'd/unknown loop point at generation time rather than silently sorting it last.
// The registry validates point names at load, so this should never fire — but if it does, the
// generator (not a confused reader) is where it must be caught.
if (!POINT_ORDER.has(p)) {
process.stderr.write(`gen-capability-matrix: WARNING — unknown loop point "${p}" (not one of the ${LOOP_POINTS.length} canonical points)\n`);
}
}
return [...set]
.sort((a, b) => (POINT_ORDER.has(a) ? POINT_ORDER.get(a) : 99) - (POINT_ORDER.has(b) ? POINT_ORDER.get(b) : 99) || a.localeCompare(b))
.map((p) => '`' + p + '`')
.join(', ');
}
function fmtKinds(set) {
if (!set || set.size === 0) return '—';
const order = { step: 0, contribution: 1, gate: 2 };
return [...set].sort((a, b) => (order[a] ?? 9) - (order[b] ?? 9)).join(', ');
}
function fmtEngines(cap) {
const g = cap && cap.engines && cap.engines.gsd;
return typeof g === 'string' && g ? '`' + g + '`' : '—';
}
/** Render one capability table (rows sorted by id) for the given role. */
function renderTable(caps, role, extByCap) {
const rows = caps
.filter((c) => c.role === role)
.sort((a, b) => a.id.localeCompare(b.id))
.map((c) => {
const ext = extByCap.get(c.id) || { points: null, kinds: null };
return `| \`${c.id}\` | ${c.role} | ${c.tier || '—'} | ${fmtEngines(c)} | ${fmtPoints(ext.points)} | ${fmtKinds(ext.kinds)} | first-party |`;
});
return [
'| id | role | tier | engines.gsd | extension points | hook kinds | source |',
'|---|---|---|---|---|---|---|',
...rows,
].join('\n');
}
function buildMatrix(registry) {
const caps = Object.values(registry.capabilities || {});
const extByCap = extensionsByCapability(registry);
const featureTable = renderTable(caps, 'feature', extByCap);
const runtimeTable = renderTable(caps, 'runtime', extByCap);
const featureCount = caps.filter((c) => c.role === 'feature').length;
const runtimeCount = caps.filter((c) => c.role === 'runtime').length;
return `# Capability matrix reference
> **Generated file — do not edit by hand.**
> This matrix is generated from the capability registry by
> \`scripts/gen-capability-matrix.cjs\` and kept honest by a drift guard
> (\`tests/capability-matrix-sync.test.cjs\` runs \`--check\`). Any manual edit is
> overwritten on the next generation run. To change a capability's declared
> metadata, edit the corresponding \`capabilities/<id>/capability.json\` and run
> \`node scripts/gen-capability-matrix.cjs --write\`.
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
[Capability manifest fields](#manifest-field-reference) —
[The capability trust model](../explanation/capability-trust-model.md)
---
## Column definitions
| Column | Description |
|---|---|
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: \`gsd-\`, \`gsd-core-\`, \`anthropic-\`. |
| **role** | \`feature\` — extends what the loop does; \`runtime\` — adapts GSD to a specific AI runtime/IDE. |
| **tier** | \`core\` — always active; \`standard\` — active when the runtime supports it; \`full\` — opt-in or runtime-specific. |
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. \`—\` means the capability declares no range. |
| **extension points** | The loop points this capability registers hooks into (from the registry's \`byLoopPoint\` index). \`—\` means it registers none (typical for runtime capabilities, whose job is surface emission). |
| **hook kinds** | Which of \`step\`, \`contribution\`, \`gate\` the capability's hooks use. \`—\` means none. |
| **source** | \`first-party\` — ships with GSD Core; \`third-party\` — installed from an external source via \`gsd capability install\`. |
> **On versions.** This matrix intentionally omits a per-capability \`version\`
> column. First-party capabilities are versioned **in lockstep** with the GSD
> Core package (their \`capability.json\` \`version\` always equals the GSD release
> version), so a per-row version would simply repeat the package version and
> churn the committed file on every release. The stable host-compatibility
> signal — \`engines.gsd\` — is shown instead. A third-party capability's exact
> version is recorded in the per-runtime ledger (\`.gsd-capabilities.json\`) at
> install time.
---
## Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD
Core package and are stamped with the package version at release (per
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
to third-party capabilities.
### Feature capabilities (role: feature) — ${featureCount}
Feature capabilities extend what the loop does — contributing research,
planning, execution, verification, or ship artefacts at the loop extension
points.
${featureTable}
### Runtime capabilities (role: runtime) — ${runtimeCount}
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are \`—\`.
${runtimeTable}
---
## Third-party capabilities
This matrix is the **first-party catalogue**: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via \`gsd capability install <spec>\` it enters the **runtime
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is \`gsd capability list\` (see the
[\`gsd capability\` command reference](gsd-capability-command.md)), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with \`source\` = \`third-party\`.
### Column values for third-party rows
| Column | Value |
|---|---|
| **id** | As declared in \`capability.json\`. Must not use reserved prefixes (\`gsd-\`, \`gsd-core-\`, \`anthropic-\`). |
| **role** | \`feature\` or \`runtime\`, as declared. |
| **tier** | \`core\`, \`standard\`, or \`full\`, as declared. |
| **engines.gsd** | Range from \`capability.json\`; verified at install and at each load. |
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
| **hook kinds** | \`step\`, \`contribution\`, and/or \`gate\` as declared. Disclosed in the consent summary at install. |
| **source** | \`third-party\` |
### Community registry
Whether GSD operates or advertises a central community registry of third-party
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
ship regardless of that decision; URL/git/npm/tarball import does not depend on
a central registry.
---
## Manifest field reference
The fields below are defined in \`capability.json\` and govern how a capability
appears in this matrix. For the full schema, see
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
and the [capability manifest reference](capability-manifest.md).
| Field | Required | Type | Purpose |
|---|---|---|---|
| \`version\` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
| \`engines.gsd\` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
| \`compatVersions\` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
| \`integrity\` | No | \`sha512-<base64>\` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
| \`provenance\` | No | \`{ sourceRepo, commit }\` | Source provenance; populated in CI for first-party/curated capabilities. |
---
## Related documents
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
- [Capability manifest reference](capability-manifest.md) — the full \`capability.json\` schema
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)
`;
}
function loadRegistry() {
delete require.cache[require.resolve(REGISTRY_PATH)];
return require(REGISTRY_PATH);
}
/** Normalize CRLF→LF + ensure a single trailing newline, for cross-platform compare. */
function normalize(s) {
return s.replace(/\r\n/g, '\n').replace(/\n+$/, '\n');
}
function main() {
const flag = process.argv[2];
const registry = loadRegistry();
const content = buildMatrix(registry);
if (flag === '--check') {
let committed;
try {
committed = fs.readFileSync(MATRIX_PATH, 'utf8');
} catch {
throw new ExitError(1, `${path.relative(ROOT, MATRIX_PATH)} is missing. Run:\n node scripts/gen-capability-matrix.cjs --write`);
}
if (normalize(committed) !== normalize(content)) {
throw new ExitError(1, `${path.relative(ROOT, MATRIX_PATH)} is stale. Run:\n node scripts/gen-capability-matrix.cjs --write`);
}
console.log(`${path.relative(ROOT, MATRIX_PATH)} is up to date.`);
return;
}
if (flag === '--write') {
fs.mkdirSync(path.dirname(MATRIX_PATH), { recursive: true });
fs.writeFileSync(MATRIX_PATH, content, 'utf8');
console.log(`Wrote ${path.relative(ROOT, MATRIX_PATH)}`);
return;
}
process.stdout.write(content);
}
if (require.main === module) runMain(main);
module.exports = { buildMatrix, extensionsByCapability };

View File

@@ -6,7 +6,7 @@
* from a crash mid-upgrade. The LEDGER WRITE is the commit point for every operation: a crash
* before it leaves the prior state fully intact; a crash after it is a completed operation.
*
* Trust invariants enforced here (see docs/explanation/the-capability-trust-model.md):
* Trust invariants enforced here (see docs/explanation/capability-trust-model.md):
* - install/upgrade never execute capability code (resolver stages copy-only; we only swap
* directories and edit JSON);
* - executable surfaces are disclosed and consent is required before anything is promoted

View File

@@ -427,7 +427,7 @@ export function loadRegistry(options: LoadRegistryOptions = {}): Registry {
// the user actually installed+consented to via the lifecycle — a bundle merely dropped on
// disk with no ledger entry provides declarative surfaces (Phase 2) but is NOT command-
// dispatchable. (Project-scope ledgers live in the repo tree and are thus only as trustworthy
// as the repo — see docs/explanation/the-capability-trust-model.md.)
// as the repo — see docs/explanation/capability-trust-model.md.)
if (families.length > 0 && committedIds.has(id)) commandRoots[id] = capDir;
}
}

View File

@@ -5,7 +5,7 @@
* never mutates the filesystem and never performs I/O beyond reading staged files to confirm
* declared executable artifacts exist. The actual consent decision (yes/no) is passed in by the
* caller — GSD has no interactive-prompt layer in lib (the runtime/CLI edge owns that), so the
* gate stays testable and side-effect-free. See docs/explanation/the-capability-trust-model.md.
* gate stays testable and side-effect-free. See docs/explanation/capability-trust-model.md.
*
* LEAF MODULE — imports ONLY: node:fs, node:path, and ./semver-compare.cjs.
*

View File

@@ -0,0 +1,59 @@
'use strict';
/**
* capability-matrix-sync.test.cjs — ADR-1244 Phase 6 (D9) drift guard.
*
* Asserts the committed docs/reference/capability-matrix.md is exactly what
* scripts/gen-capability-matrix.cjs would generate from the current registry.
* If a capability is added/removed or its tier/role/engines/extension-points/
* hook-kinds change without regenerating the matrix, this fails — the same
* pattern that keeps docs/INVENTORY-MANIFEST.json and capability-registry.cjs honest.
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const ROOT = path.resolve(__dirname, '..');
const GENERATOR = path.join(ROOT, 'scripts', 'gen-capability-matrix.cjs');
const MATRIX = path.join(ROOT, 'docs', 'reference', 'capability-matrix.md');
const { buildMatrix } = require('../scripts/gen-capability-matrix.cjs');
const registry = require('../gsd-core/bin/lib/capability-registry.cjs');
describe('capability-matrix drift guard (ADR-1244 Phase 6)', () => {
test('the committed matrix is in sync with the registry (`gen-capability-matrix.cjs --check` exits 0)', () => {
// execFileSync throws if the generator exits non-zero (i.e. the committed file is stale).
assert.doesNotThrow(() => {
execFileSync(process.execPath, [GENERATOR, '--check'], { cwd: ROOT, stdio: 'pipe' });
}, 'committed capability-matrix.md is stale — run: node scripts/gen-capability-matrix.cjs --write');
});
test('buildMatrix(registry) equals the committed file byte-for-byte (modulo line endings)', () => {
const generated = buildMatrix(registry).replace(/\r\n/g, '\n').replace(/\n+$/, '\n');
const committed = fs.readFileSync(MATRIX, 'utf8').replace(/\r\n/g, '\n').replace(/\n+$/, '\n');
assert.equal(committed, generated);
});
test('every first-party capability in the registry appears as a matrix row', () => {
const md = fs.readFileSync(MATRIX, 'utf8');
for (const cap of Object.values(registry.capabilities)) {
if (cap.role !== 'feature' && cap.role !== 'runtime') continue;
assert.ok(md.includes('`' + cap.id + '`'), `capability ${cap.id} (${cap.role}) must appear in the matrix`);
}
});
test('extension points + hook kinds reflect the registry byLoopPoint index (not placeholders)', () => {
const md = fs.readFileSync(MATRIX, 'utf8');
// The stub used "see capability.json" placeholders — the generated matrix must not.
assert.ok(!md.includes('see capability.json'), 'matrix must show real extension points, not placeholders');
// `security` registers a gate at ship:pre — a hard architectural invariant. Assert the precondition
// UNCONDITIONALLY (so this never degrades to a vacuous pass if the registry changes), then assert the
// rendered row reflects it.
const shipPreGates = (registry.byLoopPoint['ship:pre'] && registry.byLoopPoint['ship:pre'].gates) || [];
assert.ok(shipPreGates.some((g) => g.capId === 'security'), 'precondition: security registers a ship:pre gate in the registry');
const securityRow = md.split('\n').find((l) => l.includes('`security`') && l.includes('|'));
assert.ok(securityRow && securityRow.includes('`ship:pre`'), 'security row must list its real ship:pre extension point');
});
});