* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local>
11 KiB
Capability matrix reference
Generated file — do not edit by hand. This matrix is generated from the capability registry by
scripts/gen-capability-matrix.cjsand kept honest by a drift guard (tests/capability-matrix-sync.test.cjsruns--check). Any manual edit is overwritten on the next generation run. To change a capability's declared metadata, edit the correspondingcapabilities/<id>/capability.jsonand runnode scripts/gen-capability-matrix.cjs --write.
See also: ADR-1244 — Capability manifest fields — The capability trust model
Column definitions
| Column | Description |
|---|---|
| id | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: gsd-, gsd-core-, anthropic-. |
| role | feature — extends what the loop does; runtime — adapts GSD to a specific AI runtime/IDE; reviewer — declares a cross-AI reviewer lane (ADR-2782). A capability may be both a runtime and a reviewer. |
| tier | core — always active; standard — active when the runtime supports it; full — opt-in or runtime-specific. |
| engines.gsd | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. — means the capability declares no range. |
| extension points | The loop points this capability registers hooks into (from the registry's byLoopPoint index). — means it registers none (typical for runtime capabilities, whose job is surface emission). |
| hook kinds | Which of step, contribution, gate the capability's hooks use. — means none. |
| source | first-party — ships with GSD Core; third-party — installed from an external source via gsd capability install. |
On versions. This matrix intentionally omits a per-capability
versioncolumn. First-party capabilities are versioned in lockstep with the GSD Core package (theircapability.jsonversionalways equals the GSD release version), so a per-row version would simply repeat the package version and churn the committed file on every release. The stable host-compatibility signal —engines.gsd— is shown instead. A third-party capability's exact version is recorded in the per-runtime ledger (.gsd-capabilities.json) at install time.
Native (first-party) capabilities
First-party capabilities are implicitly trusted: they ship as part of the GSD Core package and are stamped with the package version at release (per ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied to third-party capabilities.
Feature capabilities (role: feature) — 22
Feature capabilities extend what the loop does — contributing research, planning, execution, verification, or ship artefacts at the loop extension points.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
ai-integration |
feature | full | >=1.6.0 |
plan:pre, verify:pre |
step, contribution, gate | first-party |
assumption-delta |
feature | full | >=1.6.0 |
plan:pre |
contribution | first-party |
audit |
feature | full | >=1.6.0 |
— | — | first-party |
broken-windows |
feature | full | >=1.7.0 |
ship:pre |
gate | first-party |
claude-orchestration |
feature | full | >=1.7.0 |
plan:post, execute:wave:pre |
contribution | first-party |
code-review |
feature | full | >=1.6.0 |
execute:post |
step | first-party |
drift |
feature | full | >=1.6.0 |
plan:pre, execute:wave:post |
gate | first-party |
external-job |
feature | full | >=1.7.0 |
plan:post, execute:wave:post |
contribution | first-party |
gap-analysis |
feature | standard | >=1.6.0 |
plan:post |
gate | first-party |
graphify |
feature | full | >=1.6.0 |
— | — | first-party |
intel |
feature | full | >=1.6.0 |
plan:pre |
step | first-party |
live-dom-uat |
feature | full | >=1.11.0 |
execute:wave:post |
step | first-party |
mempalace |
feature | full | >=1.6.0 |
discuss:pre, discuss:post, plan:pre, plan:post, execute:wave:post, verify:post, ship:post |
step, contribution | first-party |
nyquist |
feature | full | >=1.6.0 |
verify:post |
step | first-party |
pattern-mapper |
feature | full | >=1.6.0 |
plan:pre |
step | first-party |
profile-pipeline |
feature | full | >=1.6.0 |
— | — | first-party |
refactor-trigger |
feature | full | >=1.10.0 |
execute:post |
step | first-party |
research |
feature | standard | >=1.6.0 |
plan:pre |
step | first-party |
schema-gate |
feature | full | >=1.6.0 |
plan:pre |
contribution | first-party |
security |
feature | full | >=1.6.0 |
plan:pre, verify:post, ship:pre |
step, contribution, gate | first-party |
tdd |
feature | full | >=1.6.0 |
plan:pre, execute:post |
contribution, gate | first-party |
ui |
feature | full | >=1.6.0 |
plan:pre, execute:wave:post, verify:post |
step, gate | first-party |
Runtime capabilities (role: runtime) — 19
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
skills, agents, hooks configuration, and surface files for that host. They
typically register no loop hooks (their primary responsibility is surface
emission), so their extension-point and hook-kind cells are —.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
antigravity |
runtime | core | >=1.6.0 |
— | — | first-party |
augment |
runtime | core | >=1.6.0 |
— | — | first-party |
claude |
runtime | core | >=1.6.0 |
— | — | first-party |
cline |
runtime | core | >=1.6.0 |
— | — | first-party |
codebuddy |
runtime | core | >=1.6.0 |
— | — | first-party |
codex |
runtime | core | >=1.6.0 |
— | — | first-party |
copilot |
runtime | core | >=1.6.0 |
— | — | first-party |
cursor |
runtime | core | >=1.6.0 |
— | — | first-party |
hermes |
runtime | core | >=1.6.0 |
— | — | first-party |
kilo |
runtime | core | >=1.6.0 |
— | — | first-party |
kimi |
runtime | core | >=1.6.0 |
— | — | first-party |
kimi-code |
runtime | core | >=1.7.0 |
— | — | first-party |
opencode |
runtime | core | >=1.6.0 |
— | — | first-party |
pi |
runtime | core | >=1.7.0 |
— | — | first-party |
qwen |
runtime | core | >=1.6.0 |
— | — | first-party |
trae |
runtime | core | >=1.6.0 |
— | — | first-party |
vscode |
runtime | core | >=1.7.0 |
— | — | first-party |
windsurf |
runtime | core | >=1.6.0 |
— | — | first-party |
zcode |
runtime | core | >=1.6.0 |
— | — | first-party |
Reviewer capabilities (role: reviewer) — 5
Reviewer capabilities declare a cross-AI reviewer lane — one external CLI or
model endpoint /gsd-review hands a plan to (ADR-2782 D3). They are not install
targets: they emit no skills, agents, hooks or surface files, so their
extension-point and hook-kind cells are —. A host that is also a reviewer
(Claude, Codex, Cursor, OpenCode, Qwen, Antigravity) keeps one manifest and
appears under runtime above, carrying its lane alongside its runtime body;
only lanes that GSD never installs into appear here.
Because a lane receives the plan text, requirements, research findings and
CONTEXT.md decisions, it is a disclosed executable surface and is consent-gated
at install like any other — see
the trust model.
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|---|---|---|---|---|---|---|
coderabbit |
reviewer | full | >=1.8.0 |
— | — | first-party |
gemini |
reviewer | full | >=1.8.0 |
— | — | first-party |
llama-cpp |
reviewer | full | >=1.8.0 |
— | — | first-party |
lm-studio |
reviewer | full | >=1.8.0 |
— | — | first-party |
ollama |
reviewer | full | >=1.8.0 |
— | — | first-party |
Third-party capabilities
This matrix is the first-party catalogue: it is generated from the committed
registry and therefore lists only the capabilities that ship with GSD Core.
Installed third-party capabilities are NOT written into this committed file. Once a
user installs one via gsd capability install <spec> it enters the runtime
registry overlay (ADR-1244 D2); the overlay-aware view of what is installed on a
given machine is gsd capability list (see the
gsd capability command reference), which reports
first-party and installed third-party capabilities together using the same column
fields described below, with source = third-party.
Column values for third-party rows
| Column | Value |
|---|---|
| id | As declared in capability.json. Must not use reserved prefixes (gsd-, gsd-core-, anthropic-). |
| role | feature, runtime, or reviewer, as declared. |
| tier | core, standard, or full, as declared. |
| engines.gsd | Range from capability.json; verified at install and at each load. |
| extension points | The loop points the capability registers into, validated against the known 12 identifiers. |
| hook kinds | step, contribution, and/or gate as declared. Disclosed in the consent summary at install. |
| source | third-party |
Community registry
Whether GSD operates or advertises a central community registry of third-party capabilities is TBD/TBA (PRD). The matrix mechanic and all manifest fields ship regardless of that decision; URL/git/npm/tarball import does not depend on a central registry.
Manifest field reference
The fields below are defined in capability.json and govern how a capability
appears in this matrix. For the full schema, see
ADR-1244 D1
and the capability manifest reference.
| Field | Required | Type | Purpose |
|---|---|---|---|
version |
Yes | semver string | Capability version. The registry rejects manifests without it. |
engines.gsd |
Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
compatVersions |
No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
integrity |
No | sha512-<base64> |
SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
provenance |
No | { sourceRepo, commit } |
Source provenance; populated in CI for first-party/curated capabilities. |
Related documents
- ADR-1244 — Capability Ecosystem
- The capability trust model — why the trust rules are structured as they are
- The phase loop — the 12 loop extension points in context
- Capability manifest reference — the full
capability.jsonschema - ADR-857 — the original capability architecture (D7/D8 extended by ADR-1244)