* test(#3248): failing-first suite for instruction-surface disclosure 28 matrix rows from 50-test-matrix.md. Rows requiring the new Disclosure.instructionSurfaces field fail today; rows 18-20/23-25 (the ADR-2363 D4 signature invariants) pass today by construction because the current code never reads skills/agents at all, and stand as regression guards for the implementation commit. Refs #3248 * feat(#3248): disclose capability skills and agents as an instruction surface ADR-2363 D5. A capability whose only contribution was skills disclosed nothing at install: summarizeDisclosure early-returned "ships no executable surfaces (declarative only)" because hasExecutable was false, while each SKILL.md body landed verbatim in the agent's instruction context. discloseExecutableSurfaces gains a fifth, NON-executable class, instructionSurfaces, collecting declared skills/agents stems through the same safeCollect wrapper as the four existing collectors, so a hostile value degrades only this class and the function stays total for any manifest shape. Nothing existing is edited: the collectors, hasExecutable, disclosureSignature and missingArtifacts are untouched. get_impact rates the symbol CRITICAL at 196 affected, which is why the design is strictly additive. D4 is implemented by omission and pinned rather than left incidental: adding, changing or removing skills/agents leaves disclosureSignature byte-identical, so no stored consent record is perturbed and no spurious re-consent fires. ADR-2782's conditional-append trick is deliberately NOT reused - it worked because no manifest could declare a reviewer body before that class existed, whereas skills predate this one, so a conditional append would re-sign every already-consented skill-bearing capability. The renderer is extracted as summarizeInstructionSurfaces and called from BOTH branches of summarizeDisclosure. Appending only at the end would never render for skill-only capabilities - the ones that need it - since those take the early return. That branch's "declarative only" claim is now conditional on there being no instruction surface either. The renderer iterates rather than spreading into push, so an unbounded stem count cannot throw RangeError, and tolerates the bare {} the CLI edge passes via `res.disclosure || {}`. Scope note: #3248's prose says "skill stems"; ADR-2363 D3 classifies instruction surfaces as "skills, agents". Shipping skills alone would leave an ADR deliverable owned by no phase, and the epic has no Phase 2. Agents are the same shape at no extra cost. Narrowing back is a two-line change. Ratifies ADR-2363 (Proposed -> Accepted) and adds the owed ADR-1244 back-link. Closes #3248 * fix(#3248): escape consent-prompt values and narrow disclosure to skills Two review findings, both of which made the previous commit wrong. BLOCKER (isolated adversarial review). Every manifest-supplied value interpolated into a consent-prompt line was rendered unescaped. Those lines are joined with \n and written RAW to stderr on the needs-consent path (capability-command-router -> cli-exit runMain), so a stem carrying a newline forged additional lines indistinguishable from genuine GSD disclosure text, and an ANSI escape could clear or rewrite lines already printed. That defeats the informed-consent guarantee this change exists to provide, and is a prompt-injection vector against any agent that reads the stderr text to decide whether to retry with --yes. The hole was not unique to the new class - hook event/script, command family/module/router, every MCP field, and every reviewer-lane field were equally unescaped. Fixing only the new one would have created the generative-fix divergence this repo tracks, so renderValueForPrompt is applied to all five classes through one helper, guarded by a parity test that fails if a future class skips it. Escaping is identity for ordinary names, so no well-formed manifest's output changes. The disclosure OBJECT stays verbatim - only the rendered LINE is escaped - because the signature and every consumer reasoning about identity depend on the declared value. NARROWED to skills only. The previous commit also collected agents, arguing ADR-2363 D3 classifies instruction surfaces as "skills, agents". Verified against staging: stageSkillsForRuntimeAsSkills takes a registry and unions third-party skills in via readInstalledCapabilitySkill, while stageAgentsForRuntimeWithConverter takes only a source directory and has no registry-aware path. Third-party agents are never staged into the instruction context, so disclosing them would have put a false claim in a security prompt - worse than the scope creep two reviewers flagged it as. D3's classification stands; D5 now records that Phase 1 implements the skills half and that whether agents should be staged at all is an open maintainer question. Also reverts the premature ADR-2363 ratification. The previous commit flipped it to Accepted and asserted "#3248 merged" while this branch IS #3248 and is unmerged. Status returns to Proposed, and the ADR-1244 back-link - owed only on ratification - is withdrawn. Adds the fast-check property suite CLAUDE.md requires and the direct precedent (reviewer-trust-disclosure) already had: totality, D4 signature invariance, D3 hasExecutable invariance, and renderer totality over adversarial manifests. Refs #3248 * chore(#3248): correct changeset scope claim and backfill pr number The fragment was written against the pre-narrowing commit and still advertised 'skills and agents'. 4d26887e narrowed disclosure to skills only - third-party agents are never staged into the instruction context - but did not touch the fragment, so the release notes would have carried a claim the code does not implement. Also backfills pr:0 -> 3253 and names the prompt-escaping fix, which is user-visible and was absent from the original body. Changeset-only; no code or test changed, so the gsd-test pass recorded for 4d26887e still describes this tree's behavior. Refs #3248 --------- Co-authored-by: sim <sim@local>
267 lines
15 KiB
Markdown
267 lines
15 KiB
Markdown
# Develop a Capability for GSD 1.5+
|
|
|
|
This guide shows you how to add or change a first-party GSD Capability after the ADR-857 cutover. In GSD terms, the extension unit is a **Capability**. A plugin is a packaging or host-runtime term, for example a Claude Code plugin, not the unit that owns a GSD feature.
|
|
|
|
A Capability is right when the feature can be toggled as one unit and owns its own skills, agents, hooks, config keys, or command family. Keep verifier predicate contracts, the five-step loop spine, and shared infrastructure in core unless the ADRs explicitly move that boundary.
|
|
|
|
## Start from the boundary
|
|
|
|
Before writing a manifest, decide whether the work is core or a Capability.
|
|
|
|
Use a Capability when the feature:
|
|
|
|
- Can be enabled or disabled without changing the meaning of the base loop.
|
|
- Owns a stable feature name such as `research`, `ui`, `graphify`, `ai-integration`, or `pattern-mapper`.
|
|
- Adds a step, gate, or contribution at a Loop Extension Point.
|
|
- Owns one or more feature config keys.
|
|
- Owns a command family that can route through the capability registry.
|
|
|
|
Keep the work in core when the feature:
|
|
|
|
- Defines the reliability substrate of the loop, such as the verifier predicate contract described by ADR-550 and ADR-857.
|
|
- Is required for every installation profile.
|
|
- Mutates shared host workflow state rather than adding a declared hook contribution.
|
|
|
|
## Create the folder
|
|
|
|
Create one folder per Capability:
|
|
|
|
```text
|
|
capabilities/<id>/
|
|
capability.json
|
|
fragments/
|
|
plan-pre.md
|
|
```
|
|
|
|
The manifest path is `capabilities/<id>/capability.json`.
|
|
|
|
The `<id>` must match the `id` field in `capability.json`. Co-locate prompt fragments, owned skills, owned agents, and other owned artefacts under the Capability folder when the schema allows it. Shared host artefacts can be referenced by name, but ownership should remain clear in the manifest.
|
|
|
|
## Write `capability.json`
|
|
|
|
Use the existing manifests as the source of truth while the schema is still first-party:
|
|
|
|
- `capabilities/research/capability.json`
|
|
- `capabilities/ai-integration/capability.json`
|
|
- `capabilities/pattern-mapper/capability.json`
|
|
- `capabilities/ui/capability.json`
|
|
- `capabilities/graphify/capability.json`
|
|
|
|
At minimum, a feature Capability declares:
|
|
|
|
```json
|
|
{
|
|
"id": "example",
|
|
"role": "feature",
|
|
"version": "0.1.0",
|
|
"title": "Example",
|
|
"description": "Adds an example planning step.",
|
|
"tier": "standard",
|
|
"requires": [],
|
|
"runtimeCompat": { "supported": ["*"], "unsupported": [] },
|
|
"skills": [],
|
|
"agents": ["gsd-example-agent"],
|
|
"hooks": [],
|
|
"config": {},
|
|
"steps": [
|
|
{
|
|
"point": "plan:pre",
|
|
"ref": { "agent": "gsd-example-agent" },
|
|
"fragment": { "path": "fragments/plan-pre.md" },
|
|
"produces": ["EXAMPLE.md"],
|
|
"consumes": ["CONTEXT.md"],
|
|
"onError": "skip"
|
|
}
|
|
],
|
|
"contributions": [],
|
|
"gates": []
|
|
}
|
|
```
|
|
|
|
The registry generator validates the shape, ownership, and cross-capability contracts. `ref.skill` must name a skill declared by the same Capability. `ref.agent` must name an agent declared by the same Capability.
|
|
|
|
## Declare runtime compatibility
|
|
|
|
Every feature Capability must declare `runtimeCompat`. This is part of the GSD 1.5+ developer contract: a Capability says which runtime descriptors it can surface through, and the generator validates that declaration before the central registry is written.
|
|
|
|
Use the wildcard when the Capability is runtime-agnostic:
|
|
|
|
```json
|
|
"runtimeCompat": {
|
|
"supported": ["*"],
|
|
"unsupported": []
|
|
}
|
|
```
|
|
|
|
Use explicit runtime ids when the Capability is intentionally narrower:
|
|
|
|
```json
|
|
"runtimeCompat": {
|
|
"supported": ["claude", "codex"],
|
|
"unsupported": ["kilo"],
|
|
"notes": {
|
|
"kilo": "Requires a hook surface Kilo does not expose yet."
|
|
}
|
|
}
|
|
```
|
|
|
|
`supported` is required and must be non-empty. `"*"` means every descriptor-backed runtime, including future first-party runtime descriptors, is compatible unless it is listed in `unsupported`. Explicit runtime ids and `notes` keys must match runtime Capability ids such as `claude`, `codex`, `opencode`, or `kilo`; typos fail `node scripts/gen-capability-registry.cjs --check`.
|
|
|
|
## Add hooks
|
|
|
|
Loop Extension Points are the stable sites where Capabilities attach to the host loop. Phase 6 planning-time features use `plan:pre` so the core planner can ask the registry for active planning hooks instead of reading feature config directly.
|
|
|
|
Choose the hook kind that matches the behaviour:
|
|
|
|
- `steps` add a sequenced unit of work, such as running `gsd-phase-researcher` or `gsd-pattern-mapper`.
|
|
- `contributions` add labelled context to a host prompt.
|
|
- `gates` check a condition and may block when `blocking` is true.
|
|
|
|
Declare file artefact flow with `produces` and `consumes`. The registry uses those arrays to order hooks and to reject unsatisfied dependencies.
|
|
|
|
## Ship skills — and know what you are shipping
|
|
|
|
A skill your Capability declares is not an inert asset. Its `SKILL.md` body is copied **verbatim** into the user's runtime skills directory at install, where it becomes an agent-invocable instruction file. GSD does **not** scan it: there is no content inspection of any kind, at install or at any later point. What you write is what reaches the agent.
|
|
|
|
That makes a skill body an **instruction surface**, and it carries author responsibility that an inert artifact does not:
|
|
|
|
- **Write instructions you would be comfortable defending.** Your reach is bounded only by what the agent will do when told. Consent, integrity pinning and reversibility are the user's protections; content review is not among them.
|
|
- **Do not embed anything that tries to redirect the agent away from the user's task** — reframing its role, overriding host instructions, or persisting directives past the skill's own scope. That is indistinguishable from a prompt-injection payload, and the fact that it is not scanned is not permission.
|
|
- **Treat anything your skill tells the agent to read as data, not instructions.** If your skill has the agent ingest a file, a URL, or tool output, say so explicitly in the body. See [the untrusted-input boundary](../adr/1577-untrusted-input-boundary-and-injection-blocking.md).
|
|
- **The surface is disclosed.** The skills your Capability contributes are named, by stem, in their own section of the pre-install consent summary ([#3248](https://github.com/open-gsd/gsd-core/issues/3248)). Write your skill body on the assumption that a user sees it listed before they accept. Declared `agents[]` are not named there today — a third-party capability's agents are never staged into the agent's instruction context, so there is nothing to disclose — but they remain classified as an instruction surface, not an inert or safe one.
|
|
|
|
The reasoning behind this posture — including why scanning was considered and rejected — is recorded in [ADR-2363](../adr/2363-capability-instruction-surface-trust.md), and the user-facing side is [the capability trust model](../explanation/capability-trust-model.md).
|
|
|
|
## Use prompt fragments
|
|
|
|
Use `fragment.path` for prompt text longer than a short sentence:
|
|
|
|
```json
|
|
"fragment": { "path": "fragments/plan-pre.md" }
|
|
```
|
|
|
|
The generator materialises that file into `fragment.inline` in the generated registry. Paths must be relative to the Capability folder, must not be absolute, and must not contain `..`.
|
|
|
|
Use `fragment.inline` only for short, stable text:
|
|
|
|
```json
|
|
"fragment": { "inline": "Add the generated example context to the planner input." }
|
|
```
|
|
|
|
## Generate and verify the registry
|
|
|
|
After editing any Capability manifest or fragment, regenerate the committed registry:
|
|
|
|
```bash
|
|
node scripts/gen-capability-registry.cjs --write
|
|
```
|
|
|
|
Then run the drift check:
|
|
|
|
```bash
|
|
node scripts/gen-capability-registry.cjs --check
|
|
```
|
|
|
|
For a planning hook, verify the rendered output:
|
|
|
|
```bash
|
|
node gsd-core/bin/gsd-tools.cjs loop render-hooks plan:pre --raw
|
|
```
|
|
|
|
In installed workflow prose, the same resolver surface appears as `gsd-tools loop render-hooks plan:pre`.
|
|
|
|
The rendered JSON should include the active hook, the declared `ref`, and the materialised `fragment.inline`.
|
|
|
|
To verify a runtime-specific surface, pass the same config directory that the runtime installation uses:
|
|
|
|
```bash
|
|
node gsd-core/bin/gsd-tools.cjs loop render-hooks plan:pre --config-dir ~/.claude --raw
|
|
node gsd-core/bin/gsd-tools.cjs capability state --config-dir ~/.claude --raw
|
|
```
|
|
|
|
`capability state` is the diagnostic view for the same state that workflow dispatch consumes. For each Capability, `enabled` is true only when the Capability is both installed by the active profile and surfaced by the runtime surface. A hook's `configured` field reflects the `when` config key; `active` is true only when the Capability is enabled and the hook is configured on.
|
|
|
|
## Add or change a runtime descriptor
|
|
|
|
Runtime-specific facts belong in the runtime Capability declaration, not in a parallel allowlist or runtime-name branch. When adding a first-party runtime under `capabilities/<runtime>/capability.json`, declare these fields in the `runtime` object:
|
|
|
|
- `configHome`: the global config root resolver, including env overrides and probes.
|
|
- `configHome.skillsHome`: optional separate base home for runtimes whose global skills root differs from the config root.
|
|
- `artifactLayout.global` and `artifactLayout.local`: command, agent, skill, and Kimi-agent destinations.
|
|
- `hooksSurface`, `hookEvents`, and `extendedHookEvents`: hook registration surface and event dialect.
|
|
- `installSurface`, `writesSharedSettings`, and `permissionWriter`: install-time config mutation behavior.
|
|
|
|
The runtime homes, artifact layout, and install-plan resolvers read those descriptor fields directly. Adding a descriptor-backed runtime should not require editing a second list in `runtime-artifact-layout` or a fallback branch in `runtime-config-adapter-registry`.
|
|
|
|
## Own config in the Capability
|
|
|
|
Declare feature config keys in the Capability manifest:
|
|
|
|
```json
|
|
"config": {
|
|
"workflow.example_enabled": {
|
|
"type": "boolean",
|
|
"default": true,
|
|
"description": "Enable the example planning step."
|
|
}
|
|
}
|
|
```
|
|
|
|
Do not add migrated Capability keys to `gsd-core/bin/shared/config-schema.manifest.json`. The central schema remains for host/core keys; Capability-owned keys validate through the generated registry and are merged into `loadConfig` by the federated config overlay. Existing nested `.planning/config.json` values such as `{ "workflow": { "example_enabled": false } }` continue to override the Capability default.
|
|
|
|
## Wire the workflow through the registry
|
|
|
|
Host workflows should ask the resolver for active hooks and then dispatch from the resolved data. Do not add new direct `config-get workflow.<feature>` checks to a host workflow for a migrated feature.
|
|
|
|
The Phase 6 migrations follow this pattern:
|
|
|
|
- `research` registers a `plan:pre` step that invokes `gsd-phase-researcher`.
|
|
- `ai-integration` owns the AI-SPEC planning activation and config key.
|
|
- `pattern-mapper` registers a `plan:pre` step that invokes `gsd-pattern-mapper`.
|
|
- `plan-phase.md` reads the resolved `PLAN_PRE_HOOKS_JSON` and dispatches `ref.agent` or `ref.skill` from that data.
|
|
- `code-review` registers an `execute:post` step and review workflows dispatch from the active hook's `ref.skill`.
|
|
- `security` registers a `plan:pre` contribution, a `verify:post` step, and a blocking `ship:pre` gate.
|
|
- `nyquist` registers a `verify:post` step for validation coverage auditing.
|
|
|
|
This keeps "off means off" enforceable by construction: disabled Capabilities are absent from the active hook set, so the host workflow has nothing feature-specific to run.
|
|
|
|
## Test the Capability
|
|
|
|
Add focused tests before changing behaviour:
|
|
|
|
- Registry tests for manifest validation, ordering, config ownership, command family dispatch, or fragment materialisation.
|
|
- Workflow text tests for host workflow cutovers when a workflow stops reading direct feature config.
|
|
- Behavioural command tests when a command family moves behind `dispatchCapabilityCommand`.
|
|
- Documentation tests when the feature changes developer-facing behaviour.
|
|
|
|
Run the smallest affected tests first, then the full suite before opening a ready PR:
|
|
|
|
```bash
|
|
node --test tests/capability-registry.test.cjs
|
|
node --test tests/phase6-planning-capabilities.test.cjs
|
|
node --test tests/phase6-capstone-conformance.test.cjs
|
|
npm test
|
|
```
|
|
|
|
Run the Phase 6 capstone test whenever a Capability adds a `when` key or moves activation logic in a host workflow. It checks that migrated activation keys are resolved through Capability hooks or state, and that Capability-owned config keys stay out of the central schema. If a host workflow must keep core behaviour behind a `workflow.*` key, document why it is core rather than Capability-owned before adding it.
|
|
|
|
## Keep the docs with the slice
|
|
|
|
Every Phase 6 slice that changes capability behaviour must update the relevant docs in the same PR. Use this manual for developer-facing Capability authoring facts, use how-to guides for task flows, and use ADRs only for decisions and trade-offs.
|
|
|
|
## The capability ecosystem (1.6.0)
|
|
|
|
From GSD 1.6.0, capabilities are versioned (the `version` field is required in `capability.json`) and can be installed directly from a URL, a git ref, an npm package, or a local path — without modifying the core repo.
|
|
|
|
- **Tutorial** — [Build your first capability](../tutorials/build-your-first-capability.md): scaffold and install a declarative capability end-to-end in under ten minutes.
|
|
- **How-to** — [Publish a capability](../how-to/publish-a-capability.md): package and distribute a capability via a URL or registry.
|
|
- **How-to** — [Import a capability from a URL](../how-to/import-a-capability-from-a-url.md): install a third-party capability from a git URL, tarball, or npm package.
|
|
- **How-to** — [Version and update a capability](../how-to/version-a-capability.md): manage `version`, `engines.gsd`, and `compatVersions`; use `gsd capability update`.
|
|
- **How-to** — [Remove a capability](../how-to/remove-a-capability.md): uninstall cleanly with `gsd capability remove`, including the `--purge-data` option.
|
|
- **How-to** — [Ship a reviewer lane in your capability](../how-to/ship-a-reviewer-lane.md): declare a `reviewer` body (GSD 1.9.0+) so `/gsd-review` discovers and invokes your external review CLI or model endpoint.
|
|
- **How-to** — [List your reviewer lane in the registry](../how-to/list-your-reviewer-lane.md): publish a lane to the Reviewer Lane Registry (GSD 1.9.1+) so other people can find and install it.
|
|
- **Reference** — [Capability manifest](../reference/capability-manifest.md): all fields and validation rules for `capability.json`.
|
|
- **Reference** — [Capability matrix](../reference/capability-matrix.md): which first-party capabilities exist, their extension points, and their compatibility matrix.
|
|
- **Explanation** — [Capability trust model](../explanation/capability-trust-model.md): how declarative and executable capabilities are treated differently at install time.
|
|
- **ADR-1244** — `docs/adr/1244-capability-ecosystem.md`: the architectural decision that introduced the installable capability ecosystem.
|