Files
msd-core/docs/reference/capability-manifest.md
0xdhx 0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00

30 KiB
Raw Blame History

Capability Manifest Reference (capability.json)

Canonical ADRs: ADR-1244 · ADR-894 · ADR-1016 See also: How to develop a capability · Capability Command Reference

Each capability is a folder capabilities/<id>/ (or an overlay root ~/.gsd/capabilities/<id>/ / .gsd/capabilities/<id>/) containing one capability.json declaration. The file is schema-validated JSON with a common envelope plus a role-typed body (role: "feature", role: "runtime", or role: "reviewer").


Envelope fields

These fields are present for role: "feature", role: "runtime", and role: "reviewer" capabilities.

Field Type Required Description
id string (kebab-case) Yes Unique identifier; must equal the folder name. The prefix gsd-, gsd-core-, and anthropic- are reserved for first-party use.
role "feature" | "runtime" | "reviewer" Yes Discriminator that selects the body schema. "reviewer" is for lane-only capabilities that ship a reviewer body and nothing else — see Reviewer body below.
version semver string Yes (1.6.0+) Semantic version of this capability. The registry rejects a manifest without one.
title string Yes Short human-readable label. Must be a non-empty string.
description string Yes Longer summary sentence. Must be a non-empty string.
tier "core" | "standard" | "full" Yes Source of truth for install-profile membership and surface cluster assignment. tier propagates via the requires-closure; install profiles are generated from it.
requires string[] Yes Capability id values this capability depends on. Must be present as an array (use [] when there are no dependencies). Each entry must exist in the registry, be acyclic, and be tier-monotone (a core capability may not require a standard or full capability; a standard capability may not require a full capability).
engines object No Host-compatibility constraint. Sub-field: gsd — semver range string (e.g. ">=1.6.0 <3.0.0"). Acts as a hard gate at install and at load; a mismatch blocks installation and causes the overlay to be skipped with a warning at load time.
runtimeCompat object Yes (role: "feature") Declares which host runtimes this capability can surface through. Validated for every role: "feature" capability (a feature manifest without it fails validation). Sub-fields: supported — a non-empty array of kebab-case runtime ids, or the single wildcard ["*"] for a runtime-agnostic capability; unsupported — an array of kebab-case runtime ids (the wildcard is not permitted here); notes — optional object mapping a runtime id (or "*") to a non-empty explanatory string. The wildcard "*" may not be mixed with concrete ids in the same array, and the reserved names __proto__/constructor/prototype are rejected.
compatVersions object No Graceful-downgrade table mapping "<capVersion>" to "<min gsd version>". Only meaningful for sources that enumerate versions (git tags, registry, npm); a bare tarball URL carries one version and simply blocks on incompatibility.
integrity string No sha512-<base64> hash of the capability bundle. Verified before extraction when present; mismatch aborts install.
provenance object No { sourceRepo: string, commit: string }. Emitted in CI for first-party and curated capabilities.
author object No { name: string, email?: string, url?: string }.
homepage string No URL.
repository string No URL.
license string No SPDX licence identifier (e.g. "MIT").
keywords string[] No Arbitrary search tags.

Feature body (role: "feature")

Feature capabilities declare owned artefacts, lifecycle hooks, a federated configuration slice, and loop extension registrations.

skills and agents

Sub-field Type Description
skills string[] Owned skill stems. Exactly one capability may own each stem across the entire merged registry (first-party ∪ overlay).
agents string[] Owned agent stems. Same uniqueness constraint as skills.

The skills stems declared here are disclosed by name as instruction surfaces in the pre-install consent summary (ADR-2363). Bodies are installed verbatim and are not content-scanned — see the capability trust model.

agents are classified as an instruction surface too (ADR-2363 D3), but the stems declared here are not disclosed at the prompt: a third-party capability's agents[] are never staged into the agent's instruction context — the staging path that unions third-party skills into a runtime's skills directory has no equivalent for agents — so naming them would claim a surface that does not exist. This does not make agents safe or inert; it means the mechanism does not yet reach them.

hooks

Non-loop lifecycle hooks.

Sub-field Type Description
event string Hook event name (host-runtime specific).
script string Path to the hook script, relative to the capability root. The hook command written into the host settings is the realpath-confined absolute path to this script (so it always runs the bundle's own file regardless of the working directory) and is POSIX single-quoted (so an install prefix containing spaces cannot break it). For shell safety the path must contain only [A-Za-z0-9._/-] — no whitespace, no shell metacharacters (; | & $ ` ( ) < > * ? [ ] { } ! ~ # ' " \ newline), no leading -, no absolute path, and no .. segment. A script outside this allowlist fails validation and the capability is rejected.

config — federated config-key schema slice

The config field is an object whose keys are federated configuration keys contributed by this capability. Each key must be absent from the central config-schema and absent from every other capability's config object (collision fails the build gate). Each entry has the following shape:

Property Type Description
type "boolean" | "string" | "number" | "enum" Value type.
default (type-consistent) Default value; must be consistent with type.
description string Human-readable explanation of the key's effect.
values string[] enum only. Exhaustive list of permitted string values.

steps

Steps run at a loop extension point as independent units. Ordering within a point is derived from produces/consumes (topological sort; capability-id is the tiebreak).

Sub-field Type Required Description
point string Yes One of the 12 valid loop extension point identifiers (see table below).
ref object Yes The dispatch target. Exactly one of { "skill": "<stem>" }, { "agent": "<stem>" }, or { "command": "<name>" } (the three are mutually exclusive). A skill/agent stem must be declared in this capability's skills/agents array.
produces string[] Yes Artefact names this step produces. Must be present as an array (use [] when it produces none); an omitted produces fails validation. No two capability steps may produce the same artefact at the same point.
consumes string[] Yes Artefact names this step consumes. Must be present as an array (use [] when it consumes none); an omitted consumes fails validation.
onError "skip" | "halt" Yes Behaviour on failure; must be present and one of "skip" or "halt" (an omitted onError fails validation). Steps are purely additive — they never halt or redirect the host workflow on their own; a blocking precondition is expressed as a gate.
when string No Dotted config key; the step is active only when the key is truthy. Evaluated deterministically at render time; phase-context applicability is the skill's own responsibility.
fragment object No Optional inline-or-file prompt fragment attached to the step, with the same { "path": "<relative path>" } or { "inline": "<string>" } semantics as a contribution's fragment. A path is materialised (read and inlined) at load time, resolved against the capability directory and confined to it (.. traversal is rejected).

contributions

Contributions inject a fragment into a named agent role's prompt at a loop extension point. Multiple contributions into the same agent role render as ordered labelled blocks (<contribution from="<id>">…</contribution>).

Sub-field Type Required Description
point string Yes One of the 12 valid loop extension point identifiers.
into string Yes Agent role name. Must be a role published by that loop extension point in the host contract.
produces string[] Yes Artefact names this contribution produces. Use [] when it produces none.
consumes string[] Yes Artefact names this contribution reads. Use [] when it reads none.
fragment object Yes Either { "path": "<relative path>" } (file content) or { "inline": "<string>" } (literal text).
when string No Dotted config key; activates the contribution conditionally.
onError "skip" | "halt" No Behaviour on failure.

gates

Gates check a condition at a loop extension point and optionally block progression.

Sub-field Type Required Description
point string Yes One of the 12 valid loop extension point identifiers.
check object Yes One of three forms (see table below). Must be present as an object; an omitted check fails validation.
blocking boolean Yes Must be present and a boolean; an omitted blocking fails validation. When true, a failed check halts the loop at this point.
onError "skip" | "halt" Yes Behaviour when the check itself errors; must be present and one of "skip" or "halt" (an omitted onError fails validation).
when string No Dotted config key; activates the gate conditionally.

check forms:

Form Shape Blocking permitted Notes
Query { "query": "<gsd_run query>" } Yes Deterministic first-party code.
Predicate { "predicate": { "kind": "artifact-exists" | "config-equals" | …, … } } Yes Declarative; no code path.
Agent verdict { "agentVerdict": { "ref": …, "prompt": … } } No (forced advisory) LLM evaluation; non-deterministic checks may not halt the loop.

Valid point values

The 12 loop extension points are a closed, additive-only vocabulary. Every steps, contributions, and gates entry must use one of these identifiers exactly.

Point Phase Position
discuss:pre Discuss Before the discuss step executes
discuss:post Discuss After the discuss step completes
plan:pre Plan Before the plan step executes
plan:post Plan After the plan step completes
execute:pre Execute Before the execute phase begins
execute:wave:pre Execute Before each execution wave
execute:wave:post Execute After each execution wave
execute:post Execute After the execute phase completes
verify:pre Verify Before the verify step executes
verify:post Verify After the verify step completes
ship:pre Ship Before the ship step executes
ship:post Ship After the ship step completes

Runtime body (role: "runtime")

Runtime capabilities describe how GSD projects its artefacts onto one host CLI. The body is a closed 8-axis (plus 4 install-surface) vocabulary; no feature-only fields (skills, agents, steps, contributions, gates, hooks) are permitted. Full semantic specifications, the closed enum values for each axis, and the 16-runtime worked examples are in ADR-1016.

Axis Field Type summary
Config home runtime.configHome Structured object with kind (dot-home | dot-home-nested | xdg | generic-agents-root), name, optional parent, env[], probe[], probeExists, skillsHome. probeExists is an optional sub-path applied to probe candidates: for generic-agents-root it is a hard filter (a candidate qualifies only if <candidate>/<probeExists> exists); for dot-home-nested it is a preference that makes probing pick the candidate GSD owns (e.g. gsd-core/VERSION) over a bare-existing sibling before falling back — see ADR-1016 and #213/#217.
Local config dir runtime.localConfigDir Required dot-prefixed string. The runtime's local content-rewrite directory — the ./ target GSD stamps into rewritten artefact bodies (e.g. ./.claude/ → ./<localConfigDir>/) and the local install dir basename. Backs getDirName() (registry-derived, #1679). Usually .<runtime> (the runtime's home dot-dir), but three runtimes diverge because they read GSD's content from a non-home directory: copilot → .github (GitHub Copilot reads custom instructions from .github/copilot-instructions.md / .github/instructions/; see convertClaudeToCopilotContent rewrites in src/runtime-artifact-conversion.cts), antigravity → .agents (local agent/workflow dir; see the antigravity rewrites in src/runtime-artifact-conversion.cts), kimi → .kimi-code. Distinct from configHome.name (the global install home, which for these three is .copilot / antigravity / agents). Byte-parity-proven against the prior hand-maintained mapping by the golden-install-parity harness.
Config format runtime.configFormat Closed enum: settings-json | toml | markdown | markdown-dir | none.
Artefact layout runtime.artifactLayout Object with global and local arrays of ArtifactKind (kind, destSubpath, prefix, nesting, recursive, stage).
Command style runtime.commandStyle Closed enum: slash-hyphen | shell-var.
Hooks surface runtime.hooksSurface Closed enum: settings-json | codex-hooks-json | cursor-hooks-json | copilot-inline | cline-rules | kimi-hooks-toml | none.
Sandbox tier runtime.sandboxTier Closed enum: none | codex-agent-sandbox.
Support tier runtime.supportTier Integer: 1 (fully tested first-party) | 2 (shipped, lower coverage).
Install surface runtime.installSurface Closed enum: settings-json | codex-toml | copilot-instructions | cline-rules | cursor-hooks-json | profile-marker-only.
Shared settings runtime.writesSharedSettings boolean. Whether the runtime writes a shared settings.json.
Permission writer runtime.permissionWriter null | "opencode" | "kilo" | "antigravity". The finish-time permissions-sidecar writer.
Extended hook events runtime.extendedHookEvents string[] over a closed vocabulary: SubagentStop, Stop, PreCompact, FileChanged, BeforeAgent, AfterAgent, BeforeModel, SubagentStart.

hostBehaviors

runtime.hostBehaviors is a closed vocabulary of per-host behavior switches consumed directly by installer and runtime-adaptation code. A key outside the vocabulary is ignored, with a non-fatal warning naming the capability and the key; it is never a validation error, so a manifest authored against a newer GSD degrades visibly instead of failing the build of a repo that merely reads it.

Adding a key is a reviewed first-party change, which is ADR-1016's intended friction rather than an obstacle: the runtime descriptor expresses every per-host difference as a value over a closed vocabulary, and a host needing a new shape gets a named primitive rather than an open escape hatch.

History. hostBehaviors went unvalidated until #2801, and this page previously described it as a deliberate open seam sanctioned by ADR-1016. That attribution was wrong — ADR-1016 does not mention hostBehaviors at all. See the ADR-1016 amendment.

The vocabulary holds 59 keys; 39 of them are set by exactly one capability. This table is not exhaustive — it lists the keys with the widest reuse so a reader can pattern-match new ones against the same shape:

Key Capabilities declaring it
reapplyCommand 9
skipSharedHooksInstall 8
frontmatterDialect 5
hyphenNameAgentBody 3
legacyCommandsGsdInstallMigration 3
legacyCommandsGsdUninstall 3
nativePlugin 3
skipUpdateBannerCommand 3
verificationStyle 3

reviewerCli has been removed. It was a boolean that marked a runtime capability as also being a reviewer lane. ADR-2782 replaced it with the reviewer body; it survived one release (1.9.0 → 1.10.0) as a derived legacy alias and was deleted in Phase 7 (#2801). No shipped capability declares it.

If your out-of-tree manifest still sets it: nothing crashes and nothing else about your capability changes — it simply contributes no reviewer lane, and the registry reports a non-fatal warning naming the capability. The warning reaches you at build time on stderr, and at install time through the overlay loader's diagnostics. To restore the lane, declare a reviewer body; Ship a reviewer lane in your capability is the migration path, and the field reference is below.

See ADR-1016 (the runtime descriptor is a closed vocabulary; its 2026-08-09 amendment closes hostBehaviors too) and ADR-2782 (introduces the reviewer body, and D9 retires the reviewerCli alias).

For a minimal role: "runtime" example, see ADR-1016 §Decision 8.


Reviewer body (role: "reviewer", or on any role)

ADR-2782 introduces the reviewer lane: one external CLI or model endpoint that /gsd-review hands a plan to for independent review.

To declare one, follow Ship a reviewer lane in your capability. This section is the field reference behind that guide.

The reviewer body is optional and absent-safe at every layer. A capability with no reviewer body is simply not a lane — that is never a validation error. This is a normative forward/backward-compatibility invariant, not a nicety: a plugin, a runtime, or a future GSD version may omit reviewer entirely with no consequence.

The shape is hybrid:

  • A reviewer body is admissible on role: "runtime", so an existing runtime capability — codex, antigravity — keeps one manifest that is both an installable runtime and a reviewer lane.
  • A third role, role: "reviewer", exists for lane-only CLIs that GSD never installs into. There are currently 5: coderabbit, gemini, llama-cpp, lm-studio, ollama.

Current role counts across capabilities/: feature 20, runtime 19, reviewer 5.

All 12 shipped lane declarations carry all 13 fields below.

Field Type Notes
slug string Lane identity; grammar ^[a-z0-9][a-z0-9_-]*$. May use _ (lm_studio, llama_cpp) even where the capability folder id is kebab-case (lm-studio).
flags string[] User-facing CLI flags that select this lane. A lane may declare more than one — antigravity declares --antigravity and --agy. 12 lanes declare 13 flags in total.
transport closed enum spawn | openai-http.
probe object Availability check. probe.kind is a closed enum: command-exists | command-capability | http-reachable. command-capability additionally takes binary, needle, and a required timeoutMs — it exists because a bare binary name can be ambiguous (kimi is claimed by both the Kimi Code CLI and the legacy Python kimi-cli), and the timeout bound is mandatory because an unbounded --help | grep probe is this repo's named Unbounded Subprocesses defect.
invoke object Shape is selected by transport. For spawn: binary, args[], promptChannel (stdin | argv | argv-file-ref | none), outputChannel (stdout | file-arg), outputArg (required when outputChannel is file-arg), modelArg (string or null), effortChannel (none | argv | env), env (optional; an object of environment name/value pairs, string values only, merged over the inherited environment for that one spawn — keys must match the portable environment-name grammar [A-Za-z_][A-Za-z0-9_]*, which is a portability policy rather than an OS limit, and __proto__ is refused because it would be dropped before reaching the child). For openai-http: hostConfigKey, defaultHost, path, modelDiscovery (none | first-from-models-endpoint), fallbackModel, effortChannel. args supports the {{model}}, {{prompt}}, {{effort}}, and {{output}} placeholders. Every field in this object is disclosed at install and bound to the consent signature — env and defaultHost by name in the consent prompt, the rest through a residual, so any change to a declared invoke field forces re-consent. env additionally refuses execution-primitive names — PATH, NODE_OPTIONS, LD_PRELOAD, DYLD_INSERT_LIBRARIES, BASH_ENV, PYTHONPATH, PERL5OPT, RUBYOPT, GIT_SSH_COMMAND, JAVA_TOOL_OPTIONS and siblings, matched case-insensitively (Windows environment lookup is). A lane needing a specific executable declares an absolute binary rather than reshaping the child's PATH. That denylist is defence in depth and not the boundary: it cannot be complete against an arbitrary child, and disclosure runs before validation, so install-time consent — which shows every declared pair and warns on execution-primitive names — is what actually gates them.
timeoutFloorMs number Measured per-lane floor. Lane divergence here is real and correct — the descriptor's job is to declare divergence in one place, not to promise uniformity.
emptyOutput closed enum stub-with-stderr | handler-owned.
reviewsSection string The REVIEWS.md heading this lane renders under. Must be unique across the merged roster.
evidenceClass closed enum source-grounded | diff-only (diff-only findings are down-weighted in consensus).
requiresBinaries string[] Extra binaries the lane needs beyond invoke.binary.
promptBudgetKey string or null Federated config key bounding prompt size.
modelConfigKey string or null Federated config key naming the model, e.g. review.models.kimi-code.
handler closed enum or null antigravity | openai-compatible | opencode | null.

handler is a closed enum of first-party handler names, not an open escape hatch. ADR-1016 explicitly rejected "arbitrary code in the descriptor"; hard shapes are absorbed by adding a named primitive that is reviewed first-party. The consequence, stated plainly: a third-party reviewer lane is strictly data-only. A plugin can ship a lane, but not a quirky lane that needs imperative code — a lane requiring behavior beyond the closed handler set is not expressible and must be proposed and merged first-party.

Uniqueness is enforced across the merged first-party ∪ overlay set: duplicate slug, duplicate flags entry, and duplicate reviewsSection are each build-time violations. Two lanes sharing a reviewsSection heading would silently merge their output in REVIEWS.md, producing apparent consensus that does not exist.

An unknown field inside a reviewer body is a non-fatal warning on stderr, never a build failure (ADR-2782 D4), so a manifest built against a newer GSD degrades visibly rather than crashing.

Example — lane-only role: "reviewer" capability

{
  "id": "coderabbit",
  "role": "reviewer",
  "version": "1.8.0",
  "title": "CodeRabbit",
  "description": "CodeRabbit CLI — cross-AI /gsd-review reviewer lane only; not a GSD install target (no runtime body, no artifacts).",
  "tier": "full",
  "requires": [],
  "engines": { "gsd": ">=1.8.0" },
  "reviewer": {
    "slug": "coderabbit",
    "flags": ["--coderabbit"],
    "transport": "spawn",
    "probe": { "kind": "command-exists", "binary": "coderabbit" },
    "invoke": {
      "binary": "coderabbit",
      "args": ["review", "--prompt-only"],
      "promptChannel": "none",
      "outputChannel": "stdout",
      "modelArg": null,
      "effortChannel": "none"
    },
    "timeoutFloorMs": 360000,
    "emptyOutput": "stub-with-stderr",
    "reviewsSection": "CodeRabbit",
    "evidenceClass": "diff-only",
    "requiresBinaries": [],
    "promptBudgetKey": null,
    "modelConfigKey": null,
    "handler": null
  }
}

Conformance invariants

The following invariants are enforced at build time by scripts/gen-capability-registry.cjs and at install time by the runtime-callable validateCapability() / validateCrossCapability() over the merged first-party ∪ overlay set.

  • version is required. The registry rejects any manifest without a semver version field.
  • id uniqueness. No two capabilities may share an id. An overlay whose id collides with a first-party id is rejected; first-party always wins.
  • Skill and agent stem uniqueness. Exactly one capability may own each skill or agent stem across the entire merged registry.
  • requires exist and are acyclic. Every id listed in requires must exist in the registry; the dependency graph must be acyclic.
  • requires is tier-monotone. A core capability may not require a standard or full capability. A standard capability may not require a full capability.
  • point values are from the closed set. Every point in steps, contributions, and gates must be one of the 12 identifiers above.
  • contribution.into is a published agent role. The into value must be an agent role declared by the host contract for that loop extension point.
  • Config key exclusivity. A federated config key must be owned by exactly one capability and absent from the central config-schema. Presence in both is a collision; a half-migrated key fails the build gate.
  • Artefact production uniqueness per point. No two capability steps may produces the same artefact name at the same loop extension point.
  • engines.gsd is a hard gate. A capability whose engines.gsd range does not satisfy the installed GSD version is blocked at install and skipped (with a warning) at load time.
  • Path confinement. Declared module paths may not use parent-directory traversal (../); modules are require()'d only from the capability's own install root.
  • Reserved namespace. Capability id values beginning with gsd-, gsd-core-, or anthropic- are reserved; third-party capabilities using these prefixes are rejected.

Example — complete role: "feature" capability

The following is the canonical UI design-contract capability from ADR-894. It illustrates all major body sections.

{
  "id": "ui",
  "role": "feature",
  "version": "1.0.0",
  "title": "UI design contracts",
  "description": "UI-SPEC design contract and retrospective UI audit for frontend phases.",
  "tier": "standard",
  "requires": [],
  "engines": { "gsd": ">=1.6.0" },
  "runtimeCompat": { "supported": ["*"], "unsupported": [] },
  "skills": ["ui-phase", "ui-review"],
  "agents": ["gsd-ui-checker", "gsd-ui-auditor"],
  "hooks": [],
  "config": {
    "workflow.ui_phase": {
      "type": "boolean",
      "default": true,
      "description": "Enable the UI design-contract gate during planning."
    },
    "workflow.ui_review": {
      "type": "boolean",
      "default": true,
      "description": "Enable the retrospective UI audit."
    },
    "workflow.ui_safety_gate": {
      "type": "boolean",
      "default": true,
      "description": "Block execution on unmet UI-SPEC contracts."
    }
  },
  "steps": [
    {
      "point": "plan:pre",
      "ref": { "skill": "ui-phase" },
      "produces": ["UI-SPEC.md"],
      "consumes": ["CONTEXT.md"],
      "when": "workflow.ui_phase",
      "onError": "skip"
    },
    {
      "point": "verify:post",
      "ref": { "skill": "ui-review" },
      "produces": ["UI-REVIEW.md"],
      "consumes": ["UI-SPEC.md"],
      "when": "workflow.ui_review",
      "onError": "skip"
    }
  ],
  "contributions": [],
  "gates": [
    {
      "point": "execute:wave:post",
      "check": { "query": "ui.safety-gate" },
      "when": "workflow.ui_safety_gate",
      "blocking": true,
      "onError": "halt"
    }
  ]
}

Notes on this example:

  • when on each hook references its own config key; whether the phase is actually a frontend phase is decided inside ui-phase (self-gate).
  • The plan:pre step self-skips on non-frontend phases, producing no UI-SPEC.md; the execute:wave:post gate's ui.safety-gate query passes gracefully when no UI-SPEC.md exists.
  • A contribution follows this shape: { "point": "plan:pre", "into": "planner", "produces": [], "consumes": [], "fragment": { "path": "loop/threat-model.md" }, "when": "workflow.security_enforcement" } (produces and consumes are required arrays — use [] when empty).