Commit Graph

9 Commits

Author SHA1 Message Date
Tom Boucher
6badb839a0 fix(#3514): deny internal fetch hosts; disclose unverified integrity (#3516)
* test(#3514): add failing-first denylist and integrity suites

* fix(#3514): deny internal fetch hosts; disclose unverified integrity

* docs(#3514): trust-model, glossary, and changeset entries

* fix(#3514): scope v6 checks to literals; exact pin kinds in prompt

* chore(#3514): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:34:29 -04:00
Tom Boucher
e57918a648 fix(#3515): disclose the intentional mcp unconfined posture (#3517)
* test(#3515): add failing-first unconfined-mcp notice suite

* fix(#3515): disclose the intentional mcp unconfined posture

* chore(#3515): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:19:04 -04:00
Tom Boucher
f96cb44f85 enhance(#3248): disclose capability skills as an instruction surface (#3253)
* test(#3248): failing-first suite for instruction-surface disclosure

28 matrix rows from 50-test-matrix.md. Rows requiring the new
Disclosure.instructionSurfaces field fail today; rows 18-20/23-25 (the
ADR-2363 D4 signature invariants) pass today by construction because the
current code never reads skills/agents at all, and stand as regression
guards for the implementation commit.

Refs #3248

* feat(#3248): disclose capability skills and agents as an instruction surface

ADR-2363 D5. A capability whose only contribution was skills disclosed
nothing at install: summarizeDisclosure early-returned "ships no executable
surfaces (declarative only)" because hasExecutable was false, while each
SKILL.md body landed verbatim in the agent's instruction context.

discloseExecutableSurfaces gains a fifth, NON-executable class,
instructionSurfaces, collecting declared skills/agents stems through the same
safeCollect wrapper as the four existing collectors, so a hostile value
degrades only this class and the function stays total for any manifest shape.
Nothing existing is edited: the collectors, hasExecutable, disclosureSignature
and missingArtifacts are untouched. get_impact rates the symbol CRITICAL at
196 affected, which is why the design is strictly additive.

D4 is implemented by omission and pinned rather than left incidental: adding,
changing or removing skills/agents leaves disclosureSignature byte-identical,
so no stored consent record is perturbed and no spurious re-consent fires.
ADR-2782's conditional-append trick is deliberately NOT reused - it worked
because no manifest could declare a reviewer body before that class existed,
whereas skills predate this one, so a conditional append would re-sign every
already-consented skill-bearing capability.

The renderer is extracted as summarizeInstructionSurfaces and called from BOTH
branches of summarizeDisclosure. Appending only at the end would never render
for skill-only capabilities - the ones that need it - since those take the
early return. That branch's "declarative only" claim is now conditional on
there being no instruction surface either. The renderer iterates rather than
spreading into push, so an unbounded stem count cannot throw RangeError, and
tolerates the bare {} the CLI edge passes via `res.disclosure || {}`.

Scope note: #3248's prose says "skill stems"; ADR-2363 D3 classifies
instruction surfaces as "skills, agents". Shipping skills alone would leave an
ADR deliverable owned by no phase, and the epic has no Phase 2. Agents are the
same shape at no extra cost. Narrowing back is a two-line change.

Ratifies ADR-2363 (Proposed -> Accepted) and adds the owed ADR-1244 back-link.

Closes #3248

* fix(#3248): escape consent-prompt values and narrow disclosure to skills

Two review findings, both of which made the previous commit wrong.

BLOCKER (isolated adversarial review). Every manifest-supplied value
interpolated into a consent-prompt line was rendered unescaped. Those lines
are joined with \n and written RAW to stderr on the needs-consent path
(capability-command-router -> cli-exit runMain), so a stem carrying a newline
forged additional lines indistinguishable from genuine GSD disclosure text,
and an ANSI escape could clear or rewrite lines already printed. That defeats
the informed-consent guarantee this change exists to provide, and is a
prompt-injection vector against any agent that reads the stderr text to decide
whether to retry with --yes.

The hole was not unique to the new class - hook event/script, command
family/module/router, every MCP field, and every reviewer-lane field were
equally unescaped. Fixing only the new one would have created the
generative-fix divergence this repo tracks, so renderValueForPrompt is applied
to all five classes through one helper, guarded by a parity test that fails if
a future class skips it. Escaping is identity for ordinary names, so no
well-formed manifest's output changes. The disclosure OBJECT stays verbatim -
only the rendered LINE is escaped - because the signature and every consumer
reasoning about identity depend on the declared value.

NARROWED to skills only. The previous commit also collected agents, arguing
ADR-2363 D3 classifies instruction surfaces as "skills, agents". Verified
against staging: stageSkillsForRuntimeAsSkills takes a registry and unions
third-party skills in via readInstalledCapabilitySkill, while
stageAgentsForRuntimeWithConverter takes only a source directory and has no
registry-aware path. Third-party agents are never staged into the instruction
context, so disclosing them would have put a false claim in a security prompt -
worse than the scope creep two reviewers flagged it as. D3's classification
stands; D5 now records that Phase 1 implements the skills half and that
whether agents should be staged at all is an open maintainer question.

Also reverts the premature ADR-2363 ratification. The previous commit flipped
it to Accepted and asserted "#3248 merged" while this branch IS #3248 and is
unmerged. Status returns to Proposed, and the ADR-1244 back-link - owed only on
ratification - is withdrawn.

Adds the fast-check property suite CLAUDE.md requires and the direct precedent
(reviewer-trust-disclosure) already had: totality, D4 signature invariance, D3
hasExecutable invariance, and renderer totality over adversarial manifests.

Refs #3248

* chore(#3248): correct changeset scope claim and backfill pr number

The fragment was written against the pre-narrowing commit and still
advertised 'skills and agents'. 4d26887e narrowed disclosure to skills
only - third-party agents are never staged into the instruction context -
but did not touch the fragment, so the release notes would have carried a
claim the code does not implement.

Also backfills pr:0 -> 3253 and names the prompt-escaping fix, which is
user-visible and was absent from the original body.

Changeset-only; no code or test changed, so the gsd-test pass recorded for
4d26887e still describes this tree's behavior.

Refs #3248

---------

Co-authored-by: sim <sim@local>
2026-08-09 13:52:29 -04:00
Tom Boucher
bc5619dd27 docs(#3247): record the capability instruction-surface trust model (#3249)
* docs(#3247): record the capability instruction-surface trust model

ADR-2363 records the trust posture for third-party capability SKILL.md
bodies, which #2322/#2340 made agent-invocable without any content-level
control. The path-level protections that fix shipped are all present; no
content scanner exists, and external-descriptor-trust.cts never had one.
Nothing was bypassed - the control did not exist and the boundary was
never written down.

D1 records the posture: skill bodies are trusted, unscanned agent
instructions. D2 rejects content scanning on Kerckhoffs (a shipped rule
set is readable by the adversary who installs it), on threat-model
non-transfer from ADR-1577 (there, instructions are anomalous inside
data; here they are the payload's legitimate form), and on Goodhart (a
scanned-OK line displaces the judgment the consent prompt exists to
provoke). D3 replaces the executable/non-executable binary with three
classes, adding instruction surface.

D4 keeps instruction surfaces out of the v1 disclosureSignature. The
signature is NOT the activation binding - hasProjectConsent compares
contentHash only, and a global install carries no consent record at all.
What re-encoding would do is perturb the signature of every skill-bearing
capability and fire a spurious re-consent prompt on its next upgrade,
which is what ADR-2782 D4 rule 5 already forbids. If instruction surfaces
ever need to be signature-bound, that lands as a versioned v2 signature
with a migration, never an in-place re-encoding.

Corrects capability-trust-model.md, which claimed skills get lighter
consent because they do not execute code - true, and not the relevant
property, since the agent is the interpreter. Adds the author-side
boundary to develop-a-capability.md and links it from
publish-a-capability.md. Both state that per-skill disclosure at the
consent prompt lands with #3248 and does not happen today.

Docs-only. No behavior change; no consent record perturbed. D5's
mechanism is Phase 1 (#3248), which is why the ADR is Proposed.

Refs #2363

* chore(#3247): backfill changeset pr number to 3249

---------

Co-authored-by: sim <sim@local>
2026-08-09 11:47:54 -04:00
Tom Boucher
ffd5370464 fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs

Docs told readers to type the colon form, which no runtime registers -- 18 of
19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an
example got an unrecognized command. Swept 178 occurrences across 53 files,
locale mirrors included so they do not re-diverge from English.

The colon form is a source-authoring token, not a user-facing one: install-time
converters key on it to produce the hyphen form runtimes actually register. So
the sweep is scoped, and three things are deliberately left alone:

- ADRs, which are a historical record; editing their prose falsifies what was
  written at the time.
- The legacy release-notes archive, pending a maintainer decision on whether it
  follows the same historical carve-out. Excluding it keeps a later reversal
  additive rather than a revert.
- Source artifacts under commands, workflows and agents, where the colon form is
  load-bearing. Rewriting those would break the installed-skill guarantee across
  every runtime -- the single largest hazard here.

The plugin namespace form is a real, separate token and survives untouched.

Adds a lint enforcing exactly that boundary, since the correct form genuinely
differs by directory and nothing previously caught the drift.

Also fixes a hardcoded colon form in the capability-matrix generator. The sweep
alone would have left the generated matrix disagreeing with the template that
produces it, so the fix is at the source and the output regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): stop the sweep misquoting source frontmatter

Adversarial review caught three lines where the sweep rewrote a citation of the
literal YAML name: key from a source command file. That key genuinely is the
colon form -- this change's own carve-out logic says source-authoring tokens keep
it -- so the docs ended up misquoting the real files. One of the three is an
acceptance-checklist assertion, which the sweep turned into a false statement.

Restored the three citations to match their sources verbatim, surgically: where a
line carried both a name: citation and a real reader-facing slash command, only
the citation reverted and the command stayed corrected.

The guard needed the same distinction, or it would have flagged the restoration
and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a
source token and is now permitted. The exemption is deliberately narrow -- a bare
gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness.

Also makes the detection case-insensitive. Review found /GSD:next slipped through
silently; no such casing exists in the tree today, so this closes a latent gap
rather than fixing a live one.

Swept the whole tree for further corrupted citations: none beyond the three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): retire the stale-next invariant and sweep next like every other command

Maintainer decision on a genuine conflict between two contracts.

Invariant #3054 banned the literal /gsd-next from user-facing docs because it
named a retired workflow-advance command. But commands/gsd/next.md is a live
command -- the state-aware smart-entry launcher -- and this issue requires docs
to use the hyphen form every runtime actually registers. Both could not hold for
this one command, so docs had been sidestepping the ban by keeping the colon
form, which is exactly the defect this issue exists to remove.

FEATURES.md already recorded the reassignment: the hyphen form "is not the
retired workflow-advance command; it is reserved for the state-aware smart-entry
launcher. Workflow advancement remains under /gsd-progress --next." With that
reassignment the invariant's premise is obsolete and the guard now contradicts
the documented command form, so it is retired with a comment recording why
rather than deleted silently.

next is now swept like every other command, and the earlier exemption added to
the new guard is removed so nothing is special-cased.

Four citations of the literal name: frontmatter key stay in colon form, because
the source file really does carry name: gsd:next and a doc quoting it must
reproduce it verbatim. Two of those lines were reworded to say which side is the
frontmatter key and which is the slash command, since they previously conflated
the two.

Verified the retired scan would now genuinely fail against this tree -- the
conflict was real and resolved, not dodged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2903): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 13:23:44 -04:00
Tom Boucher
69dbf28ca7 feat(#2796): reviewer lane as a fourth trust-disclosure class (#2826)
* feat(#2796): reviewer lane as a fourth trust-disclosure class

Phase 3 of epic #2782, delivering ADR-2782 D5. A reviewer lane is piped the plan
text, requirements, research findings and CONTEXT.md decisions, and its output is
read back into REVIEWS.md -- an egress channel for the most sensitive artifacts
GSD produces. Making lanes pluggable WITHOUT a disclosure class would open a
data-exfiltration path behind a manifest field, which is why this gates the
feature rather than following it.

- discloseExecutableSurfaces was cyclomatic 51 / cognitive 99 / 110 lines with
  risk_level critical. Rather than grow it, it is now a short orchestrator over
  four extracted collectors (hooks, commands, mcp -- behaviour-preserving -- plus
  the new lane collector), each independently testable. That is also what makes
  the 80% mutation threshold survivable: 51 branches in one function cannot be
  mutation-covered by whole-function tests.

- A spawn lane discloses its binary AND its full declared args, in rendered and
  raw form. Binary-only disclosure would be insufficient and not hypothetically:
  a lane declaring python3 with innocuous args could later change them to
  ['-c', '<program>'] without the binary changing. That is the bug class #1459
  already fixed for MCP servers.

- An openai-http lane has no binary, so it discloses the destination host and the
  config key naming it. A localhost destination is disclosed and distinguished
  from a remote one. Both forms name the egress payload classes.

THE CONSTRAINT THAT SHAPED THE DESIGN: the lane element is appended to the
disclosure signature ONLY when at least one lane is declared. signatureForManifest
is the consent key both the loader and the lifecycle compare, so appending
unconditionally would have changed every installed capability's signature and
re-prompted every user for every capability on their next upgrade -- for a feature
they do not use. Two pre-change goldens are asserted byte-for-byte as the tripwire.

The resolved host is deliberately NOT in the signature. The loader has no config
resolver, so including it would make the loader and the lifecycle compute
different signatures for the same manifest and produce a permanent false-mismatch
loop. It is disclosed and recorded instead; Phase 5b re-resolves and compares at
invocation, which is D5 rule 4's own placement.

reviewsSection and timeoutFloorMs are also excluded from the signature: a cosmetic
change must not force re-consent, because a prompt carrying no security
information is how users learn to click through.

A lane's binary is NOT existence-checked against the staged bundle. It is a PATH
tool, never a bundle artifact; treating it like a hook script would add every lane
to missingArtifacts and block every lane install.

Two defects fixed beyond the fourth class:
- isLocalHostValue mis-parsed a scheme-less host: new URL('localhost:1234') does
  NOT throw, it reads 'localhost' as the URL scheme and yields an empty hostname,
  so a bare host:port would have been reported as non-local. Now falls back on an
  empty hostname rather than only on a caught throw.
- The orchestrator's safeCollect closes a PRE-EXISTING totality gap in the other
  three classes: a null manifest, or one with a throwing getter or Proxy trap,
  previously threw out of disclosure -- which runs on an UNVALIDATED manifest at
  install time. No well-formed input changes; all 51 existing trust tests pass.

Closes #2796

* fix(#2796): close four disclosure gaps found by the isolated security review

All four were REPRODUCED by execution against the shipped module, and all four
passed the existing 41-test suite while live -- each exists because the matrix
did not think to ask.

B (MEDIUM, reachable via plain JSON). Non-string argv members were folded into
the consent SIGNATURE but dropped from the human-facing text, because the summary
rendered the string-filtered args rather than the raw declared array. A manifest
declaring args ['--json', 7, {mode:'exfiltrate-everything'}, true] printed as
'--json' alone -- the host still receives the rest, so the user consented to a
surface never shown. That directly contradicts this design's own Kerckhoffs claim
that nothing about a lane is hidden. The summary now renders the raw array, with
non-strings shown in a visible form, and never throws on a circular or BigInt
member.

F (MEDIUM, reachable). The [local] flag is design-load-bearing, and it was
dropped for every loopback form except the dotted quad and the bare hostname.
Bracketed IPv6 was mangled by splitting on the address's own colons ([::1]:8080
became '['), and legacy IPv4 encodings were not recognised at all. A browser,
curl and the OS resolver all treat 127.1, 2130706433, 0x7f000001 and 0177.0.0.1
as loopback. isLocalHostValue now handles bracketed and bare IPv6, IPv4-mapped
loopback, and inet_aton shorthand/decimal/hex/octal. The dangerous direction was
already clean and is now pinned by tests: localhost.evil.com,
http://user@localhost@evil.com and friends stay REMOTE.

C (LOW). An empty reviewer body flipped hasExecutable true and perturbed the
disclosure signature, producing a re-consent prompt whose only content was
'(no binary declared)'. A prompt carrying no security information is the
click-through-training harm this design explicitly refuses for reviewsSection and
timeoutFloorMs; refusing it there and permitting it here was inconsistent. A body
declaring nothing recognised is no longer a lane. The test is deliberately broad
-- any ONE recognised field suffices -- because requiring specifically a binary,
or specifically a slug, would let a lane declaring only the other slip through
unconsented, which is the far worse failure. Pinned in both directions.

D (LOW-MEDIUM). Disclosure runs BEFORE validation, so a mis-cased or unrecognised
transport reaches this code. Keying on an exact string sent a lane that plainly
declares a hostConfigKey down the spawn branch, printing '(no binary declared)'
for a lane egressing to a live remote host, and left its resolvedHost blank --
which reads as 'no destination', the precise thing the design forbids. Both the
collector and the summary now branch on the declared SHAPE, so such a lane
discloses its key and either a resolved host or the explicit unresolved marker.

Two further findings were reproduced but confirmed NOT reachable through the real
pipeline and are recorded as known limits rather than fixed: a selective-throw
Proxy blanking a whole lane, and NaN/Infinity/undefined colliding to 'null' in a
signature. Every production manifest reaches disclosure through
readManifestBounded's strict JSON.parse, which cannot produce a Proxy, a getter,
a BigInt, a circular reference, NaN or Infinity. The 0/-0 sub-case IS reachable
via valid JSON but is inert -- String(0) === String(-0), so a spawned process
receives identical argv.

9 regression tests added (50 total in this file, up from 41).

* chore(#2796): backfill changeset pr number to 2826
2026-07-29 11:21:49 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
f52a7a5f77 feat(#1245): add capability ecosystem ADR, PRD, and developer documentation (#1248)
Phase 0 of the Capability Ecosystem epic (#1244): the design record and the
third-party-author documentation set, with no runtime or code changes.

- docs/adr/1244-capability-ecosystem.md — architecture decision record
  (amends/extends ADR-857 Decisions 7 & 8)
- docs/prd/1244-capability-ecosystem.md — product requirements
- Diataxis docs: tutorial, how-to (publish/import/version/remove),
  reference (manifest schema, /gsd:capability command, capability matrix),
  explanation (trust model); cross-links added to develop-a-capability.md

Refs #1244

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 17:51:43 -04:00