c19d3d7bdae02aaf836c79f27982d8be0451f507
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
23a65c4a3d |
fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem surfaced (#2045) but no SKILL.md is ever written to disk. The other four are controls that must keep holding: first-party-wins collision, profile-tier filter, nested-router layout unperturbed, and absent/malformed capability must not throw. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): materialize installed third-party capability skills A capability could report installed:true, surfaced:true, active:true and still never exist as an invocable command. #2045 fixed the registry layer — resolveSurface unions registry.capabilityClusters into the resolved skill set — but the materialization layer never got the matching fix. stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled commands/gsd/*.md and silently skipped any stem it couldn't find there, so a third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/ was never copied. Registry said surfaced; disk had nothing. Installed capability skills are now staged alongside the first-party ones, copied verbatim (they are authored complete for their target runtime and need no converter). First-party stems always win a collision, the profile filter still applies, and an absent or malformed capability degrades rather than throwing. Security: capability.json's skills[] entries are validated only as non-empty non-reserved strings (capability-validator.cjs:503-514) — no path shape is enforced upstream — so stems are sanitized (rejecting separators, '..', absolute paths, NUL) with an independent isPathConfined check on both the read and write paths. A '../../evil' stem writes nothing outside the capability's own dir. Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability skill dir has no first-party manifest entry, so every apply logged "preserving (user-owned or unknown)" for a live GSD-managed dir. The retained check now precedes the manifest gate; no deletion outcome changes, and genuinely unknown gsd-* dirs still warn and are preserved. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): address security review — bind skills to declaring capability, fix full profile An independent security review BLOCKED the first pass. Both blockers were mine. BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir and returned the first sorted match, never checking that a capability DECLARES the stem — ownership was inferred from attacker-controlled filesystem layout. Since install copies the whole bundle and the validator only checks DECLARED entries, a capability declaring `skills: []` could ship an undeclared skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the agent-invocable instructions the user believed came from the registered capability. Stems are now bound to their owning capId via registry.capabilityClusters, and only that capability's dir is read. BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that applySurface materializes `full` into a concrete Set. True for applySurface — false for the installer, which is the default path: resolveProfile returns the '*' sentinel and bin/install.js passes it straight to staging. So #2322 survived on the default `full` profile, i.e. the fix didn't fix the reported bug. The registry is now plumbed to staging, and '*' stages all capability-cluster stems. Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary install path) never threaded its registry either, which would have silently defeated the fix on the real default install. HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the first-party manifest, so uninstalling a capability left its instructions live in the agent's context forever. Staged skills now carry a marker making them GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved. MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over the whole stage dir. The tests asserted byte-equality and passed only because their fixtures contained no rewrite triggers. Claim dropped; tests now assert the rewrite against triggering content. LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently unreachable because install rejects symlinks) — comment corrected. The validator does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not a second layer — comment corrected and it now has traversal/NUL/absolute/empty test coverage (previously mutating it to `return true` left every test green). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2322): pin that the imperative adapter forwards a capability registry The delegation-args test deep-equalled the exact argv to installRuntimeArtifacts, so threading the composed capability registry through the ADR-1239 imperative adapter (required for #2322 — without it the default `full` install path never materializes third-party capability skills) failed it. The contract legitimately gained a parameter, so this is a stale-test correction, not a regression. Rather than deep-equalling the whole composed registry (brittle — it embeds the full agent/profile map), the test pins the leading args exactly and asserts only that a registry-shaped value is forwarded. That still fails if the adapter stops threading it, which is the regression the test exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2322): backfill PR number 2340 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7c93d9e222 |
feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e7855bc217 | fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) | ||
|
|
c866ac1b24 | fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) | ||
|
|
34bc096ec2 |
feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the six subcommands, dispatching to the existing lifecycle/ledger: - install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]… - update [<id>|--all] [--scope] [--yes] [--shared-file] (re-resolves recorded source) - remove <id> [--purge-data] [--scope] (first-party rejected) - list [--json] (first-party + overlay, both scopes, JSON array) - disable|enable <id> (activation-state alias of capability set --off/--on) Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home, project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json). Consent is non-interactive: --yes grants; without it an executable install aborts after printing the disclosure and writes nothing. Best-effort reconcile before each mutation. Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs, GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip, remove round-trip + first-party guard, disable/enable, unknown subcommand. Docs: docs/reference/gsd-capability-command.md reconciled to the real surface (ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned); docs/COMMANDS.md gains the gsd capability entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug Adversarial-review (Codex) fixes: - capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate fail-closes on it (was silently downgrading to permissive) - installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't act on a different id if the recorded source was retargeted - capability update: prints the consent disclosure, exits non-zero on --all partial failure, no longer masks the resolved id - capability remove: ledger-first ordering so an overlay is removable even if it shadows a first-party name; first-party guard only fires for ids not in the ledger - gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle not yet wired through this path) Silent-output bug (root cause, not waved off as pre-existing): - captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of stdout. Now it flushes the captured buffer before re-throwing (exit code preserved). - cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws ExitError so the wrapper flushes — matches the repo's no-process-exit architecture. - Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout. Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed - confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it. - mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or another capability's) server is skipped, so install/remove can't silently clobber user MCP config (hooks already append; the map-keyed mcpServers path was the gap). - capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead of silently downgrading the strict_known_registries policy to permissive. - Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry preserved; unparseable config blocks an external install. capability suite 83/83, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc - install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per the lifecycle contract; latent today, hardened for future status additions). - Clarify capResolveScope comment (project scope === already-resolved cwd) and document that strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide allowlist) in gsd-capability-command.md. - Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1451): backfill changeset PR number → #1457 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |