Commit Graph

563 Commits

Author SHA1 Message Date
Tom Boucher
35478b615e refactor(#1646): route capability routers through Command Routing Hub per ADR-959 (#1647)
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959

Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.

src/cjs-command-router-adapter.cts
  * Imported ERROR_REASON from io.cjs.
  * UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
    as the second arg to error() — additive for host routers (their
    existing one-arg error callbacks ignore the second arg), required
    for capability routers whose tests assert reason === 'sdk_unknown_command'
    on the JSON-error envelope.

src/graphify-command-router.cts
  * Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing/invalid --budget) now
    return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
    instead of calling error() directly (Q2=C, Q4=ii from grilling).
  * Success handlers keep direct output() calls.
  * Subcommands array is alphabetical for byte-identical 'Available:'
    text in the unknown-subcommand message.
  * The unknown-subcommand path is now owned by the Hub's manifest
    check (the adapter passes SDK_UNKNOWN_COMMAND).

src/intel-command-router.cts
  * Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing filePath for patch-meta
    and extract-exports) return makeInvalidArgs Results.
  * Preserved the timeAgo mutation in the non-raw status handler.
  * Preserved the lazy require('./intel.cjs') inside the route function.

src/audit-command-router.cts
  * routeAuditUat: routes through the Hub with a synthetic 'run'
    defaultSubcommand (no real subcommands). Gives uniform observability.
  * routeAuditOpen: captures --json in a closure, strips it from args
    before Hub dispatch (so it isn't mistaken for a subcommand by the
    manifest check), then branches on wantJson inside the handler to
    preserve the formatAuditReport success-path quirk.

docs/CONFIGURATION.md
  * Observability section: noted capability commands (graphify, intel,
    audit-uat, audit-open) now emit DispatchEvent records since #1646.

.changeset/capability-routers-via-hub.md
  * Changed fragment describing the user-visible audit-trail expansion.
    pr:0 placeholder will be backfilled after gh pr create returns the
    real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).

Verification
  * graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
    error path, JSON-errors, and registry assertions)
  * intel cutover tests: 39/39 pass
  * audit cutover tests: 24/24 pass
  * bug-974-graphify-budget-missing-value regression test: pass
  * npm run test:unit (full suite): 2384 tests, 0 fail
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.

* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
2026-06-23 23:21:06 -04:00
Tom Boucher
bcc5a6d1ba fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command

Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.

- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
  absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
  .sh and others keep the bare quoted path (unchanged).

Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.

Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.

* chore(#1634): backfill changeset pr:1638

* fix(#1634): resolve lint and windows CI failures

- validator: replace the control-character range regex with a char-code loop.
  The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
  codes are equally precise and lint-clean. Behavior unchanged (still rejects
  matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
  honor POSIX write modes (a 0o644 write reads back as 0o666), so the
  precondition is meaningless there and failed the windows-latest lane. The
  node-prefix assertion — the actual fix — is platform-independent and still
  runs everywhere.

* docs(#1634): amend ADR-894 for optional lifecycle hook matcher

The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.

* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md

Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
2026-06-23 21:49:43 -04:00
Tom Boucher
61be100a02 docs(#1642): amend ADR-0174 §5 — reconcile Result type + add exitReason? (#1643)
Two changes to ADR-0174 §5 (Sync dispatch with tight-typed Result<T>):

1. Reconcile the documented Result type to the as-built code. The
   original §5 text planned 'Unknown' / 'BadArgs' / 'ValidationFailed'
   / 'NotImplemented' / 'HandlerFailed'; the SDK retirement migration
   kept the ADR-0012 names (UnknownCommand / InvalidArgs / HandlerFailure)
   and never added the planned ValidationFailed or NotImplemented
   variants. HandlerRefusal was added during implementation but never
   back-filled into this ADR. The ADR now documents what consumers
   actually depend on.

2. Add the optional exitReason?: string field on the InvalidArgs variant
   (and update the makeInvalidArgs factory signature). This carries an
   ERROR_REASON enum value separately from the existing reason
   explanation text, so capability routers migrating from direct
   error(msg, ERROR_REASON.USAGE) calls to makeInvalidArgs(...) Results
   preserve ERROR_REASON granularity through the Hub Result →
   error(msg, exitReason) translation. The field is additive and
   backward-compatible.

Also tightens the amendment requirement: 'Adding a new variant OR
adding a field to an existing variant requires amending this ADR.'

Phase 0 of parent #1641. No code changes; pure ADR amendment.
2026-06-23 21:42:16 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Tom Boucher
f9d9dfb4bc fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.

- New reference gsd-core/references/security-asvs-levels.md defines L1
  (opportunistic), L2 (standard), L3 (comprehensive) for both planner
  threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
  hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
  L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
  by extracting the goal-backward worked example to planner-guidance.md.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 18:48:58 -04:00
Tom Boucher
481ca1b18e Merge origin/next into fix/1615-workflows-agents-not-installing-for-wind
Resolved conflict in docs/reference/skill-mapping-matrix.md: kept next's Antigravity row (flat layout per #1614, doc-fixed by #1617) AND kept HEAD's Windsurf row (workflows layout per #1615). Also updated the Structural Facts section counts that the original #1615 PR left stale: '15 skill-bearing runtimes' → '14' (Windsurf no longer skill-bearing), 'Nine runtimes stay flat' → 'Eight', removed windsurf from the 'FLAT (unconfirmed)' row in the loader-verification table.
2026-06-23 15:01:16 -04:00
Tom Boucher
e32eac56a6 docs(#1618): correct Claude skill layout from nested to flat 2026-06-23 14:11:22 -04:00
Tom Boucher
08dfcad5f1 docs(#1617): correct Antigravity skill layout from nested to flat 2026-06-23 14:05:49 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
c28cccbf85 Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
2026-06-23 10:56:29 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Tom Boucher
1985d4a4de docs(#1603): add opencode host-plugin binding to adr-1239 2026-06-22 22:29:48 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Tom Boucher
d108932136 docs(#1600): Phase C+D — per-platform plugin skill model assessment (all N/A)
Resolves TBD entries in the skill-mapping matrix provision/consumption
table. All 15 non-Claude platforms assessed as N/A — none have a
documented plugin skill-provision model. File-copy install path
handles skill provision for all runtimes. Updates claude row to
reflect Phase B-provide merged status. Closes epic #1258 acceptance.

Closes #1600
2026-06-22 19:22:18 -04:00
Tom Boucher
da4a86d8c1 feat(#1596): ship GSD skills via .claude-plugin/plugin.json
Phase B-provide of epic #1258. Adds a build-generated skills/ dir +
a skills manifest field so plugin-installed GSD exposes gsd-core:<skill>
the native Claude Code way. Closes the gap where plugin-only installs
lacked the skill surface because bin/install.js never ran.

- scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md
  to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill
- .claude-plugin/plugin.json: add "skills": "./skills/"
- package.json: add skills to files, gen:plugin-skills to build chain
- tests/issue-766-plugin-manifest.test.cjs: Section H conformance
  (manifest field + dir + frontmatter + count parity) + C2 skills symlink
- docs/adr/766-*.md: dated amendment adding skills surface row
- .changeset/rapid-bears-hum.md: type Added
- skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed)

Closes #1596
2026-06-22 18:07:35 -04:00
Tom Boucher
860179ee5d docs(#1593): ADR for cross-runtime skill mapping + converter methodology
Phase A of epic #1258. Adds the single authoritative ADR documenting
the per-runtime skill mapping + converter transform-contract catalog
that ADR-3660 (layout) and ADR-1508 (module ownership) each carry a
third of. Ships a companion reference matrix projecting all 16
runtimes from their capability descriptors.

- docs/adr/1593-skill-mapping-converter-methodology.md (new ADR)
- docs/reference/skill-mapping-matrix.md (new reference page)
- docs/adr/README.md (index row + status fixes: 3660, 1016 Proposed->Accepted)
- docs/adr/1016-runtime-capability-descriptor.md (header Proposed->Accepted;
  the ConverterName enum is already code-enforced per ADR-857 phase 5e)

Doc-only. No code, no behavior change. Closes #1593.
2026-06-22 16:27:41 -04:00
Tom Boucher
3a06b4888a Merge branch 'next' into feat/1173-wire-agent-converters-descriptor 2026-06-22 13:04:46 -04:00
Joe
e12a2abfd8 feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit

Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on

- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
  seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
  slug, status, scope, trigger_when, planted, title. Optional case-insensitive
  status filter. User-controlled content is sanitized (sanitizeForDisplay) and
  every path validated (requireSafePath); read-only. Independent of
  audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
  renders the seed table.

Closes #441

* chore(#441): point changeset fragment at PR #722

* test(#441): allowlist list-seeds test in prompt-injection scan

The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.

* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow

Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.

* docs(#441): sync help full.md + INVENTORY for --list-seeds

Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.

* docs(#441): add --list-seeds how-to + drop phantom statuses

Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):

- USER-GUIDE.md Seeds section (how-to): extend the task to cover
  auditing parked seeds on demand via --list-seeds, including the
  status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
  from the list-seeds filter vocabulary; the system only produces
  dormant|active|triggered (src/audit.cts scanSeeds). Reference must
  be factually accurate and complete.

* fix(#441): guard non-scalar status frontmatter in cmdListSeeds

A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.

Adds regression coverage for empty and array `status:` and non-scalar fields.

Refs #441

* docs(#441): align list-seeds workflow status vocabulary

The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.

Refs #441

* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds

Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.

Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).

* test(#441): add fast-check property coverage and count=1 boundary for list-seeds

Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).

Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).

* chore(#441): sync runtime launcher snippet into list-seeds workflow

Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.

* test(#441): record list-seeds.md in workflow size baseline (#1074)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:59:41 -04:00
Behruz Nassre Esfahani
ac40f070ef feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source

/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.

Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): add changeset for reviewer source-grounding

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt

Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1318): harden build_prompt fence extraction to be fence-run-aware

Addresses maintainer review on PR #1421 (required-before-merge).

The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.

Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.

Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:47:49 -04:00
Tom Boucher
405ae9b3b7 refactor(#1557): add runtime artifact install plan module (#1560) 2026-06-21 20:40:13 -04:00
Rezolv
195d356d7c Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio 2026-06-21 17:41:37 -04:00
Behruz Nassre Esfahani
faac9331f2 feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298
2026-06-21 15:38:44 -04:00
Tom Boucher
21fe9d3627 chore(#1544): test-quality + doc cleanups from the #1507 conversion-module review (#1547)
Follow-up hygiene from the #1507 epic review (release blocker fixed in #1537):

- path-replacement.test.cjs: delete the hand-reimplemented computePathPrefix copy
  (ADR-1508 said to); route all cases (incl. Windows + outside-home) through the
  real _computePathPrefix so the test can no longer drift from the function.
- enh-1511: add an isWindowsHost no-op tripwire characterization test.
- enh-1511: add a deterministic (fs-method monkeypatch, root/OS-independent)
  error-path test for applyRuntimeContentRewritesForCommandsInPlace — asserts the
  temp dir is rm'd on a read failure with no orphaned gsd-cmd-rewrites-* leak.
- runtime-artifact-conversion.cts: @internal note on rewriteStagedCommandBodies
  (deep-seam companion to rewriteStagedSkillBodies; no production caller today).
- ADR-1508 + CONTEXT.md: qualify the "single owner" claim with the deliberate
  opencode/kilo applyOpencodeFamilyPathPrefix pre-conversion carve-out (#784).

No user-facing behavior change (tests + internal docs + one code comment).
Refs #1507.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-21 15:20:46 -04:00
Behruz Nassre Esfahani
2b6107d46f refactor(#1173): defer agents-kind descriptor declarations (option a)
Adopt maintainer-recommended option (a): keep the convertedAgentsKind /
stageAgentsForRuntimeWithConverter scope-threading plumbing, but DEFER the
8 runtimes' capability.json `agents`-kind declarations to a follow-up that
first ships the ADR-1235 §0 byte-for-byte parity harness.

The declarations were a live regression: the second `layout.kinds` consumer,
applySurface / `/gsd:surface` / `--materialize` (src/surface.cts), does not
mirror the legacy agent pipeline (copilot `.agent.md` rename, path-prefix
rewrite + attribution, stale cleanup), so a `/gsd:surface` toggle deleted
installed copilot `gsd-*.agent.md` and path-unrewrote the other 7 runtimes.
trek-e + davesienkowski both flagged this.

- Revert the agents-kind entries from the 8 capability.json files and
  regenerate capability-registry.cjs (now matches next; 0 converted agents
  kinds declared).
- Revert the declaration-driven kind-count test bumps
  (runtime-artifact-layout, descriptor-drive, bug-782, enh-789, enh-790).
- Keep the synthetic-descriptor seam tests for convertedAgentsKind dispatch;
  add a synthetic scope-threading test so the kept isGlobal plumbing stays
  covered without depending on real declarations.
- Update the convertedAgentsKind doc comment to state declarations are
  deferred pending the ADR-1235 §0 parity harness.
- Move ADR-1235 to Accepted.
- Reword the changeset to plumbing-only (docs-exempt now honest: no runtime
  declares the kind, legacy loop authoritative, installed output unchanged).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:29:50 -07:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Dave
c09c13f295 enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).

Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.

Closes #1346

Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
2026-06-21 10:44:48 -04:00
Tom Boucher
b65892e6a2 docs(#1508): add ADR for Runtime Artifact Conversion Module content-rewrite ownership
Phase 0 of epic #1507. Records the decision to make the Runtime Artifact
Conversion Module the single owner of per-runtime content rewriting, flip the
dependency direction to installer/layout -> conversion, and close the
surface.cts -> bin/install.js getInstallExports relay. Doc-only.

Resolves ADR-3660 Initial-Scope deferral; distinct from epic #1258.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
2026-06-20 20:54:32 -04:00
Tom Boucher
b63a500fad fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to
  gsd-core/references/execute-phase-context-guard.md (@-ref lazy load),
  bringing execute-phase.md back under the ADR-857 phase-6 size ceiling
  (92914 < 93166 bytes)
- Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule
- Update feat-1452 tests to check the reference file for extracted content
- Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and
  INVENTORY.md Workflow References section
- Regenerate workflow-size-baseline.json after file shrinkage

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:38:20 -04:00
Tom Boucher
a77f3c4b3a docs(#1452): document workflow.context_guard_mode in CONFIGURATION.md
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:12:03 -04:00
Tom Boucher
c330f70f65 feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command and plan_chunked in planning-config.md (#1500)
* feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command, plan_chunked, mvp_mode in planning-config.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: backfill PR number 1500 in changeset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 15:32:07 -04:00
Tom Boucher
f5276b36b3 fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase (#1492)
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase

Two compounding issues caused wave N+1 worktrees to fork from the stale
pre-wave-N commit, immediately tripping the worktree_branch_check FATAL
guard in every executor:

1. worktree.base-check auto-degrade only ran once at initialize time.
   After wave N merges advanced orchestrator HEAD past origin/HEAD, new
   worktrees were still forked from origin/HEAD (Claude Code "fresh" base).

2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1
   reused the consumed wave-N manifest file, which would have blocked the
   step 5.5 manifest guard (#3384) on subsequent waves.

Fix: add two safeguards in execute-phase.md —
- Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades
  USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD.
- Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1
  creates a fresh per-wave manifest; re-asserts worktree.set-baseref
  (idempotent) and re-evaluates base degradation after wave merges land.

17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): rename test to fix-NNN convention; update workflow size baseline

Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs
to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix).

Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393
(LF-normalized byte count after adding step 0.5 inter-wave base re-check and
step 7c between-wave manifest reset).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap

Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave
dependency check + between-wave manifest reset/base refresh) added by this
PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6
architectural mandate that host-loop bodies remain strictly below the
pre-phase-6 baseline of 93166 bytes.

Extract both new blocks into dedicated reference files:
- gsd-core/references/execute-phase-wave-guard.md (step 0.5)
- gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c)

Replace inline prose with @-reference pointers. File now measures 92851
bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate.

Also update tests/workflow-size-baseline.json to 92851 and add both new
reference files to docs/INVENTORY-MANIFEST.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): update regression tests to read from extracted reference files

Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857
size cap on execute-phase.md. Tests now check @-reference pointers in the
workflow for ordering and read content assertions from the reference files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 14:22:59 -04:00
Tom Boucher
3a3b2135c2 chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.

Closes #1073
2026-06-20 13:36:57 -04:00
Tom Boucher
02c6491e61 docs(#1464): fix ADR-1244 capability doc set — followable tutorials, overlay-model + install tutorial, set/fragment/runtimeCompat reference, accuracy fixes 2026-06-20 13:02:24 -04:00
Tom Boucher
7c93d9e222 feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 10:50:12 -04:00
Tom Boucher
08d1c57d6e fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 09:10:04 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
81d4b6d573 docs(#1010): document projects-sync as the first reference third-party capability (#1470)
Add a worked example to the URL-import how-to showing how to install and
enable the projects-sync ADR-1244 ecosystem capability (published at
The-Artificer-of-Ciphers-LLC/projects-sync-capability). The capability ships
as an external installable repo; GSD Core gains the worked example only.

Closes #1010
2026-06-19 21:03:26 -04:00
Tom Boucher
c866ac1b24 fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) 2026-06-19 19:55:10 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
34bc096ec2 feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI

ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs
install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a
user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the
six subcommands, dispatching to the existing lifecycle/ledger:

- install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]…
- update [<id>|--all] [--scope] [--yes] [--shared-file]  (re-resolves recorded source)
- remove <id> [--purge-data] [--scope]  (first-party rejected)
- list [--json]  (first-party + overlay, both scopes, JSON array)
- disable|enable <id>  (activation-state alias of capability set --off/--on)

Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home,
project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json).
Consent is non-interactive: --yes grants; without it an executable install aborts after
printing the disclosure and writes nothing. Best-effort reconcile before each mutation.

Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs,
GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip,
remove round-trip + first-party guard, disable/enable, unknown subcommand.
Docs: docs/reference/gsd-capability-command.md reconciled to the real surface
(ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned);
docs/COMMANDS.md gains the gsd capability entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug

Adversarial-review (Codex) fixes:
- capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate
  fail-closes on it (was silently downgrading to permissive)
- installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection
  (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't
  act on a different id if the recorded source was retargeted
- capability update: prints the consent disclosure, exits non-zero on --all partial failure,
  no longer masks the resolved id
- capability remove: ledger-first ordering so an overlay is removable even if it shadows a
  first-party name; first-party guard only fires for ids not in the ledger
- gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle
  not yet wired through this path)

Silent-output bug (root cause, not waved off as pre-existing):
- captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw
  command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of
  stdout. Now it flushes the captured buffer before re-throwing (exit code preserved).
- cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws
  ExitError so the wrapper flushes — matches the repo's no-process-exit architecture.
- Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout.

Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed

- confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors
  safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it.
- mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or
  another capability's) server is skipped, so install/remove can't silently clobber user MCP config
  (hooks already append; the map-keyed mcpServers path was the gap).
- capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead
  of silently downgrading the strict_known_registries policy to permissive.
- Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry
  preserved; unparseable config blocks an external install. capability suite 83/83, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc

- install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never
  fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per
  the lifecycle contract; latent today, hardened for future status additions).
- Clarify capResolveScope comment (project scope === already-resolved cwd) and document that
  strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide
  allowlist) in gsd-capability-command.md.
- Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that
  looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1451): backfill changeset PR number → #1457

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:02:36 -04:00
Tom Boucher
1abebbf4fd feat(#1434): registry-driven dispatch for third-party capabilities (ADR-1244 Phase 5) (#1450)
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.

Closes #1434.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:37:43 -04:00
Rezolv
dcceb1a004 fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing

resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first
bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also
had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy
dir (probed first). Regression from #217 — pre-#217 returned
~/.gemini/antigravity unconditionally.

Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested
descriptor: a two-pass probe prefers the candidate GSD installed into, then
bare existence, then probe[0]. Behavior is byte-identical when probeMarker is
absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for
installer/operator guidance on already-misinstalled users (auto-relocation
ruled out per ADR-0008's single-configDir migration bound).

Regression tests fail before / pass after: coexistence + marker-priority
cases, end-to-end through the registry descriptor.

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB

* chore(#1441): add changeset for antigravity resolver fix

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
2026-06-18 21:47:18 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
1abb0d3427 feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3)

ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability
install command + consent gate are Phase 4). NEW modules only; not yet wired into
bin/install.js.

- src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one
  adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/
  git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install),
  tarball (https download + sha512 integrity-before-extraction + tar), registry (stub).
  SECURITY: install never executes capability code (copy/extract only, --ignore-scripts);
  integrity verified before staging; symlink members rejected at interior, source-root,
  AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction;
  npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell);
  git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging
  (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd
  pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet.
- src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest:
  readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent,
  prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against
  hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir).
- 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex
  adversarial rounds; all 12 findings fixed + regression-tested.
- .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored
  (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen.

Closes #1432

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1432): add changeset for capability source resolver + ledger

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 16:23:20 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
2421cf1b4a feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1)

Make the capability manifest versioned — the data substrate the Capability
Ecosystem (ADR-1244) keys off:

- capability.json gains a REQUIRED semver `version` plus the optional
  ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`,
  `provenance`); the build-time conformance validator enforces them via a new
  `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay).
- All 32 native capabilities stamped with `version` (= package version,
  lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep
  that keeps them in sync, and the issue-844 regression guard is extended.
- Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in
  version strings; range/integrity fields are shape-validated (satisfaction
  and the load-time gate are deferred to Phase 2/4).
- Capability rel-paths emitted forward-slash for cross-platform git correctness.

Closes #1430

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1430): add changeset for versioned capability manifest

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:45:53 -04:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
6f2547610e docs(#1412): ADR-1411 — Resolution must report provenance, not fall open silently (#1413)
Decision record for the Resolution Provenance epic (#1411, P0). Establishes that
context resolution (config loading, project-root anchoring, workstream resolution)
must report its provenance rather than fall open silently to defaults — the
resolution-side analog of ADR-227. Binds the Config Loader Module, Project-Root
Resolution Module, and I/O Module. Adds the ADR, the index/seam-map entry, and the
CONTEXT.md glossary term.

Part of #1411
Closes #1412

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 23:42:17 -04:00