diff --git a/.changeset/1143-claude-orchestration-capability.md b/.changeset/1143-claude-orchestration-capability.md new file mode 100644 index 000000000..d87aae625 --- /dev/null +++ b/.changeset/1143-claude-orchestration-capability.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 2044 +--- +**A default-off, BETA, claude-only "Claude orchestration" capability** — adopts Claude Code's Workflow tool (`/effort ultracode`, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, restoring the wave parallelism + plan-checker + verifier that the #853 backgrounded-agent nesting limitation forces inline on Claude Code, and folding the existing `gsd-ultraplan-phase` plan-offload under the same runtime gate. When `claude_orchestration.enabled` is on AND the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets the floor (`claude_orchestration.min_agent_sdk_version`, default `0.3.149`), `execute-phase` emits a generated Workflow script (`waves → parallel() barriers`, `plans → agent({ agentType: 'gsd-executor', isolation: 'worktree' })`, `files_modified overlap → separate sequential stages`, `resumeFromRunId` wired to the phase run id, shared `budget` pool) that composes the SAME executor agent + worktree isolation the inline path uses, so artifacts/commits are produced identically. Detection is pure and fail-closed (any miss → inline), so on any runtime lacking the Workflow tool behaviour is byte-identical to today. Adds a pure module `gsd-core/bin/lib/claude-orchestration.cjs` (`detectWorkflowBackend`, `emitWorkflowScript`), the `capabilities/claude-orchestration/` declaration with two gated loop contributions (`execute:wave:post`, `plan:post`) and a `claude-orchestration` command family (`gsd-tools claude-orchestration detect-backend|emit-workflow`), federated config keys, and an ADR-1143 implementation amendment. (#1143) diff --git a/.changeset/gentle-badgers-roar.md b/.changeset/gentle-badgers-roar.md new file mode 100644 index 000000000..a20eb5765 --- /dev/null +++ b/.changeset/gentle-badgers-roar.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 2048 +--- +**`model_overrides` Claude model IDs now resolve to Agent-tool aliases on the claude runtime** — a full Claude model ID (e.g. `claude-sonnet-5`) in `model_overrides` was returned verbatim and silently dropped by the Claude Agent tool (whose `model` parameter documents only tier aliases), causing the spawned subagent to inherit the parent session model instead of the configured one. It now maps to the tier alias (`sonnet`/`opus`/`haiku`/`fable`), consistent with the `model_policy` path (#1144). Bare aliases, non-Claude values, and non-Claude runtimes are unchanged; a Claude ID with no alias warns once and falls through to tier resolution. (#2041) diff --git a/.changeset/plucky-jays-dart.md b/.changeset/plucky-jays-dart.md new file mode 100644 index 000000000..0a183d3c1 --- /dev/null +++ b/.changeset/plucky-jays-dart.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 2040 +--- +**`/gsd:surface` and `--materialize` now produce byte-identical agent output to a fresh install** — surface-path agents for descriptor-driven runtimes (cursor, windsurf, augment, trae, codebuddy, copilot, antigravity) now receive the same path-prefix rewrite, Co-Authored-By attribution, runtime-specific conversion, and body normalization as the install path. Copilot and Antigravity agents are now installed via the descriptor-driven path (copilot agents get the `.agent.md` filename rename). Cline remains on the inline loop (rules-only local branch). (#1575) diff --git a/.changeset/steady-ibex-run.md b/.changeset/steady-ibex-run.md new file mode 100644 index 000000000..bf1b37a3b --- /dev/null +++ b/.changeset/steady-ibex-run.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 2049 +--- +**Skill-bearing capabilities now surface correctly on flat command-layout installs** — on an install using the flat `commands/gsd-.md` source layout (e.g. a Claude Code local project install with no `commands/gsd/` subdir), every skill-bearing capability (`nyquist`, `code-review`, `security`, `ui`, `mempalace`, `ai-integration`, `profile-pipeline`) was silently reported `surfaced:false`/`enabled:false`/`active:false`, so their loop hooks (`verify:post`, `execute:post`, etc.) never fired even with the corresponding `workflow.*` toggle on. The skill-manifest resolver now detects the flat layout and produces the same stems the nested `commands/gsd/*.md` loader does. (#1858) diff --git a/CONTEXT.md b/CONTEXT.md index 0b191c568..02e92c938 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -220,6 +220,8 @@ ADR-1244 Phase 4 (D5+D6) orchestration seam (`gsd-core/bin/lib/capability-lifecy ### Capability Command Dispatch ADR-1244 Phase 5 (D7) registry-driven dispatch of capability command families. First-party families (`graphify`/`intel`/`audit`, shipped in `bin/lib/`) dispatch via `dispatchCapabilityCommand` (`gsd-core/bin/gsd-tools.cjs`) against the FROZEN `capability-registry.cjs` `commandFamilies` (confined to `bin/lib/`) — unchanged. Third-party (installed overlay) families dispatch via `dispatchOverlayCapabilityCommand`: after the first-party path returns false, it calls `loadRegistry({ includeInstalled, cwd })` and dispatches a family iff its `capId` is in `_overlay.commandRoots` — which `capability-loader.cjs` populates ONLY for accepted overlay capabilities that declare `commands` AND pass the loader's activation gate (a **committed** ledger entry, present and non-`_pending`, PLUS — for PROJECT scope — a matching user consent record in the Capability Consent Store; GLOBAL scope needs no consent record). A bundle dropped on disk with no install (no ledger entry) or no on-this-machine consent is NOT command-dispatchable. The router module is `require()`'d FROM the capability's install root via `defaultRequireFromInstallRoot` (bare-`.cjs` basename + `realpath` containment, rejecting `..` traversal and symlink escape); same own-property/function/sync-only guards as the first-party path. Wired into the `runCommand` default arm before "Unknown command". A repo-planted project ledger no longer activates anything on its own (#1459) — see `docs/explanation/capability-trust-model.md` "project-scope trust boundary". +### Claude Orchestration Capability +Default-off, BETA, claude-only Capability (`capabilities/claude-orchestration/`, `role: feature`, `runtimeCompat.supported: ["claude"]`, `tier: full`, `activationKey: claude_orchestration.enabled`) adopting Claude Code's Workflow tool (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, and folding the `gsd-ultraplan-phase` plan-offload under the same runtime gate (#1143; ADR-1143). Pure, fail-closed core in `gsd-core/bin/lib/claude-orchestration.cjs` (generated from `src/claude-orchestration.cts`): `detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) → { available, backend:'workflow'|'inline', reason }` (gate ladder: enabled → Claude → execution_backend ≠ inline → host dispatch nested+background → valid Agent SDK → SDK ≥ floor; every miss degrades to `inline`, never throws); `emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) → { ok, script, summary }` mapping waves → `parallel()` stage barriers, plans → `agent({ agentType:'gsd-executor', isolation:'worktree' })`, `files_modified` overlap → separate sequential stages (greedy first-fit), `resumeFromRunId` wired to the run id, shared `budget(tokens)`; all interpolated identifiers validated script-safe (no `"`,`\`,control chars) and briefs JSON-quoted (review anti-injection). Registers two loop contributions at WIRED points only (execute:wave:pre/execute:pre are declared but not rendered, same constraint external-job documents): `execute:wave:post into:executor` (Workflow-backend guidance) and `plan:post into:planner` (ultraplan ownership declaration), both `when: claude_orchestration.enabled`, `onError: skip`. Federated config keys (`claude_orchestration.enabled` default false, `execution_backend` enum auto|workflow|inline default auto, `min_agent_sdk_version` string default "0.3.149") live only in the registry — uninstall removes them cleanly. Pre-release versions of the floor compare below GA (SemVer precedence). Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code; on any runtime lacking the Workflow tool, behaviour is byte-identical to today. BETA v1 ships detection + emission + declarative ultraplan ownership + a `claude-orchestration` command family (`gsd-tools claude-orchestration detect-backend|emit-workflow`, router `gsd-core/bin/lib/claude-orchestration-command-router.cjs` from `src/claude-orchestration-command-router.cts`); full install-profile migration of the ultraplan skill into `skills[]` is a follow-up (CLUSTERS/profile gate). Test anchors: `tests/claude-orchestration.test.cjs`, `tests/claude-orchestration-command-router.test.cjs`. ### Loop Extension Point A named, stable site on a host loop step (per-step `pre`/`post` plus per-wave in Execute; 12 total) where Capabilities register hooks. Three hook kinds: `step` (runs as its own sequenced unit), `contribution` (injects into the core step's prompt/context), and `gate` (checks and optionally blocks via a declared `blocking` flag). Each hook declares the artifacts it produces and consumes; hook order is derived by topological sort of that produces/consumes graph (capability-id tiebreak), which also defines data flow — file-artifact based, surviving `/clear` and fresh executor contexts. Hooks are surfaced by runtime resolution with concrete projection: the workflow calls a query that resolves the active hooks and returns fully-rendered, ordered markdown for the executor. Failure is default-resilient — a non-gate hook that errors is skipped with a warning; a hook may opt into `onError: halt`. Part of the Capability system. ADR-857 phase 3c ships the registry-consuming query layer: `gsd-core/bin/lib/loop-resolver.cjs` exposes `resolveLoopHooks({ point, registry, config })` (pure, no I/O), `renderLoopHooks(resolved)` (pure markdown renderer), and `cmdLoopRenderHooks(cwd, point, raw, opts)` (I/O entry point); activated via `gsd-tools loop render-hooks ` which emits `{ point, activeHooks[], rendered }`. Activation is driven by `when` (dotted config key resolved against `loadConfig`), with inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard. The first phase-6 cutovers wiring workflows to this query have landed — ui-phase at `plan:pre` and ui-review at `verify:post` (in `plan-phase.md`/`autonomous.md`); further per-feature cutovers are ongoing. diff --git a/bin/install.js b/bin/install.js index 1159443e9..589ca0218 100755 --- a/bin/install.js +++ b/bin/install.js @@ -8992,9 +8992,11 @@ function install(isGlobal, runtime = 'claude', options = {}) { // (by installRuntimeArtifacts at line 8912), which also performs its own // stale-file prune pass. The inline stale-removal + inline loop both skip them. // Trivial group (cursor/windsurf/augment/trae/codebuddy) cut over together. - // cline is excluded: it takes a rules-only local branch and has a local/global - // complication that the descriptor-driven path does not handle correctly. - const _DESCRIPTOR_AGENTS_RUNTIMES = new Set(['cursor', 'windsurf', 'augment', 'trae', 'codebuddy']); + // #1575: copilot and antigravity cut over — copilot gets .agent.md filename + // rename via _copyStaged(runtime); antigravity uses scope-aware converter. + // cline remains excluded: rules-only local branch + local/global complication + // that the descriptor-driven path does not handle correctly. + const _DESCRIPTOR_AGENTS_RUNTIMES = new Set(['cursor', 'windsurf', 'augment', 'trae', 'codebuddy', 'copilot', 'antigravity']); // Always remove stale gsd-* agents first so re-installing with // `--minimal` actually shrinks a previously-full install. diff --git a/capabilities/antigravity/capability.json b/capabilities/antigravity/capability.json index 3a47cef3d..96b69c82f 100644 --- a/capabilities/antigravity/capability.json +++ b/capabilities/antigravity/capability.json @@ -35,6 +35,14 @@ "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ], "local": [ @@ -45,6 +53,14 @@ "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ] }, diff --git a/capabilities/claude-orchestration/capability.json b/capabilities/claude-orchestration/capability.json new file mode 100644 index 000000000..95474d492 --- /dev/null +++ b/capabilities/claude-orchestration/capability.json @@ -0,0 +1,85 @@ +{ + "id": "claude-orchestration", + "role": "feature", + "version": "1.7.0-rc.3", + "title": "Claude orchestration (Workflow backend)", + "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", + "tier": "full", + "requires": [], + "engines": { + "gsd": ">=1.7.0" + }, + "runtimeCompat": { + "supported": [ + "claude" + ], + "unsupported": [] + }, + "skills": [], + "agents": [], + "hooks": [], + "commands": [ + { + "family": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand", + "subcommands": [ + "detect-backend", + "emit-workflow" + ] + } + ], + "activationKey": "claude_orchestration.enabled", + "config": { + "claude_orchestration.enabled": { + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "type": "enum", + "values": [ + "auto", + "workflow", + "inline" + ], + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability." + }, + "claude_orchestration.min_agent_sdk_version": { + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + } + }, + "steps": [], + "contributions": [ + { + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, + { + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + } + ], + "gates": [] +} diff --git a/capabilities/claude-orchestration/fragments/execute-wave-post.md b/capabilities/claude-orchestration/fragments/execute-wave-post.md new file mode 100644 index 000000000..db0e76d5a --- /dev/null +++ b/capabilities/claude-orchestration/fragments/execute-wave-post.md @@ -0,0 +1,64 @@ +# Claude orchestration — Workflow execution backend (BETA) + +> Injected at `execute:wave:post` `into: executor` only when +> `claude_orchestration.enabled` is true. Default-off; `onError: skip`. + +## When this contribution is active + +The Claude orchestration capability is **default-off and BETA**. It activates only +when ALL of the following hold: + +1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND +2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent + SDK-specific), AND +3. `claude_orchestration.execution_backend` resolves to `workflow` — either + explicitly, or via `auto` — **and** the Agent SDK version is + `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK + floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release + or older SDK never activates the preview backend). + +Detection is fail-closed: any miss degrades to **inline, manual, one-agent-per- +message dispatch** — exactly today's behaviour. On a non-Claude runtime this +contribution is a no-op. + +## What the executor does when the Workflow backend is active + +Instead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor, +isolation=worktree, run_in_background=true)` per message (which on Claude Code +cannot nest further subagents — #853 — and so degrades to sequential inline +execution), execute-phase **emits a generated Workflow script** and lets the main +loop orchestrate it: + +- **waves → one or more sequential `parallel()` barriers** — each wave is a + barrier group; when plans within a wave share `files_modified`, they are split + into separate sequential stages within that wave's barrier (the next wave + still waits for the previous wave to complete). +- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`** + — the SAME executor agent and worktree isolation the inline path uses, so the + produced `SUMMARY.md` and commits are identical. +- **`files_modified` overlap → separate sequential stages** — two plans that + touch the same file are placed in different stages within the wave (the same + overlap rule execute-phase already applies inline). +- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase + resumes without re-running completed plans. +- **`budget(tokens)`** — a shared token pool across the whole phase when the + orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a + function parameter, not a config key; the orchestrator decides the budget). + +The emitter is a pure function exposed through the capability command surface: +`gsd-tools claude-orchestration emit-workflow --waves --run-id +[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript` +directly). It maps the phase's wave/plan manifest to the Workflow script string +and never invokes the Workflow tool itself; the orchestrator runs the emitted +script. Detection is resolved by the orchestrator calling the pure +`detectWorkflowBackend` with the LIVE host descriptor (the CLI +`gsd-tools claude-orchestration detect-backend` is a simulation harness that +assumes a capable host unless `--no-nested-dispatch` is passed — it does not probe +the real runtime; the orchestrator supplies the real descriptor). + +## Fallback contract + +If detection resolves to `inline` (tool absent, SDK too old, runtime not Claude, +or the capability disabled), execute-phase MUST proceed with the standard inline +wave dispatch. The executor MUST NOT assume parallelism, a shared budget, or +resume-from-run-id semantics in that mode. diff --git a/capabilities/claude-orchestration/fragments/plan-post.md b/capabilities/claude-orchestration/fragments/plan-post.md new file mode 100644 index 000000000..bec9b01fa --- /dev/null +++ b/capabilities/claude-orchestration/fragments/plan-post.md @@ -0,0 +1,28 @@ +# Claude orchestration — ultraplan plan-offload ownership (BETA) + +> Injected at `plan:post` `into: planner` only when +> `claude_orchestration.enabled` is true. Default-off; `onError: skip`. + +## Ownership declaration + +The `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to +Claude Code's ultraplan cloud) is **owned by this capability**, not by a +standalone BETA skill. Both surfaces share one runtime gate +(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only +detection seam. + +## When the planner should consider ultraplan offload + +When this contribution is active (capability enabled, Claude Code runtime), the +planner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local +`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is +advisory, not mandatory — the stable local planner remains the default. + +## Fallback contract + +If the capability is disabled, or the runtime is not Claude Code, ultraplan +offload is **not surfaced** and the planner proceeds with the standard local +`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed +(its own runtime gate already no-ops on non-Claude runtimes); this contribution +only governs whether the capability manifest advertises it as part of the +orchestration surface. diff --git a/capabilities/copilot/capability.json b/capabilities/copilot/capability.json index aba2ab991..b4c5df53a 100644 --- a/capabilities/copilot/capability.json +++ b/capabilities/copilot/capability.json @@ -29,6 +29,14 @@ "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ], "local": [ @@ -39,6 +47,14 @@ "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ] }, diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 2c94b3bc9..0bc7c3b28 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -306,6 +306,8 @@ "capability-writer.cjs", "check-command-router.cjs", "cjs-command-router-adapter.cjs", + "claude-orchestration-command-router.cjs", + "claude-orchestration.cjs", "cli-exit.cjs", "cli-skew-check.cjs", "clock.cjs", diff --git a/docs/adr/1143-claude-orchestration-capability.md b/docs/adr/1143-claude-orchestration-capability.md index 128a40909..4541bce43 100644 --- a/docs/adr/1143-claude-orchestration-capability.md +++ b/docs/adr/1143-claude-orchestration-capability.md @@ -86,3 +86,39 @@ These existing multi-model features (`execute-phase` `cross_ai_delegation`, the - **Neutral:** no effect on non-Claude runtimes by construction; no behavior change until explicitly enabled. > **Governance note:** This ADR is a *draft design* accompanying feature request #1143. Per CONTRIBUTING, it is PR'd only after the issue receives `approved-feature`, and the capability is implemented only after #857 is released. + +## Amendment (2026-07-06): BETA v1 implementation landed + +#857 is **released** (CLOSED); the capability infrastructure is live. The BETA v1 +of this capability has shipped as `capabilities/claude-orchestration/` with the +scope agreed in the Decision, refined to the lowest-risk first slice: + +- **Detection + emission** live as pure, fail-closed functions in + `gsd-core/bin/lib/claude-orchestration.cjs` (source `src/claude-orchestration.cts`): + `detectWorkflowBackend` (gate ladder: enabled → Claude runtime → + execution_backend ≠ inline → host dispatch nested+background → valid Agent SDK + → SDK ≥ `claude_orchestration.min_agent_sdk_version`, default `0.3.149`) and + `emitWorkflowScript` (waves → `parallel()` stage barriers, plans → + `agent({ agentType: 'gsd-executor', isolation: 'worktree' })`, `files_modified` + overlap → separate sequential stages, `resumeFromRunId` wired to the phase run + id, shared `budget(tokens)` pool). All interpolated values are validated as + script-safe identifiers or JSON-quoted (review Finding 1). +- **Loop registration** is at the two **wired** points the loop host contract + actually renders: `execute:wave:post into:executor` (Workflow-backend guidance) + and `plan:post into:planner` (ultraplan ownership declaration). `execute:wave:pre` + and `execute:pre` are declared in the contract but **not wired** today, so the + capability registers at `wave:post` (the constraint `external-job` also documents). +- **Config** is federated (`claude_orchestration.enabled` default false / + `activationKey`, `execution_backend` enum `auto|workflow|inline` default `auto`, + `min_agent_sdk_version`); the keys live only in the registry, so uninstall + removes them cleanly. +- **ultraplan ownership** is declared in the manifest (`plan:post` contribution); + full install-profile migration of the `gsd-ultraplan-phase` skill into the + capability's `skills[]` is deferred to a follow-up (it triggers the CLUSTERS / + profile membership gate and is a heavier, install-machinery change). + +Status remains **Proposed** — the BETA is default-off and the end-to-end Workflow +execution path (actual orchestration via the Workflow tool inside Claude Code) is +not verifiable outside that runtime. The capability is structurally complete and +tested at the contract level; flipping to Accepted follows maintainer sign-off on +the E2E behaviour once exercised on Claude Code with the Workflow tool present. diff --git a/docs/adr/1235-descriptor-driven-agent-conversion-migration.md b/docs/adr/1235-descriptor-driven-agent-conversion-migration.md index efff11bc5..e41f10daf 100644 --- a/docs/adr/1235-descriptor-driven-agent-conversion-migration.md +++ b/docs/adr/1235-descriptor-driven-agent-conversion-migration.md @@ -80,6 +80,14 @@ Cross-cutting steps (a, b, c) are applied by the descriptor pipeline for the app 5. **Codex** — fold the `.toml` sidecar into the descriptor (or declare it an explicit companion artifact); the `.md` + `.toml` must both reach parity. 6. **Delete the inline loop** once every runtime is green; remove the now-dead `isKimi`/minimal special-casing that referenced it. +### Cutover progress (#1575) + +- **Step 0 (parity harness):** shipped in `tests/issue-1575-agent-descriptor-parity.test.cjs`. Asserts `applySurface` output is byte-identical to `installRuntimeArtifacts` for all descriptor-driven runtimes. Covers stale-cleanup convergence (pre-existing legacy `.agent.md` pruned correctly). +- **Step 1 (trivial converters):** cursor, windsurf, augment, trae, codebuddy — install-path cutover complete (PR #1438); surface-path parity shipped (#1575: `applySurface` now builds `agentCtx` and passes it to `kind.stage()` for agents, applying path-rewrite + attribution + converter + normalize). +- **Step 2 (scope-aware):** copilot and antigravity — cutover complete (#1575: declared `agents` kind in `capability.json`, added to `_DESCRIPTOR_AGENTS_RUNTIMES`, copilot `.agent.md` rename handled in both `_copyStaged` and `_syncGsdDir`). +- **Cline:** deferred — rules-only local branch + local/global complication not handled by the descriptor-driven path. +- **Remaining:** steps 3–6 (config-reading, no-converter, codex, inline-loop deletion). + ## Risks / trade-offs - **Silent install regression** across ~15 runtimes is the dominant risk; the byte-for-byte golden gate is the mitigation, and per-runtime sequencing bounds the blast radius of any single step. diff --git a/docs/explanation/claude-orchestration-capability.md b/docs/explanation/claude-orchestration-capability.md new file mode 100644 index 000000000..e044a4e1f --- /dev/null +++ b/docs/explanation/claude-orchestration-capability.md @@ -0,0 +1,94 @@ +# Claude orchestration capability (BETA) + +> **Explanation** — *why this capability exists and how it fits the loop.* For the +> step-by-step, see the [capability reference](../reference/capability-matrix.md); +> for the design record, see [ADR-1143](../adr/1143-claude-orchestration-capability.md). + +## The problem + +GSD's `execute-phase` is wave-based: plans carry a wave number, waves run +sequentially, and plans *within* a wave run in parallel when their +`files_modified` sets don't overlap. On most runtimes GSD realizes that by +fanning out one backgrounded `gsd-executor` agent (in a worktree) per plan. + +On **Claude Code** that fan-out degrades. Backgrounded agents on Claude Code have +no `Agent`/`Task` tool, so they cannot nest subagents ([#853]). The autonomous +loop therefore falls back to **inline sequential execution** — and with it +silently drops wave parallelism, the plan-checker, and the verifier — on the one +runtime most GSD users run. + +Claude Code ships an orchestration primitive that sidesteps exactly this: the +**Workflow tool** (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149). +A Workflow script *is* the orchestrator — it runs from the main loop and spawns +subagents itself via `agent()`, `parallel()` (barrier), `pipeline()`, and +`phase()`, with `isolation: 'worktree'`, a shared token `budget`, and +`resumeFromRunId`. + +## The capability + +`claude-orchestration` is a **default-off, BETA, claude-only** capability that +adopts the Workflow tool as an optional, runtime-gated parallel-execution +backend, and folds the existing `gsd-ultraplan-phase` plan-offload under the same +gate. It is blocked-on-nothing now that the ADR-857 capability system is released. + +- **`role: feature`**, `runtimeCompat.supported: ["claude"]`, `tier: full`. +- **`activationKey: claude_orchestration.enabled`** — default `false`. Nothing + changes until you opt in. +- Registers at two **wired** loop points: `execute:wave:post` (into the executor) + and `plan:post` (into the planner). Both are `onError: skip` and gated by the + `enabled` key. + +## How it decides whether to activate + +Detection is a pure, **fail-closed** function — `detectWorkflowBackend`. The +Workflow backend activates only when *every* gate passes; any miss degrades to +`inline` (today's behaviour): + +1. `claude_orchestration.enabled` is true. +2. The runtime is Claude (the Workflow tool is Claude / Agent SDK-specific). +3. `claude_orchestration.execution_backend` is `auto` or `workflow` (not `inline`). +4. The host descriptor advertises `dispatch.nested` **and** `dispatch.background` + (the nesting-capable Claude-Code shape — a proxy for Workflow-tool presence, + meaningful only after gate 2). +5. The Agent SDK reports a valid semver version. +6. That version is `>= claude_orchestration.min_agent_sdk_version` + (default `0.3.149`). A pre-release of the floor (e.g. `0.3.149-rc.1`) compares + *below* the GA release per SemVer, so the preview backend stays off. + +## What the executor runs when the backend is active + +`emitWorkflowScript` maps the phase's wave/plan model onto Workflow primitives: + +| GSD concept | Workflow primitive | +|---|---| +| Wave | `parallel()` stage barrier | +| Plan | `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })` | +| `files_modified` overlap | forces the plans into separate sequential stages | +| Phase run id | `resumeFromRunId("")` | +| Phase token cap | `budget()` | + +Because the emitted script composes the **same** `gsd-executor` agent and +**worktree isolation** the inline path uses, it produces the same `SUMMARY.md` +artifacts and commits — the only difference is the execution vehicle. + +## The fallback contract + +On any runtime lacking the Workflow tool — or when the capability is disabled, +the SDK is too old, or detection fails for any reason — execute-phase proceeds +with the standard inline wave dispatch. This is a release gate, not a nicety: a +regression test asserts the inline fallback on every non-capable combination, so +the capability is default-off and low-risk by construction. + +## BETA scope (v1) + +The first slice ships **detection + emission + declarative ultraplan ownership**. +The emitter is exercised at the contract level (structure, overlap splitting, +resume, budget, anti-injection). End-to-end execution through the Workflow tool +is verifiable only inside Claude Code with the tool present. Full install-profile +migration of the `gsd-ultraplan-phase` skill into the capability's `skills[]` +array is a follow-up (it touches the cluster/profile machinery); for v1 the +manifest *declares* ultraplan ownership at `plan:post` and the existing skill's +own runtime gate continues to no-op on non-Claude runtimes. + +[#853]: https://github.com/open-gsd/gsd-core/issues/853 +[#1143]: https://github.com/open-gsd/gsd-core/issues/1143 diff --git a/docs/how-to/enable-claude-orchestration-workflow-backend.md b/docs/how-to/enable-claude-orchestration-workflow-backend.md new file mode 100644 index 000000000..a8fea05e1 --- /dev/null +++ b/docs/how-to/enable-claude-orchestration-workflow-backend.md @@ -0,0 +1,171 @@ +# How to enable and use the Claude orchestration backend (BETA) + +Run GSD's execute-phase waves through Claude Code's Workflow tool (`/effort ultracode`, Agent SDK ≥ v0.3.149) instead of the default one-agent-per-message dispatch, and fold the `gsd-ultraplan-phase` plan-offload under the same gate. On Claude Code this restores the wave parallelism that backgrounded-agent nesting (#853) otherwise forces inline. + +> **BETA.** This capability tracks a Claude Code preview surface. It is default-off, fail-closed, and Claude-only. Every detection miss degrades silently to today's inline behaviour — enabling it can never break the loop. See the [explanation doc](../explanation/claude-orchestration-capability.md) for the why, and [ADR-1143](../adr/1143-claude-orchestration-capability.md) for the design. + +**What you need:** +- GSD installed with the `full` profile (the capability is `tier: full`). +- **Claude Code** with the Workflow tool available (Agent SDK ≥ `0.3.149`). On any other runtime the capability is an explicit no-op — you can flip the switch safely, nothing happens. +- A GSD project with at least one planned phase (you need a wave/plan manifest to emit a script for). + +--- + +## Step 1 — Enable the capability + +The capability ships disabled. Turn on the master switch inside your GSD project: + +```bash +gsd-tools query config-set claude_orchestration.enabled true +``` + +That single key gates everything — both the Workflow-backend hook at `execute:wave:post` and the ultraplan ownership declaration at `plan:post`. All other `claude_orchestration.*` keys are optional refinements. + +Verify it took: + +```bash +gsd-tools query config-get claude_orchestration.enabled +# → true +``` + +--- + +## Step 2 — Check whether your runtime qualifies + +Detection is fail-closed: the Workflow backend activates only when **every** gate opens. Before relying on it, confirm your runtime reports as capable: + +```bash +gsd-tools claude-orchestration detect-backend \ + --runtime claude \ + --agent-sdk-version 1.2.0 +``` + +You will get one of two results: + +| `backend` | `available` | Meaning | +|-----------|-------------|---------| +| `workflow` | `true` | Every gate passed — the emitter will produce a Workflow script the orchestrator can run. | +| `inline` | `false` | A gate failed. The `reason` field tells you which: `capability_disabled`, `runtime_not_claude`, `backend_inline`, `workflow_tool_unavailable`, `agent_sdk_version_unknown`, or `agent_sdk_version_below_floor`. | + +> **The CLI is a simulation harness, not a probe.** `detect-backend` assumes a capable host descriptor unless you pass `--no-nested-dispatch`. It exists so you (and the orchestrator) can ask "given these facts, would the backend activate?" The real detection the loop uses is the pure `detectWorkflowBackend` function, called with the live host descriptor. + +### If detection returns `inline` + +Work through the `reason`: + +- **`runtime_not_claude`** — you are on Codex / Cursor / opencode / etc. The Workflow tool is Claude-specific; there is nothing to enable here. Your loop is unchanged. +- **`agent_sdk_version_below_floor`** — upgrade Claude Code / the Agent SDK to at least `claude_orchestration.min_agent_sdk_version` (default `0.3.149`). A pre-release of the floor (e.g. `0.3.149-rc.1`) compares *below* the GA release and will not activate. +- **`workflow_tool_unavailable`** — your host descriptor does not advertise nested + background dispatch. This is unusual on Claude Code; if you see it, the Workflow tool is not present in this session. +- **`agent_sdk_version_unknown`** — the version could not be determined. Supply it explicitly via `--agent-sdk-version`. + +### Pin a higher floor (optional) + +If you want to gate the BETA behind a newer Agent SDK than the default: + +```bash +gsd-tools query config-set claude_orchestration.min_agent_sdk_version 1.0.0 +``` + +--- + +## Step 3 — Choose the execution backend + +`claude_orchestration.execution_backend` controls how aggressively the backend is used once detection passes: + +| Value | Behaviour | +|-------|-----------| +| `auto` (default) | Use the Workflow backend **if** detection passes; otherwise inline. The safe, recommended value. | +| `workflow` | Force the Workflow backend when the tool is present (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). | +| `inline` | Force today's manual one-agent-per-message dispatch, even on a capable Claude Code runtime. Use this to A/B compare or to temporarily retire the BETA. | + +Switch with: + +```bash +gsd-tools query config-set claude_orchestration.execution_backend workflow +``` + +--- + +## Step 4 — Emit a Workflow script for a phase + +With the capability enabled and detection passing, generate the Workflow script for a phase's wave/plan manifest. The manifest is the wave/plan model execute-phase already builds: + +```json +{ + "waves": [ + { + "id": "w1", + "plans": [ + { "id": "p1", "brief": "Implement the foo module", "files_modified": ["src/foo.cts"] }, + { "id": "p2", "brief": "Wire the bar seam", "files_modified": ["src/bar.cts"] } + ] + } + ] +} +``` + +Emit the script: + +```bash +gsd-tools claude-orchestration emit-workflow \ + --waves .planning/phases/01-foo/waves.json \ + --run-id phase-01-foo \ + --phase-dir .planning/phases/01-foo \ + --budget 500000 +``` + +The output is a generated Workflow script that maps GSD's model 1:1 onto Workflow primitives: + +- **waves → sequential `parallel()` barriers** (split into separate stages within a wave when `files_modified` overlap), +- **plans → `agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses, +- **`resumeFromRunId("")`** wired to the phase run id, +- **`budget()`** — a shared token pool across the whole phase (omit `--budget` to skip). + +Because the script composes the same `gsd-executor` agent + worktree isolation + `SUMMARY.md` artifact as the inline path, the artifacts and commits it produces are identical — only the execution vehicle differs. + +### Run the emitted script + +Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation). The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, waves barrier between each other, and `resumeFromRunId` lets an interrupted phase resume without re-running completed plans. + +--- + +## Step 5 — Ultraplan plan-offload + +Enabling the capability also folds `gsd-ultraplan-phase` under the same runtime gate. When the capability is on, the planner may offer the `/gsd-ultraplan-phase` path (offload plan-phase to Claude Code's ultraplan cloud) as an alternative to local `/gsd-plan-phase`. This is advisory — the stable local planner remains the default. + +If the capability is off, or the runtime is not Claude Code, ultraplan offload is not surfaced and `/gsd-plan-phase` runs as normal. + +--- + +## Disabling + +To turn the capability off and return to byte-identical inline behaviour: + +```bash +gsd-tools query config-set claude_orchestration.enabled false +``` + +Or force inline dispatch while leaving the capability otherwise on: + +```bash +gsd-tools query config-set claude_orchestration.execution_backend inline +``` + +Either step is sufficient — no uninstall or resurface needed. The federated config keys live only in the capability registry, so they vanish cleanly if the capability is ever removed. + +--- + +## What is and is not wired in BETA v1 + +**Working today:** +- Detection (`detectWorkflowBackend` / `gsd-tools claude-orchestration detect-backend`) — fail-closed, tested across every gate. +- Emission (`emitWorkflowScript` / `gsd-tools claude-orchestration emit-workflow`) — waves→barriers, overlap→stages, resume, budget, anti-injection. +- The contribution fragments at `execute:wave:post` and `plan:post` (gated, `onError: skip`). +- Inline fallback on every non-capable combination (regression-tested). + +**Not yet wired (follow-ups):** +- `execute-phase.md` does not yet auto-branch to emit-and-run the Workflow script. Today you emit the script explicitly (Step 4) and run it via the Workflow tool. Automatic dispatch inside the loop is the next milestone. +- The plan-checker and verifier still run inline — this capability delivers the parallel-execution backend, not those gates. +- Full install-profile migration of the `gsd-ultraplan-phase` skill into the capability's `skills[]` (it is currently declared in the manifest; the skill's own runtime gate continues to no-op on non-Claude runtimes). + +If a preview-API change breaks detection, the capability degrades to inline; it cannot destabilise the core loop. diff --git a/docs/reference/capability-matrix.md b/docs/reference/capability-matrix.md index 4efee626d..c74755eac 100644 --- a/docs/reference/capability-matrix.md +++ b/docs/reference/capability-matrix.md @@ -44,7 +44,7 @@ Core package and are stamped with the package version at release (per ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied to third-party capabilities. -### Feature capabilities (role: feature) — 18 +### Feature capabilities (role: feature) — 19 Feature capabilities extend what the loop does — contributing research, planning, execution, verification, or ship artefacts at the loop extension @@ -55,6 +55,7 @@ points. | `ai-integration` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party | | `assumption-delta` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party | | `audit` | feature | full | `>=1.6.0` | — | — | first-party | +| `claude-orchestration` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party | | `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party | | `drift` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post` | gate | first-party | | `external-job` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party | diff --git a/eslint.config.mjs b/eslint.config.mjs index 580c90822..d382077df 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -56,6 +56,8 @@ export default tseslint.config( 'coverage/**', '**/*.generated.cjs', // ADR-457: tsc-generated runtime artifact — lint the src/*.cts source, not the emitted .cjs. + 'gsd-core/bin/lib/claude-orchestration.cjs', + 'gsd-core/bin/lib/claude-orchestration-command-router.cjs', 'gsd-core/bin/lib/semver-compare.cjs', 'gsd-core/bin/lib/host-integration.cjs', 'gsd-core/bin/lib/handshake-serialized.cjs', diff --git a/gsd-core/bin/lib/capability-registry.cjs b/gsd-core/bin/lib/capability-registry.cjs index 3b6a26d41..33e4b1ce1 100644 --- a/gsd-core/bin/lib/capability-registry.cjs +++ b/gsd-core/bin/lib/capability-registry.cjs @@ -97,6 +97,14 @@ const capabilities = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ], "local": [ @@ -107,6 +115,14 @@ const capabilities = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ] }, @@ -408,6 +424,93 @@ const capabilities = { } } }, + "claude-orchestration": { + "id": "claude-orchestration", + "role": "feature", + "version": "1.7.0-rc.3", + "title": "Claude orchestration (Workflow backend)", + "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", + "tier": "full", + "requires": [], + "engines": { + "gsd": ">=1.7.0" + }, + "runtimeCompat": { + "supported": [ + "claude" + ], + "unsupported": [] + }, + "skills": [], + "agents": [], + "hooks": [], + "commands": [ + { + "family": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand", + "subcommands": [ + "detect-backend", + "emit-workflow" + ] + } + ], + "activationKey": "claude_orchestration.enabled", + "config": { + "claude_orchestration.enabled": { + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "type": "enum", + "values": [ + "auto", + "workflow", + "inline" + ], + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability." + }, + "claude_orchestration.min_agent_sdk_version": { + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + } + }, + "steps": [], + "contributions": [ + { + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md", + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:post` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## What the executor does when the Workflow backend is active\n\nInstead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor,\nisolation=worktree, run_in_background=true)` per message (which on Claude Code\ncannot nest further subagents — #853 — and so degrades to sequential inline\nexecution), execute-phase **emits a generated Workflow script** and lets the main\nloop orchestrate it:\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier (the next wave\n still waits for the previous wave to complete).\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n — the SAME executor agent and worktree isolation the inline path uses, so the\n produced `SUMMARY.md` and commits are identical.\n- **`files_modified` overlap → separate sequential stages** — two plans that\n touch the same file are placed in different stages within the wave (the same\n overlap rule execute-phase already applies inline).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n- **`budget(tokens)`** — a shared token pool across the whole phase when the\n orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a\n function parameter, not a config key; the orchestrator decides the budget).\n\nThe emitter is a pure function exposed through the capability command surface:\n`gsd-tools claude-orchestration emit-workflow --waves --run-id \n[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript`\ndirectly). It maps the phase's wave/plan manifest to the Workflow script string\nand never invokes the Workflow tool itself; the orchestrator runs the emitted\nscript. Detection is resolved by the orchestrator calling the pure\n`detectWorkflowBackend` with the LIVE host descriptor (the CLI\n`gsd-tools claude-orchestration detect-backend` is a simulation harness that\nassumes a capable host unless `--no-nested-dispatch` is passed — it does not probe\nthe real runtime; the orchestrator supplies the real descriptor).\n\n## Fallback contract\n\nIf detection resolves to `inline` (tool absent, SDK too old, runtime not Claude,\nor the capability disabled), execute-phase MUST proceed with the standard inline\nwave dispatch. The executor MUST NOT assume parallelism, a shared budget, or\nresume-from-run-id semantics in that mode.\n" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, + { + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md", + "inline": "# Claude orchestration — ultraplan plan-offload ownership (BETA)\n\n> Injected at `plan:post` `into: planner` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## Ownership declaration\n\nThe `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to\nClaude Code's ultraplan cloud) is **owned by this capability**, not by a\nstandalone BETA skill. Both surfaces share one runtime gate\n(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only\ndetection seam.\n\n## When the planner should consider ultraplan offload\n\nWhen this contribution is active (capability enabled, Claude Code runtime), the\nplanner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local\n`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is\nadvisory, not mandatory — the stable local planner remains the default.\n\n## Fallback contract\n\nIf the capability is disabled, or the runtime is not Claude Code, ultraplan\noffload is **not surfaced** and the planner proceeds with the standard local\n`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed\n(its own runtime gate already no-ops on non-Claude runtimes); this contribution\nonly governs whether the capability manifest advertises it as part of the\norchestration surface.\n" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + } + ], + "gates": [] + }, "cline": { "id": "cline", "role": "runtime", @@ -735,6 +838,14 @@ const capabilities = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ], "local": [ @@ -745,6 +856,14 @@ const capabilities = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ] }, @@ -2813,6 +2932,21 @@ const byLoopPoint = { } ], "contributions": [ + { + "capId": "claude-orchestration", + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md", + "inline": "# Claude orchestration — ultraplan plan-offload ownership (BETA)\n\n> Injected at `plan:post` `into: planner` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## Ownership declaration\n\nThe `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to\nClaude Code's ultraplan cloud) is **owned by this capability**, not by a\nstandalone BETA skill. Both surfaces share one runtime gate\n(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only\ndetection seam.\n\n## When the planner should consider ultraplan offload\n\nWhen this contribution is active (capability enabled, Claude Code runtime), the\nplanner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local\n`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is\nadvisory, not mandatory — the stable local planner remains the default.\n\n## Fallback contract\n\nIf the capability is disabled, or the runtime is not Claude Code, ultraplan\noffload is **not surfaced** and the planner proceeds with the standard local\n`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed\n(its own runtime gate already no-ops on non-Claude runtimes); this contribution\nonly governs whether the capability manifest advertises it as part of the\norchestration surface.\n" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, { "capId": "external-job", "point": "plan:post", @@ -2853,6 +2987,21 @@ const byLoopPoint = { "execute:wave:post": { "steps": [], "contributions": [ + { + "capId": "claude-orchestration", + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md", + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:post` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## What the executor does when the Workflow backend is active\n\nInstead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor,\nisolation=worktree, run_in_background=true)` per message (which on Claude Code\ncannot nest further subagents — #853 — and so degrades to sequential inline\nexecution), execute-phase **emits a generated Workflow script** and lets the main\nloop orchestrate it:\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier (the next wave\n still waits for the previous wave to complete).\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n — the SAME executor agent and worktree isolation the inline path uses, so the\n produced `SUMMARY.md` and commits are identical.\n- **`files_modified` overlap → separate sequential stages** — two plans that\n touch the same file are placed in different stages within the wave (the same\n overlap rule execute-phase already applies inline).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n- **`budget(tokens)`** — a shared token pool across the whole phase when the\n orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a\n function parameter, not a config key; the orchestrator decides the budget).\n\nThe emitter is a pure function exposed through the capability command surface:\n`gsd-tools claude-orchestration emit-workflow --waves --run-id \n[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript`\ndirectly). It maps the phase's wave/plan manifest to the Workflow script string\nand never invokes the Workflow tool itself; the orchestrator runs the emitted\nscript. Detection is resolved by the orchestrator calling the pure\n`detectWorkflowBackend` with the LIVE host descriptor (the CLI\n`gsd-tools claude-orchestration detect-backend` is a simulation harness that\nassumes a capable host unless `--no-nested-dispatch` is passed — it does not probe\nthe real runtime; the orchestrator supplies the real descriptor).\n\n## Fallback contract\n\nIf detection resolves to `inline` (tool absent, SDK too old, runtime not Claude,\nor the capability disabled), execute-phase MUST proceed with the standard inline\nwave dispatch. The executor MUST NOT assume parallelism, a shared budget, or\nresume-from-run-id semantics in that mode.\n" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, { "capId": "external-job", "point": "execute:wave:post", @@ -3063,6 +3212,9 @@ const byLoopPoint = { const configKeys = { "workflow.ai_integration_phase": "ai-integration", "workflow.assumption_delta": "assumption-delta", + "claude_orchestration.enabled": "claude-orchestration", + "claude_orchestration.execution_backend": "claude-orchestration", + "claude_orchestration.min_agent_sdk_version": "claude-orchestration", "workflow.code_review": "code-review", "workflow.code_review_depth": "code-review", "workflow.drift_threshold": "drift", @@ -3114,6 +3266,29 @@ const configSchema = { "default": true, "description": "Enable the assumption-delta architecture checkpoint during planning. When a pluralization/optional/chosen signal is detected in the phase scope, the planner is prompted to re-ask whether the primary key / identity model still names the right thing. Advisory (non-blocking)." }, + "claude_orchestration.enabled": { + "owner": "claude-orchestration", + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "owner": "claude-orchestration", + "type": "enum", + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability.", + "values": [ + "auto", + "workflow", + "inline" + ] + }, + "claude_orchestration.min_agent_sdk_version": { + "owner": "claude-orchestration", + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + }, "workflow.code_review": { "owner": "code-review", "type": "boolean", @@ -3394,6 +3569,14 @@ const runtimes = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ], "local": [ @@ -3404,6 +3587,14 @@ const runtimes = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToAntigravitySkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToAntigravityAgent" } ] }, @@ -3888,6 +4079,14 @@ const runtimes = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ], "local": [ @@ -3898,6 +4097,14 @@ const runtimes = { "nesting": "flat", "recursive": false, "converter": "convertClaudeCommandToCopilotSkill" + }, + { + "kind": "agents", + "destSubpath": "agents", + "prefix": "gsd-", + "nesting": "flat", + "recursive": false, + "converter": "convertClaudeAgentToCopilotAgent" } ] }, @@ -4713,6 +4920,11 @@ const commandFamilies = { "module": "audit-command-router.cjs", "router": "routeAuditUat" }, + "claude-orchestration": { + "capId": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand" + }, "extract-messages": { "capId": "profile-pipeline", "module": "profile-pipeline-command-router.cjs", @@ -4852,6 +5064,7 @@ const _requiresGraph = { "audit": [], "augment": [], "claude": [], + "claude-orchestration": [], "cline": [], "code-review": [], "codebuddy": [], diff --git a/gsd-core/bin/lib/claude-orchestration-command-router.cjs b/gsd-core/bin/lib/claude-orchestration-command-router.cjs new file mode 100644 index 000000000..a07ad64c2 --- /dev/null +++ b/gsd-core/bin/lib/claude-orchestration-command-router.cjs @@ -0,0 +1,136 @@ +"use strict"; +/** + * Claude orchestration command router — CLI dispatcher for + * `gsd-tools claude-orchestration `. + * + * #1143 — thin CLI adapter over the pure `claude-orchestration.cjs` module. + * Lets execute-phase (or any orchestrator) invoke the Workflow-backend + * detection and the Workflow-script emitter through the standard capability + * command surface (ADR-959) instead of a bare `require()`. + * + * Router signature: { args, cwd, raw, error } — identical to the other host + * routers; discovered by dispatchCapabilityCommand via the registry's + * commandFamilies index. + * + * Subcommands: + * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] + * Resolves whether the Workflow backend should activate. `--runtime` + * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * `claude_orchestration.*` keys from .planning/config.json. Emits + * { available, backend, reason }. + * + * emit-workflow --waves --run-id [--phase-dir ] [--budget ] + * Reads a wave/plan manifest JSON file and emits the generated Workflow + * script + summary. The manifest shape matches emitWorkflowScript's input: + * { waves: [{ id, plans: [{ id, brief, files_modified: string[] }] }] }. + */ +var __importDefault = (this && this.__importDefault) || function (mod) { + return (mod && mod.__esModule) ? mod : { "default": mod }; +}; +const node_fs_1 = __importDefault(require("node:fs")); +const node_path_1 = __importDefault(require("node:path")); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const io = require("./io.cjs"); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const core = require("./claude-orchestration.cjs"); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const configLoader = require("./config-loader.cjs"); +const { output } = io; +const { detectWorkflowBackend, emitWorkflowScript } = core; +const CAPABLE_HOST = { dispatch: { nested: true, background: true } }; +function usage(error) { + error('Usage: gsd-tools claude-orchestration [...]\n' + + ' detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch]\n' + + ' emit-workflow --waves --run-id [--phase-dir ] [--budget ]'); +} +function argValue(args, flag) { + const i = args.indexOf(flag); + return i !== -1 && i + 1 < args.length ? args[i + 1] : undefined; +} +/** + * Detect whether the Workflow backend should activate for the current/given + * runtime. Reads `claude_orchestration.*` from the project config; runtime and + * SDK version come from flags (the orchestrator already knows these) or env. + */ +function cmdDetectBackend(args, cwd, raw) { + const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; + const agentSdkVersion = argValue(args, '--agent-sdk-version'); + const noNested = args.includes('--no-nested-dispatch'); + const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; + // Resolve the claude_orchestration.* slice from the project config (federated + // keys are merged by loadConfig as a nested object). A config read failure + // degrades to inline — it must not break the core loop. + let claudeSlice = {}; + try { + const loaded = configLoader.loadConfig(cwd); + const slice = loaded['claude_orchestration']; + if (slice && typeof slice === 'object' && !Array.isArray(slice)) { + claudeSlice = slice; + } + } + catch { + claudeSlice = {}; + } + // Flatten the nested slice into the dotted-key shape detectWorkflowBackend expects. + const flatConfig = {}; + for (const k of Object.keys(claudeSlice)) { + flatConfig['claude_orchestration.' + k] = claudeSlice[k]; + } + const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); + output(result, raw); +} +/** + * Emit a Workflow script from a wave/plan manifest file. + */ +function cmdEmitWorkflow(args, _cwd, raw, error) { + const wavesPath = argValue(args, '--waves'); + const runId = argValue(args, '--run-id'); + const phaseDir = argValue(args, '--phase-dir') || '.planning/phases/current'; + const budgetRaw = argValue(args, '--budget'); + if (!wavesPath) { + error('emit-workflow requires --waves '); + return; + } + if (!runId) { + error('emit-workflow requires --run-id '); + return; + } + let waves; + try { + const content = node_fs_1.default.readFileSync(node_path_1.default.resolve(wavesPath), 'utf8'); + const parsed = JSON.parse(content); + waves = parsed['waves']; + } + catch (e) { + error('emit-workflow: could not read/parse --waves file "' + wavesPath + '": ' + (e instanceof Error ? e.message : String(e))); + return; + } + const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; + const budget = (typeof budgetTokens === 'number' && !Number.isNaN(budgetTokens)) ? budgetTokens : undefined; + const result = emitWorkflowScript({ + phaseDir, + runId, + waves: waves, + budgetTokens: budget, + }); + if (!result.ok) { + error('emit-workflow: ' + result.reason); + return; + } + output({ script: result.script, summary: result.summary }, raw); +} +function routeClaudeOrchestrationCommand(opts) { + const { args, cwd, raw, error } = opts; + // args[0] is the family ('claude-orchestration'); the subcommand is args[1]. + const subcommand = args[1]; + if (subcommand === 'detect-backend') { + cmdDetectBackend(args, cwd, raw); + } + else if (subcommand === 'emit-workflow') { + cmdEmitWorkflow(args, cwd, raw, error); + } + else { + usage(error); + } +} +module.exports = { routeClaudeOrchestrationCommand }; diff --git a/gsd-core/bin/lib/claude-orchestration.cjs b/gsd-core/bin/lib/claude-orchestration.cjs new file mode 100644 index 000000000..956fdc2ed --- /dev/null +++ b/gsd-core/bin/lib/claude-orchestration.cjs @@ -0,0 +1,404 @@ +"use strict"; +/** + * Claude Orchestration Capability — Workflow-tool backend detection + emitter + * + * #1143 — adopts Claude Code's Workflow tool (the engine behind `/effort ultracode`) + * as an optional, runtime-gated parallel-execution backend for the GSD loop. + * + * This module is the pure, testable core of the capability. It owns two seams: + * + * detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) + * → { available: boolean, backend: 'workflow'|'inline', reason: string } + * Fail-closed: every miss degrades to `inline` (today's behaviour), so the + * core loop is byte-identical unless every gate opens. This is criteria 3 + 6. + * + * emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) + * → { ok:true, script, summary } | { ok:false, reason } + * Maps GSD's wave/plan model 1:1 onto Workflow primitives: + * wave → sequential `parallel()` stage barriers, + * plan → `agent(brief, { agentType:'gsd-executor', isolation:'worktree' })`, + * files_modified overlap → forces plans into separate sequential stages + * (the same overlap rule execute-phase already applies inline), + * resumeFromRunId → wired to the phase run id, + * budgetTokens → a shared token pool. + * The emitted script composes the SAME gsd-executor agent and worktree + * isolation the inline path uses, so it produces the same artifacts/commits + * (criterion 2). It is a generated string consumed by the orchestrator; this + * module never invokes the Workflow tool itself. + * + * Design laws: + * - Gall's Law: ship a small working slice that composes existing primitives + * (gsd-executor + worktree isolation) rather than reinventing them. + * - Greenspun's Tenth Rule (cited in #1143): adopt the Workflow tool's + * barrier/pipeline/budget/resume semantics instead of hand-rolling them. + * - Postel's Law: liberal in input (missing fields → inline), conservative in + * output (workflow only when every gate opens). + * - Fail-closed: an unknown version, a missing descriptor, or a disabled + * toggle all resolve to `inline`, never to `workflow`. + * + * Zero external dependencies. Pure functions. Never throws on bad input. + */ +// ─── Constants ──────────────────────────────────────────────────────────────── +/** + * The Agent SDK version that introduced the Workflow tool (#1143 prior art). + * Used as the default floor when config does not override it. A runtime reporting + * an agentSdkVersion below this cannot host the Workflow backend. + */ +const WORKFLOW_TOOL_FLOOR_VERSION = '0.3.149'; +/** Closed enum for the `claude_orchestration.execution_backend` config key. */ +const BACKEND_VALUES = new Set(['auto', 'workflow', 'inline']); +/** Only this runtime can host the Workflow tool (Claude Code / Agent SDK). */ +const WORKFLOW_RUNTIME = 'claude'; +// ─── Semver helpers ─────────────────────────────────────────────────────────── +/** Official-ish strict SemVer 2.0.0 numeric triple (+ optional pre/build). */ +const SEMVER_RE = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$/; +/** True for a syntactically valid semver string. */ +function isValidSemver(s) { + return typeof s === 'string' && SEMVER_RE.test(s); +} +/** + * Compare two semver strings. + * Returns -1/0/1 in the usual sense. Garbage in either position → -1 (fail-closed: + * an unparseable version is treated as "less than" any real floor, so detection + * never accidentally enables the preview backend on an unknown SDK). + * + * Pre-release/build metadata are ignored for the comparison — only the numeric + * major.minor.patch triple participates, matching how the Workflow-tool floor is + * specified (a plain "0.3.149"). + */ +function compareSemver(a, b) { + if (!isValidSemver(a) || !isValidSemver(b)) + return -1; + // Split numeric triple from pre-release/build metadata. + const parseTriple = (s) => { + const core = s.split('-')[0].split('+')[0].split('.'); + return [parseInt(core[0], 10), parseInt(core[1], 10), parseInt(core[2], 10)]; + }; + const hasPre = (s) => s.indexOf('-') !== -1; + const preIdentifiers = (s) => (s.split('-')[1] || '').split('+')[0].split('.').filter((x) => x.length > 0); + const am = parseTriple(a); + const bm = parseTriple(b); + for (let i = 0; i < 3; i++) { + if (am[i] < bm[i]) + return -1; + if (am[i] > bm[i]) + return 1; + } + // Numeric triple is equal. SemVer 2.0.0 §11 precedence: + // - a version WITH a pre-release tag is LOWER than the same triple WITHOUT one + // (keeps the floor fail-closed for pre-release builds of the GA floor); + // - two pre-releases of the same triple are ordered by their dot-separated + // identifiers (numeric < alphanumeric; numeric compared numerically, + // alphanumeric lexically; fewer identifiers < more). + const aPre = hasPre(a); + const bPre = hasPre(b); + if (aPre && !bPre) + return -1; + if (!aPre && bPre) + return 1; + if (aPre && bPre) { + const ai = preIdentifiers(a); + const bi = preIdentifiers(b); + const len = Math.min(ai.length, bi.length); + for (let i = 0; i < len; i++) { + const ax = ai[i]; + const bx = bi[i]; + const aNum = /^\d+$/.test(ax); + const bNum = /^\d+$/.test(bx); + if (aNum && bNum) { + const an = parseInt(ax, 10); + const bn = parseInt(bx, 10); + if (an < bn) + return -1; + if (an > bn) + return 1; + } + else if (aNum && !bNum) { + return -1; // numeric identifiers always lower than alphanumeric + } + else if (!aNum && bNum) { + return 1; + } + else { + if (ax < bx) + return -1; + if (ax > bx) + return 1; + } + } + if (ai.length < bi.length) + return -1; + if (ai.length > bi.length) + return 1; + } + return 0; +} +/** Inline result shorthand. */ +function inline(reason, available = false) { + return { available, backend: 'inline', reason }; +} +/** + * Resolve whether the Workflow-tool backend should activate. + * + * Gate ladder (all must pass for `workflow`; first miss wins, fail-closed): + * 1. capability enabled (claude_orchestration.enabled truthy) + * 2. runtime is Claude (the only runtime that exposes the Workflow tool) + * 3. execution_backend !== 'inline' + * 4. host descriptor signals nested+background dispatch (Workflow-tool capable) + * 5. agentSdkVersion is a known, valid semver + * 6. agentSdkVersion >= the configured floor (default WORKFLOW_TOOL_FLOOR_VERSION) + * 7. execution_backend === 'workflow' OR 'auto' (both reach here; 'inline' exited at 3) + * + * Never throws. Destructures defensively. + */ +function detectWorkflowBackend(input) { + if (input === null || input === undefined || typeof input !== 'object') { + return inline('capability_disabled'); + } + const cfg = (input.config !== null && input.config !== undefined && typeof input.config === 'object') + ? input.config + : {}; + // 1. capability must be opted in (default-off — ships disabled). + if (!cfg['claude_orchestration.enabled']) { + return inline('capability_disabled'); + } + // 2. only Claude can host the Workflow tool. + if (input.runtimeId !== WORKFLOW_RUNTIME) { + return inline('runtime_not_claude'); + } + // 3. explicit inline opt-out short-circuits. + let backendRaw = cfg['claude_orchestration.execution_backend']; + if (typeof backendRaw !== 'string' || !BACKEND_VALUES.has(backendRaw)) { + backendRaw = 'auto'; + } + if (backendRaw === 'inline') { + return inline('backend_inline'); + } + // 4. the host dispatch descriptor must be the nesting-capable Claude-Code shape + // (a proxy for Workflow-tool presence). This is Claude-specific and already + // gated at step 2; `background:true` alone is true on several non-Claude hosts, + // so the proxy is only meaningful after the runtime check above. Note: this is + // NOT the canonical `shouldFlattenDispatch` rule (which keys on + // `backgroundDispatch`); the Workflow backend works precisely because a single + // tool-call orchestrates internally, sidestepping the backgroundDispatch:false + // limitation. Missing/false/foreign descriptor → fail-closed. + const hi = input.hostIntegration; + if (hi === null || hi === undefined || typeof hi !== 'object' || Array.isArray(hi)) { + return inline('workflow_tool_unavailable'); + } + const dispatch = hi.dispatch; + if (typeof dispatch !== 'object' || dispatch === null || Array.isArray(dispatch)) { + return inline('workflow_tool_unavailable'); + } + const nested = dispatch['nested']; + const background = dispatch['background']; + if (nested !== true || background !== true) { + return inline('workflow_tool_unavailable'); + } + // 5. an unknown agentSdkVersion cannot be trusted to meet the floor. + if (!isValidSemver(input.agentSdkVersion)) { + return inline('agent_sdk_version_unknown'); + } + // 6. version floor (config override > default constant). + const floorRaw = cfg['claude_orchestration.min_agent_sdk_version']; + const floor = typeof floorRaw === 'string' && isValidSemver(floorRaw) ? floorRaw : WORKFLOW_TOOL_FLOOR_VERSION; + if (compareSemver(input.agentSdkVersion, floor) < 0) { + return inline('agent_sdk_version_below_floor'); + } + // 7. auto/workflow both reach the workflow backend once every gate passes. + return { available: true, backend: 'workflow', reason: 'workflow_backend_active' }; +} +/** + * Partition a wave's plans into a near-minimal number of sequential stages (via + * greedy first-fit — not guaranteed optimal for arbitrary overlap graphs, but + * correct: no two plans sharing a file ever cohabit a stage) such that no two + * plans in the same stage share a modified file. Each plan goes into the earliest + * stage where it does not overlap any plan already there. + * + * A plan with an EMPTY files_modified set declares no files; it overlaps nothing + * and coalesces into stage 0 (same behavior as the inline path, which also cannot + * guard against undeclared concurrent writes — declare filesModified accurately). + * + * This is the same overlap rule execute-phase applies inline — the only difference + * is the execution vehicle (Workflow `parallel()` vs one-agent-per-message). + */ +function partitionStages(plans) { + const stages = []; + for (const plan of plans) { + const fileSet = new Set(plan.files_modified); + let placed = false; + for (const stage of stages) { + let overlap = false; + for (const f of fileSet) { + if (stage.files.has(f)) { + overlap = true; + break; + } + } + if (!overlap) { + stage.plans.push(plan); + for (const f of fileSet) + stage.files.add(f); + placed = true; + break; + } + } + if (!placed) { + stages.push({ plans: [plan], files: new Set(fileSet) }); + } + } + return stages.map((s) => s.plans.map((p) => p.id)); +} +/** + * Quote a free-text value for safe embedding as a JavaScript/Workflow double-quoted + * string literal. Uses JSON.stringify so every JS-relevant escape (backslash, quote, + * newline, tab, NUL, U+2028/U+2029, all control chars) is handled by the language + * itself — there is no hand-rolled escape table to drift. Returns the value already + * wrapped in its surrounding quotes. + */ +function quoteString(s) { + return JSON.stringify(s); +} +/** + * True if `s` is a safe identifier/path token to interpolate into the generated + * script WITHOUT requiring a string-literal context — i.e. it contains no + * character that could terminate a comment line (`\n`/`\r`), break out of a + * string literal (`"` / `\`), or smuggle a NUL/control sequence. Used for + * `phaseDir`, `runId`, `wave.id`, and `plan.id`, which are identifiers/paths and + * must never legitimately contain such characters. Rejecting them at validation + * (rather than silently flattening) keeps the emitted script faithful to input. + */ +const UNSCRIPTABLE_CHAR_RE = /[\r\n"\\\x00-\x1f\x7f\u2028\u2029]/; +function isScriptableIdentifier(s) { + if (typeof s !== 'string' || s.length === 0) + return false; + return !UNSCRIPTABLE_CHAR_RE.test(s); +} +/** + * Emit a Workflow script mapping the phase's wave/plan model onto Workflow + * primitives. Pure and deterministic: identical input yields an identical string. + * + * Returns ok:false (never throws) on invalid input — empty waves, missing runId, + * a wave with no plans, etc. + */ +function emitWorkflowScript(input) { + if (input === null || input === undefined || typeof input !== 'object') { + return { ok: false, reason: 'invalid_input' }; + } + const { phaseDir, waves, runId } = input; + // Identifiers/paths interpolated into the generated script must be free of any + // character that could terminate a comment, break out of a string literal, or + // smuggle control bytes — reject up front (security: #1143 review Finding 1). + if (!isScriptableIdentifier(phaseDir)) { + return { ok: false, reason: 'phaseDir must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!isScriptableIdentifier(runId)) { + return { ok: false, reason: 'runId must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(waves) || waves.length === 0) { + return { ok: false, reason: 'waves must be a non-empty array' }; + } + for (let i = 0; i < waves.length; i++) { + const w = waves[i]; + if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { + return { ok: false, reason: 'waves[' + i + '] must be { id, plans: non-empty[] }' }; + } + if (!isScriptableIdentifier(w.id)) { + return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(w.plans) || w.plans.length === 0) { + return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; + } + const seenIds = new Set(); + for (let j = 0; j < w.plans.length; j++) { + const p = w.plans[j]; + if (p === null || typeof p !== 'object' || typeof p.id !== 'string' || typeof p.brief !== 'string' || !Array.isArray(p.files_modified)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '] must be { id, brief, files_modified[] }' }; + } + if (!isScriptableIdentifier(p.id)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (seenIds.has(p.id)) { + return { ok: false, reason: 'waves[' + i + '] has duplicate plan id "' + p.id + '"' }; + } + seenIds.add(p.id); + for (const f of p.files_modified) { + if (typeof f !== 'string' || f.length === 0) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].files_modified entries must be non-empty strings' }; + } + } + } + } + const budgetTokens = (typeof input.budgetTokens === 'number' && Number.isFinite(input.budgetTokens) && input.budgetTokens > 0) + ? Math.floor(input.budgetTokens) + : null; + const lines = []; + lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); + lines.push('// phase: ' + phaseDir); + lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); + lines.push('// Composes the SAME gsd-executor agent + worktree isolation as the inline path,'); + lines.push('// so artifacts (SUMMARY.md) and commits are produced identically.'); + lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + if (budgetTokens !== null) { + lines.push('budget(' + budgetTokens + ')'); + } + lines.push(''); + const stagesByWave = []; + let totalPlans = 0; + for (let wi = 0; wi < waves.length; wi++) { + const wave = waves[wi]; + const stages = partitionStages(wave.plans); + stagesByWave.push(stages); + totalPlans += wave.plans.length; + lines.push('// Wave ' + wave.id); + for (let si = 0; si < stages.length; si++) { + const stagePlanIds = stages[si]; + // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). + const stagePlans = stagePlanIds.map((id) => wave.plans.find((p) => p.id === id)); + if (stages.length > 1) { + lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); + } + if (stagePlans.length === 1) { + const p = stagePlans[0]; + lines.push('parallel('); + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" })'); + lines.push(')'); + } + else { + lines.push('parallel('); + for (const p of stagePlans) { + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" }),'); + } + // Replace trailing comma on the last agent line with nothing. + const lastIdx = lines.length - 1; + lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); + lines.push(')'); + } + } + if (wi < waves.length - 1) + lines.push(''); + } + lines.push('// Each agent writes SUMMARY.md on its worktree branch; commits land there'); + lines.push('// and are merged by the orchestrator exactly as in inline wave dispatch.'); + const script = lines.join('\n'); + return { + ok: true, + script, + summary: { + waves: waves.length, + plans: totalPlans, + stagesByWave, + resumeRunId: runId, + budgetTokens, + }, + }; +} +module.exports = { + detectWorkflowBackend, + emitWorkflowScript, + compareSemver, + isValidSemver, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + WORKFLOW_RUNTIME, +}; diff --git a/scripts/run-tests.cjs b/scripts/run-tests.cjs index b4df63d3c..e52c7fe92 100644 --- a/scripts/run-tests.cjs +++ b/scripts/run-tests.cjs @@ -570,11 +570,13 @@ function main() { // progress (verified: no leaked handle / hang; --test-force-exit exits leaks // cleanly, so the timeout was pure slowness, NOT the leak the kill message guesses). // The per-chunk timeout is sized for a "healthy chunk (~4-5 min)"; keep chunks at - // roughly half a shard so each gets its own fresh 600s budget and a fresh node - // process (also relieving per-process memory pressure from 170+ files at once). + // roughly a third of a shard so each gets its own fresh 600s budget and a fresh + // node process (also relieving per-process memory pressure from 170+ files at once). + // Lowered from 90 to 60 after #1575 — macOS Node 22 shard 2/3 chunk 2 (~80 files + // including state.test.cjs, perf-*, worktree-cleanup) exceeded 600s with 90. const MAX_FILES_PER_CHUNK = process.env.RUN_TESTS_MAX_FILES_PER_CHUNK ? Number(process.env.RUN_TESTS_MAX_FILES_PER_CHUNK) - : 90; + : 60; // node:test does not exit until the event loop drains. A unit test that leaks // an open handle (un-terminated Worker, un-killed child_process, ref'd timer) diff --git a/src/capability-state.cts b/src/capability-state.cts index a43b16f3b..9527e930a 100644 --- a/src/capability-state.cts +++ b/src/capability-state.cts @@ -46,7 +46,7 @@ const { loadConfig } = configLoaderMod; // eslint-disable-next-line @typescript-eslint/no-require-imports import installProfilesMod = require('./install-profiles.cjs'); -const { readActiveProfile, loadSkillsManifest, resolveProfile, parseRequires } = installProfilesMod; +const { readActiveProfile, loadSkillsManifest, resolveProfile, parseRequires, parseCallsAgents } = installProfilesMod; // eslint-disable-next-line @typescript-eslint/no-require-imports import surfaceMod = require('./surface.cjs'); @@ -384,22 +384,84 @@ function _loadInstalledSkillsManifest(configDir: string): Map return manifest; } +/** + * #1858 — Build a skill dependency manifest from a FLAT commands/gsd-.md + * source layout (the Claude local project install shape, where the `gsd-` + * prefix is baked into each filename at the commands/ level and there is no + * commands/gsd/ subdir). Strips the `gsd-` prefix so stems match the nested + * loader's output (gsd-validate-phase.md → validate-phase, same as nested + * validate-phase.md). + * + * Map shape is identical to loadSkillsManifest: each stem maps to its + * `requires` deps (parsed via the same shared parseRequires) and carries a + * companion `_calls_agents_` key (parsed via parseCallsAgents) so the + * flat and nested paths cannot drift. + * + * Returns an empty Map when the parent directory does not exist or contains + * no gsd-*.md files (so _resolveManifest can use size>0 as the "flat layout + * present" signal and fall through to the installed-skills branch otherwise). + */ +function _loadFlatCommandsGsdManifest(commandsParentDir: string): Map { + const manifest = new Map(); + let entries: fs.Dirent[]; + try { + entries = fs.readdirSync(commandsParentDir, { withFileTypes: true }); + } catch { + return manifest; + } + for (const entry of entries) { + if (!entry.isFile()) continue; + if (!entry.name.startsWith('gsd-')) continue; + if (!entry.name.endsWith('.md')) continue; + // Strip 'gsd-' prefix (4 chars) and '.md' suffix (3 chars) → stem. + const stem = entry.name.slice(4, -3); + if (!stem) continue; + // Mirror loadSkillsManifest's try/catch structure exactly: wrap read + + // parse + set together so an unreadable file OR a thrown parser degrades + // both keys to [] (parity; closes the latent catch-scope drift a reviewer + // flagged — both parsers are non-throwing today, but the structural + // match future-proofs the "identical Map shape" contract). + try { + const content = fs.readFileSync(path.join(commandsParentDir, entry.name), 'utf8'); + manifest.set(stem, parseRequires(content)); + manifest.set(`_calls_agents_${stem}`, parseCallsAgents(content)); + } catch { + manifest.set(stem, []); + manifest.set(`_calls_agents_${stem}`, []); + } + } + return manifest; +} + /** * Resolve the skill dependency manifest for capability-state resolution. * - * Resolution order (fixes #1160 — installed-runtime capability surface): - * 1. If commandsGsdDir exists, load from source (repo-checkout behavior). - * 2. Otherwise, fall back to installed skills at configDir/skills/gsd-[stem]/SKILL.md. + * Resolution order: + * 1. If commandsGsdDir exists, load from the nested source layout + * (repo-checkout behavior: /commands/gsd/*.md). + * 2. #1858 — otherwise, if the flat source layout is present (gsd-.md + * files in dirname(commandsGsdDir)), load from there. This is the Claude + * local project install shape where commands/gsd/ does not exist but + * commands/gsd-.md files do. + * 3. #1160 — otherwise, fall back to installed skills at + * configDir/skills/gsd-[stem]/SKILL.md. * - * In an installed runtime the commands/gsd source tree is absent; only the - * skills/ layout exists. Returning an empty manifest caused resolveSurface to - * materialise the full-sentinel to an empty Set, making every capability appear - * unsurfaced even when the skill was physically installed. + * In an installed runtime both source trees are absent; only the skills/ + * layout exists. Returning an empty manifest caused resolveSurface to + * materialise the full-sentinel to an empty Set, making every skill-bearing + * capability appear unsurfaced even when the skill was physically installed + * (#1160) or authored as a flat command file (#1858). */ function _resolveManifest(commandsGsdDir: string, configDir: string): Map { if (fs.existsSync(commandsGsdDir)) { return loadSkillsManifest(commandsGsdDir); } + // #1858: flat source layout — gsd-.md files at dirname(commandsGsdDir). + // Only claim the flat branch when it actually has gsd-*.md files; otherwise + // fall through to the installed-skills branch (a commands/ dir with no gsd + // files must not shadow an installed skills/ tree). + const flat = _loadFlatCommandsGsdManifest(path.dirname(commandsGsdDir)); + if (flat.size > 0) return flat; return _loadInstalledSkillsManifest(configDir); } @@ -619,6 +681,7 @@ export = { // Exported for tests _resolveCommandsGsdDir, _loadInstalledSkillsManifest, + _loadFlatCommandsGsdManifest, _resolveManifest, _isSafePropKey, }; diff --git a/src/capability-writer.cts b/src/capability-writer.cts index f9a51579c..7605614f8 100644 --- a/src/capability-writer.cts +++ b/src/capability-writer.cts @@ -92,7 +92,7 @@ interface DesiredCapability { } interface SetCapabilityStateOptions { - materialize?: { runtime: string; scope: string }; + materialize?: { runtime: string; scope: string; resolveAttribution?: (runtime: string) => string | null | undefined }; } /** @@ -352,8 +352,18 @@ function setCapabilityState( const layout = runtimeArtifactLayout.resolveRuntimeArtifactLayout(runtime, resolvedConfigDir, scope); const commandsGsdDir = _resolveCommandsGsdDir(); const manifest = _resolveManifest(commandsGsdDir, resolvedConfigDir); + // #1575: applySurface now accepts opts.resolveAttribution so surface-path + // agents get the same Co-Authored-By trailer as the install path. The + // resolver is not threaded here yet — the CLI command handler does not have + // access to getCommitAttribution (which lives in bin/install.js). Until that + // is refactored into a shared module, surface-path agents for descriptor- + // driven runtimes will lack the Co-Authored-By trailer that install adds. + // Parity is proven when resolveAttribution IS provided (see + // tests/issue-1575-agent-descriptor-parity.test.cjs). // eslint-disable-next-line @typescript-eslint/no-unsafe-argument - applySurface(resolvedConfigDir, layout, manifest, undefined, registry); + applySurface(resolvedConfigDir, layout, manifest, undefined, registry, opts?.materialize?.resolveAttribution + ? { resolveAttribution: opts.materialize.resolveAttribution } + : undefined); } catch (err: unknown) { const msg = err instanceof Error ? err.message : String(err); // Fix C: materialise was explicitly requested — a failure is an error (non-zero exit), diff --git a/src/claude-orchestration-command-router.cts b/src/claude-orchestration-command-router.cts new file mode 100644 index 000000000..cabc96435 --- /dev/null +++ b/src/claude-orchestration-command-router.cts @@ -0,0 +1,159 @@ +/** + * Claude orchestration command router — CLI dispatcher for + * `gsd-tools claude-orchestration `. + * + * #1143 — thin CLI adapter over the pure `claude-orchestration.cjs` module. + * Lets execute-phase (or any orchestrator) invoke the Workflow-backend + * detection and the Workflow-script emitter through the standard capability + * command surface (ADR-959) instead of a bare `require()`. + * + * Router signature: { args, cwd, raw, error } — identical to the other host + * routers; discovered by dispatchCapabilityCommand via the registry's + * commandFamilies index. + * + * Subcommands: + * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] + * Resolves whether the Workflow backend should activate. `--runtime` + * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * `claude_orchestration.*` keys from .planning/config.json. Emits + * { available, backend, reason }. + * + * emit-workflow --waves --run-id [--phase-dir ] [--budget ] + * Reads a wave/plan manifest JSON file and emits the generated Workflow + * script + summary. The manifest shape matches emitWorkflowScript's input: + * { waves: [{ id, plans: [{ id, brief, files_modified: string[] }] }] }. + */ + +import fs from 'node:fs'; +import path from 'node:path'; +// eslint-disable-next-line @typescript-eslint/no-require-imports +import io = require('./io.cjs'); +// eslint-disable-next-line @typescript-eslint/no-require-imports +import core = require('./claude-orchestration.cjs'); +// eslint-disable-next-line @typescript-eslint/no-require-imports +import configLoader = require('./config-loader.cjs'); + +const { output } = io; +const { detectWorkflowBackend, emitWorkflowScript } = core; + +const CAPABLE_HOST = { dispatch: { nested: true, background: true } }; + +interface RouterOpts { + args: string[]; + cwd: string; + raw: boolean; + error: (msg: string, reason?: string) => void; +} + +function usage(error: (msg: string, reason?: string) => void): void { + error( + 'Usage: gsd-tools claude-orchestration [...]\n' + + ' detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch]\n' + + ' emit-workflow --waves --run-id [--phase-dir ] [--budget ]', + ); +} + +function argValue(args: string[], flag: string): string | undefined { + const i = args.indexOf(flag); + return i !== -1 && i + 1 < args.length ? args[i + 1] : undefined; +} + +/** + * Detect whether the Workflow backend should activate for the current/given + * runtime. Reads `claude_orchestration.*` from the project config; runtime and + * SDK version come from flags (the orchestrator already knows these) or env. + */ +function cmdDetectBackend(args: string[], cwd: string, raw: boolean): void { + const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; + const agentSdkVersion = argValue(args, '--agent-sdk-version'); + const noNested = args.includes('--no-nested-dispatch'); + const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; + + // Resolve the claude_orchestration.* slice from the project config (federated + // keys are merged by loadConfig as a nested object). A config read failure + // degrades to inline — it must not break the core loop. + let claudeSlice: Record = {}; + try { + const loaded = configLoader.loadConfig(cwd); + const slice = loaded['claude_orchestration']; + if (slice && typeof slice === 'object' && !Array.isArray(slice)) { + claudeSlice = slice as Record; + } + } catch { + claudeSlice = {}; + } + + // Flatten the nested slice into the dotted-key shape detectWorkflowBackend expects. + const flatConfig: Record = {}; + for (const k of Object.keys(claudeSlice)) { + flatConfig['claude_orchestration.' + k] = claudeSlice[k]; + } + + const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); + output(result, raw); +} + +/** + * Emit a Workflow script from a wave/plan manifest file. + */ +function cmdEmitWorkflow(args: string[], _cwd: string, raw: boolean, error: (msg: string, reason?: string) => void): void { + const wavesPath = argValue(args, '--waves'); + const runId = argValue(args, '--run-id'); + const phaseDir = argValue(args, '--phase-dir') || '.planning/phases/current'; + const budgetRaw = argValue(args, '--budget'); + + if (!wavesPath) { + error('emit-workflow requires --waves '); + return; + } + if (!runId) { + error('emit-workflow requires --run-id '); + return; + } + + let waves: unknown; + try { + const content = fs.readFileSync(path.resolve(wavesPath), 'utf8'); + const parsed = JSON.parse(content) as Record; + waves = parsed['waves']; + } catch (e) { + error('emit-workflow: could not read/parse --waves file "' + wavesPath + '": ' + (e instanceof Error ? e.message : String(e))); + return; + } + + const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; + const budget = (typeof budgetTokens === 'number' && !Number.isNaN(budgetTokens)) ? budgetTokens : undefined; + + const result = emitWorkflowScript({ + phaseDir, + runId, + waves: waves as EmitInput['waves'], + budgetTokens: budget, + }); + + if (!result.ok) { + error('emit-workflow: ' + result.reason); + return; + } + output({ script: result.script, summary: result.summary }, raw); +} + +// Re-declared minimal input type for the cast above (avoids importing private types). +interface EmitInput { + waves: Array<{ id: string; plans: Array<{ id: string; brief: string; files_modified: string[] }> }>; +} + +function routeClaudeOrchestrationCommand(opts: RouterOpts): void { + const { args, cwd, raw, error } = opts; + // args[0] is the family ('claude-orchestration'); the subcommand is args[1]. + const subcommand = args[1]; + if (subcommand === 'detect-backend') { + cmdDetectBackend(args, cwd, raw); + } else if (subcommand === 'emit-workflow') { + cmdEmitWorkflow(args, cwd, raw, error); + } else { + usage(error); + } +} + +export = { routeClaudeOrchestrationCommand }; diff --git a/src/claude-orchestration.cts b/src/claude-orchestration.cts new file mode 100644 index 000000000..8e877edbf --- /dev/null +++ b/src/claude-orchestration.cts @@ -0,0 +1,485 @@ +/** + * Claude Orchestration Capability — Workflow-tool backend detection + emitter + * + * #1143 — adopts Claude Code's Workflow tool (the engine behind `/effort ultracode`) + * as an optional, runtime-gated parallel-execution backend for the GSD loop. + * + * This module is the pure, testable core of the capability. It owns two seams: + * + * detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) + * → { available: boolean, backend: 'workflow'|'inline', reason: string } + * Fail-closed: every miss degrades to `inline` (today's behaviour), so the + * core loop is byte-identical unless every gate opens. This is criteria 3 + 6. + * + * emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) + * → { ok:true, script, summary } | { ok:false, reason } + * Maps GSD's wave/plan model 1:1 onto Workflow primitives: + * wave → sequential `parallel()` stage barriers, + * plan → `agent(brief, { agentType:'gsd-executor', isolation:'worktree' })`, + * files_modified overlap → forces plans into separate sequential stages + * (the same overlap rule execute-phase already applies inline), + * resumeFromRunId → wired to the phase run id, + * budgetTokens → a shared token pool. + * The emitted script composes the SAME gsd-executor agent and worktree + * isolation the inline path uses, so it produces the same artifacts/commits + * (criterion 2). It is a generated string consumed by the orchestrator; this + * module never invokes the Workflow tool itself. + * + * Design laws: + * - Gall's Law: ship a small working slice that composes existing primitives + * (gsd-executor + worktree isolation) rather than reinventing them. + * - Greenspun's Tenth Rule (cited in #1143): adopt the Workflow tool's + * barrier/pipeline/budget/resume semantics instead of hand-rolling them. + * - Postel's Law: liberal in input (missing fields → inline), conservative in + * output (workflow only when every gate opens). + * - Fail-closed: an unknown version, a missing descriptor, or a disabled + * toggle all resolve to `inline`, never to `workflow`. + * + * Zero external dependencies. Pure functions. Never throws on bad input. + */ + +// ─── Constants ──────────────────────────────────────────────────────────────── + +/** + * The Agent SDK version that introduced the Workflow tool (#1143 prior art). + * Used as the default floor when config does not override it. A runtime reporting + * an agentSdkVersion below this cannot host the Workflow backend. + */ +const WORKFLOW_TOOL_FLOOR_VERSION = '0.3.149'; + +/** Closed enum for the `claude_orchestration.execution_backend` config key. */ +const BACKEND_VALUES = new Set(['auto', 'workflow', 'inline']); + +/** Only this runtime can host the Workflow tool (Claude Code / Agent SDK). */ +const WORKFLOW_RUNTIME = 'claude'; + +// ─── Semver helpers ─────────────────────────────────────────────────────────── + +/** Official-ish strict SemVer 2.0.0 numeric triple (+ optional pre/build). */ +const SEMVER_RE = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$/; + +/** True for a syntactically valid semver string. */ +function isValidSemver(s: unknown): s is string { + return typeof s === 'string' && SEMVER_RE.test(s); +} + +/** + * Compare two semver strings. + * Returns -1/0/1 in the usual sense. Garbage in either position → -1 (fail-closed: + * an unparseable version is treated as "less than" any real floor, so detection + * never accidentally enables the preview backend on an unknown SDK). + * + * Pre-release/build metadata are ignored for the comparison — only the numeric + * major.minor.patch triple participates, matching how the Workflow-tool floor is + * specified (a plain "0.3.149"). + */ +function compareSemver(a: string, b: string): number { + if (!isValidSemver(a) || !isValidSemver(b)) return -1; + // Split numeric triple from pre-release/build metadata. + const parseTriple = (s: string): number[] => { + const core = s.split('-')[0].split('+')[0].split('.'); + return [parseInt(core[0], 10), parseInt(core[1], 10), parseInt(core[2], 10)]; + }; + const hasPre = (s: string): boolean => s.indexOf('-') !== -1; + const preIdentifiers = (s: string): string[] => (s.split('-')[1] || '').split('+')[0].split('.').filter((x) => x.length > 0); + const am = parseTriple(a); + const bm = parseTriple(b); + for (let i = 0; i < 3; i++) { + if (am[i] < bm[i]) return -1; + if (am[i] > bm[i]) return 1; + } + // Numeric triple is equal. SemVer 2.0.0 §11 precedence: + // - a version WITH a pre-release tag is LOWER than the same triple WITHOUT one + // (keeps the floor fail-closed for pre-release builds of the GA floor); + // - two pre-releases of the same triple are ordered by their dot-separated + // identifiers (numeric < alphanumeric; numeric compared numerically, + // alphanumeric lexically; fewer identifiers < more). + const aPre = hasPre(a); + const bPre = hasPre(b); + if (aPre && !bPre) return -1; + if (!aPre && bPre) return 1; + if (aPre && bPre) { + const ai = preIdentifiers(a); + const bi = preIdentifiers(b); + const len = Math.min(ai.length, bi.length); + for (let i = 0; i < len; i++) { + const ax = ai[i]; + const bx = bi[i]; + const aNum = /^\d+$/.test(ax); + const bNum = /^\d+$/.test(bx); + if (aNum && bNum) { + const an = parseInt(ax, 10); + const bn = parseInt(bx, 10); + if (an < bn) return -1; + if (an > bn) return 1; + } else if (aNum && !bNum) { + return -1; // numeric identifiers always lower than alphanumeric + } else if (!aNum && bNum) { + return 1; + } else { + if (ax < bx) return -1; + if (ax > bx) return 1; + } + } + if (ai.length < bi.length) return -1; + if (ai.length > bi.length) return 1; + } + return 0; +} + +// ─── detectWorkflowBackend ──────────────────────────────────────────────────── + +interface HostIntegration { + dispatch?: { + nested?: boolean; + background?: boolean; + backgroundDispatch?: boolean; + [k: string]: unknown; + }; + [k: string]: unknown; +} + +interface BackendConfig { + 'claude_orchestration.enabled'?: unknown; + 'claude_orchestration.execution_backend'?: unknown; + 'claude_orchestration.min_agent_sdk_version'?: unknown; + [k: string]: unknown; +} + +interface DetectInput { + runtimeId?: string; + hostIntegration?: HostIntegration | null; + config?: BackendConfig | null; + agentSdkVersion?: string; +} + +interface DetectResult { + available: boolean; + backend: 'workflow' | 'inline'; + reason: string; +} + +/** Inline result shorthand. */ +function inline(reason: string, available = false): DetectResult { + return { available, backend: 'inline', reason }; +} + +/** + * Resolve whether the Workflow-tool backend should activate. + * + * Gate ladder (all must pass for `workflow`; first miss wins, fail-closed): + * 1. capability enabled (claude_orchestration.enabled truthy) + * 2. runtime is Claude (the only runtime that exposes the Workflow tool) + * 3. execution_backend !== 'inline' + * 4. host descriptor signals nested+background dispatch (Workflow-tool capable) + * 5. agentSdkVersion is a known, valid semver + * 6. agentSdkVersion >= the configured floor (default WORKFLOW_TOOL_FLOOR_VERSION) + * 7. execution_backend === 'workflow' OR 'auto' (both reach here; 'inline' exited at 3) + * + * Never throws. Destructures defensively. + */ +function detectWorkflowBackend(input: DetectInput | null | undefined): DetectResult { + if (input === null || input === undefined || typeof input !== 'object') { + return inline('capability_disabled'); + } + + const cfg: BackendConfig = + (input.config !== null && input.config !== undefined && typeof input.config === 'object') + ? input.config + : {}; + + // 1. capability must be opted in (default-off — ships disabled). + if (!cfg['claude_orchestration.enabled']) { + return inline('capability_disabled'); + } + + // 2. only Claude can host the Workflow tool. + if (input.runtimeId !== WORKFLOW_RUNTIME) { + return inline('runtime_not_claude'); + } + + // 3. explicit inline opt-out short-circuits. + let backendRaw = cfg['claude_orchestration.execution_backend']; + if (typeof backendRaw !== 'string' || !BACKEND_VALUES.has(backendRaw)) { + backendRaw = 'auto'; + } + if (backendRaw === 'inline') { + return inline('backend_inline'); + } + + // 4. the host dispatch descriptor must be the nesting-capable Claude-Code shape + // (a proxy for Workflow-tool presence). This is Claude-specific and already + // gated at step 2; `background:true` alone is true on several non-Claude hosts, + // so the proxy is only meaningful after the runtime check above. Note: this is + // NOT the canonical `shouldFlattenDispatch` rule (which keys on + // `backgroundDispatch`); the Workflow backend works precisely because a single + // tool-call orchestrates internally, sidestepping the backgroundDispatch:false + // limitation. Missing/false/foreign descriptor → fail-closed. + const hi = input.hostIntegration; + if (hi === null || hi === undefined || typeof hi !== 'object' || Array.isArray(hi)) { + return inline('workflow_tool_unavailable'); + } + const dispatch = (hi as { dispatch?: Record }).dispatch; + if (typeof dispatch !== 'object' || dispatch === null || Array.isArray(dispatch)) { + return inline('workflow_tool_unavailable'); + } + const nested = dispatch['nested']; + const background = dispatch['background']; + if (nested !== true || background !== true) { + return inline('workflow_tool_unavailable'); + } + + // 5. an unknown agentSdkVersion cannot be trusted to meet the floor. + if (!isValidSemver(input.agentSdkVersion)) { + return inline('agent_sdk_version_unknown'); + } + + // 6. version floor (config override > default constant). + const floorRaw = cfg['claude_orchestration.min_agent_sdk_version']; + const floor = typeof floorRaw === 'string' && isValidSemver(floorRaw) ? floorRaw : WORKFLOW_TOOL_FLOOR_VERSION; + if (compareSemver(input.agentSdkVersion, floor) < 0) { + return inline('agent_sdk_version_below_floor'); + } + + // 7. auto/workflow both reach the workflow backend once every gate passes. + return { available: true, backend: 'workflow', reason: 'workflow_backend_active' }; +} + +// ─── emitWorkflowScript ─────────────────────────────────────────────────────── + +interface Plan { + id: string; + brief: string; + files_modified: string[]; +} + +interface Wave { + id: string; + plans: Plan[]; +} + +interface EmitInput { + phaseDir: string; + waves: Wave[]; + runId: string; + budgetTokens?: number; +} + +interface EmitOk { + ok: true; + script: string; + summary: { + waves: number; + plans: number; + stagesByWave: string[][][]; // wave → stage → planId[] + resumeRunId: string; + budgetTokens: number | null; + }; +} + +interface EmitErr { + ok: false; + reason: string; +} + +/** + * Partition a wave's plans into a near-minimal number of sequential stages (via + * greedy first-fit — not guaranteed optimal for arbitrary overlap graphs, but + * correct: no two plans sharing a file ever cohabit a stage) such that no two + * plans in the same stage share a modified file. Each plan goes into the earliest + * stage where it does not overlap any plan already there. + * + * A plan with an EMPTY files_modified set declares no files; it overlaps nothing + * and coalesces into stage 0 (same behavior as the inline path, which also cannot + * guard against undeclared concurrent writes — declare filesModified accurately). + * + * This is the same overlap rule execute-phase applies inline — the only difference + * is the execution vehicle (Workflow `parallel()` vs one-agent-per-message). + */ +function partitionStages(plans: Plan[]): string[][] { + const stages: { plans: Plan[]; files: Set }[] = []; + for (const plan of plans) { + const fileSet = new Set(plan.files_modified); + let placed = false; + for (const stage of stages) { + let overlap = false; + for (const f of fileSet) { + if (stage.files.has(f)) { overlap = true; break; } + } + if (!overlap) { + stage.plans.push(plan); + for (const f of fileSet) stage.files.add(f); + placed = true; + break; + } + } + if (!placed) { + stages.push({ plans: [plan], files: new Set(fileSet) }); + } + } + return stages.map((s) => s.plans.map((p) => p.id)); +} + +/** + * Quote a free-text value for safe embedding as a JavaScript/Workflow double-quoted + * string literal. Uses JSON.stringify so every JS-relevant escape (backslash, quote, + * newline, tab, NUL, U+2028/U+2029, all control chars) is handled by the language + * itself — there is no hand-rolled escape table to drift. Returns the value already + * wrapped in its surrounding quotes. + */ +function quoteString(s: string): string { + return JSON.stringify(s); +} + +/** + * True if `s` is a safe identifier/path token to interpolate into the generated + * script WITHOUT requiring a string-literal context — i.e. it contains no + * character that could terminate a comment line (`\n`/`\r`), break out of a + * string literal (`"` / `\`), or smuggle a NUL/control sequence. Used for + * `phaseDir`, `runId`, `wave.id`, and `plan.id`, which are identifiers/paths and + * must never legitimately contain such characters. Rejecting them at validation + * (rather than silently flattening) keeps the emitted script faithful to input. + */ +const UNSCRIPTABLE_CHAR_RE = /[\r\n"\\\x00-\x1f\x7f\u2028\u2029]/; +function isScriptableIdentifier(s: unknown): boolean { + if (typeof s !== 'string' || s.length === 0) return false; + return !UNSCRIPTABLE_CHAR_RE.test(s); +} + +/** + * Emit a Workflow script mapping the phase's wave/plan model onto Workflow + * primitives. Pure and deterministic: identical input yields an identical string. + * + * Returns ok:false (never throws) on invalid input — empty waves, missing runId, + * a wave with no plans, etc. + */ +function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitErr { + if (input === null || input === undefined || typeof input !== 'object') { + return { ok: false, reason: 'invalid_input' }; + } + const { phaseDir, waves, runId } = input; + // Identifiers/paths interpolated into the generated script must be free of any + // character that could terminate a comment, break out of a string literal, or + // smuggle control bytes — reject up front (security: #1143 review Finding 1). + if (!isScriptableIdentifier(phaseDir)) { + return { ok: false, reason: 'phaseDir must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!isScriptableIdentifier(runId)) { + return { ok: false, reason: 'runId must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(waves) || waves.length === 0) { + return { ok: false, reason: 'waves must be a non-empty array' }; + } + for (let i = 0; i < waves.length; i++) { + const w = waves[i]; + if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { + return { ok: false, reason: 'waves[' + i + '] must be { id, plans: non-empty[] }' }; + } + if (!isScriptableIdentifier(w.id)) { + return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(w.plans) || w.plans.length === 0) { + return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; + } + const seenIds = new Set(); + for (let j = 0; j < w.plans.length; j++) { + const p = w.plans[j]; + if (p === null || typeof p !== 'object' || typeof p.id !== 'string' || typeof p.brief !== 'string' || !Array.isArray(p.files_modified)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '] must be { id, brief, files_modified[] }' }; + } + if (!isScriptableIdentifier(p.id)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (seenIds.has(p.id)) { + return { ok: false, reason: 'waves[' + i + '] has duplicate plan id "' + p.id + '"' }; + } + seenIds.add(p.id); + for (const f of p.files_modified) { + if (typeof f !== 'string' || f.length === 0) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].files_modified entries must be non-empty strings' }; + } + } + } + } + + const budgetTokens = (typeof input.budgetTokens === 'number' && Number.isFinite(input.budgetTokens) && input.budgetTokens > 0) + ? Math.floor(input.budgetTokens) + : null; + + const lines: string[] = []; + lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); + lines.push('// phase: ' + phaseDir); + lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); + lines.push('// Composes the SAME gsd-executor agent + worktree isolation as the inline path,'); + lines.push('// so artifacts (SUMMARY.md) and commits are produced identically.'); + lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + if (budgetTokens !== null) { + lines.push('budget(' + budgetTokens + ')'); + } + lines.push(''); + + const stagesByWave: string[][][] = []; + let totalPlans = 0; + + for (let wi = 0; wi < waves.length; wi++) { + const wave = waves[wi]; + const stages = partitionStages(wave.plans); + stagesByWave.push(stages); + totalPlans += wave.plans.length; + + lines.push('// Wave ' + wave.id); + for (let si = 0; si < stages.length; si++) { + const stagePlanIds = stages[si]; + // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). + const stagePlans = stagePlanIds.map((id) => wave.plans.find((p) => p.id === id) as Plan); + if (stages.length > 1) { + lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); + } + if (stagePlans.length === 1) { + const p = stagePlans[0]; + lines.push('parallel('); + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" })'); + lines.push(')'); + } else { + lines.push('parallel('); + for (const p of stagePlans) { + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" }),'); + } + // Replace trailing comma on the last agent line with nothing. + const lastIdx = lines.length - 1; + lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); + lines.push(')'); + } + } + if (wi < waves.length - 1) lines.push(''); + } + + lines.push('// Each agent writes SUMMARY.md on its worktree branch; commits land there'); + lines.push('// and are merged by the orchestrator exactly as in inline wave dispatch.'); + + const script = lines.join('\n'); + + return { + ok: true, + script, + summary: { + waves: waves.length, + plans: totalPlans, + stagesByWave, + resumeRunId: runId, + budgetTokens, + }, + }; +} + +// ─── Exports ────────────────────────────────────────────────────────────────── + +export = { + detectWorkflowBackend, + emitWorkflowScript, + compareSemver, + isValidSemver, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + WORKFLOW_RUNTIME, +}; diff --git a/src/install-engine.cts b/src/install-engine.cts index 9cf83f660..f7f6a3baa 100644 --- a/src/install-engine.cts +++ b/src/install-engine.cts @@ -244,7 +244,7 @@ function migrateLegacyDevPreferencesToSkill(targetDir: string, saved: Map(); +function warnModelOverrideUnmappable(agentType: string, overrideValue: string): void { + const key = `${agentType}::${overrideValue}`; + if (_modelOverrideUnmappableWarned.has(key)) return; + _modelOverrideUnmappableWarned.add(key); + // Cap emission length so an oversized or secret-shaped value cannot leak in + // full to stderr/logs (#2041 security review). MUST go to stderr — resolve- + // model's JSON result is parsed from stdout. + const safe = overrideValue.length > 64 ? overrideValue.slice(0, 64) + '…' : overrideValue; + process.stderr.write( + `gsd: warning — model_overrides value "${safe}" for ${agentType} ` + + `has no Claude agent alias; falling through to tier resolution.\n`, + ); +} + +// Test-only: reset the model_overrides warn-dedupe cache between cases (#2041). +function _resetModelOverrideWarningCacheForTests(): void { + _modelOverrideUnmappableWarned.clear(); +} + +/** + * #2041 — Map a `model_overrides` value to its Claude Agent-tool alias on the + * claude runtime, mirroring the `model_policy` path (#1144). Claude Code's + * Agent tool `model` parameter documents only tier aliases (opus/sonnet/haiku/ + * fable); a full Claude model ID returned verbatim is silently dropped by the + * spawner. Returns the value to return verbatim, or null to signal "fall + * through to normal tier/dynamic-routing resolution" (used when a Claude full + * ID has no alias — matches model_policy's warn-and-fall-through). Non-Claude + * runtimes and non-Claude values always pass through verbatim. + * + * Hardening (code+security review): a `typeof` guard preserves the pre-fix + * no-crash behavior if a malformed config surfaces a non-string value, and an + * `Object.hasOwn` lookup defeats `__proto__`/`constructor` lookups on the plain + * object literal so those reserved keys cannot return a truthy non-string. + */ +function mapClaudeOverrideForRuntime( + override: string, + configRuntime: string | null | undefined, + agentType: string, +): string | null { + // Defensive: model_overrides is typed Record but a malformed + // config could surface a non-string; pass through verbatim (preserving the + // pre-fix no-crash behaviour) and let the downstream Agent tool reject it. + if (typeof override !== 'string') return override; + const onClaude = !configRuntime || configRuntime === 'claude'; + if (!onClaude) return override; + // Object.hasOwn guards against __proto__/constructor returning a truthy + // non-string from the plain object literal (#2041 security review). + if (Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, override)) { + return CLAUDE_POLICY_ID_TO_ALIAS[override]; + } + if (CLAUDE_AGENT_ALIASES.has(override)) return override; + if (override.startsWith('claude-')) { + warnModelOverrideUnmappable(agentType, override); + return null; + } + return override; +} + /** * #49 — Provider-neutral model policy preset resolution. */ @@ -159,11 +219,15 @@ function resolveModelPolicy(policy: Record | null | undefined, function resolveModelInternal(cwd: string, agentType: string): string { const config = loadConfig(cwd); - // 1. Per-agent override + // 1. Per-agent override (#2041: map Claude full IDs → Agent-tool aliases on + // the claude runtime, mirroring the model_policy path #1144; non-Claude + // runtimes and non-Claude values pass through verbatim). const modelOverrides = config['model_overrides'] as Record | null | undefined; const override = modelOverrides?.[agentType]; if (override) { - return override; + const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType); + if (mapped !== null) return mapped; + // Unmappable Claude ID — fall through to tier resolution (matches model_policy). } // 2. Compute the tier @@ -287,7 +351,11 @@ function resolveModelForTier(cwd: string, agentType: string, attempt?: number): const modelOverrides = config['model_overrides'] as Record | null | undefined; const override = modelOverrides?.[agentType]; - if (override) return override; + if (override) { + const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType); + if (mapped !== null) return mapped; + // Unmappable Claude ID — fall through to dynamic_routing / model_policy resolution. + } if (config['model_policy'] && config['runtime'] && config['runtime'] !== 'claude') { return resolveModelInternal(cwd, agentType); @@ -508,6 +576,7 @@ export = { resolveModelPolicy, resolveModelInternal, _resetModelPolicyWarningCacheForTests, + _resetModelOverrideWarningCacheForTests, VALID_GRANULARITIES, resolveGranularityInternal, assertValidGranularityOverride, diff --git a/src/surface.cts b/src/surface.cts index caa4bcc34..6c93c61d0 100644 --- a/src/surface.cts +++ b/src/surface.cts @@ -29,6 +29,7 @@ */ import fs from 'node:fs'; +import os from 'node:os'; import path from 'node:path'; import { platformWriteSync } from './shell-command-projection.cjs'; // eslint-disable-next-line @typescript-eslint/no-require-imports @@ -55,11 +56,17 @@ const SURFACE_FILE_NAME = '.gsd-surface.json'; // Types // --------------------------------------------------------------------------- +interface AgentCtx { + runtime: string; + pathPrefix: string; + attribution: string | null | undefined; +} + interface ArtifactKind { kind: string; destSubpath: string; prefix: string; - stage: (resolvedProfile: { name: string; skills: Set | '*'; agents: Set }) => string; + stage: (resolvedProfile: { name: string; skills: Set | '*'; agents: Set }, agentCtx?: AgentCtx) => string; } interface Layout { @@ -69,6 +76,12 @@ interface Layout { kinds: ArtifactKind[]; } +interface ApplySurfaceOptions { + resolveAttribution?: (runtime: string) => string | null | undefined; + homedir?: () => string; + platform?: string; +} + // --------------------------------------------------------------------------- // State IO // --------------------------------------------------------------------------- @@ -301,32 +314,54 @@ function resolveSurface(runtimeConfigDir: string, manifest: Map | object, clusterMap?: ClusterMap | Record, registry?: { capabilityClusters?: Record; profileMembership?: Record }): { name: string; skills: Set; agents: Set } { +function applySurface(runtimeConfigDir: string, layout: Layout, manifest: Map | object, clusterMap?: ClusterMap | Record, registry?: { capabilityClusters?: Record; profileMembership?: Record }, opts?: ApplySurfaceOptions): { name: string; skills: Set; agents: Set } { if (path.resolve(runtimeConfigDir) !== path.resolve(layout.configDir)) { throw new TypeError('applySurface runtimeConfigDir must match layout.configDir'); } const skillManifest = normalizeSkillManifest(layout.configDir, manifest); const resolved = resolveSurface(layout.configDir, skillManifest, clusterMap, registry); - // Mirror installRuntimeArtifacts: skills kinds get per-runtime path rewrites - // so SKILL.md bodies reference the install target (pathPrefix), not the - // converter's default ~/.claude paths (#813). Delegated to the conversion - // module's deep seam (ADR-1508 / #1511 Phase 2) — no attribution resolver - // needed here (proven: Co-Authored-By never appears in staged content; see - // brief PROVEN KEY FACT). No getInstallExports() call required. - // #1615 adversarial review (PR #1622): commands kind was previously skipped, - // leaving raw @~/.claude/... references in Windsurf workflow bodies after a - // /gsd-surface profile change. Same gap affected any runtime with commands - // kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini). - // - // Asymmetry note: rewriteStagedSkillBodies mutates in place (returns void), - // but rewriteStagedCommandBodies copies to a fresh mkdtemp dir and returns - // its path (commands .md files are flat; mutating the staged source would - // corrupt the package source on full-profile runs). Caller MUST sync from - // the returned dir and clean it up. + // #1575: agents kind now mirrors createRuntimeArtifactInstallPlan — build + // agentCtx (pathPrefix + attribution) and pass it to kind.stage() so + // stageAgentsForRuntimeWithConverter applies the full inline-loop pipeline + // (pathRewrites -> attribution -> converter -> normalize). Without this, + // surface-path agents lack path-prefix rewrites and Co-Authored-By trailers, + // diverging from a fresh install. + const _homedirFn: () => string = opts?.homedir ?? (() => os.homedir()); + const _resolvedTarget = path.resolve(layout.configDir).replace(/\\/g, '/'); + const _homeDir = _homedirFn().replace(/\\/g, '/'); + const _isGlobal = (layout.scope ?? 'global') === 'global'; + const _isOpencode = layout.runtime === 'opencode'; + const _isWindowsHost = (opts?.platform ?? process.platform) === 'win32'; + const _pathPrefix = runtimeArtifactConversion._computePathPrefix({ isGlobal: _isGlobal, isOpencode: _isOpencode, isWindowsHost: _isWindowsHost, resolvedTarget: _resolvedTarget, homeDir: _homeDir }); + const _attribution = opts?.resolveAttribution ? opts.resolveAttribution(layout.runtime) : undefined; + const agentCtx: AgentCtx = { runtime: layout.runtime, pathPrefix: _pathPrefix, attribution: _attribution }; + const tempDirsToClean: string[] = []; + // #1575: When the surface has no state modifications AND the base profile is + // 'full', pass the '*' sentinel for agents staging so ALL agents are staged — + // matching the install path which uses { skills: '*' }. Without this, agents + // not referenced by any skill's _calls_agents_ manifest entry would be silently + // dropped from the surface path. For tiered profiles (core/standard) or when + // surface mods exist, pass the resolved set so only the filtered subset stages. + const _surfaceState = readSurface(layout.configDir); + const _baseProfileName = (_surfaceState && _surfaceState.baseProfile) + ? _surfaceState.baseProfile + : (readActiveProfile(layout.configDir) || 'full'); + const _hasSurfaceMods = !!_surfaceState && ( + _surfaceState.disabledClusters.length > 0 || + _surfaceState.explicitAdds.length > 0 || + _surfaceState.explicitRemoves.length > 0 + ); + const _isUnmodifiedFull = _baseProfileName === 'full' && !_hasSurfaceMods; try { for (const kind of layout.kinds) { - let staged: string = kind.stage(resolved); + let staged: string; + if (kind.kind === 'agents') { + const agentProfile = _isUnmodifiedFull ? { ...resolved, skills: '*' as const } : resolved; + staged = kind.stage(agentProfile, agentCtx); + } else { + staged = kind.stage(resolved); + } if (kind.kind === 'skills') { runtimeArtifactConversion.rewriteStagedSkillBodies(staged, { runtime: layout.runtime, @@ -345,7 +380,7 @@ function applySurface(runtimeConfigDir: string, layout: Layout, manifest: Map, prefix: s * user-owned dirs. GSD-owned = stem in manifest; removal targets = in manifest AND * not in staged set. User-owned (not in manifest) are always preserved. */ -function _syncGsdDir(stagedDir: string, destDir: string, kind: ArtifactKind | string, manifest?: Map): void { +function _syncGsdDir(stagedDir: string, destDir: string, kind: ArtifactKind | string, manifest?: Map, runtime?: string): void { if (!fs.existsSync(stagedDir)) return; fs.mkdirSync(destDir, { recursive: true }); @@ -459,6 +494,11 @@ function _syncGsdDir(stagedDir: string, destDir: string, kind: ArtifactKind | st const kindName = (typeof kind === 'string') ? kind : kind.kind; const kindPrefix = (typeof kind === 'object' && kind !== null) ? kind.prefix : 'gsd-'; + // #1575: copilot agents are renamed .md -> .agent.md at copy time, mirroring + // the inline agent loop in bin/install.js (line ~9118). Other runtimes keep + // the staged filename verbatim. + const isCopilotAgents = runtime === 'copilot' && kindName === 'agents'; + if (kindName === 'skills') { // Skills kind: work with directories, not files. // Each staged entry is a directory named ${prefix}${stem}. @@ -497,24 +537,20 @@ function _syncGsdDir(stagedDir: string, destDir: string, kind: ArtifactKind | st const stagedFiles = fs.readdirSync(stagedDir).filter(f => f.endsWith('.md')); const stagedDestNames = new Set(); for (const file of stagedFiles) { - const destName = (kindName === 'agents' || namespacedByDir) - ? file - : `${kindPrefix}${file.slice(0, -3)}.md`; + const destName = isCopilotAgents + ? file.replace(/\.md$/, '.agent.md') + : (kindName === 'agents' || namespacedByDir) + ? file + : `${kindPrefix}${file.slice(0, -3)}.md`; fs.copyFileSync(path.join(stagedDir, file), path.join(destDir, destName)); stagedDestNames.add(destName); } // Prune stale GSD-owned files not in the staged set, preserving user-owned files // (mirrors install's prefix-scoped _removeGsdEntries): - // - agents: only gsd-* are GSD-owned + // - agents: only gsd-* are GSD-owned (copilot: gsd-*.agent.md) // - flat command dirs: only `${kindPrefix}`-prefixed are GSD-owned // - namespaced command dirs: the whole dir is GSD-owned - // - // Manifest gate (#2018): when the manifest is empty/absent (e.g. an unresolvable - // install source root yields an empty staged dir), the staged set is untrustworthy. - // Skills are guarded by pruneSkillDirs' manifest-membership check; agents must be - // guarded here — skip the prune loop entirely so an empty manifest never deletes - // every gsd-* agent. Copying (above) still runs so genuinely new agents are added. const shouldPruneAgents = !(kindName === 'agents' && (!manifest || manifest.size === 0)); if (shouldPruneAgents) { for (const file of fs.readdirSync(destDir).filter(f => f.endsWith('.md'))) { diff --git a/tests/capability-state.test.cjs b/tests/capability-state.test.cjs index 66bba4e9c..456e8ee84 100644 --- a/tests/capability-state.test.cjs +++ b/tests/capability-state.test.cjs @@ -22,6 +22,7 @@ const { isCapabilityActive, _isSafePropKey, _loadInstalledSkillsManifest, + _loadFlatCommandsGsdManifest, _resolveManifest, } = require('../gsd-core/bin/lib/capability-state.cjs'); @@ -823,6 +824,223 @@ describe('cmdCapabilityState — end-to-end via gsd-tools CLI', () => { // FAIL before the fix and PASS after, regardless of whether commands/gsd // happens to exist in the current checkout. +describe('regressions: flat commands/gsd-.md layout (#1858)', () => { + // Flat command layout (Claude local project install shape): skills live at + // /commands/gsd-.md — the gsd- prefix is baked into the filename + // and there is NO commands/gsd/ subdir. _resolveManifest must detect this + // layout, strip the gsd- prefix, and produce the same stems the nested + // loader (commands/gsd/.md) would, or every skill-bearing capability + // is silently reported surfaced:false / enabled:false / active:false. + function makeFlatCommandMd(stem, requires) { + const req = requires ? `requires: [${requires.join(', ')}]` : 'requires: [phase]'; + return [ + '---', + `name: gsd:${stem}`, + `description: ${stem} skill`, + 'argument-hint: "[phase number]"', + 'allowed-tools:', + ' - Read', + req, + '---', + 'Execute end-to-end.', + ].join('\n') + '\n'; + } + + // ── Unit tests for _loadFlatCommandsGsdManifest ───────────────────────────── + + test('_loadFlatCommandsGsdManifest: returns empty map when parent dir absent', () => { + const missing = path.join(os.tmpdir(), 'cap-flat-missing-' + Date.now()); + const manifest = _loadFlatCommandsGsdManifest(missing); + assert.ok(manifest instanceof Map, 'should return a Map'); + assert.strictEqual(manifest.size, 0, 'should be empty when parent dir absent'); + }); + + test('_loadFlatCommandsGsdManifest: scans gsd-.md and strips the gsd- prefix', () => { + const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-flat-scan-')); + try { + fs.writeFileSync(path.join(tmpDir, 'gsd-validate-phase.md'), makeFlatCommandMd('validate-phase'), 'utf8'); + fs.writeFileSync(path.join(tmpDir, 'gsd-secure-phase.md'), makeFlatCommandMd('secure-phase'), 'utf8'); + // Non-gsd file must be ignored + fs.writeFileSync(path.join(tmpDir, 'random-doc.md'), '# not a skill\n', 'utf8'); + // Non-markdown gsd file must be ignored + fs.writeFileSync(path.join(tmpDir, 'gsd-notskill.txt'), 'nope\n', 'utf8'); + + const manifest = _loadFlatCommandsGsdManifest(tmpDir); + assert.ok(manifest.has('validate-phase'), 'flat gsd-validate-phase.md -> stem validate-phase'); + assert.ok(manifest.has('secure-phase'), 'flat gsd-secure-phase.md -> stem secure-phase'); + assert.ok(!manifest.has('gsd-validate-phase'), 'must NOT keep the gsd- prefix on the stem'); + assert.ok(!manifest.has('random-doc'), 'non-gsd file must be ignored'); + assert.ok(!manifest.has('notskill'), 'non-.md gsd file must be ignored'); + } finally { + cleanup(tmpDir); + } + }); + + test('_loadFlatCommandsGsdManifest: parses requires via shared parseRequires (no drift)', () => { + const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-flat-req-')); + try { + fs.writeFileSync( + path.join(tmpDir, 'gsd-my-skill.md'), + makeFlatCommandMd('my-skill', ['dep-a', 'dep-b']), + 'utf8', + ); + const manifest = _loadFlatCommandsGsdManifest(tmpDir); + assert.deepStrictEqual(manifest.get('my-skill'), ['dep-a', 'dep-b']); + } finally { + cleanup(tmpDir); + } + }); + + test('_loadFlatCommandsGsdManifest: companion _calls_agents_ key present (parity with nested loader)', () => { + const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-flat-agents-')); + try { + fs.writeFileSync(path.join(tmpDir, 'gsd-validate-phase.md'), makeFlatCommandMd('validate-phase'), 'utf8'); + const manifest = _loadFlatCommandsGsdManifest(tmpDir); + assert.ok(manifest.has('_calls_agents_validate-phase'), + 'flat loader must emit the companion _calls_agents_ key (same Map shape as loadSkillsManifest)'); + } finally { + cleanup(tmpDir); + } + }); + + // Low-1 (review): boundary — empty stem (gsd-.md) skipped, single-char stem kept. + test('_loadFlatCommandsGsdManifest: skips gsd-.md (empty stem) and keeps single-char stem (slice boundary)', () => { + const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-flat-edge-')); + try { + fs.writeFileSync(path.join(tmpDir, 'gsd-.md'), makeFlatCommandMd(''), 'utf8'); + fs.writeFileSync(path.join(tmpDir, 'gsd-x.md'), makeFlatCommandMd('x'), 'utf8'); + const manifest = _loadFlatCommandsGsdManifest(tmpDir); + assert.ok(!manifest.has(''), 'gsd-.md must NOT register an empty-string stem'); + assert.ok(!manifest.has('_calls_agents_'), 'no companion key for an empty stem'); + assert.ok(manifest.has('x'), 'gsd-x.md -> single-char stem "x" (slice(4,-3) boundary)'); + } finally { + cleanup(tmpDir); + } + }); + + // Low-2 (review): unreadable file degrades both keys to [] (parity with nested + // loader's catch). POSIX-only AND must not run as root — root bypasses POSIX + // read permission bits, so chmod 0o000 would NOT make the file unreadable and + // the test would assert [] against the real parsed deps (false failure). + // Skip on win32 (DEFECT.WINDOWS-POSIX-MODE-BIT-ASSERT) and when getuid()==0. + const _skipUnreadable = process.platform === 'win32' || (typeof process.getuid === 'function' && process.getuid() === 0); + test('_loadFlatCommandsGsdManifest: unreadable file degrades to empty deps + agents (parity)', { skip: _skipUnreadable }, () => { + const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-flat-unread-')); + let madeUnreadable = false; + try { + const skillPath = path.join(tmpDir, 'gsd-secure.md'); + fs.writeFileSync(skillPath, '---\nname: gsd:secure\nrequires: [phase]\n---\nbody\n', { mode: 0o644 }); + fs.chmodSync(skillPath, 0o000); + madeUnreadable = true; + const manifest = _loadFlatCommandsGsdManifest(tmpDir); + assert.deepStrictEqual(manifest.get('secure'), [], 'unreadable file -> empty requires (parity with nested catch)'); + assert.deepStrictEqual(manifest.get('_calls_agents_secure'), [], 'unreadable file -> empty agents (parity with nested catch)'); + } finally { + // Restore writability so cleanup() can rm the tmp tree. + if (madeUnreadable) { + try { fs.chmodSync(path.join(tmpDir, 'gsd-secure.md'), 0o644); } catch { /* best effort */ } + } + cleanup(tmpDir); + } + }); + + // ── _resolveManifest picks the flat branch when nested is absent ──────────── + + test('_resolveManifest: detects flat commands/gsd-.md layout when nested commands/gsd/ is absent (#1858)', () => { + const tmpRepo = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-rm-flat-repo-')); + try { + // Flat source layout: /commands/gsd-.md + // NO commands/gsd/ subdir, NO skills/ dir. + const commandsDir = path.join(tmpRepo, 'commands'); + fs.mkdirSync(commandsDir, { recursive: true }); + fs.writeFileSync(path.join(commandsDir, 'gsd-validate-phase.md'), makeFlatCommandMd('validate-phase'), 'utf8'); + fs.writeFileSync(path.join(commandsDir, 'gsd-secure-phase.md'), makeFlatCommandMd('secure-phase'), 'utf8'); + + // commandsGsdDir = /commands/gsd (nested — does NOT exist). + // dirname(commandsGsdDir) = /commands (where the flat files live). + const commandsGsdDir = path.join(commandsDir, 'gsd'); + // configDir = a separate empty tmp dir (no skills/ → installed fallback empty). + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-rm-flat-cfg-')); + try { + const manifest = _resolveManifest(commandsGsdDir, configDir); + assert.ok(manifest.has('validate-phase'), + 'flat layout must populate validate-phase stem (was empty pre-fix → all skill caps unsurfaced)'); + assert.ok(manifest.has('secure-phase'), 'flat layout must populate secure-phase stem'); + assert.ok(!manifest.has('gsd-validate-phase'), 'stem must have gsd- prefix stripped'); + } finally { + cleanup(configDir); + } + } finally { + cleanup(tmpRepo); + } + }); + + test('_resolveManifest: flat branch does NOT shadow a populated installed skills dir when no flat files exist', () => { + // Precedence: nested > flat-source > installed. If the flat parent dir has + // NO gsd-*.md files, the flat loader returns an empty Map and _resolveManifest + // must fall through to the installed-skills branch (not return empty). + const tmpRepo = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-rm-precedence-')); + try { + const commandsDir = path.join(tmpRepo, 'commands'); + fs.mkdirSync(commandsDir, { recursive: true }); + // No gsd-*.md files in commands/ — flat loader yields empty. + // Installed skills present under configDir: + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-rm-precedence-cfg-')); + try { + const secureDir = path.join(configDir, 'skills', 'gsd-secure-phase'); + fs.mkdirSync(secureDir, { recursive: true }); + fs.writeFileSync(path.join(secureDir, 'SKILL.md'), + '---\nname: gsd:secure-phase\nrequires: [phase]\n---\nbody\n', 'utf8'); + + const commandsGsdDir = path.join(commandsDir, 'gsd'); // nested absent + const manifest = _resolveManifest(commandsGsdDir, configDir); + assert.ok(manifest.has('secure-phase'), + 'when flat dir is empty, installed-skills fallback must still work (precedence flat > installed only when flat non-empty)'); + } finally { + cleanup(configDir); + } + } finally { + cleanup(tmpRepo); + } + }); + + // ── Parity: flat loader produces the same stems as the nested loader ──────── + // (DEFECT.GENERATIVE-FIX — guards against silent divergence between the two + // parallel manifest-builders.) + + test('flat loader and nested loader produce identical stems for the same command set (parity)', () => { + const realCommandsGsdDir = path.resolve(__dirname, '..', 'commands', 'gsd'); + if (!fs.existsSync(realCommandsGsdDir)) return; // skip outside a repo checkout + // Build a flat mirror of the real nested commands/gsd/.md as + // commands/gsd-.md in a temp dir, then compare stem sets. + const tmpFlat = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-parity-flat-')); + try { + const nested = require('../gsd-core/bin/lib/install-profiles.cjs').loadSkillsManifest(realCommandsGsdDir); + for (const [stem] of nested) { + if (stem.startsWith('_calls_agents_')) continue; + const src = path.join(realCommandsGsdDir, stem + '.md'); + if (!fs.existsSync(src)) continue; + fs.writeFileSync(path.join(tmpFlat, 'gsd-' + stem + '.md'), fs.readFileSync(src, 'utf8'), 'utf8'); + } + const flat = _loadFlatCommandsGsdManifest(tmpFlat); + const nestedStems = [...nested.keys()].filter((k) => !k.startsWith('_calls_agents_')).sort(); + const flatStems = [...flat.keys()].filter((k) => !k.startsWith('_calls_agents_')).sort(); + assert.deepStrictEqual(flatStems, nestedStems, + 'flat loader stem set must match nested loader stem set for the real command tree'); + // Nit-2 (review): also compare _calls_agents_ VALUES, not just the + // stem set — proves the shared parseCallsAgents output is identical. + for (const stem of flatStems) { + assert.deepStrictEqual(flat.get(`_calls_agents_${stem}`), nested.get(`_calls_agents_${stem}`), + `agent refs for stem "${stem}" must match between flat and nested loaders`); + assert.deepStrictEqual(flat.get(stem), nested.get(stem), + `requires for stem "${stem}" must match between flat and nested loaders`); + } + } finally { + cleanup(tmpFlat); + } + }); +}); + describe('regressions: installed-runtime capability surface (#1160)', () => { // Minimal valid SKILL.md content (frontmatter only — matches what install emits) function makeSkillMd(stem) { diff --git a/tests/check-gap-analysis-plan-post-e2e.test.cjs b/tests/check-gap-analysis-plan-post-e2e.test.cjs index 4a4a7930a..62b4bf87a 100644 --- a/tests/check-gap-analysis-plan-post-e2e.test.cjs +++ b/tests/check-gap-analysis-plan-post-e2e.test.cjs @@ -493,7 +493,7 @@ describe('resolveLoopHooks plan:post — pure function against real registry', ( assert.strictEqual(result.activeHooks[0].capId, 'gap-analysis'); }); - test('[happy] real registry byLoopPoint plan:post has 1 step (mempalace), 1 contribution (external-job planner fragment), and 1 gate (gap-analysis)', () => { + test('[happy] real registry byLoopPoint plan:post has 1 step (mempalace), 2 contributions (external-job planner + claude-orchestration ultraplan ownership), and 1 gate (gap-analysis)', () => { const entry = realRegistry.byLoopPoint['plan:post']; assert.ok(entry, 'plan:post must exist in byLoopPoint'); assert.ok(Array.isArray(entry.steps), 'steps must be an array'); @@ -501,8 +501,12 @@ describe('resolveLoopHooks plan:post — pure function against real registry', ( assert.ok(Array.isArray(entry.gates), 'gates must be an array'); assert.strictEqual(entry.steps.length, 1, 'plan:post must have 1 step (mempalace capture)'); assert.strictEqual(entry.steps[0].capId, 'mempalace', 'plan:post step must be from mempalace'); - assert.strictEqual(entry.contributions.length, 1, 'plan:post must have 1 contribution (external-job planner runtime-budget fragment)'); - assert.strictEqual(entry.contributions[0].capId, 'external-job', 'plan:post contribution must be from external-job'); + // #1143: claude-orchestration registers a plan:post contribution declaring + // ultraplan plan-offload ownership under its runtime gate (default-off). + assert.strictEqual(entry.contributions.length, 2, 'plan:post must have 2 contributions (external-job planner + claude-orchestration ultraplan ownership)'); + const capIds = entry.contributions.map(c => c.capId).sort(); + assert.deepStrictEqual(capIds, ['claude-orchestration', 'external-job'], + `plan:post contributions must be external-job + claude-orchestration; got ${capIds.join(',')}`); assert.strictEqual(entry.gates.length, 1, 'plan:post must have exactly one gate'); assert.strictEqual(entry.gates[0].capId, 'gap-analysis'); }); diff --git a/tests/claude-orchestration-command-router.test.cjs b/tests/claude-orchestration-command-router.test.cjs new file mode 100644 index 000000000..e0a408f1b --- /dev/null +++ b/tests/claude-orchestration-command-router.test.cjs @@ -0,0 +1,169 @@ +'use strict'; + +/** + * claude-orchestration-command-router.test.cjs — end-to-end tests for the + * `gsd-tools claude-orchestration` command surface (#1143). + * + * Exercises the full dispatch path: gsd-tools → dispatchCapabilityCommand → + * routeClaudeOrchestrationCommand → the pure detect/emitter functions. + * + * NOTE: runGsdTools() returns { success, output, exitCode, error } — it does + * NOT throw on a non-zero exit (helpers.cjs). Tests assert `.success` and parse + * `.output`, and check `.error` on the failure path. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs'); + +// ─── Fixtures ───────────────────────────────────────────────────────────────── + +const WAVES_MANIFEST = { + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] }, + { id: 'p2', brief: 'Wire the bar seam', files_modified: ['src/bar.cts'] }, + ], + }, + ], +}; + +function writeManifest(tmpDir) { + const manifestPath = path.join(tmpDir, 'waves.json'); + fs.writeFileSync(manifestPath, JSON.stringify(WAVES_MANIFEST), 'utf8'); + return manifestPath; +} + +/** Run the command and assert it succeeded, returning the parsed JSON output. */ +function runAndParse(args, cwd) { + const res = runGsdTools(args, cwd); + assert.strictEqual(res.success, true, 'command should succeed; stderr: ' + (res.error || '')); + assert.ok(res.output.length > 0, 'command should emit output'); + return JSON.parse(res.output); +} + +// ─── emit-workflow ──────────────────────────────────────────────────────────── + +describe('claude-orchestration emit-workflow (CLI)', () => { + test('emits a Workflow script with parallel barriers, gsd-executor + worktree, and resumeFromRunId', () => { + const tmp = createTempProject('claw-emit-'); + try { + const manifestPath = writeManifest(tmp); + const parsed = runAndParse([ + 'claude-orchestration', 'emit-workflow', + '--waves', manifestPath, + '--run-id', 'run-cli-1143', + '--phase-dir', '.planning/phases/01-foo', + ], tmp); + assert.ok(typeof parsed.script === 'string' && parsed.script.length > 0); + assert.ok(parsed.script.includes('parallel('), 'parallel() barrier emitted'); + assert.ok(parsed.script.includes('gsd-executor'), 'gsd-executor agentType'); + assert.ok(parsed.script.includes('worktree'), 'worktree isolation'); + assert.ok(parsed.script.includes('resumeFromRunId'), 'resumeFromRunId wired'); + assert.ok(parsed.script.includes('run-cli-1143'), 'carries the run id'); + assert.strictEqual(parsed.summary.resumeRunId, 'run-cli-1143'); + assert.strictEqual(parsed.summary.waves, 1); + assert.strictEqual(parsed.summary.plans, 2); + } finally { + cleanup(tmp); + } + }); + + test('budget flag threads a shared token pool into the script', () => { + const tmp = createTempProject('claw-budget-'); + try { + const manifestPath = writeManifest(tmp); + const parsed = runAndParse([ + 'claude-orchestration', 'emit-workflow', + '--waves', manifestPath, + '--run-id', 'r', + '--budget', '750000', + ], tmp); + assert.ok(parsed.script.includes('budget('), 'budget() pool emitted'); + assert.ok(parsed.script.includes('750000')); + } finally { + cleanup(tmp); + } + }); + + test('missing --waves -> non-zero exit with a diagnostic', () => { + const tmp = createTempProject('claw-noargs-'); + try { + const res = runGsdTools([ + 'claude-orchestration', 'emit-workflow', '--run-id', 'r', + ], tmp); + assert.strictEqual(res.success, false, 'missing --waves must fail'); + assert.ok(res.exitCode !== 0, 'non-zero exit'); + assert.match(res.error || '', /--waves/); + } finally { + cleanup(tmp); + } + }); +}); + +// ─── detect-backend ─────────────────────────────────────────────────────────── + +describe('claude-orchestration detect-backend (CLI)', () => { + test('default config (capability off) -> inline, even on Claude', () => { + const tmp = createTempProject('claw-detect-'); + try { + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'claude', + '--agent-sdk-version', '1.0.0', + ], tmp); + assert.strictEqual(parsed.backend, 'inline'); + assert.strictEqual(parsed.available, false); + assert.match(parsed.reason, /disabled/); + } finally { + cleanup(tmp); + } + }); + + test('enabled + claude + capable + new-enough SDK -> workflow', () => { + const tmp = createTempProject('claw-detect-on-'); + try { + fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(tmp, '.planning', 'config.json'), + JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }), + 'utf8', + ); + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'claude', + '--agent-sdk-version', '1.2.0', + ], tmp); + assert.strictEqual(parsed.backend, 'workflow'); + assert.strictEqual(parsed.available, true); + } finally { + cleanup(tmp); + } + }); + + test('non-Claude runtime -> inline (criterion 6)', () => { + const tmp = createTempProject('claw-detect-nonclaude-'); + try { + fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(tmp, '.planning', 'config.json'), + JSON.stringify({ claude_orchestration: { enabled: true } }), + 'utf8', + ); + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'codex', + '--agent-sdk-version', '1.0.0', + ], tmp); + assert.strictEqual(parsed.backend, 'inline'); + assert.match(parsed.reason, /claude/i); + } finally { + cleanup(tmp); + } + }); +}); diff --git a/tests/claude-orchestration.test.cjs b/tests/claude-orchestration.test.cjs new file mode 100644 index 000000000..808501e1e --- /dev/null +++ b/tests/claude-orchestration.test.cjs @@ -0,0 +1,615 @@ +'use strict'; + +/** + * claude-orchestration.test.cjs — Behavioral tests for the Claude orchestration + * capability (#1143): Workflow-tool backend detection, Workflow-script emission, + * capability-declaration validation, registry integration, and inline-fallback parity. + * + * The capability is default-off + BETA + claude-only. On any runtime lacking the + * Workflow tool it must be a byte-identical no-op. These tests encode that contract. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const fc = require('fast-check'); + +const { + detectWorkflowBackend, + emitWorkflowScript, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + compareSemver, +} = require('../gsd-core/bin/lib/claude-orchestration.cjs'); + +const { + validateCapability, + validateAgainstContract, + loadAndValidate, + buildRegistry, + serializeRegistry, + normalizeLineEndings, + stripGeneratedComment, +} = require('../scripts/gen-capability-registry.cjs'); + +const ROOT = path.resolve(__dirname, '..'); +const CAP_PATH = path.join(ROOT, 'capabilities', 'claude-orchestration', 'capability.json'); +const REGISTRY_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs'); + +// ─── Fixtures ───────────────────────────────────────────────────────────────── + +/** A host-integration descriptor whose dispatch axis signals Workflow-tool capability. */ +const CAPABLE_HOST = { + dispatch: { namedDispatch: true, nested: true, background: true, backgroundDispatch: false }, +}; + +/** Read the real capability declaration (data file — not a source grep). */ +function loadCap() { + return JSON.parse(fs.readFileSync(CAP_PATH, 'utf8')); +} + +/** A minimal single-plan wave manifest. */ +function singleWaveManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] }, + ], + }, + ], + }; +} + +/** Two plans in one wave that DO NOT overlap (parallel-safe in a single stage). */ +function nonOverlappingManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Plan A', files_modified: ['src/a.cts'] }, + { id: 'p2', brief: 'Plan B', files_modified: ['src/b.cts'] }, + ], + }, + ], + }; +} + +/** Two plans in one wave that DO overlap on files_modified (must split into stages). */ +function overlappingManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Plan A', files_modified: ['src/shared.cts', 'src/a.cts'] }, + { id: 'p2', brief: 'Plan B', files_modified: ['src/shared.cts', 'src/b.cts'] }, + ], + }, + ], + }; +} + +// ─── 1. detectWorkflowBackend ───────────────────────────────────────────────── + +describe('detectWorkflowBackend', () => { + + test('capability disabled (default-off) -> inline, even on Claude with the tool', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': false }, + }); + assert.strictEqual(r.available, false); + assert.strictEqual(r.backend, 'inline'); + assert.match(r.reason, /disabled/); + }); + + test('non-Claude runtime -> inline (criterion 6: no change to non-Claude loop)', () => { + for (const runtimeId of ['codex', 'cursor', 'opencode', 'copilot', ' Windsurf'.trim()]) { + const r = detectWorkflowBackend({ + runtimeId, + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'workflow' }, + }); + assert.strictEqual(r.backend, 'inline', runtimeId + ' should be inline'); + assert.strictEqual(r.available, false, runtimeId + ' should be unavailable'); + assert.match(r.reason, /claude/i, runtimeId + ' reason should mention claude'); + } + }); + + test('Claude + auto + capable host + new-enough SDK -> workflow', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.2.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'workflow'); + assert.strictEqual(r.available, true); + }); + + test('Claude + execution_backend:"workflow" forces workflow when tool is capable', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'workflow' }, + }); + assert.strictEqual(r.backend, 'workflow'); + assert.strictEqual(r.available, true); + }); + + test('Claude + execution_backend:"inline" -> inline even when tool is capable', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'inline' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.match(r.reason, /inline/); + }); + + test('Claude + auto + host lacking nested dispatch -> inline (fail-closed)', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: { dispatch: { nested: false, background: true } }, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.strictEqual(r.available, false); + }); + + test('Claude + unknown agentSdkVersion -> inline fail-closed (criterion 3 fallback)', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: undefined, + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.strictEqual(r.available, false); + assert.match(r.reason, /version|sdk|unknown/i); + }); + + test('agent SDK version boundary: floor-1 -> inline, floor -> workflow, floor+patch -> workflow', () => { + const floor = WORKFLOW_TOOL_FLOOR_VERSION; + const [maj, min, pat] = floor.split('.').map((n) => parseInt(n, 10)); + // Robust "below" derivation with full borrow chain (works for .0.0 floors too). + let below; + if (pat > 0) below = `${maj}.${min}.${pat - 1}`; + else if (min > 0) below = `${maj}.${min - 1}.999`; + else if (maj > 0) below = `${maj - 1}.999.999`; + else { assert.ok(false, 'cannot derive below for 0.0.0 floor'); return; } + const above = `${maj}.${min}.${pat + 1}`; + // Sanity: confirm below really is below per the comparator under test. + assert.ok(compareSemver(below, floor) < 0, below + ' must compare below ' + floor); + + const cfg = { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }; + + const rBelow = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: below, config: cfg }); + assert.strictEqual(rBelow.backend, 'inline', below + ' (floor-1) must be inline'); + + const rAt = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: floor, config: cfg }); + assert.strictEqual(rAt.backend, 'workflow', floor + ' (exact floor) must be workflow'); + + const rAbove = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: above, config: cfg }); + assert.strictEqual(rAbove.backend, 'workflow', above + ' (floor+patch) must be workflow'); + }); + + test('config-level min_agent_sdk_version override raises/lowers the floor', () => { + const cfg = { + 'claude_orchestration.enabled': true, + 'claude_orchestration.execution_backend': 'auto', + 'claude_orchestration.min_agent_sdk_version': '2.0.0', + }; + const r1 = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '1.9.9', config: cfg }); + assert.strictEqual(r1.backend, 'inline', 'below raised floor -> inline'); + const r2 = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '2.0.0', config: cfg }); + assert.strictEqual(r2.backend, 'workflow', 'at raised floor -> workflow'); + }); + + test('execution_backend:"workflow" + SDK below floor -> inline (M-1: floor applies in both modes)', () => { + const cfg = { + 'claude_orchestration.enabled': true, + 'claude_orchestration.execution_backend': 'workflow', + }; + const r = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '0.3.0', config: cfg }); + assert.strictEqual(r.backend, 'inline', 'workflow mode must still honor the SDK floor (fail-closed)'); + assert.strictEqual(r.available, false); + assert.match(r.reason, /floor|version/); + }); + + test('pre-release of the floor (0.3.149-rc.1) -> inline (pre-release < GA per SemVer)', () => { + const cfg = { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }; + // Explicitly assert the precedence rule: a pre-release tag is below the GA release. + assert.ok(compareSemver('0.3.149-rc.1', '0.3.149') < 0, 'pre-release must compare below GA'); + const r = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '0.3.149-rc.1', config: cfg }); + assert.strictEqual(r.backend, 'inline', 'pre-release of the floor must not activate the BETA backend'); + assert.strictEqual(r.available, false); + }); + + test('two pre-releases of the same triple order by their identifiers (SemVer §11)', () => { + assert.ok(compareSemver('0.3.149-rc.0', '0.3.149-rc.1') < 0, 'rc.0 < rc.1'); + assert.ok(compareSemver('1.0.0-alpha.1', '1.0.0-alpha.2') < 0, 'alpha.1 < alpha.2'); + assert.ok(compareSemver('1.0.0-rc.1', '1.0.0-rc.2') < 0, 'rc.1 < rc.2'); + // numeric < alphanumeric at the same position + assert.ok(compareSemver('1.0.0-1', '1.0.0-alpha') < 0, 'numeric identifier < alphanumeric'); + }); + + test('missing/empty input -> inline, never throws (Postel: liberal-in-input)', () => { + assert.strictEqual(detectWorkflowBackend({}).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend(null).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend(undefined).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend({ runtimeId: 'claude' }).backend, 'inline'); + }); + + test('BACKEND_VALUES exposes the closed enum', () => { + assert.deepStrictEqual([...BACKEND_VALUES].sort(), ['auto', 'inline', 'workflow']); + }); + + test('property: pure & deterministic (same input -> same output)', () => { + fc.assert(fc.property( + fc.record({ + runtimeId: fc.constantFrom('claude', 'codex', 'cursor', 'opencode'), + sdk: fc.option(fc.string({ minLength: 1, maxLength: 8 }).filter((s) => /^\d/.test(s)), { nil: undefined }), + backend: fc.constantFrom('auto', 'workflow', 'inline'), + enabled: fc.boolean(), + }), + (input) => { + const cfg = { + 'claude_orchestration.enabled': input.enabled, + 'claude_orchestration.execution_backend': input.backend, + }; + const a = detectWorkflowBackend({ runtimeId: input.runtimeId, hostIntegration: CAPABLE_HOST, agentSdkVersion: input.sdk, config: cfg }); + const b = detectWorkflowBackend({ runtimeId: input.runtimeId, hostIntegration: CAPABLE_HOST, agentSdkVersion: input.sdk, config: cfg }); + assert.deepStrictEqual(a, b); + assert.ok(['workflow', 'inline'].includes(a.backend)); + }, + )); + }); +}); + +// ─── 2. compareSemver helper ────────────────────────────────────────────────── + +describe('compareSemver', () => { + test('ordering', () => { + assert.ok(compareSemver('1.0.0', '0.9.9') > 0); + assert.ok(compareSemver('1.0.0', '1.0.0') === 0); + assert.ok(compareSemver('1.0.0', '1.0.1') < 0); + assert.ok(compareSemver('2.0.0', '1.9.9') > 0); + }); + test('garbage versions compare as -1 (fail-closed)', () => { + assert.strictEqual(compareSemver('garbage', '1.0.0'), -1); + assert.strictEqual(compareSemver('1.0.0', ''), -1); + }); +}); + +// ─── 3. emitWorkflowScript ──────────────────────────────────────────────────── + +describe('emitWorkflowScript', () => { + + test('single-wave single-plan -> one parallel barrier, one agent, executor+worktree', () => { + const { ok, script, summary } = emitWorkflowScript(singleWaveManifest()); + assert.strictEqual(ok, true); + assert.ok(typeof script === 'string' && script.length > 0); + + const parallelCount = (script.match(/parallel\s*\(/g) || []).length; + assert.ok(parallelCount >= 1, 'at least one parallel() barrier'); + assert.ok(script.includes('agent('), 'agent() call per plan'); + assert.ok(script.includes('gsd-executor'), 'uses gsd-executor agentType'); + assert.ok(script.includes('worktree'), 'uses worktree isolation'); + assert.ok(script.includes('SUMMARY.md'), 'produces SUMMARY.md (same artifact as inline path)'); + + assert.deepStrictEqual(summary.waves, 1); + assert.deepStrictEqual(summary.plans, 1); + }); + + test('multi-wave -> one parallel() barrier per wave (sequential barriers)', () => { + const r = emitWorkflowScript({ + phaseDir: '.planning/phases/01-foo', + runId: 'run-multi', + waves: [ + { id: 'w1', plans: [{ id: 'p1', brief: 'A', files_modified: ['src/a.cts'] }] }, + { id: 'w2', plans: [{ id: 'p2', brief: 'B', files_modified: ['src/b.cts'] }] }, + { id: 'w3', plans: [{ id: 'p3', brief: 'C', files_modified: ['src/c.cts'] }] }, + ], + }); + assert.strictEqual(r.ok, true); + const parallelCount = (r.script.match(/parallel\s*\(/g) || []).length; + assert.strictEqual(parallelCount, 3, 'one parallel() per wave'); + assert.strictEqual(r.summary.waves, 3); + assert.strictEqual(r.summary.plans, 3); + }); + + test('overlapping files_modified -> plans split into separate sequential stages (criterion 2)', () => { + const r = emitWorkflowScript(overlappingManifest()); + assert.strictEqual(r.ok, true); + // Two plans sharing src/shared.cts must NOT be in the same stage. + const stages = r.summary.stagesByWave[0]; // wave w1 + assert.ok(Array.isArray(stages), 'stagesByWave present'); + assert.strictEqual(stages.length, 2, 'overlapping plans split into 2 stages'); + const stagePlanSets = stages.map((s) => s.slice().sort()); + const allPlans = stagePlanSets.flat().sort(); + assert.deepStrictEqual(allPlans, ['p1', 'p2']); + // p1 and p2 must be in different stages + assert.ok(stages[0].length === 1 && stages[1].length === 1, 'one plan per stage when they overlap'); + }); + + test('non-overlapping plans -> coalesced into a single parallel stage', () => { + const r = emitWorkflowScript(nonOverlappingManifest()); + assert.strictEqual(r.ok, true); + const stages = r.summary.stagesByWave[0]; + assert.strictEqual(stages.length, 1, 'non-overlapping plans share one stage'); + assert.deepStrictEqual(stages[0].slice().sort(), ['p1', 'p2']); + }); + + test('resumeFromRunId wired to the provided runId (criterion 4)', () => { + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(r.script.includes('resumeFromRunId'), 'references resumeFromRunId'); + assert.ok(r.script.includes('run-abc-1143'), 'carries the run id'); + assert.strictEqual(r.summary.resumeRunId, 'run-abc-1143'); + }); + + test('shared budget pool emitted when budgetTokens provided', () => { + const r = emitWorkflowScript({ ...singleWaveManifest(), budgetTokens: 500000 }); + assert.ok(r.script.includes('budget('), 'emits budget() pool'); + assert.ok(r.script.includes('500000')); + }); + + test('no budget() emitted when budgetTokens omitted', () => { + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(!r.script.includes('budget('), 'no budget() when unset'); + }); + + test('invalid input -> ok:false with a reason, never throws', () => { + const empty = emitWorkflowScript({ phaseDir: '.p', runId: 'r', waves: [] }); + assert.strictEqual(empty.ok, false); + assert.ok(typeof empty.reason === 'string' && empty.reason.length > 0); + + const noRun = emitWorkflowScript({ phaseDir: '.p', runId: '', waves: singleWaveManifest().waves }); + assert.strictEqual(noRun.ok, false); + + const noPhase = emitWorkflowScript({ phaseDir: '', runId: 'r', waves: singleWaveManifest().waves }); + assert.strictEqual(noPhase.ok, false); + + const badWave = emitWorkflowScript({ phaseDir: '.p', runId: 'r', waves: [{ id: 'w1', plans: [] }] }); + assert.strictEqual(badWave.ok, false); + }); + + test('SECURITY: runId/phaseDir/wave.id/plan.id with injection chars -> ok:false (never reach the script)', () => { + // runId is interpolated inside resumeFromRunId("...") — a quote/backslash/newline + // could break out of the call. Identifier validation must reject it. + const injectRun = emitWorkflowScript({ phaseDir: '.p', runId: 'x");evil("y', waves: singleWaveManifest().waves }); + assert.strictEqual(injectRun.ok, false); + assert.match(injectRun.reason, /runId/i); + + const newlineRun = emitWorkflowScript({ phaseDir: '.p', runId: 'r\nbreakout', waves: singleWaveManifest().waves }); + assert.strictEqual(newlineRun.ok, false); + + const injectPhase = emitWorkflowScript({ phaseDir: '.p"; drop table', runId: 'r', waves: singleWaveManifest().waves }); + assert.strictEqual(injectPhase.ok, false); + + const injectWave = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1\nagent("evil")', plans: [{ id: 'p1', brief: 'b', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(injectWave.ok, false); + + const injectPlan = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1";x("y', brief: 'b', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(injectPlan.ok, false); + }); + + test('SECURITY: a brief containing quotes/backslash/newlines is neutralised (never breaks the string literal)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1', brief: 'he said "hi" \\ then \n newline', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(r.ok, true); + // The emitted script must not contain a raw unescaped quote that closes the + // agent() string literal, nor a raw newline inside the brief. + assert.ok(!r.script.includes('he said "hi" \\\\'), 'no unescaped breakout'); + // The full brief text never appears verbatim with its dangerous chars intact. + assert.ok(!r.script.includes('"hi"'), 'the inner quote must be JSON-escaped, not raw'); + }); + + test('duplicate plan id within a wave -> ok:false (L-5: no silent brief loss)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [ + { id: 'p1', brief: 'first', files_modified: ['a.cts'] }, + { id: 'p1', brief: 'second', files_modified: ['b.cts'] }, + ] }], + }); + assert.strictEqual(r.ok, false); + assert.match(r.reason, /duplicate/i); + }); + + test('non-string files_modified entries -> ok:false (L-7: strict element typing)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1', brief: 'b', files_modified: ['ok.cts', 42, { path: 'x' }] }] }], + }); + assert.strictEqual(r.ok, false); + assert.match(r.reason, /files_modified/); + }); + + test('property: deterministic (same input -> identical script)', () => { + fc.assert(fc.property( + fc.record({ + runId: fc.string({ minLength: 1, maxLength: 12 }).filter((s) => /^[a-zA-Z0-9-]+$/.test(s)), + nPlans: fc.integer({ min: 1, max: 5 }), + }), + ({ runId, nPlans }) => { + const waves = [{ + id: 'w1', + plans: Array.from({ length: nPlans }, (_, i) => ({ + id: 'p' + i, + brief: 'brief ' + i, + files_modified: ['src/file' + i + '.cts'], + })), + }]; + const a = emitWorkflowScript({ phaseDir: '.planning/phases/01-x', runId, waves }); + const b = emitWorkflowScript({ phaseDir: '.planning/phases/01-x', runId, waves }); + assert.strictEqual(a.script, b.script); + assert.deepStrictEqual(a.summary, b.summary); + }, + )); + }); +}); + +// ─── 4. Capability declaration validation ───────────────────────────────────── + +describe('capability declaration (capabilities/claude-orchestration/capability.json)', () => { + + test('file exists and parses', () => { + const cap = loadCap(); + assert.strictEqual(cap.id, 'claude-orchestration'); + }); + + test('passes per-file validateCapability', () => { + const errors = validateCapability(loadCap(), 'claude-orchestration'); + assert.deepEqual(errors, [], 'Expected no validation errors: ' + JSON.stringify(errors)); + }); + + test('passes contract validation (contribution.into roles, when references)', () => { + const errors = validateAgainstContract(loadCap(), 'claude-orchestration'); + assert.deepEqual(errors, [], 'Expected no contract errors: ' + JSON.stringify(errors)); + }); + + test('default-off: activationKey default is false and points at the enabled key', () => { + const cap = loadCap(); + assert.strictEqual(cap.activationKey, 'claude_orchestration.enabled'); + assert.strictEqual(cap.config['claude_orchestration.enabled'].default, false); + assert.strictEqual(cap.config['claude_orchestration.enabled'].type, 'boolean'); + }); + + test('runtimeCompat is claude-only (criterion 6)', () => { + const cap = loadCap(); + assert.deepStrictEqual(cap.runtimeCompat.supported, ['claude']); + assert.deepStrictEqual(cap.runtimeCompat.unsupported, []); + }); + + test('BETA posture: tier full, role feature', () => { + const cap = loadCap(); + assert.strictEqual(cap.role, 'feature'); + assert.strictEqual(cap.tier, 'full'); + }); + + test('execution_backend is an enum with auto|workflow|inline defaulting to auto', () => { + const slice = loadCap().config['claude_orchestration.execution_backend']; + assert.strictEqual(slice.type, 'enum'); + assert.deepStrictEqual(slice.values, ['auto', 'workflow', 'inline']); + assert.strictEqual(slice.default, 'auto'); + }); + + test('registers at WIRED points only (execute:wave:post, plan:post)', () => { + const cap = loadCap(); + const points = cap.contributions.map((c) => c.point); + for (const p of points) { + assert.ok( + ['discuss:pre', 'discuss:post', 'plan:pre', 'plan:post', 'execute:post', 'execute:wave:post', 'verify:post', 'ship:pre', 'ship:post'].includes(p), + 'contribution point ' + p + ' must be a wired point', + ); + } + assert.ok(points.includes('execute:wave:post'), 'registers the execute wave hook'); + assert.ok(points.includes('plan:post'), 'declares plan:* ownership for ultraplan (criterion 5)'); + }); + + test('all contributions gated by the enabled key + onError:skip (default-resilient)', () => { + const cap = loadCap(); + for (const c of cap.contributions) { + assert.strictEqual(c.when, 'claude_orchestration.enabled', 'every contribution gated by enabled'); + assert.strictEqual(c.onError, 'skip', 'every contribution onError:skip'); + } + }); +}); + +// ─── 5. Registry integration ────────────────────────────────────────────────── + +describe('registry integration', () => { + + test('loadAndValidate includes claude-orchestration with no errors', () => { + const { capMap, errors } = loadAndValidate(new Set()); // empty central keys = no collision noise + // Filter errors to only those touching our capability. + const ours = errors.filter((e) => e.includes('claude-orchestration')); + assert.deepEqual(ours, [], 'our capability produced errors: ' + JSON.stringify(ours)); + assert.ok(capMap.has('claude-orchestration'), 'capMap includes claude-orchestration'); + }); + + test('buildRegistry surfaces the federated config keys in configSchema', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + assert.ok(registry.configSchema['claude_orchestration.enabled'], 'enabled key federated'); + assert.ok(registry.configSchema['claude_orchestration.execution_backend'], 'execution_backend key federated'); + assert.strictEqual(registry.configSchema['claude_orchestration.enabled'].owner, 'claude-orchestration'); + assert.strictEqual(registry.configSchema['claude_orchestration.execution_backend'].default, 'auto'); + }); + + test('byLoopPoint[execute:wave:post].contributions includes our capability', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + const contribs = registry.byLoopPoint['execute:wave:post'].contributions; + const ours = contribs.find((c) => c.capId === 'claude-orchestration'); + assert.ok(ours, 'our execute:wave:post contribution is registered'); + assert.strictEqual(ours.into, 'executor'); + }); + + test('committed registry is in sync (gen-capability-registry --check)', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + const live = serializeRegistry(registry, capMap); + const committed = fs.readFileSync(REGISTRY_PATH, 'utf8'); + assert.strictEqual( + normalizeLineEndings(stripGeneratedComment(committed)), + normalizeLineEndings(stripGeneratedComment(live)), + 'registry is stale — run: node scripts/gen-capability-registry.cjs --write', + ); + }); +}); + +// ─── 6. Inline-fallback parity (criterion 3 + 6) ────────────────────────────── + +describe('inline-fallback parity', () => { + + test('default config (capability off) -> inline on every runtime, including Claude', () => { + // The capability ships default-off; with no user opt-in the backend is always inline. + const defaultCfg = {}; // nothing set + for (const runtimeId of ['claude', 'codex', 'cursor', 'opencode']) { + const r = detectWorkflowBackend({ + runtimeId, + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: defaultCfg, + }); + assert.strictEqual(r.backend, 'inline', runtimeId + ' default must be inline'); + assert.strictEqual(r.available, false, runtimeId + ' default must be unavailable'); + } + }); + + test('generated Workflow script preserves the inline-path contract (same agent + isolation + artifact)', () => { + // Criterion 2: the emitted Workflow composes the SAME gsd-executor agent and worktree + // isolation the inline path uses, and produces the same SUMMARY.md artifact. + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(r.script.includes('gsd-executor'), 'same executor agent as inline dispatch'); + assert.ok(r.script.includes('worktree'), 'same worktree isolation as inline dispatch'); + assert.ok(r.script.includes('SUMMARY.md'), 'same SUMMARY.md artifact as inline dispatch'); + }); +}); diff --git a/tests/execute-wave-post-gate-pipeline-e2e.test.cjs b/tests/execute-wave-post-gate-pipeline-e2e.test.cjs index 99fdbeec9..c01f1903d 100644 --- a/tests/execute-wave-post-gate-pipeline-e2e.test.cjs +++ b/tests/execute-wave-post-gate-pipeline-e2e.test.cjs @@ -629,15 +629,17 @@ describe('F. Real registry execute:wave:post shape — guard against accidental `ui.safety-gate onError must be 'halt'; got ${uiGate.onError}`); }); - test('[happy] real registry: execute:wave:post has no steps and 2 contributions (mempalace capture-problems + external-job executor fragment)', () => { + test('[happy] real registry: execute:wave:post has no steps and 3 contributions (claude-orchestration executor + external-job executor + mempalace capture-problems)', () => { const point = realRegistry.byLoopPoint['execute:wave:post']; assert.strictEqual(point.steps.length, 0, `execute:wave:post steps must be empty; got ${point.steps.length}`); - assert.strictEqual(point.contributions.length, 2, - `execute:wave:post must have 2 contributions (mempalace + external-job); got ${point.contributions.length}`); + // #1143: claude-orchestration registers an execute:wave:post contribution + // providing the Workflow-tool backend guidance (default-off, claude-only). + assert.strictEqual(point.contributions.length, 3, + `execute:wave:post must have 3 contributions (claude-orchestration + external-job + mempalace); got ${point.contributions.length}`); const capIds = point.contributions.map(c => c.capId).sort(); - assert.deepStrictEqual(capIds, ['external-job', 'mempalace'], - `execute:wave:post contributions must be mempalace + external-job; got ${capIds.join(',')}`); + assert.deepStrictEqual(capIds, ['claude-orchestration', 'external-job', 'mempalace'], + `execute:wave:post contributions must be claude-orchestration + external-job + mempalace; got ${capIds.join(',')}`); }); }); diff --git a/tests/helpers.cjs b/tests/helpers.cjs index 3538a26a5..234bc8f99 100644 --- a/tests/helpers.cjs +++ b/tests/helpers.cjs @@ -423,6 +423,7 @@ function resetRuntimeWarningCaches() { const modelResolver = require('../gsd-core/bin/lib/model-resolver.cjs'); configLoader._resetRuntimeWarningCacheForTests(); modelResolver._resetModelPolicyWarningCacheForTests(); + modelResolver._resetModelOverrideWarningCacheForTests(); } module.exports = { runGsdTools, createTempDir, createTempProject, createTempGitProject, cleanup, parseFrontmatter, isUsageOutput, captureConsole, toPosixPath, runNpm, isolatedNpmEnv, withIsolatedProcessState, delay, waitFor, resetRuntimeWarningCaches, TOOLS_PATH }; diff --git a/tests/issue-1575-agent-descriptor-parity.test.cjs b/tests/issue-1575-agent-descriptor-parity.test.cjs new file mode 100644 index 000000000..3557391b7 --- /dev/null +++ b/tests/issue-1575-agent-descriptor-parity.test.cjs @@ -0,0 +1,176 @@ +'use strict'; + +// #1575 — Golden-parity harness (ADR-1235 §0). +// +// Asserts that the surface path (applySurface) produces byte-for-byte identical +// agent output to the install path (installRuntimeArtifacts) for every +// descriptor-driven runtime. Both paths run against the SAME configDir so +// pathPrefix, attribution, and converter outputs match. +// +// The harness: +// 1. installRuntimeArtifacts(runtime, configDir, 'global', profile, resolveAttribution) +// 2. Snapshot every gsd-* agent file in configDir/agents/ +// 3. applySurface(configDir, layout, manifest, ..., opts) +// 4. Compare every agent file byte-for-byte: snapshot === current + +process.env.GSD_TEST_MODE = '1'; + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const os = require('node:os'); + +const ROOT = path.join(__dirname, '..'); +const COMMANDS_GSD = path.join(ROOT, 'commands', 'gsd'); + +const { installRuntimeArtifacts } = require('../gsd-core/bin/lib/install-engine.cjs'); +const { applySurface } = require('../gsd-core/bin/lib/surface.cjs'); +const { loadSkillsManifest, resolveProfile } = require('../gsd-core/bin/lib/install-profiles.cjs'); +const { resolveRuntimeArtifactLayout } = require('../gsd-core/bin/lib/runtime-artifact-layout.cjs'); +const { cleanup } = require('./helpers.cjs'); + +// The 7 descriptor-driven agent runtimes (cline deferred per code comment: +// rules-only local branch + local/global complication). +const DESCRIPTOR_RUNTIMES = [ + 'cursor', + 'windsurf', + 'augment', + 'trae', + 'codebuddy', + 'copilot', + 'antigravity', +]; + +function snapshotAgents(agentsDir) { + const snap = new Map(); + if (!fs.existsSync(agentsDir)) return snap; + for (const name of fs.readdirSync(agentsDir)) { + if (!name.startsWith('gsd-')) continue; + if (!name.endsWith('.md') && !name.endsWith('.agent.md')) continue; + snap.set(name, fs.readFileSync(path.join(agentsDir, name), 'utf8')); + } + return snap; +} + +// Shared manifest + profile so both paths see the same source agents. +const manifest = loadSkillsManifest(COMMANDS_GSD); +const profile = resolveProfile({ modes: ['full'], manifest }); +// Same attribution resolver for both paths (undefined → no Co-Authored-By mutation). +const resolveAttribution = () => undefined; + +describe('#1575 — golden-parity: surface path matches install path for descriptor-driven agents', () => { + + for (const runtime of DESCRIPTOR_RUNTIMES) { + test(`${runtime}: surface agents byte-identical to install agents`, (t) => { + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), `gsd-1575-${runtime}-`)); + t.after(() => { try { cleanup(configDir); } catch { /* best-effort */ } }); + + // Step 1: install path writes agents + installRuntimeArtifacts(runtime, configDir, 'global', profile, resolveAttribution); + + // Step 2: snapshot agent files + const agentsDir = path.join(configDir, 'agents'); + const installSnap = snapshotAgents(agentsDir); + assert.ok(installSnap.size > 0, `${runtime}: install must produce at least one gsd-* agent`); + + // Step 3: surface path re-materializes into the SAME configDir + const layout = resolveRuntimeArtifactLayout(runtime, configDir, 'global'); + applySurface(configDir, layout, manifest, undefined, undefined, { resolveAttribution }); + + // Step 4: compare byte-for-byte + const surfaceSnap = snapshotAgents(agentsDir); + + // File lists must match + const installFiles = [...installSnap.keys()].sort(); + const surfaceFiles = [...surfaceSnap.keys()].sort(); + assert.deepEqual( + surfaceFiles, + installFiles, + `${runtime}: file lists must match after surface. Install: [${installFiles.join(', ')}] Surface: [${surfaceFiles.join(', ')}]`, + ); + + // Content must match byte-for-byte + for (const [fileName, installContent] of installSnap) { + const surfaceContent = surfaceSnap.get(fileName); + assert.strictEqual( + surfaceContent, + installContent, + `${runtime}/${fileName}: surface content must be byte-identical to install content`, + ); + } + + test('cursor with non-undefined attribution: surface agents byte-identical to install agents (M2 coverage)', (t) => { + // M2 regression guard: verify parity holds when resolveAttribution returns + // a real value. Source agents don't carry Co-Authored-By, so processAttribution + // is a no-op (it replaces existing lines, doesn't add new ones). But this test + // proves the agentCtx threading is correct for both paths regardless. + const attrResolver = () => 'Test Bot '; + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-1575-attr-')); + t.after(() => { try { cleanup(configDir); } catch { /* best-effort */ } }); + + installRuntimeArtifacts('cursor', configDir, 'global', profile, attrResolver); + + const agentsDir = path.join(configDir, 'agents'); + const installSnap = snapshotAgents(agentsDir); + assert.ok(installSnap.size > 0, 'install must produce agents'); + + const layout = resolveRuntimeArtifactLayout('cursor', configDir, 'global'); + applySurface(configDir, layout, manifest, undefined, undefined, { resolveAttribution: attrResolver }); + + const surfaceSnap = snapshotAgents(agentsDir); + for (const [fileName, installContent] of installSnap) { + assert.strictEqual(surfaceSnap.get(fileName), installContent, + `cursor/${fileName}: content must be byte-identical with non-undefined attribution`); + } + }); + }); + } + + test('copilot: agents installed as .agent.md (filename rename parity)', (t) => { + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-1575-copilot-rename-')); + t.after(() => { try { cleanup(configDir); } catch { /* best-effort */ } }); + + installRuntimeArtifacts('copilot', configDir, 'global', profile, resolveAttribution); + + const agentsDir = path.join(configDir, 'agents'); + assert.ok(fs.existsSync(agentsDir), 'copilot agents dir must exist'); + const agentFiles = fs.readdirSync(agentsDir).filter((f) => f.startsWith('gsd-')); + assert.ok(agentFiles.length > 0, 'copilot must have installed agents'); + assert.ok( + agentFiles.every((f) => f.endsWith('.agent.md')), + `copilot agents must be .agent.md, got: [${agentFiles.slice(0, 3).join(', ')}]`, + ); + }); +}); + +describe('#1575 — surface path: no prune data-loss over pre-existing legacy agents', () => { + test('pre-existing gsd-* agents not in staged set are pruned; user agents preserved', (t) => { + const configDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-1575-prune-')); + t.after(() => { try { cleanup(configDir); } catch { /* best-effort */ } }); + + // Seed a pre-existing legacy .agent.md (simulating a prior install) + const agentsDir = path.join(configDir, 'agents'); + fs.mkdirSync(agentsDir, { recursive: true }); + fs.writeFileSync(path.join(agentsDir, 'gsd-old-defunct.agent.md'), '# Old\n'); + fs.writeFileSync(path.join(agentsDir, 'user-custom.md'), '# User\n'); + + // Install (should prune stale gsd-*, preserve user agents) + installRuntimeArtifacts('copilot', configDir, 'global', profile, resolveAttribution); + + const afterInstall = fs.readdirSync(agentsDir); + assert.ok(!afterInstall.includes('gsd-old-defunct.agent.md'), 'stale gsd-* agent must be pruned'); + assert.ok(afterInstall.includes('user-custom.md'), 'user agent must be preserved'); + + // Now surface over the install — must converge to the same state + const layout = resolveRuntimeArtifactLayout('copilot', configDir, 'global'); + applySurface(configDir, layout, manifest, undefined, undefined, { resolveAttribution }); + + const afterSurface = fs.readdirSync(agentsDir); + // Same set of agent files as after install + const installAgents = afterInstall.filter((f) => f.startsWith('gsd-')).sort(); + const surfaceAgents = afterSurface.filter((f) => f.startsWith('gsd-')).sort(); + assert.deepEqual(surfaceAgents, installAgents, 'surface must converge to same agent set as install'); + assert.ok(afterSurface.includes('user-custom.md'), 'user agent still preserved after surface'); + }); +}); diff --git a/tests/model-resolver.test.cjs b/tests/model-resolver.test.cjs index 9fc7fa30e..3e2d06ddf 100644 --- a/tests/model-resolver.test.cjs +++ b/tests/model-resolver.test.cjs @@ -3770,6 +3770,190 @@ describe('#49 resolveModelPolicy: prototype-pollution guards', () => { }); }); +// ─── #2041: model_overrides Claude full ID → Agent-tool alias on claude runtime ─ +// +// Mirrors the #1133 model_policy alias-mapping tests (above) for the +// model_overrides path. Bug: a full Claude model ID in model_overrides +// (e.g. "claude-sonnet-5") was returned VERBATIM on the claude runtime and +// handed to the Claude Agent tool, whose typed `model` parameter documents only +// tier aliases (opus/sonnet/haiku/fable). The model_policy path already maps +// full IDs → aliases via CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides +// skipped that mapping entirely. The fix mirrors #1144 on the override path. +// Non-Claude runtimes and non-Claude values pass through verbatim (parity). + +describe('#2041 model_overrides: Claude full ID → alias on claude runtime', () => { + let tmpDir; + beforeEach(() => { + tmpDir = makeTmp('2041'); + resetRuntimeWarningCaches(); + }); + afterEach(() => { + rmr(tmpDir); + resetRuntimeWarningCaches(); + }); + + // AC1 + AC2: mappable Claude full IDs resolve to their aliases on claude runtime + test('model_overrides claude-sonnet-5 → "sonnet" on runtime:claude (resolveModelInternal)', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-executor': 'claude-sonnet-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'sonnet'); + }); + + test('model_overrides claude-opus-4-8 → "opus" on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-planner': 'claude-opus-4-8' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + test('model_overrides claude-haiku-4-5 → "haiku" on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-codebase-mapper': 'claude-haiku-4-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-codebase-mapper'), 'haiku'); + }); + + test('model_overrides claude-fable-5 → "fable" on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-planner': 'claude-fable-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'fable'); + }); + + // AC3: bare aliases pass through verbatim + test('model_overrides bare "sonnet" alias passes through verbatim on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-executor': 'sonnet' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'sonnet'); + }); + + test('model_overrides bare "fable" alias passes through verbatim on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-planner': 'fable' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'fable'); + }); + + // AC1 (implicit claude): mapping fires when runtime key is absent (defaults to claude) + test('model_overrides claude-sonnet-5 → "sonnet" with implicit claude runtime (no runtime key)', () => { + writeConfig(tmpDir, { + model_overrides: { 'gsd-executor': 'claude-sonnet-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'sonnet'); + }); + + // AC4: non-claude runtimes keep full IDs verbatim (parity with model_policy path) + test('model_overrides claude-sonnet-5 → verbatim ID on non-claude runtime (opencode)', () => { + writeConfig(tmpDir, { + runtime: 'opencode', + model_overrides: { 'gsd-executor': 'claude-sonnet-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'claude-sonnet-5'); + }); + + // AC5: unmappable Claude full ID warns once + falls through to tier alias + test('model_overrides unmappable claude ID (claude-opus-4-5) falls through to tier alias on claude', () => { + resetRuntimeWarningCaches(); + writeConfig(tmpDir, { + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-planner': 'claude-opus-4-5' }, + }); + // gsd-planner balanced → opus tier; claude-opus-4-5 has no alias → warn + fall through → 'opus' + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'opus'); + }); + + test('model_overrides unmappable claude ID emits a stderr warning exactly once (dedupe)', () => { + resetRuntimeWarningCaches(); + writeConfig(tmpDir, { + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-planner': 'claude-opus-4-5' }, + }); + const writes = []; + const original = process.stderr.write.bind(process.stderr); + process.stderr.write = (chunk) => { writes.push(String(chunk)); return true; }; + try { + resolveModelInternal(tmpDir, 'gsd-planner'); + resolveModelInternal(tmpDir, 'gsd-planner'); // second call — dedupe must suppress + } finally { + process.stderr.write = original; + } + const warnings = writes.filter((w) => w.includes('model_overrides') && w.includes('claude-opus-4-5')); + assert.strictEqual(warnings.length, 1, + `expected exactly one override warning, got ${warnings.length}: ${JSON.stringify(writes)}`); + }); + + // AC6: resolveModelForTier (escalation / --attempt path) maps the same way + test('resolveModelForTier maps claude-sonnet-5 → "sonnet" on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-executor': 'claude-sonnet-5' }, + }); + assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-executor', 0), 'sonnet'); + }); + + test('resolveModelForTier keeps full ID verbatim on non-claude runtime', () => { + writeConfig(tmpDir, { + runtime: 'opencode', + model_overrides: { 'gsd-executor': 'claude-sonnet-5' }, + }); + assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-executor', 0), 'claude-sonnet-5'); + }); + + // MEDIUM-1 (review): exercise the unmappable-override fall-through branch in + // resolveModelForTier (closes the mutation-score gap — a future refactor that + // accidentally returned the verbatim override instead of falling through + // would otherwise survive the suite). + test('resolveModelForTier unmappable claude ID falls through to tier alias on claude', () => { + resetRuntimeWarningCaches(); + writeConfig(tmpDir, { + runtime: 'claude', + model_profile: 'balanced', + model_overrides: { 'gsd-planner': 'claude-opus-4-5' }, + }); + // unmappable override → fall through → no dynamic_routing → resolveModelInternal → 'opus' + assert.strictEqual(resolveModelForTier(tmpDir, 'gsd-planner', 0), 'opus'); + }); + + // LOW-2 (review): pin the case-sensitive contract — a case-variant like + // "Claude-Sonnet-5" is NOT mapped (alias keys are case-sensitive, matching + // the model_policy path and the Claude API). + test('model_overrides case-variant "Claude-Sonnet-5" passes through verbatim (case-sensitive contract)', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-executor': 'Claude-Sonnet-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'Claude-Sonnet-5'); + }); + + // Regression guard: non-Claude custom / vendor values still pass through verbatim + // on the claude runtime (the fix must NOT touch values that aren't Claude IDs). + test('model_overrides non-Claude custom model passes through verbatim on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-planner': 'my-custom-model' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-planner'), 'my-custom-model'); + }); + + test('model_overrides non-Claude vendor ID (openai/gpt-5) passes through verbatim on runtime:claude', () => { + writeConfig(tmpDir, { + runtime: 'claude', + model_overrides: { 'gsd-executor': 'openai/gpt-5' }, + }); + assert.strictEqual(resolveModelInternal(tmpDir, 'gsd-executor'), 'openai/gpt-5'); + }); +}); + // ─── resolveModelForTier: model_policy beats dynamic_routing ───────────────── describe('#49 resolveModelForTier: model_policy beats dynamic_routing', () => { diff --git a/tests/repo-layout.test.cjs b/tests/repo-layout.test.cjs index eb6465df9..61a95c8d0 100644 --- a/tests/repo-layout.test.cjs +++ b/tests/repo-layout.test.cjs @@ -20,22 +20,36 @@ const test = require('node:test'); const assert = require('node:assert/strict'); const fs = require('fs'); const path = require('path'); +const { execFileSync } = require('child_process'); const ROOT = path.resolve(__dirname, '..'); -test('repo-layout: root AGENTS.md is absent — no ad-hoc AI instruction file committed alongside CONTEXT.md', () => { - const agentsMdPath = path.join(ROOT, 'AGENTS.md'); +test('repo-layout: root AGENTS.md is not git-tracked — no ad-hoc AI instruction file committed alongside CONTEXT.md', () => { + // The installer legitimately writes AGENTS.md to process.cwd() when + // `gsd install copilot` runs inside a repo checkout (issue #786). The file + // may exist on disk — that's expected after a local install. What must NOT + // happen is committing it to git, where editors and AI tools would silently + // pick up the installer-generated stub instead of CONTEXT.md. + let tracked; + try { + execFileSync('git', ['ls-files', '--error-unmatch', 'AGENTS.md'], { + cwd: ROOT, encoding: 'utf8', stdio: 'pipe', + }); + tracked = true; + } catch { + tracked = false; + } assert.equal( - fs.existsSync(agentsMdPath), + tracked, false, [ - 'root AGENTS.md must not be committed.', + 'root AGENTS.md must not be git-tracked.', 'This file is written by `gsd install copilot` (bin/install.js, local Copilot path, issue #786)', - 'when the installer runs inside a repo checkout.', + 'when the installer runs inside a repo checkout — its presence on disk is fine,', + 'but it must never be committed.', 'The repository source of truth for architecture and contributor guidance is', 'CONTEXT.md and docs/adr/ — not an installer-generated instruction stub.', - 'Run `gsd uninstall copilot` to remove the artefact, then verify it is gitignored', - 'before re-running the install in this checkout.', + 'Run `git rm --cached AGENTS.md` to untrack it if accidentally staged.', ].join(' '), ); }); diff --git a/tests/runtime-artifact-layout-descriptor-drive.test.cjs b/tests/runtime-artifact-layout-descriptor-drive.test.cjs index e31e064fe..686529649 100644 --- a/tests/runtime-artifact-layout-descriptor-drive.test.cjs +++ b/tests/runtime-artifact-layout-descriptor-drive.test.cjs @@ -83,20 +83,26 @@ const GOLDEN = { // ── copilot ────────────────────────────────────────────────────────────────── // Old switch: no scope branch → local == global. 5b backfill restores this. + // #1575: agents kind added (copilot cutover — .agent.md rename handled by _copyStaged). 'copilot/global': [ { kind: 'skills', destSubpath: 'skills', prefix: 'gsd-' }, + { kind: 'agents', destSubpath: 'agents', prefix: 'gsd-' }, ], 'copilot/local': [ { kind: 'skills', destSubpath: 'skills', prefix: 'gsd-' }, + { kind: 'agents', destSubpath: 'agents', prefix: 'gsd-' }, ], // ── antigravity ────────────────────────────────────────────────────────────── // Old switch: no scope branch → local == global. 5b backfill restores this. + // #1575: agents kind added (antigravity cutover — scope-aware converter). 'antigravity/global': [ { kind: 'skills', destSubpath: 'skills', prefix: 'gsd-' }, + { kind: 'agents', destSubpath: 'agents', prefix: 'gsd-' }, ], 'antigravity/local': [ { kind: 'skills', destSubpath: 'skills', prefix: 'gsd-' }, + { kind: 'agents', destSubpath: 'agents', prefix: 'gsd-' }, ], // ── windsurf ───────────────────────────────────────────────────────────────── diff --git a/tests/runtime-artifact-layout.test.cjs b/tests/runtime-artifact-layout.test.cjs index 1bb726368..e0df1ac38 100644 --- a/tests/runtime-artifact-layout.test.cjs +++ b/tests/runtime-artifact-layout.test.cjs @@ -108,11 +108,16 @@ describe('resolveRuntimeArtifactLayout — copilot', () => { const layout = resolveRuntimeArtifactLayout('copilot', FAKE_DIR); assert.strictEqual(layout.runtime, 'copilot'); assert.strictEqual(layout.configDir, FAKE_DIR); - assert.strictEqual(layout.kinds.length, 1); + // #1575: agents kind added (copilot cutover) + assert.strictEqual(layout.kinds.length, 2); assert.strictEqual(layout.kinds[0].kind, 'skills'); assert.strictEqual(layout.kinds[0].destSubpath, 'skills'); assert.strictEqual(layout.kinds[0].prefix, 'gsd-'); assert.strictEqual(typeof layout.kinds[0].stage, 'function'); + assert.strictEqual(layout.kinds[1].kind, 'agents'); + assert.strictEqual(layout.kinds[1].destSubpath, 'agents'); + assert.strictEqual(layout.kinds[1].prefix, 'gsd-'); + assert.strictEqual(typeof layout.kinds[1].stage, 'function'); }); }); @@ -121,11 +126,16 @@ describe('resolveRuntimeArtifactLayout — antigravity', () => { const layout = resolveRuntimeArtifactLayout('antigravity', FAKE_DIR); assert.strictEqual(layout.runtime, 'antigravity'); assert.strictEqual(layout.configDir, FAKE_DIR); - assert.strictEqual(layout.kinds.length, 1); + // #1575: agents kind added (antigravity cutover) + assert.strictEqual(layout.kinds.length, 2); assert.strictEqual(layout.kinds[0].kind, 'skills'); assert.strictEqual(layout.kinds[0].destSubpath, 'skills'); assert.strictEqual(layout.kinds[0].prefix, 'gsd-'); assert.strictEqual(typeof layout.kinds[0].stage, 'function'); + assert.strictEqual(layout.kinds[1].kind, 'agents'); + assert.strictEqual(layout.kinds[1].destSubpath, 'agents'); + assert.strictEqual(layout.kinds[1].prefix, 'gsd-'); + assert.strictEqual(typeof layout.kinds[1].stage, 'function'); }); });