diff --git a/.changeset/1143-claude-orchestration-capability.md b/.changeset/1143-claude-orchestration-capability.md new file mode 100644 index 000000000..d87aae625 --- /dev/null +++ b/.changeset/1143-claude-orchestration-capability.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 2044 +--- +**A default-off, BETA, claude-only "Claude orchestration" capability** — adopts Claude Code's Workflow tool (`/effort ultracode`, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, restoring the wave parallelism + plan-checker + verifier that the #853 backgrounded-agent nesting limitation forces inline on Claude Code, and folding the existing `gsd-ultraplan-phase` plan-offload under the same runtime gate. When `claude_orchestration.enabled` is on AND the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets the floor (`claude_orchestration.min_agent_sdk_version`, default `0.3.149`), `execute-phase` emits a generated Workflow script (`waves → parallel() barriers`, `plans → agent({ agentType: 'gsd-executor', isolation: 'worktree' })`, `files_modified overlap → separate sequential stages`, `resumeFromRunId` wired to the phase run id, shared `budget` pool) that composes the SAME executor agent + worktree isolation the inline path uses, so artifacts/commits are produced identically. Detection is pure and fail-closed (any miss → inline), so on any runtime lacking the Workflow tool behaviour is byte-identical to today. Adds a pure module `gsd-core/bin/lib/claude-orchestration.cjs` (`detectWorkflowBackend`, `emitWorkflowScript`), the `capabilities/claude-orchestration/` declaration with two gated loop contributions (`execute:wave:post`, `plan:post`) and a `claude-orchestration` command family (`gsd-tools claude-orchestration detect-backend|emit-workflow`), federated config keys, and an ADR-1143 implementation amendment. (#1143) diff --git a/CONTEXT.md b/CONTEXT.md index 0b191c568..02e92c938 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -220,6 +220,8 @@ ADR-1244 Phase 4 (D5+D6) orchestration seam (`gsd-core/bin/lib/capability-lifecy ### Capability Command Dispatch ADR-1244 Phase 5 (D7) registry-driven dispatch of capability command families. First-party families (`graphify`/`intel`/`audit`, shipped in `bin/lib/`) dispatch via `dispatchCapabilityCommand` (`gsd-core/bin/gsd-tools.cjs`) against the FROZEN `capability-registry.cjs` `commandFamilies` (confined to `bin/lib/`) — unchanged. Third-party (installed overlay) families dispatch via `dispatchOverlayCapabilityCommand`: after the first-party path returns false, it calls `loadRegistry({ includeInstalled, cwd })` and dispatches a family iff its `capId` is in `_overlay.commandRoots` — which `capability-loader.cjs` populates ONLY for accepted overlay capabilities that declare `commands` AND pass the loader's activation gate (a **committed** ledger entry, present and non-`_pending`, PLUS — for PROJECT scope — a matching user consent record in the Capability Consent Store; GLOBAL scope needs no consent record). A bundle dropped on disk with no install (no ledger entry) or no on-this-machine consent is NOT command-dispatchable. The router module is `require()`'d FROM the capability's install root via `defaultRequireFromInstallRoot` (bare-`.cjs` basename + `realpath` containment, rejecting `..` traversal and symlink escape); same own-property/function/sync-only guards as the first-party path. Wired into the `runCommand` default arm before "Unknown command". A repo-planted project ledger no longer activates anything on its own (#1459) — see `docs/explanation/capability-trust-model.md` "project-scope trust boundary". +### Claude Orchestration Capability +Default-off, BETA, claude-only Capability (`capabilities/claude-orchestration/`, `role: feature`, `runtimeCompat.supported: ["claude"]`, `tier: full`, `activationKey: claude_orchestration.enabled`) adopting Claude Code's Workflow tool (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149) as an optional parallel-execution backend for the GSD loop, and folding the `gsd-ultraplan-phase` plan-offload under the same runtime gate (#1143; ADR-1143). Pure, fail-closed core in `gsd-core/bin/lib/claude-orchestration.cjs` (generated from `src/claude-orchestration.cts`): `detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) → { available, backend:'workflow'|'inline', reason }` (gate ladder: enabled → Claude → execution_backend ≠ inline → host dispatch nested+background → valid Agent SDK → SDK ≥ floor; every miss degrades to `inline`, never throws); `emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) → { ok, script, summary }` mapping waves → `parallel()` stage barriers, plans → `agent({ agentType:'gsd-executor', isolation:'worktree' })`, `files_modified` overlap → separate sequential stages (greedy first-fit), `resumeFromRunId` wired to the run id, shared `budget(tokens)`; all interpolated identifiers validated script-safe (no `"`,`\`,control chars) and briefs JSON-quoted (review anti-injection). Registers two loop contributions at WIRED points only (execute:wave:pre/execute:pre are declared but not rendered, same constraint external-job documents): `execute:wave:post into:executor` (Workflow-backend guidance) and `plan:post into:planner` (ultraplan ownership declaration), both `when: claude_orchestration.enabled`, `onError: skip`. Federated config keys (`claude_orchestration.enabled` default false, `execution_backend` enum auto|workflow|inline default auto, `min_agent_sdk_version` string default "0.3.149") live only in the registry — uninstall removes them cleanly. Pre-release versions of the floor compare below GA (SemVer precedence). Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code; on any runtime lacking the Workflow tool, behaviour is byte-identical to today. BETA v1 ships detection + emission + declarative ultraplan ownership + a `claude-orchestration` command family (`gsd-tools claude-orchestration detect-backend|emit-workflow`, router `gsd-core/bin/lib/claude-orchestration-command-router.cjs` from `src/claude-orchestration-command-router.cts`); full install-profile migration of the ultraplan skill into `skills[]` is a follow-up (CLUSTERS/profile gate). Test anchors: `tests/claude-orchestration.test.cjs`, `tests/claude-orchestration-command-router.test.cjs`. ### Loop Extension Point A named, stable site on a host loop step (per-step `pre`/`post` plus per-wave in Execute; 12 total) where Capabilities register hooks. Three hook kinds: `step` (runs as its own sequenced unit), `contribution` (injects into the core step's prompt/context), and `gate` (checks and optionally blocks via a declared `blocking` flag). Each hook declares the artifacts it produces and consumes; hook order is derived by topological sort of that produces/consumes graph (capability-id tiebreak), which also defines data flow — file-artifact based, surviving `/clear` and fresh executor contexts. Hooks are surfaced by runtime resolution with concrete projection: the workflow calls a query that resolves the active hooks and returns fully-rendered, ordered markdown for the executor. Failure is default-resilient — a non-gate hook that errors is skipped with a warning; a hook may opt into `onError: halt`. Part of the Capability system. ADR-857 phase 3c ships the registry-consuming query layer: `gsd-core/bin/lib/loop-resolver.cjs` exposes `resolveLoopHooks({ point, registry, config })` (pure, no I/O), `renderLoopHooks(resolved)` (pure markdown renderer), and `cmdLoopRenderHooks(cwd, point, raw, opts)` (I/O entry point); activated via `gsd-tools loop render-hooks ` which emits `{ point, activeHooks[], rendered }`. Activation is driven by `when` (dotted config key resolved against `loadConfig`), with inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard. The first phase-6 cutovers wiring workflows to this query have landed — ui-phase at `plan:pre` and ui-review at `verify:post` (in `plan-phase.md`/`autonomous.md`); further per-feature cutovers are ongoing. diff --git a/capabilities/claude-orchestration/capability.json b/capabilities/claude-orchestration/capability.json new file mode 100644 index 000000000..95474d492 --- /dev/null +++ b/capabilities/claude-orchestration/capability.json @@ -0,0 +1,85 @@ +{ + "id": "claude-orchestration", + "role": "feature", + "version": "1.7.0-rc.3", + "title": "Claude orchestration (Workflow backend)", + "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", + "tier": "full", + "requires": [], + "engines": { + "gsd": ">=1.7.0" + }, + "runtimeCompat": { + "supported": [ + "claude" + ], + "unsupported": [] + }, + "skills": [], + "agents": [], + "hooks": [], + "commands": [ + { + "family": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand", + "subcommands": [ + "detect-backend", + "emit-workflow" + ] + } + ], + "activationKey": "claude_orchestration.enabled", + "config": { + "claude_orchestration.enabled": { + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "type": "enum", + "values": [ + "auto", + "workflow", + "inline" + ], + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability." + }, + "claude_orchestration.min_agent_sdk_version": { + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + } + }, + "steps": [], + "contributions": [ + { + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, + { + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + } + ], + "gates": [] +} diff --git a/capabilities/claude-orchestration/fragments/execute-wave-post.md b/capabilities/claude-orchestration/fragments/execute-wave-post.md new file mode 100644 index 000000000..db0e76d5a --- /dev/null +++ b/capabilities/claude-orchestration/fragments/execute-wave-post.md @@ -0,0 +1,64 @@ +# Claude orchestration — Workflow execution backend (BETA) + +> Injected at `execute:wave:post` `into: executor` only when +> `claude_orchestration.enabled` is true. Default-off; `onError: skip`. + +## When this contribution is active + +The Claude orchestration capability is **default-off and BETA**. It activates only +when ALL of the following hold: + +1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND +2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent + SDK-specific), AND +3. `claude_orchestration.execution_backend` resolves to `workflow` — either + explicitly, or via `auto` — **and** the Agent SDK version is + `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK + floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release + or older SDK never activates the preview backend). + +Detection is fail-closed: any miss degrades to **inline, manual, one-agent-per- +message dispatch** — exactly today's behaviour. On a non-Claude runtime this +contribution is a no-op. + +## What the executor does when the Workflow backend is active + +Instead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor, +isolation=worktree, run_in_background=true)` per message (which on Claude Code +cannot nest further subagents — #853 — and so degrades to sequential inline +execution), execute-phase **emits a generated Workflow script** and lets the main +loop orchestrate it: + +- **waves → one or more sequential `parallel()` barriers** — each wave is a + barrier group; when plans within a wave share `files_modified`, they are split + into separate sequential stages within that wave's barrier (the next wave + still waits for the previous wave to complete). +- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`** + — the SAME executor agent and worktree isolation the inline path uses, so the + produced `SUMMARY.md` and commits are identical. +- **`files_modified` overlap → separate sequential stages** — two plans that + touch the same file are placed in different stages within the wave (the same + overlap rule execute-phase already applies inline). +- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase + resumes without re-running completed plans. +- **`budget(tokens)`** — a shared token pool across the whole phase when the + orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a + function parameter, not a config key; the orchestrator decides the budget). + +The emitter is a pure function exposed through the capability command surface: +`gsd-tools claude-orchestration emit-workflow --waves --run-id +[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript` +directly). It maps the phase's wave/plan manifest to the Workflow script string +and never invokes the Workflow tool itself; the orchestrator runs the emitted +script. Detection is resolved by the orchestrator calling the pure +`detectWorkflowBackend` with the LIVE host descriptor (the CLI +`gsd-tools claude-orchestration detect-backend` is a simulation harness that +assumes a capable host unless `--no-nested-dispatch` is passed — it does not probe +the real runtime; the orchestrator supplies the real descriptor). + +## Fallback contract + +If detection resolves to `inline` (tool absent, SDK too old, runtime not Claude, +or the capability disabled), execute-phase MUST proceed with the standard inline +wave dispatch. The executor MUST NOT assume parallelism, a shared budget, or +resume-from-run-id semantics in that mode. diff --git a/capabilities/claude-orchestration/fragments/plan-post.md b/capabilities/claude-orchestration/fragments/plan-post.md new file mode 100644 index 000000000..bec9b01fa --- /dev/null +++ b/capabilities/claude-orchestration/fragments/plan-post.md @@ -0,0 +1,28 @@ +# Claude orchestration — ultraplan plan-offload ownership (BETA) + +> Injected at `plan:post` `into: planner` only when +> `claude_orchestration.enabled` is true. Default-off; `onError: skip`. + +## Ownership declaration + +The `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to +Claude Code's ultraplan cloud) is **owned by this capability**, not by a +standalone BETA skill. Both surfaces share one runtime gate +(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only +detection seam. + +## When the planner should consider ultraplan offload + +When this contribution is active (capability enabled, Claude Code runtime), the +planner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local +`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is +advisory, not mandatory — the stable local planner remains the default. + +## Fallback contract + +If the capability is disabled, or the runtime is not Claude Code, ultraplan +offload is **not surfaced** and the planner proceeds with the standard local +`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed +(its own runtime gate already no-ops on non-Claude runtimes); this contribution +only governs whether the capability manifest advertises it as part of the +orchestration surface. diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 2c94b3bc9..0bc7c3b28 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -306,6 +306,8 @@ "capability-writer.cjs", "check-command-router.cjs", "cjs-command-router-adapter.cjs", + "claude-orchestration-command-router.cjs", + "claude-orchestration.cjs", "cli-exit.cjs", "cli-skew-check.cjs", "clock.cjs", diff --git a/docs/adr/1143-claude-orchestration-capability.md b/docs/adr/1143-claude-orchestration-capability.md index 128a40909..4541bce43 100644 --- a/docs/adr/1143-claude-orchestration-capability.md +++ b/docs/adr/1143-claude-orchestration-capability.md @@ -86,3 +86,39 @@ These existing multi-model features (`execute-phase` `cross_ai_delegation`, the - **Neutral:** no effect on non-Claude runtimes by construction; no behavior change until explicitly enabled. > **Governance note:** This ADR is a *draft design* accompanying feature request #1143. Per CONTRIBUTING, it is PR'd only after the issue receives `approved-feature`, and the capability is implemented only after #857 is released. + +## Amendment (2026-07-06): BETA v1 implementation landed + +#857 is **released** (CLOSED); the capability infrastructure is live. The BETA v1 +of this capability has shipped as `capabilities/claude-orchestration/` with the +scope agreed in the Decision, refined to the lowest-risk first slice: + +- **Detection + emission** live as pure, fail-closed functions in + `gsd-core/bin/lib/claude-orchestration.cjs` (source `src/claude-orchestration.cts`): + `detectWorkflowBackend` (gate ladder: enabled → Claude runtime → + execution_backend ≠ inline → host dispatch nested+background → valid Agent SDK + → SDK ≥ `claude_orchestration.min_agent_sdk_version`, default `0.3.149`) and + `emitWorkflowScript` (waves → `parallel()` stage barriers, plans → + `agent({ agentType: 'gsd-executor', isolation: 'worktree' })`, `files_modified` + overlap → separate sequential stages, `resumeFromRunId` wired to the phase run + id, shared `budget(tokens)` pool). All interpolated values are validated as + script-safe identifiers or JSON-quoted (review Finding 1). +- **Loop registration** is at the two **wired** points the loop host contract + actually renders: `execute:wave:post into:executor` (Workflow-backend guidance) + and `plan:post into:planner` (ultraplan ownership declaration). `execute:wave:pre` + and `execute:pre` are declared in the contract but **not wired** today, so the + capability registers at `wave:post` (the constraint `external-job` also documents). +- **Config** is federated (`claude_orchestration.enabled` default false / + `activationKey`, `execution_backend` enum `auto|workflow|inline` default `auto`, + `min_agent_sdk_version`); the keys live only in the registry, so uninstall + removes them cleanly. +- **ultraplan ownership** is declared in the manifest (`plan:post` contribution); + full install-profile migration of the `gsd-ultraplan-phase` skill into the + capability's `skills[]` is deferred to a follow-up (it triggers the CLUSTERS / + profile membership gate and is a heavier, install-machinery change). + +Status remains **Proposed** — the BETA is default-off and the end-to-end Workflow +execution path (actual orchestration via the Workflow tool inside Claude Code) is +not verifiable outside that runtime. The capability is structurally complete and +tested at the contract level; flipping to Accepted follows maintainer sign-off on +the E2E behaviour once exercised on Claude Code with the Workflow tool present. diff --git a/docs/explanation/claude-orchestration-capability.md b/docs/explanation/claude-orchestration-capability.md new file mode 100644 index 000000000..e044a4e1f --- /dev/null +++ b/docs/explanation/claude-orchestration-capability.md @@ -0,0 +1,94 @@ +# Claude orchestration capability (BETA) + +> **Explanation** — *why this capability exists and how it fits the loop.* For the +> step-by-step, see the [capability reference](../reference/capability-matrix.md); +> for the design record, see [ADR-1143](../adr/1143-claude-orchestration-capability.md). + +## The problem + +GSD's `execute-phase` is wave-based: plans carry a wave number, waves run +sequentially, and plans *within* a wave run in parallel when their +`files_modified` sets don't overlap. On most runtimes GSD realizes that by +fanning out one backgrounded `gsd-executor` agent (in a worktree) per plan. + +On **Claude Code** that fan-out degrades. Backgrounded agents on Claude Code have +no `Agent`/`Task` tool, so they cannot nest subagents ([#853]). The autonomous +loop therefore falls back to **inline sequential execution** — and with it +silently drops wave parallelism, the plan-checker, and the verifier — on the one +runtime most GSD users run. + +Claude Code ships an orchestration primitive that sidesteps exactly this: the +**Workflow tool** (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149). +A Workflow script *is* the orchestrator — it runs from the main loop and spawns +subagents itself via `agent()`, `parallel()` (barrier), `pipeline()`, and +`phase()`, with `isolation: 'worktree'`, a shared token `budget`, and +`resumeFromRunId`. + +## The capability + +`claude-orchestration` is a **default-off, BETA, claude-only** capability that +adopts the Workflow tool as an optional, runtime-gated parallel-execution +backend, and folds the existing `gsd-ultraplan-phase` plan-offload under the same +gate. It is blocked-on-nothing now that the ADR-857 capability system is released. + +- **`role: feature`**, `runtimeCompat.supported: ["claude"]`, `tier: full`. +- **`activationKey: claude_orchestration.enabled`** — default `false`. Nothing + changes until you opt in. +- Registers at two **wired** loop points: `execute:wave:post` (into the executor) + and `plan:post` (into the planner). Both are `onError: skip` and gated by the + `enabled` key. + +## How it decides whether to activate + +Detection is a pure, **fail-closed** function — `detectWorkflowBackend`. The +Workflow backend activates only when *every* gate passes; any miss degrades to +`inline` (today's behaviour): + +1. `claude_orchestration.enabled` is true. +2. The runtime is Claude (the Workflow tool is Claude / Agent SDK-specific). +3. `claude_orchestration.execution_backend` is `auto` or `workflow` (not `inline`). +4. The host descriptor advertises `dispatch.nested` **and** `dispatch.background` + (the nesting-capable Claude-Code shape — a proxy for Workflow-tool presence, + meaningful only after gate 2). +5. The Agent SDK reports a valid semver version. +6. That version is `>= claude_orchestration.min_agent_sdk_version` + (default `0.3.149`). A pre-release of the floor (e.g. `0.3.149-rc.1`) compares + *below* the GA release per SemVer, so the preview backend stays off. + +## What the executor runs when the backend is active + +`emitWorkflowScript` maps the phase's wave/plan model onto Workflow primitives: + +| GSD concept | Workflow primitive | +|---|---| +| Wave | `parallel()` stage barrier | +| Plan | `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })` | +| `files_modified` overlap | forces the plans into separate sequential stages | +| Phase run id | `resumeFromRunId("")` | +| Phase token cap | `budget()` | + +Because the emitted script composes the **same** `gsd-executor` agent and +**worktree isolation** the inline path uses, it produces the same `SUMMARY.md` +artifacts and commits — the only difference is the execution vehicle. + +## The fallback contract + +On any runtime lacking the Workflow tool — or when the capability is disabled, +the SDK is too old, or detection fails for any reason — execute-phase proceeds +with the standard inline wave dispatch. This is a release gate, not a nicety: a +regression test asserts the inline fallback on every non-capable combination, so +the capability is default-off and low-risk by construction. + +## BETA scope (v1) + +The first slice ships **detection + emission + declarative ultraplan ownership**. +The emitter is exercised at the contract level (structure, overlap splitting, +resume, budget, anti-injection). End-to-end execution through the Workflow tool +is verifiable only inside Claude Code with the tool present. Full install-profile +migration of the `gsd-ultraplan-phase` skill into the capability's `skills[]` +array is a follow-up (it touches the cluster/profile machinery); for v1 the +manifest *declares* ultraplan ownership at `plan:post` and the existing skill's +own runtime gate continues to no-op on non-Claude runtimes. + +[#853]: https://github.com/open-gsd/gsd-core/issues/853 +[#1143]: https://github.com/open-gsd/gsd-core/issues/1143 diff --git a/docs/how-to/enable-claude-orchestration-workflow-backend.md b/docs/how-to/enable-claude-orchestration-workflow-backend.md new file mode 100644 index 000000000..a8fea05e1 --- /dev/null +++ b/docs/how-to/enable-claude-orchestration-workflow-backend.md @@ -0,0 +1,171 @@ +# How to enable and use the Claude orchestration backend (BETA) + +Run GSD's execute-phase waves through Claude Code's Workflow tool (`/effort ultracode`, Agent SDK ≥ v0.3.149) instead of the default one-agent-per-message dispatch, and fold the `gsd-ultraplan-phase` plan-offload under the same gate. On Claude Code this restores the wave parallelism that backgrounded-agent nesting (#853) otherwise forces inline. + +> **BETA.** This capability tracks a Claude Code preview surface. It is default-off, fail-closed, and Claude-only. Every detection miss degrades silently to today's inline behaviour — enabling it can never break the loop. See the [explanation doc](../explanation/claude-orchestration-capability.md) for the why, and [ADR-1143](../adr/1143-claude-orchestration-capability.md) for the design. + +**What you need:** +- GSD installed with the `full` profile (the capability is `tier: full`). +- **Claude Code** with the Workflow tool available (Agent SDK ≥ `0.3.149`). On any other runtime the capability is an explicit no-op — you can flip the switch safely, nothing happens. +- A GSD project with at least one planned phase (you need a wave/plan manifest to emit a script for). + +--- + +## Step 1 — Enable the capability + +The capability ships disabled. Turn on the master switch inside your GSD project: + +```bash +gsd-tools query config-set claude_orchestration.enabled true +``` + +That single key gates everything — both the Workflow-backend hook at `execute:wave:post` and the ultraplan ownership declaration at `plan:post`. All other `claude_orchestration.*` keys are optional refinements. + +Verify it took: + +```bash +gsd-tools query config-get claude_orchestration.enabled +# → true +``` + +--- + +## Step 2 — Check whether your runtime qualifies + +Detection is fail-closed: the Workflow backend activates only when **every** gate opens. Before relying on it, confirm your runtime reports as capable: + +```bash +gsd-tools claude-orchestration detect-backend \ + --runtime claude \ + --agent-sdk-version 1.2.0 +``` + +You will get one of two results: + +| `backend` | `available` | Meaning | +|-----------|-------------|---------| +| `workflow` | `true` | Every gate passed — the emitter will produce a Workflow script the orchestrator can run. | +| `inline` | `false` | A gate failed. The `reason` field tells you which: `capability_disabled`, `runtime_not_claude`, `backend_inline`, `workflow_tool_unavailable`, `agent_sdk_version_unknown`, or `agent_sdk_version_below_floor`. | + +> **The CLI is a simulation harness, not a probe.** `detect-backend` assumes a capable host descriptor unless you pass `--no-nested-dispatch`. It exists so you (and the orchestrator) can ask "given these facts, would the backend activate?" The real detection the loop uses is the pure `detectWorkflowBackend` function, called with the live host descriptor. + +### If detection returns `inline` + +Work through the `reason`: + +- **`runtime_not_claude`** — you are on Codex / Cursor / opencode / etc. The Workflow tool is Claude-specific; there is nothing to enable here. Your loop is unchanged. +- **`agent_sdk_version_below_floor`** — upgrade Claude Code / the Agent SDK to at least `claude_orchestration.min_agent_sdk_version` (default `0.3.149`). A pre-release of the floor (e.g. `0.3.149-rc.1`) compares *below* the GA release and will not activate. +- **`workflow_tool_unavailable`** — your host descriptor does not advertise nested + background dispatch. This is unusual on Claude Code; if you see it, the Workflow tool is not present in this session. +- **`agent_sdk_version_unknown`** — the version could not be determined. Supply it explicitly via `--agent-sdk-version`. + +### Pin a higher floor (optional) + +If you want to gate the BETA behind a newer Agent SDK than the default: + +```bash +gsd-tools query config-set claude_orchestration.min_agent_sdk_version 1.0.0 +``` + +--- + +## Step 3 — Choose the execution backend + +`claude_orchestration.execution_backend` controls how aggressively the backend is used once detection passes: + +| Value | Behaviour | +|-------|-----------| +| `auto` (default) | Use the Workflow backend **if** detection passes; otherwise inline. The safe, recommended value. | +| `workflow` | Force the Workflow backend when the tool is present (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). | +| `inline` | Force today's manual one-agent-per-message dispatch, even on a capable Claude Code runtime. Use this to A/B compare or to temporarily retire the BETA. | + +Switch with: + +```bash +gsd-tools query config-set claude_orchestration.execution_backend workflow +``` + +--- + +## Step 4 — Emit a Workflow script for a phase + +With the capability enabled and detection passing, generate the Workflow script for a phase's wave/plan manifest. The manifest is the wave/plan model execute-phase already builds: + +```json +{ + "waves": [ + { + "id": "w1", + "plans": [ + { "id": "p1", "brief": "Implement the foo module", "files_modified": ["src/foo.cts"] }, + { "id": "p2", "brief": "Wire the bar seam", "files_modified": ["src/bar.cts"] } + ] + } + ] +} +``` + +Emit the script: + +```bash +gsd-tools claude-orchestration emit-workflow \ + --waves .planning/phases/01-foo/waves.json \ + --run-id phase-01-foo \ + --phase-dir .planning/phases/01-foo \ + --budget 500000 +``` + +The output is a generated Workflow script that maps GSD's model 1:1 onto Workflow primitives: + +- **waves → sequential `parallel()` barriers** (split into separate stages within a wave when `files_modified` overlap), +- **plans → `agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses, +- **`resumeFromRunId("")`** wired to the phase run id, +- **`budget()`** — a shared token pool across the whole phase (omit `--budget` to skip). + +Because the script composes the same `gsd-executor` agent + worktree isolation + `SUMMARY.md` artifact as the inline path, the artifacts and commits it produces are identical — only the execution vehicle differs. + +### Run the emitted script + +Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation). The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, waves barrier between each other, and `resumeFromRunId` lets an interrupted phase resume without re-running completed plans. + +--- + +## Step 5 — Ultraplan plan-offload + +Enabling the capability also folds `gsd-ultraplan-phase` under the same runtime gate. When the capability is on, the planner may offer the `/gsd-ultraplan-phase` path (offload plan-phase to Claude Code's ultraplan cloud) as an alternative to local `/gsd-plan-phase`. This is advisory — the stable local planner remains the default. + +If the capability is off, or the runtime is not Claude Code, ultraplan offload is not surfaced and `/gsd-plan-phase` runs as normal. + +--- + +## Disabling + +To turn the capability off and return to byte-identical inline behaviour: + +```bash +gsd-tools query config-set claude_orchestration.enabled false +``` + +Or force inline dispatch while leaving the capability otherwise on: + +```bash +gsd-tools query config-set claude_orchestration.execution_backend inline +``` + +Either step is sufficient — no uninstall or resurface needed. The federated config keys live only in the capability registry, so they vanish cleanly if the capability is ever removed. + +--- + +## What is and is not wired in BETA v1 + +**Working today:** +- Detection (`detectWorkflowBackend` / `gsd-tools claude-orchestration detect-backend`) — fail-closed, tested across every gate. +- Emission (`emitWorkflowScript` / `gsd-tools claude-orchestration emit-workflow`) — waves→barriers, overlap→stages, resume, budget, anti-injection. +- The contribution fragments at `execute:wave:post` and `plan:post` (gated, `onError: skip`). +- Inline fallback on every non-capable combination (regression-tested). + +**Not yet wired (follow-ups):** +- `execute-phase.md` does not yet auto-branch to emit-and-run the Workflow script. Today you emit the script explicitly (Step 4) and run it via the Workflow tool. Automatic dispatch inside the loop is the next milestone. +- The plan-checker and verifier still run inline — this capability delivers the parallel-execution backend, not those gates. +- Full install-profile migration of the `gsd-ultraplan-phase` skill into the capability's `skills[]` (it is currently declared in the manifest; the skill's own runtime gate continues to no-op on non-Claude runtimes). + +If a preview-API change breaks detection, the capability degrades to inline; it cannot destabilise the core loop. diff --git a/docs/reference/capability-matrix.md b/docs/reference/capability-matrix.md index 4efee626d..c74755eac 100644 --- a/docs/reference/capability-matrix.md +++ b/docs/reference/capability-matrix.md @@ -44,7 +44,7 @@ Core package and are stamped with the package version at release (per ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied to third-party capabilities. -### Feature capabilities (role: feature) — 18 +### Feature capabilities (role: feature) — 19 Feature capabilities extend what the loop does — contributing research, planning, execution, verification, or ship artefacts at the loop extension @@ -55,6 +55,7 @@ points. | `ai-integration` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party | | `assumption-delta` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party | | `audit` | feature | full | `>=1.6.0` | — | — | first-party | +| `claude-orchestration` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party | | `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party | | `drift` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post` | gate | first-party | | `external-job` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party | diff --git a/eslint.config.mjs b/eslint.config.mjs index 580c90822..d382077df 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -56,6 +56,8 @@ export default tseslint.config( 'coverage/**', '**/*.generated.cjs', // ADR-457: tsc-generated runtime artifact — lint the src/*.cts source, not the emitted .cjs. + 'gsd-core/bin/lib/claude-orchestration.cjs', + 'gsd-core/bin/lib/claude-orchestration-command-router.cjs', 'gsd-core/bin/lib/semver-compare.cjs', 'gsd-core/bin/lib/host-integration.cjs', 'gsd-core/bin/lib/handshake-serialized.cjs', diff --git a/gsd-core/bin/lib/capability-registry.cjs b/gsd-core/bin/lib/capability-registry.cjs index a4706bfa3..33e4b1ce1 100644 --- a/gsd-core/bin/lib/capability-registry.cjs +++ b/gsd-core/bin/lib/capability-registry.cjs @@ -424,6 +424,93 @@ const capabilities = { } } }, + "claude-orchestration": { + "id": "claude-orchestration", + "role": "feature", + "version": "1.7.0-rc.3", + "title": "Claude orchestration (Workflow backend)", + "description": "Default-off, BETA, claude-only capability that adopts Claude Code's Workflow tool (the engine behind /effort ultracode) as an optional parallel-execution backend for the GSD loop. When the runtime exposes the Workflow tool and claude_orchestration.execution_backend resolves to 'workflow', execute-phase emits a generated Workflow script (waves -> parallel() barriers, plans -> agent({ agentType: 'gsd-executor', isolation: 'worktree' }), files_modified overlap -> separate sequential stages, resumeFromRunId wired to the phase run id, shared token budget) that composes the SAME gsd-executor agent and worktree isolation the inline path uses, restoring the wave parallelism the #853 backgrounded-agent nesting limitation forces inline on Claude Code. (The plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates.) Also folds the ultraplan plan-offload under one runtime gate (plan:* surface). On any runtime lacking the Workflow tool, or when the capability is disabled, behaviour is byte-identical to today (inline/manual dispatch). Detection + emission live in gsd-core/bin/lib/claude-orchestration.cjs (pure, fail-closed). Mirrors the existing gsd-ultraplan-phase BETA-isolation posture.", + "tier": "full", + "requires": [], + "engines": { + "gsd": ">=1.7.0" + }, + "runtimeCompat": { + "supported": [ + "claude" + ], + "unsupported": [] + }, + "skills": [], + "agents": [], + "hooks": [], + "commands": [ + { + "family": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand", + "subcommands": [ + "detect-backend", + "emit-workflow" + ] + } + ], + "activationKey": "claude_orchestration.enabled", + "config": { + "claude_orchestration.enabled": { + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "type": "enum", + "values": [ + "auto", + "workflow", + "inline" + ], + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability." + }, + "claude_orchestration.min_agent_sdk_version": { + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + } + }, + "steps": [], + "contributions": [ + { + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md", + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:post` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## What the executor does when the Workflow backend is active\n\nInstead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor,\nisolation=worktree, run_in_background=true)` per message (which on Claude Code\ncannot nest further subagents — #853 — and so degrades to sequential inline\nexecution), execute-phase **emits a generated Workflow script** and lets the main\nloop orchestrate it:\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier (the next wave\n still waits for the previous wave to complete).\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n — the SAME executor agent and worktree isolation the inline path uses, so the\n produced `SUMMARY.md` and commits are identical.\n- **`files_modified` overlap → separate sequential stages** — two plans that\n touch the same file are placed in different stages within the wave (the same\n overlap rule execute-phase already applies inline).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n- **`budget(tokens)`** — a shared token pool across the whole phase when the\n orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a\n function parameter, not a config key; the orchestrator decides the budget).\n\nThe emitter is a pure function exposed through the capability command surface:\n`gsd-tools claude-orchestration emit-workflow --waves --run-id \n[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript`\ndirectly). It maps the phase's wave/plan manifest to the Workflow script string\nand never invokes the Workflow tool itself; the orchestrator runs the emitted\nscript. Detection is resolved by the orchestrator calling the pure\n`detectWorkflowBackend` with the LIVE host descriptor (the CLI\n`gsd-tools claude-orchestration detect-backend` is a simulation harness that\nassumes a capable host unless `--no-nested-dispatch` is passed — it does not probe\nthe real runtime; the orchestrator supplies the real descriptor).\n\n## Fallback contract\n\nIf detection resolves to `inline` (tool absent, SDK too old, runtime not Claude,\nor the capability disabled), execute-phase MUST proceed with the standard inline\nwave dispatch. The executor MUST NOT assume parallelism, a shared budget, or\nresume-from-run-id semantics in that mode.\n" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, + { + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md", + "inline": "# Claude orchestration — ultraplan plan-offload ownership (BETA)\n\n> Injected at `plan:post` `into: planner` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## Ownership declaration\n\nThe `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to\nClaude Code's ultraplan cloud) is **owned by this capability**, not by a\nstandalone BETA skill. Both surfaces share one runtime gate\n(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only\ndetection seam.\n\n## When the planner should consider ultraplan offload\n\nWhen this contribution is active (capability enabled, Claude Code runtime), the\nplanner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local\n`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is\nadvisory, not mandatory — the stable local planner remains the default.\n\n## Fallback contract\n\nIf the capability is disabled, or the runtime is not Claude Code, ultraplan\noffload is **not surfaced** and the planner proceeds with the standard local\n`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed\n(its own runtime gate already no-ops on non-Claude runtimes); this contribution\nonly governs whether the capability manifest advertises it as part of the\norchestration surface.\n" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + } + ], + "gates": [] + }, "cline": { "id": "cline", "role": "runtime", @@ -2845,6 +2932,21 @@ const byLoopPoint = { } ], "contributions": [ + { + "capId": "claude-orchestration", + "point": "plan:post", + "into": "planner", + "fragment": { + "path": "fragments/plan-post.md", + "inline": "# Claude orchestration — ultraplan plan-offload ownership (BETA)\n\n> Injected at `plan:post` `into: planner` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## Ownership declaration\n\nThe `gsd-ultraplan-phase` plan-offload surface (offloading GSD's plan phase to\nClaude Code's ultraplan cloud) is **owned by this capability**, not by a\nstandalone BETA skill. Both surfaces share one runtime gate\n(`claude_orchestration.enabled`), one BETA boundary, and one Claude-Code-only\ndetection seam.\n\n## When the planner should consider ultraplan offload\n\nWhen this contribution is active (capability enabled, Claude Code runtime), the\nplanner MAY offer the `/gsd-ultraplan-phase` path as an alternative to local\n`/gsd-plan-phase` for phases where cloud-assisted planning adds value. This is\nadvisory, not mandatory — the stable local planner remains the default.\n\n## Fallback contract\n\nIf the capability is disabled, or the runtime is not Claude Code, ultraplan\noffload is **not surfaced** and the planner proceeds with the standard local\n`/gsd-plan-phase`. The `gsd-ultraplan-phase` command itself remains installed\n(its own runtime gate already no-ops on non-Claude runtimes); this contribution\nonly governs whether the capability manifest advertises it as part of the\norchestration surface.\n" + }, + "produces": [], + "consumes": [ + "CONTEXT.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, { "capId": "external-job", "point": "plan:post", @@ -2885,6 +2987,21 @@ const byLoopPoint = { "execute:wave:post": { "steps": [], "contributions": [ + { + "capId": "claude-orchestration", + "point": "execute:wave:post", + "into": "executor", + "fragment": { + "path": "fragments/execute-wave-post.md", + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:post` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## What the executor does when the Workflow backend is active\n\nInstead of the orchestrator fanning out one `Agent(subagent_type=gsd-executor,\nisolation=worktree, run_in_background=true)` per message (which on Claude Code\ncannot nest further subagents — #853 — and so degrades to sequential inline\nexecution), execute-phase **emits a generated Workflow script** and lets the main\nloop orchestrate it:\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier (the next wave\n still waits for the previous wave to complete).\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n — the SAME executor agent and worktree isolation the inline path uses, so the\n produced `SUMMARY.md` and commits are identical.\n- **`files_modified` overlap → separate sequential stages** — two plans that\n touch the same file are placed in different stages within the wave (the same\n overlap rule execute-phase already applies inline).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n- **`budget(tokens)`** — a shared token pool across the whole phase when the\n orchestrator passes a `budgetTokens` value to `emitWorkflowScript` (it is a\n function parameter, not a config key; the orchestrator decides the budget).\n\nThe emitter is a pure function exposed through the capability command surface:\n`gsd-tools claude-orchestration emit-workflow --waves --run-id \n[--phase-dir ] [--budget ]` (or `require('gsd-core/bin/lib/claude-orchestration.cjs').emitWorkflowScript`\ndirectly). It maps the phase's wave/plan manifest to the Workflow script string\nand never invokes the Workflow tool itself; the orchestrator runs the emitted\nscript. Detection is resolved by the orchestrator calling the pure\n`detectWorkflowBackend` with the LIVE host descriptor (the CLI\n`gsd-tools claude-orchestration detect-backend` is a simulation harness that\nassumes a capable host unless `--no-nested-dispatch` is passed — it does not probe\nthe real runtime; the orchestrator supplies the real descriptor).\n\n## Fallback contract\n\nIf detection resolves to `inline` (tool absent, SDK too old, runtime not Claude,\nor the capability disabled), execute-phase MUST proceed with the standard inline\nwave dispatch. The executor MUST NOT assume parallelism, a shared budget, or\nresume-from-run-id semantics in that mode.\n" + }, + "produces": [], + "consumes": [ + "PLAN.md" + ], + "when": "claude_orchestration.enabled", + "onError": "skip" + }, { "capId": "external-job", "point": "execute:wave:post", @@ -3095,6 +3212,9 @@ const byLoopPoint = { const configKeys = { "workflow.ai_integration_phase": "ai-integration", "workflow.assumption_delta": "assumption-delta", + "claude_orchestration.enabled": "claude-orchestration", + "claude_orchestration.execution_backend": "claude-orchestration", + "claude_orchestration.min_agent_sdk_version": "claude-orchestration", "workflow.code_review": "code-review", "workflow.code_review_depth": "code-review", "workflow.drift_threshold": "drift", @@ -3146,6 +3266,29 @@ const configSchema = { "default": true, "description": "Enable the assumption-delta architecture checkpoint during planning. When a pluralization/optional/chosen signal is detected in the phase scope, the planner is prompted to re-ask whether the primary key / identity model still names the right thing. Advisory (non-blocking)." }, + "claude_orchestration.enabled": { + "owner": "claude-orchestration", + "type": "boolean", + "default": false, + "description": "Master toggle for the Claude orchestration capability. Default-off + BETA: the Workflow-tool execution backend and the ultraplan plan-offload surface are inert unless this is true. When false, loop behaviour is byte-identical to a non-Claude runtime (inline/manual dispatch)." + }, + "claude_orchestration.execution_backend": { + "owner": "claude-orchestration", + "type": "enum", + "default": "auto", + "description": "Which execute-phase dispatch backend to use when the capability is enabled. 'auto' (default) activates the Workflow backend only when the runtime is Claude AND the Workflow tool is detected AND the Agent SDK meets claude_orchestration.min_agent_sdk_version; otherwise it falls back to inline. 'workflow' forces the Workflow backend when the tool is present AND the Agent SDK meets the floor (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). 'inline' forces today's manual one-agent-per-message dispatch regardless of tool availability.", + "values": [ + "auto", + "workflow", + "inline" + ] + }, + "claude_orchestration.min_agent_sdk_version": { + "owner": "claude-orchestration", + "type": "string", + "default": "0.3.149", + "description": "Minimum Agent SDK version required to activate the Workflow backend under execution_backend='auto'. Defaults to 0.3.149 (the release that introduced the Workflow tool). Raise to pin a higher floor; the detection seam fails closed to inline for any runtime reporting an older or unknown version." + }, "workflow.code_review": { "owner": "code-review", "type": "boolean", @@ -4777,6 +4920,11 @@ const commandFamilies = { "module": "audit-command-router.cjs", "router": "routeAuditUat" }, + "claude-orchestration": { + "capId": "claude-orchestration", + "module": "claude-orchestration-command-router.cjs", + "router": "routeClaudeOrchestrationCommand" + }, "extract-messages": { "capId": "profile-pipeline", "module": "profile-pipeline-command-router.cjs", @@ -4916,6 +5064,7 @@ const _requiresGraph = { "audit": [], "augment": [], "claude": [], + "claude-orchestration": [], "cline": [], "code-review": [], "codebuddy": [], diff --git a/gsd-core/bin/lib/claude-orchestration-command-router.cjs b/gsd-core/bin/lib/claude-orchestration-command-router.cjs new file mode 100644 index 000000000..a07ad64c2 --- /dev/null +++ b/gsd-core/bin/lib/claude-orchestration-command-router.cjs @@ -0,0 +1,136 @@ +"use strict"; +/** + * Claude orchestration command router — CLI dispatcher for + * `gsd-tools claude-orchestration `. + * + * #1143 — thin CLI adapter over the pure `claude-orchestration.cjs` module. + * Lets execute-phase (or any orchestrator) invoke the Workflow-backend + * detection and the Workflow-script emitter through the standard capability + * command surface (ADR-959) instead of a bare `require()`. + * + * Router signature: { args, cwd, raw, error } — identical to the other host + * routers; discovered by dispatchCapabilityCommand via the registry's + * commandFamilies index. + * + * Subcommands: + * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] + * Resolves whether the Workflow backend should activate. `--runtime` + * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * `claude_orchestration.*` keys from .planning/config.json. Emits + * { available, backend, reason }. + * + * emit-workflow --waves --run-id [--phase-dir ] [--budget ] + * Reads a wave/plan manifest JSON file and emits the generated Workflow + * script + summary. The manifest shape matches emitWorkflowScript's input: + * { waves: [{ id, plans: [{ id, brief, files_modified: string[] }] }] }. + */ +var __importDefault = (this && this.__importDefault) || function (mod) { + return (mod && mod.__esModule) ? mod : { "default": mod }; +}; +const node_fs_1 = __importDefault(require("node:fs")); +const node_path_1 = __importDefault(require("node:path")); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const io = require("./io.cjs"); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const core = require("./claude-orchestration.cjs"); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const configLoader = require("./config-loader.cjs"); +const { output } = io; +const { detectWorkflowBackend, emitWorkflowScript } = core; +const CAPABLE_HOST = { dispatch: { nested: true, background: true } }; +function usage(error) { + error('Usage: gsd-tools claude-orchestration [...]\n' + + ' detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch]\n' + + ' emit-workflow --waves --run-id [--phase-dir ] [--budget ]'); +} +function argValue(args, flag) { + const i = args.indexOf(flag); + return i !== -1 && i + 1 < args.length ? args[i + 1] : undefined; +} +/** + * Detect whether the Workflow backend should activate for the current/given + * runtime. Reads `claude_orchestration.*` from the project config; runtime and + * SDK version come from flags (the orchestrator already knows these) or env. + */ +function cmdDetectBackend(args, cwd, raw) { + const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; + const agentSdkVersion = argValue(args, '--agent-sdk-version'); + const noNested = args.includes('--no-nested-dispatch'); + const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; + // Resolve the claude_orchestration.* slice from the project config (federated + // keys are merged by loadConfig as a nested object). A config read failure + // degrades to inline — it must not break the core loop. + let claudeSlice = {}; + try { + const loaded = configLoader.loadConfig(cwd); + const slice = loaded['claude_orchestration']; + if (slice && typeof slice === 'object' && !Array.isArray(slice)) { + claudeSlice = slice; + } + } + catch { + claudeSlice = {}; + } + // Flatten the nested slice into the dotted-key shape detectWorkflowBackend expects. + const flatConfig = {}; + for (const k of Object.keys(claudeSlice)) { + flatConfig['claude_orchestration.' + k] = claudeSlice[k]; + } + const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); + output(result, raw); +} +/** + * Emit a Workflow script from a wave/plan manifest file. + */ +function cmdEmitWorkflow(args, _cwd, raw, error) { + const wavesPath = argValue(args, '--waves'); + const runId = argValue(args, '--run-id'); + const phaseDir = argValue(args, '--phase-dir') || '.planning/phases/current'; + const budgetRaw = argValue(args, '--budget'); + if (!wavesPath) { + error('emit-workflow requires --waves '); + return; + } + if (!runId) { + error('emit-workflow requires --run-id '); + return; + } + let waves; + try { + const content = node_fs_1.default.readFileSync(node_path_1.default.resolve(wavesPath), 'utf8'); + const parsed = JSON.parse(content); + waves = parsed['waves']; + } + catch (e) { + error('emit-workflow: could not read/parse --waves file "' + wavesPath + '": ' + (e instanceof Error ? e.message : String(e))); + return; + } + const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; + const budget = (typeof budgetTokens === 'number' && !Number.isNaN(budgetTokens)) ? budgetTokens : undefined; + const result = emitWorkflowScript({ + phaseDir, + runId, + waves: waves, + budgetTokens: budget, + }); + if (!result.ok) { + error('emit-workflow: ' + result.reason); + return; + } + output({ script: result.script, summary: result.summary }, raw); +} +function routeClaudeOrchestrationCommand(opts) { + const { args, cwd, raw, error } = opts; + // args[0] is the family ('claude-orchestration'); the subcommand is args[1]. + const subcommand = args[1]; + if (subcommand === 'detect-backend') { + cmdDetectBackend(args, cwd, raw); + } + else if (subcommand === 'emit-workflow') { + cmdEmitWorkflow(args, cwd, raw, error); + } + else { + usage(error); + } +} +module.exports = { routeClaudeOrchestrationCommand }; diff --git a/gsd-core/bin/lib/claude-orchestration.cjs b/gsd-core/bin/lib/claude-orchestration.cjs new file mode 100644 index 000000000..956fdc2ed --- /dev/null +++ b/gsd-core/bin/lib/claude-orchestration.cjs @@ -0,0 +1,404 @@ +"use strict"; +/** + * Claude Orchestration Capability — Workflow-tool backend detection + emitter + * + * #1143 — adopts Claude Code's Workflow tool (the engine behind `/effort ultracode`) + * as an optional, runtime-gated parallel-execution backend for the GSD loop. + * + * This module is the pure, testable core of the capability. It owns two seams: + * + * detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) + * → { available: boolean, backend: 'workflow'|'inline', reason: string } + * Fail-closed: every miss degrades to `inline` (today's behaviour), so the + * core loop is byte-identical unless every gate opens. This is criteria 3 + 6. + * + * emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) + * → { ok:true, script, summary } | { ok:false, reason } + * Maps GSD's wave/plan model 1:1 onto Workflow primitives: + * wave → sequential `parallel()` stage barriers, + * plan → `agent(brief, { agentType:'gsd-executor', isolation:'worktree' })`, + * files_modified overlap → forces plans into separate sequential stages + * (the same overlap rule execute-phase already applies inline), + * resumeFromRunId → wired to the phase run id, + * budgetTokens → a shared token pool. + * The emitted script composes the SAME gsd-executor agent and worktree + * isolation the inline path uses, so it produces the same artifacts/commits + * (criterion 2). It is a generated string consumed by the orchestrator; this + * module never invokes the Workflow tool itself. + * + * Design laws: + * - Gall's Law: ship a small working slice that composes existing primitives + * (gsd-executor + worktree isolation) rather than reinventing them. + * - Greenspun's Tenth Rule (cited in #1143): adopt the Workflow tool's + * barrier/pipeline/budget/resume semantics instead of hand-rolling them. + * - Postel's Law: liberal in input (missing fields → inline), conservative in + * output (workflow only when every gate opens). + * - Fail-closed: an unknown version, a missing descriptor, or a disabled + * toggle all resolve to `inline`, never to `workflow`. + * + * Zero external dependencies. Pure functions. Never throws on bad input. + */ +// ─── Constants ──────────────────────────────────────────────────────────────── +/** + * The Agent SDK version that introduced the Workflow tool (#1143 prior art). + * Used as the default floor when config does not override it. A runtime reporting + * an agentSdkVersion below this cannot host the Workflow backend. + */ +const WORKFLOW_TOOL_FLOOR_VERSION = '0.3.149'; +/** Closed enum for the `claude_orchestration.execution_backend` config key. */ +const BACKEND_VALUES = new Set(['auto', 'workflow', 'inline']); +/** Only this runtime can host the Workflow tool (Claude Code / Agent SDK). */ +const WORKFLOW_RUNTIME = 'claude'; +// ─── Semver helpers ─────────────────────────────────────────────────────────── +/** Official-ish strict SemVer 2.0.0 numeric triple (+ optional pre/build). */ +const SEMVER_RE = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$/; +/** True for a syntactically valid semver string. */ +function isValidSemver(s) { + return typeof s === 'string' && SEMVER_RE.test(s); +} +/** + * Compare two semver strings. + * Returns -1/0/1 in the usual sense. Garbage in either position → -1 (fail-closed: + * an unparseable version is treated as "less than" any real floor, so detection + * never accidentally enables the preview backend on an unknown SDK). + * + * Pre-release/build metadata are ignored for the comparison — only the numeric + * major.minor.patch triple participates, matching how the Workflow-tool floor is + * specified (a plain "0.3.149"). + */ +function compareSemver(a, b) { + if (!isValidSemver(a) || !isValidSemver(b)) + return -1; + // Split numeric triple from pre-release/build metadata. + const parseTriple = (s) => { + const core = s.split('-')[0].split('+')[0].split('.'); + return [parseInt(core[0], 10), parseInt(core[1], 10), parseInt(core[2], 10)]; + }; + const hasPre = (s) => s.indexOf('-') !== -1; + const preIdentifiers = (s) => (s.split('-')[1] || '').split('+')[0].split('.').filter((x) => x.length > 0); + const am = parseTriple(a); + const bm = parseTriple(b); + for (let i = 0; i < 3; i++) { + if (am[i] < bm[i]) + return -1; + if (am[i] > bm[i]) + return 1; + } + // Numeric triple is equal. SemVer 2.0.0 §11 precedence: + // - a version WITH a pre-release tag is LOWER than the same triple WITHOUT one + // (keeps the floor fail-closed for pre-release builds of the GA floor); + // - two pre-releases of the same triple are ordered by their dot-separated + // identifiers (numeric < alphanumeric; numeric compared numerically, + // alphanumeric lexically; fewer identifiers < more). + const aPre = hasPre(a); + const bPre = hasPre(b); + if (aPre && !bPre) + return -1; + if (!aPre && bPre) + return 1; + if (aPre && bPre) { + const ai = preIdentifiers(a); + const bi = preIdentifiers(b); + const len = Math.min(ai.length, bi.length); + for (let i = 0; i < len; i++) { + const ax = ai[i]; + const bx = bi[i]; + const aNum = /^\d+$/.test(ax); + const bNum = /^\d+$/.test(bx); + if (aNum && bNum) { + const an = parseInt(ax, 10); + const bn = parseInt(bx, 10); + if (an < bn) + return -1; + if (an > bn) + return 1; + } + else if (aNum && !bNum) { + return -1; // numeric identifiers always lower than alphanumeric + } + else if (!aNum && bNum) { + return 1; + } + else { + if (ax < bx) + return -1; + if (ax > bx) + return 1; + } + } + if (ai.length < bi.length) + return -1; + if (ai.length > bi.length) + return 1; + } + return 0; +} +/** Inline result shorthand. */ +function inline(reason, available = false) { + return { available, backend: 'inline', reason }; +} +/** + * Resolve whether the Workflow-tool backend should activate. + * + * Gate ladder (all must pass for `workflow`; first miss wins, fail-closed): + * 1. capability enabled (claude_orchestration.enabled truthy) + * 2. runtime is Claude (the only runtime that exposes the Workflow tool) + * 3. execution_backend !== 'inline' + * 4. host descriptor signals nested+background dispatch (Workflow-tool capable) + * 5. agentSdkVersion is a known, valid semver + * 6. agentSdkVersion >= the configured floor (default WORKFLOW_TOOL_FLOOR_VERSION) + * 7. execution_backend === 'workflow' OR 'auto' (both reach here; 'inline' exited at 3) + * + * Never throws. Destructures defensively. + */ +function detectWorkflowBackend(input) { + if (input === null || input === undefined || typeof input !== 'object') { + return inline('capability_disabled'); + } + const cfg = (input.config !== null && input.config !== undefined && typeof input.config === 'object') + ? input.config + : {}; + // 1. capability must be opted in (default-off — ships disabled). + if (!cfg['claude_orchestration.enabled']) { + return inline('capability_disabled'); + } + // 2. only Claude can host the Workflow tool. + if (input.runtimeId !== WORKFLOW_RUNTIME) { + return inline('runtime_not_claude'); + } + // 3. explicit inline opt-out short-circuits. + let backendRaw = cfg['claude_orchestration.execution_backend']; + if (typeof backendRaw !== 'string' || !BACKEND_VALUES.has(backendRaw)) { + backendRaw = 'auto'; + } + if (backendRaw === 'inline') { + return inline('backend_inline'); + } + // 4. the host dispatch descriptor must be the nesting-capable Claude-Code shape + // (a proxy for Workflow-tool presence). This is Claude-specific and already + // gated at step 2; `background:true` alone is true on several non-Claude hosts, + // so the proxy is only meaningful after the runtime check above. Note: this is + // NOT the canonical `shouldFlattenDispatch` rule (which keys on + // `backgroundDispatch`); the Workflow backend works precisely because a single + // tool-call orchestrates internally, sidestepping the backgroundDispatch:false + // limitation. Missing/false/foreign descriptor → fail-closed. + const hi = input.hostIntegration; + if (hi === null || hi === undefined || typeof hi !== 'object' || Array.isArray(hi)) { + return inline('workflow_tool_unavailable'); + } + const dispatch = hi.dispatch; + if (typeof dispatch !== 'object' || dispatch === null || Array.isArray(dispatch)) { + return inline('workflow_tool_unavailable'); + } + const nested = dispatch['nested']; + const background = dispatch['background']; + if (nested !== true || background !== true) { + return inline('workflow_tool_unavailable'); + } + // 5. an unknown agentSdkVersion cannot be trusted to meet the floor. + if (!isValidSemver(input.agentSdkVersion)) { + return inline('agent_sdk_version_unknown'); + } + // 6. version floor (config override > default constant). + const floorRaw = cfg['claude_orchestration.min_agent_sdk_version']; + const floor = typeof floorRaw === 'string' && isValidSemver(floorRaw) ? floorRaw : WORKFLOW_TOOL_FLOOR_VERSION; + if (compareSemver(input.agentSdkVersion, floor) < 0) { + return inline('agent_sdk_version_below_floor'); + } + // 7. auto/workflow both reach the workflow backend once every gate passes. + return { available: true, backend: 'workflow', reason: 'workflow_backend_active' }; +} +/** + * Partition a wave's plans into a near-minimal number of sequential stages (via + * greedy first-fit — not guaranteed optimal for arbitrary overlap graphs, but + * correct: no two plans sharing a file ever cohabit a stage) such that no two + * plans in the same stage share a modified file. Each plan goes into the earliest + * stage where it does not overlap any plan already there. + * + * A plan with an EMPTY files_modified set declares no files; it overlaps nothing + * and coalesces into stage 0 (same behavior as the inline path, which also cannot + * guard against undeclared concurrent writes — declare filesModified accurately). + * + * This is the same overlap rule execute-phase applies inline — the only difference + * is the execution vehicle (Workflow `parallel()` vs one-agent-per-message). + */ +function partitionStages(plans) { + const stages = []; + for (const plan of plans) { + const fileSet = new Set(plan.files_modified); + let placed = false; + for (const stage of stages) { + let overlap = false; + for (const f of fileSet) { + if (stage.files.has(f)) { + overlap = true; + break; + } + } + if (!overlap) { + stage.plans.push(plan); + for (const f of fileSet) + stage.files.add(f); + placed = true; + break; + } + } + if (!placed) { + stages.push({ plans: [plan], files: new Set(fileSet) }); + } + } + return stages.map((s) => s.plans.map((p) => p.id)); +} +/** + * Quote a free-text value for safe embedding as a JavaScript/Workflow double-quoted + * string literal. Uses JSON.stringify so every JS-relevant escape (backslash, quote, + * newline, tab, NUL, U+2028/U+2029, all control chars) is handled by the language + * itself — there is no hand-rolled escape table to drift. Returns the value already + * wrapped in its surrounding quotes. + */ +function quoteString(s) { + return JSON.stringify(s); +} +/** + * True if `s` is a safe identifier/path token to interpolate into the generated + * script WITHOUT requiring a string-literal context — i.e. it contains no + * character that could terminate a comment line (`\n`/`\r`), break out of a + * string literal (`"` / `\`), or smuggle a NUL/control sequence. Used for + * `phaseDir`, `runId`, `wave.id`, and `plan.id`, which are identifiers/paths and + * must never legitimately contain such characters. Rejecting them at validation + * (rather than silently flattening) keeps the emitted script faithful to input. + */ +const UNSCRIPTABLE_CHAR_RE = /[\r\n"\\\x00-\x1f\x7f\u2028\u2029]/; +function isScriptableIdentifier(s) { + if (typeof s !== 'string' || s.length === 0) + return false; + return !UNSCRIPTABLE_CHAR_RE.test(s); +} +/** + * Emit a Workflow script mapping the phase's wave/plan model onto Workflow + * primitives. Pure and deterministic: identical input yields an identical string. + * + * Returns ok:false (never throws) on invalid input — empty waves, missing runId, + * a wave with no plans, etc. + */ +function emitWorkflowScript(input) { + if (input === null || input === undefined || typeof input !== 'object') { + return { ok: false, reason: 'invalid_input' }; + } + const { phaseDir, waves, runId } = input; + // Identifiers/paths interpolated into the generated script must be free of any + // character that could terminate a comment, break out of a string literal, or + // smuggle control bytes — reject up front (security: #1143 review Finding 1). + if (!isScriptableIdentifier(phaseDir)) { + return { ok: false, reason: 'phaseDir must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!isScriptableIdentifier(runId)) { + return { ok: false, reason: 'runId must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(waves) || waves.length === 0) { + return { ok: false, reason: 'waves must be a non-empty array' }; + } + for (let i = 0; i < waves.length; i++) { + const w = waves[i]; + if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { + return { ok: false, reason: 'waves[' + i + '] must be { id, plans: non-empty[] }' }; + } + if (!isScriptableIdentifier(w.id)) { + return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(w.plans) || w.plans.length === 0) { + return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; + } + const seenIds = new Set(); + for (let j = 0; j < w.plans.length; j++) { + const p = w.plans[j]; + if (p === null || typeof p !== 'object' || typeof p.id !== 'string' || typeof p.brief !== 'string' || !Array.isArray(p.files_modified)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '] must be { id, brief, files_modified[] }' }; + } + if (!isScriptableIdentifier(p.id)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (seenIds.has(p.id)) { + return { ok: false, reason: 'waves[' + i + '] has duplicate plan id "' + p.id + '"' }; + } + seenIds.add(p.id); + for (const f of p.files_modified) { + if (typeof f !== 'string' || f.length === 0) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].files_modified entries must be non-empty strings' }; + } + } + } + } + const budgetTokens = (typeof input.budgetTokens === 'number' && Number.isFinite(input.budgetTokens) && input.budgetTokens > 0) + ? Math.floor(input.budgetTokens) + : null; + const lines = []; + lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); + lines.push('// phase: ' + phaseDir); + lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); + lines.push('// Composes the SAME gsd-executor agent + worktree isolation as the inline path,'); + lines.push('// so artifacts (SUMMARY.md) and commits are produced identically.'); + lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + if (budgetTokens !== null) { + lines.push('budget(' + budgetTokens + ')'); + } + lines.push(''); + const stagesByWave = []; + let totalPlans = 0; + for (let wi = 0; wi < waves.length; wi++) { + const wave = waves[wi]; + const stages = partitionStages(wave.plans); + stagesByWave.push(stages); + totalPlans += wave.plans.length; + lines.push('// Wave ' + wave.id); + for (let si = 0; si < stages.length; si++) { + const stagePlanIds = stages[si]; + // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). + const stagePlans = stagePlanIds.map((id) => wave.plans.find((p) => p.id === id)); + if (stages.length > 1) { + lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); + } + if (stagePlans.length === 1) { + const p = stagePlans[0]; + lines.push('parallel('); + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" })'); + lines.push(')'); + } + else { + lines.push('parallel('); + for (const p of stagePlans) { + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" }),'); + } + // Replace trailing comma on the last agent line with nothing. + const lastIdx = lines.length - 1; + lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); + lines.push(')'); + } + } + if (wi < waves.length - 1) + lines.push(''); + } + lines.push('// Each agent writes SUMMARY.md on its worktree branch; commits land there'); + lines.push('// and are merged by the orchestrator exactly as in inline wave dispatch.'); + const script = lines.join('\n'); + return { + ok: true, + script, + summary: { + waves: waves.length, + plans: totalPlans, + stagesByWave, + resumeRunId: runId, + budgetTokens, + }, + }; +} +module.exports = { + detectWorkflowBackend, + emitWorkflowScript, + compareSemver, + isValidSemver, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + WORKFLOW_RUNTIME, +}; diff --git a/src/claude-orchestration-command-router.cts b/src/claude-orchestration-command-router.cts new file mode 100644 index 000000000..cabc96435 --- /dev/null +++ b/src/claude-orchestration-command-router.cts @@ -0,0 +1,159 @@ +/** + * Claude orchestration command router — CLI dispatcher for + * `gsd-tools claude-orchestration `. + * + * #1143 — thin CLI adapter over the pure `claude-orchestration.cjs` module. + * Lets execute-phase (or any orchestrator) invoke the Workflow-backend + * detection and the Workflow-script emitter through the standard capability + * command surface (ADR-959) instead of a bare `require()`. + * + * Router signature: { args, cwd, raw, error } — identical to the other host + * routers; discovered by dispatchCapabilityCommand via the registry's + * commandFamilies index. + * + * Subcommands: + * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] + * Resolves whether the Workflow backend should activate. `--runtime` + * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * `claude_orchestration.*` keys from .planning/config.json. Emits + * { available, backend, reason }. + * + * emit-workflow --waves --run-id [--phase-dir ] [--budget ] + * Reads a wave/plan manifest JSON file and emits the generated Workflow + * script + summary. The manifest shape matches emitWorkflowScript's input: + * { waves: [{ id, plans: [{ id, brief, files_modified: string[] }] }] }. + */ + +import fs from 'node:fs'; +import path from 'node:path'; +// eslint-disable-next-line @typescript-eslint/no-require-imports +import io = require('./io.cjs'); +// eslint-disable-next-line @typescript-eslint/no-require-imports +import core = require('./claude-orchestration.cjs'); +// eslint-disable-next-line @typescript-eslint/no-require-imports +import configLoader = require('./config-loader.cjs'); + +const { output } = io; +const { detectWorkflowBackend, emitWorkflowScript } = core; + +const CAPABLE_HOST = { dispatch: { nested: true, background: true } }; + +interface RouterOpts { + args: string[]; + cwd: string; + raw: boolean; + error: (msg: string, reason?: string) => void; +} + +function usage(error: (msg: string, reason?: string) => void): void { + error( + 'Usage: gsd-tools claude-orchestration [...]\n' + + ' detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch]\n' + + ' emit-workflow --waves --run-id [--phase-dir ] [--budget ]', + ); +} + +function argValue(args: string[], flag: string): string | undefined { + const i = args.indexOf(flag); + return i !== -1 && i + 1 < args.length ? args[i + 1] : undefined; +} + +/** + * Detect whether the Workflow backend should activate for the current/given + * runtime. Reads `claude_orchestration.*` from the project config; runtime and + * SDK version come from flags (the orchestrator already knows these) or env. + */ +function cmdDetectBackend(args: string[], cwd: string, raw: boolean): void { + const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; + const agentSdkVersion = argValue(args, '--agent-sdk-version'); + const noNested = args.includes('--no-nested-dispatch'); + const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; + + // Resolve the claude_orchestration.* slice from the project config (federated + // keys are merged by loadConfig as a nested object). A config read failure + // degrades to inline — it must not break the core loop. + let claudeSlice: Record = {}; + try { + const loaded = configLoader.loadConfig(cwd); + const slice = loaded['claude_orchestration']; + if (slice && typeof slice === 'object' && !Array.isArray(slice)) { + claudeSlice = slice as Record; + } + } catch { + claudeSlice = {}; + } + + // Flatten the nested slice into the dotted-key shape detectWorkflowBackend expects. + const flatConfig: Record = {}; + for (const k of Object.keys(claudeSlice)) { + flatConfig['claude_orchestration.' + k] = claudeSlice[k]; + } + + const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); + output(result, raw); +} + +/** + * Emit a Workflow script from a wave/plan manifest file. + */ +function cmdEmitWorkflow(args: string[], _cwd: string, raw: boolean, error: (msg: string, reason?: string) => void): void { + const wavesPath = argValue(args, '--waves'); + const runId = argValue(args, '--run-id'); + const phaseDir = argValue(args, '--phase-dir') || '.planning/phases/current'; + const budgetRaw = argValue(args, '--budget'); + + if (!wavesPath) { + error('emit-workflow requires --waves '); + return; + } + if (!runId) { + error('emit-workflow requires --run-id '); + return; + } + + let waves: unknown; + try { + const content = fs.readFileSync(path.resolve(wavesPath), 'utf8'); + const parsed = JSON.parse(content) as Record; + waves = parsed['waves']; + } catch (e) { + error('emit-workflow: could not read/parse --waves file "' + wavesPath + '": ' + (e instanceof Error ? e.message : String(e))); + return; + } + + const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; + const budget = (typeof budgetTokens === 'number' && !Number.isNaN(budgetTokens)) ? budgetTokens : undefined; + + const result = emitWorkflowScript({ + phaseDir, + runId, + waves: waves as EmitInput['waves'], + budgetTokens: budget, + }); + + if (!result.ok) { + error('emit-workflow: ' + result.reason); + return; + } + output({ script: result.script, summary: result.summary }, raw); +} + +// Re-declared minimal input type for the cast above (avoids importing private types). +interface EmitInput { + waves: Array<{ id: string; plans: Array<{ id: string; brief: string; files_modified: string[] }> }>; +} + +function routeClaudeOrchestrationCommand(opts: RouterOpts): void { + const { args, cwd, raw, error } = opts; + // args[0] is the family ('claude-orchestration'); the subcommand is args[1]. + const subcommand = args[1]; + if (subcommand === 'detect-backend') { + cmdDetectBackend(args, cwd, raw); + } else if (subcommand === 'emit-workflow') { + cmdEmitWorkflow(args, cwd, raw, error); + } else { + usage(error); + } +} + +export = { routeClaudeOrchestrationCommand }; diff --git a/src/claude-orchestration.cts b/src/claude-orchestration.cts new file mode 100644 index 000000000..8e877edbf --- /dev/null +++ b/src/claude-orchestration.cts @@ -0,0 +1,485 @@ +/** + * Claude Orchestration Capability — Workflow-tool backend detection + emitter + * + * #1143 — adopts Claude Code's Workflow tool (the engine behind `/effort ultracode`) + * as an optional, runtime-gated parallel-execution backend for the GSD loop. + * + * This module is the pure, testable core of the capability. It owns two seams: + * + * detectWorkflowBackend({ runtimeId, hostIntegration, config, agentSdkVersion }) + * → { available: boolean, backend: 'workflow'|'inline', reason: string } + * Fail-closed: every miss degrades to `inline` (today's behaviour), so the + * core loop is byte-identical unless every gate opens. This is criteria 3 + 6. + * + * emitWorkflowScript({ phaseDir, waves, runId, budgetTokens? }) + * → { ok:true, script, summary } | { ok:false, reason } + * Maps GSD's wave/plan model 1:1 onto Workflow primitives: + * wave → sequential `parallel()` stage barriers, + * plan → `agent(brief, { agentType:'gsd-executor', isolation:'worktree' })`, + * files_modified overlap → forces plans into separate sequential stages + * (the same overlap rule execute-phase already applies inline), + * resumeFromRunId → wired to the phase run id, + * budgetTokens → a shared token pool. + * The emitted script composes the SAME gsd-executor agent and worktree + * isolation the inline path uses, so it produces the same artifacts/commits + * (criterion 2). It is a generated string consumed by the orchestrator; this + * module never invokes the Workflow tool itself. + * + * Design laws: + * - Gall's Law: ship a small working slice that composes existing primitives + * (gsd-executor + worktree isolation) rather than reinventing them. + * - Greenspun's Tenth Rule (cited in #1143): adopt the Workflow tool's + * barrier/pipeline/budget/resume semantics instead of hand-rolling them. + * - Postel's Law: liberal in input (missing fields → inline), conservative in + * output (workflow only when every gate opens). + * - Fail-closed: an unknown version, a missing descriptor, or a disabled + * toggle all resolve to `inline`, never to `workflow`. + * + * Zero external dependencies. Pure functions. Never throws on bad input. + */ + +// ─── Constants ──────────────────────────────────────────────────────────────── + +/** + * The Agent SDK version that introduced the Workflow tool (#1143 prior art). + * Used as the default floor when config does not override it. A runtime reporting + * an agentSdkVersion below this cannot host the Workflow backend. + */ +const WORKFLOW_TOOL_FLOOR_VERSION = '0.3.149'; + +/** Closed enum for the `claude_orchestration.execution_backend` config key. */ +const BACKEND_VALUES = new Set(['auto', 'workflow', 'inline']); + +/** Only this runtime can host the Workflow tool (Claude Code / Agent SDK). */ +const WORKFLOW_RUNTIME = 'claude'; + +// ─── Semver helpers ─────────────────────────────────────────────────────────── + +/** Official-ish strict SemVer 2.0.0 numeric triple (+ optional pre/build). */ +const SEMVER_RE = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?$/; + +/** True for a syntactically valid semver string. */ +function isValidSemver(s: unknown): s is string { + return typeof s === 'string' && SEMVER_RE.test(s); +} + +/** + * Compare two semver strings. + * Returns -1/0/1 in the usual sense. Garbage in either position → -1 (fail-closed: + * an unparseable version is treated as "less than" any real floor, so detection + * never accidentally enables the preview backend on an unknown SDK). + * + * Pre-release/build metadata are ignored for the comparison — only the numeric + * major.minor.patch triple participates, matching how the Workflow-tool floor is + * specified (a plain "0.3.149"). + */ +function compareSemver(a: string, b: string): number { + if (!isValidSemver(a) || !isValidSemver(b)) return -1; + // Split numeric triple from pre-release/build metadata. + const parseTriple = (s: string): number[] => { + const core = s.split('-')[0].split('+')[0].split('.'); + return [parseInt(core[0], 10), parseInt(core[1], 10), parseInt(core[2], 10)]; + }; + const hasPre = (s: string): boolean => s.indexOf('-') !== -1; + const preIdentifiers = (s: string): string[] => (s.split('-')[1] || '').split('+')[0].split('.').filter((x) => x.length > 0); + const am = parseTriple(a); + const bm = parseTriple(b); + for (let i = 0; i < 3; i++) { + if (am[i] < bm[i]) return -1; + if (am[i] > bm[i]) return 1; + } + // Numeric triple is equal. SemVer 2.0.0 §11 precedence: + // - a version WITH a pre-release tag is LOWER than the same triple WITHOUT one + // (keeps the floor fail-closed for pre-release builds of the GA floor); + // - two pre-releases of the same triple are ordered by their dot-separated + // identifiers (numeric < alphanumeric; numeric compared numerically, + // alphanumeric lexically; fewer identifiers < more). + const aPre = hasPre(a); + const bPre = hasPre(b); + if (aPre && !bPre) return -1; + if (!aPre && bPre) return 1; + if (aPre && bPre) { + const ai = preIdentifiers(a); + const bi = preIdentifiers(b); + const len = Math.min(ai.length, bi.length); + for (let i = 0; i < len; i++) { + const ax = ai[i]; + const bx = bi[i]; + const aNum = /^\d+$/.test(ax); + const bNum = /^\d+$/.test(bx); + if (aNum && bNum) { + const an = parseInt(ax, 10); + const bn = parseInt(bx, 10); + if (an < bn) return -1; + if (an > bn) return 1; + } else if (aNum && !bNum) { + return -1; // numeric identifiers always lower than alphanumeric + } else if (!aNum && bNum) { + return 1; + } else { + if (ax < bx) return -1; + if (ax > bx) return 1; + } + } + if (ai.length < bi.length) return -1; + if (ai.length > bi.length) return 1; + } + return 0; +} + +// ─── detectWorkflowBackend ──────────────────────────────────────────────────── + +interface HostIntegration { + dispatch?: { + nested?: boolean; + background?: boolean; + backgroundDispatch?: boolean; + [k: string]: unknown; + }; + [k: string]: unknown; +} + +interface BackendConfig { + 'claude_orchestration.enabled'?: unknown; + 'claude_orchestration.execution_backend'?: unknown; + 'claude_orchestration.min_agent_sdk_version'?: unknown; + [k: string]: unknown; +} + +interface DetectInput { + runtimeId?: string; + hostIntegration?: HostIntegration | null; + config?: BackendConfig | null; + agentSdkVersion?: string; +} + +interface DetectResult { + available: boolean; + backend: 'workflow' | 'inline'; + reason: string; +} + +/** Inline result shorthand. */ +function inline(reason: string, available = false): DetectResult { + return { available, backend: 'inline', reason }; +} + +/** + * Resolve whether the Workflow-tool backend should activate. + * + * Gate ladder (all must pass for `workflow`; first miss wins, fail-closed): + * 1. capability enabled (claude_orchestration.enabled truthy) + * 2. runtime is Claude (the only runtime that exposes the Workflow tool) + * 3. execution_backend !== 'inline' + * 4. host descriptor signals nested+background dispatch (Workflow-tool capable) + * 5. agentSdkVersion is a known, valid semver + * 6. agentSdkVersion >= the configured floor (default WORKFLOW_TOOL_FLOOR_VERSION) + * 7. execution_backend === 'workflow' OR 'auto' (both reach here; 'inline' exited at 3) + * + * Never throws. Destructures defensively. + */ +function detectWorkflowBackend(input: DetectInput | null | undefined): DetectResult { + if (input === null || input === undefined || typeof input !== 'object') { + return inline('capability_disabled'); + } + + const cfg: BackendConfig = + (input.config !== null && input.config !== undefined && typeof input.config === 'object') + ? input.config + : {}; + + // 1. capability must be opted in (default-off — ships disabled). + if (!cfg['claude_orchestration.enabled']) { + return inline('capability_disabled'); + } + + // 2. only Claude can host the Workflow tool. + if (input.runtimeId !== WORKFLOW_RUNTIME) { + return inline('runtime_not_claude'); + } + + // 3. explicit inline opt-out short-circuits. + let backendRaw = cfg['claude_orchestration.execution_backend']; + if (typeof backendRaw !== 'string' || !BACKEND_VALUES.has(backendRaw)) { + backendRaw = 'auto'; + } + if (backendRaw === 'inline') { + return inline('backend_inline'); + } + + // 4. the host dispatch descriptor must be the nesting-capable Claude-Code shape + // (a proxy for Workflow-tool presence). This is Claude-specific and already + // gated at step 2; `background:true` alone is true on several non-Claude hosts, + // so the proxy is only meaningful after the runtime check above. Note: this is + // NOT the canonical `shouldFlattenDispatch` rule (which keys on + // `backgroundDispatch`); the Workflow backend works precisely because a single + // tool-call orchestrates internally, sidestepping the backgroundDispatch:false + // limitation. Missing/false/foreign descriptor → fail-closed. + const hi = input.hostIntegration; + if (hi === null || hi === undefined || typeof hi !== 'object' || Array.isArray(hi)) { + return inline('workflow_tool_unavailable'); + } + const dispatch = (hi as { dispatch?: Record }).dispatch; + if (typeof dispatch !== 'object' || dispatch === null || Array.isArray(dispatch)) { + return inline('workflow_tool_unavailable'); + } + const nested = dispatch['nested']; + const background = dispatch['background']; + if (nested !== true || background !== true) { + return inline('workflow_tool_unavailable'); + } + + // 5. an unknown agentSdkVersion cannot be trusted to meet the floor. + if (!isValidSemver(input.agentSdkVersion)) { + return inline('agent_sdk_version_unknown'); + } + + // 6. version floor (config override > default constant). + const floorRaw = cfg['claude_orchestration.min_agent_sdk_version']; + const floor = typeof floorRaw === 'string' && isValidSemver(floorRaw) ? floorRaw : WORKFLOW_TOOL_FLOOR_VERSION; + if (compareSemver(input.agentSdkVersion, floor) < 0) { + return inline('agent_sdk_version_below_floor'); + } + + // 7. auto/workflow both reach the workflow backend once every gate passes. + return { available: true, backend: 'workflow', reason: 'workflow_backend_active' }; +} + +// ─── emitWorkflowScript ─────────────────────────────────────────────────────── + +interface Plan { + id: string; + brief: string; + files_modified: string[]; +} + +interface Wave { + id: string; + plans: Plan[]; +} + +interface EmitInput { + phaseDir: string; + waves: Wave[]; + runId: string; + budgetTokens?: number; +} + +interface EmitOk { + ok: true; + script: string; + summary: { + waves: number; + plans: number; + stagesByWave: string[][][]; // wave → stage → planId[] + resumeRunId: string; + budgetTokens: number | null; + }; +} + +interface EmitErr { + ok: false; + reason: string; +} + +/** + * Partition a wave's plans into a near-minimal number of sequential stages (via + * greedy first-fit — not guaranteed optimal for arbitrary overlap graphs, but + * correct: no two plans sharing a file ever cohabit a stage) such that no two + * plans in the same stage share a modified file. Each plan goes into the earliest + * stage where it does not overlap any plan already there. + * + * A plan with an EMPTY files_modified set declares no files; it overlaps nothing + * and coalesces into stage 0 (same behavior as the inline path, which also cannot + * guard against undeclared concurrent writes — declare filesModified accurately). + * + * This is the same overlap rule execute-phase applies inline — the only difference + * is the execution vehicle (Workflow `parallel()` vs one-agent-per-message). + */ +function partitionStages(plans: Plan[]): string[][] { + const stages: { plans: Plan[]; files: Set }[] = []; + for (const plan of plans) { + const fileSet = new Set(plan.files_modified); + let placed = false; + for (const stage of stages) { + let overlap = false; + for (const f of fileSet) { + if (stage.files.has(f)) { overlap = true; break; } + } + if (!overlap) { + stage.plans.push(plan); + for (const f of fileSet) stage.files.add(f); + placed = true; + break; + } + } + if (!placed) { + stages.push({ plans: [plan], files: new Set(fileSet) }); + } + } + return stages.map((s) => s.plans.map((p) => p.id)); +} + +/** + * Quote a free-text value for safe embedding as a JavaScript/Workflow double-quoted + * string literal. Uses JSON.stringify so every JS-relevant escape (backslash, quote, + * newline, tab, NUL, U+2028/U+2029, all control chars) is handled by the language + * itself — there is no hand-rolled escape table to drift. Returns the value already + * wrapped in its surrounding quotes. + */ +function quoteString(s: string): string { + return JSON.stringify(s); +} + +/** + * True if `s` is a safe identifier/path token to interpolate into the generated + * script WITHOUT requiring a string-literal context — i.e. it contains no + * character that could terminate a comment line (`\n`/`\r`), break out of a + * string literal (`"` / `\`), or smuggle a NUL/control sequence. Used for + * `phaseDir`, `runId`, `wave.id`, and `plan.id`, which are identifiers/paths and + * must never legitimately contain such characters. Rejecting them at validation + * (rather than silently flattening) keeps the emitted script faithful to input. + */ +const UNSCRIPTABLE_CHAR_RE = /[\r\n"\\\x00-\x1f\x7f\u2028\u2029]/; +function isScriptableIdentifier(s: unknown): boolean { + if (typeof s !== 'string' || s.length === 0) return false; + return !UNSCRIPTABLE_CHAR_RE.test(s); +} + +/** + * Emit a Workflow script mapping the phase's wave/plan model onto Workflow + * primitives. Pure and deterministic: identical input yields an identical string. + * + * Returns ok:false (never throws) on invalid input — empty waves, missing runId, + * a wave with no plans, etc. + */ +function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitErr { + if (input === null || input === undefined || typeof input !== 'object') { + return { ok: false, reason: 'invalid_input' }; + } + const { phaseDir, waves, runId } = input; + // Identifiers/paths interpolated into the generated script must be free of any + // character that could terminate a comment, break out of a string literal, or + // smuggle control bytes — reject up front (security: #1143 review Finding 1). + if (!isScriptableIdentifier(phaseDir)) { + return { ok: false, reason: 'phaseDir must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!isScriptableIdentifier(runId)) { + return { ok: false, reason: 'runId must be a non-empty string without newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(waves) || waves.length === 0) { + return { ok: false, reason: 'waves must be a non-empty array' }; + } + for (let i = 0; i < waves.length; i++) { + const w = waves[i]; + if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { + return { ok: false, reason: 'waves[' + i + '] must be { id, plans: non-empty[] }' }; + } + if (!isScriptableIdentifier(w.id)) { + return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (!Array.isArray(w.plans) || w.plans.length === 0) { + return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; + } + const seenIds = new Set(); + for (let j = 0; j < w.plans.length; j++) { + const p = w.plans[j]; + if (p === null || typeof p !== 'object' || typeof p.id !== 'string' || typeof p.brief !== 'string' || !Array.isArray(p.files_modified)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '] must be { id, brief, files_modified[] }' }; + } + if (!isScriptableIdentifier(p.id)) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].id must not contain newlines/quotes/backslash/control chars' }; + } + if (seenIds.has(p.id)) { + return { ok: false, reason: 'waves[' + i + '] has duplicate plan id "' + p.id + '"' }; + } + seenIds.add(p.id); + for (const f of p.files_modified) { + if (typeof f !== 'string' || f.length === 0) { + return { ok: false, reason: 'waves[' + i + '].plans[' + j + '].files_modified entries must be non-empty strings' }; + } + } + } + } + + const budgetTokens = (typeof input.budgetTokens === 'number' && Number.isFinite(input.budgetTokens) && input.budgetTokens > 0) + ? Math.floor(input.budgetTokens) + : null; + + const lines: string[] = []; + lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); + lines.push('// phase: ' + phaseDir); + lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); + lines.push('// Composes the SAME gsd-executor agent + worktree isolation as the inline path,'); + lines.push('// so artifacts (SUMMARY.md) and commits are produced identically.'); + lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + if (budgetTokens !== null) { + lines.push('budget(' + budgetTokens + ')'); + } + lines.push(''); + + const stagesByWave: string[][][] = []; + let totalPlans = 0; + + for (let wi = 0; wi < waves.length; wi++) { + const wave = waves[wi]; + const stages = partitionStages(wave.plans); + stagesByWave.push(stages); + totalPlans += wave.plans.length; + + lines.push('// Wave ' + wave.id); + for (let si = 0; si < stages.length; si++) { + const stagePlanIds = stages[si]; + // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). + const stagePlans = stagePlanIds.map((id) => wave.plans.find((p) => p.id === id) as Plan); + if (stages.length > 1) { + lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); + } + if (stagePlans.length === 1) { + const p = stagePlans[0]; + lines.push('parallel('); + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" })'); + lines.push(')'); + } else { + lines.push('parallel('); + for (const p of stagePlans) { + lines.push(' agent(' + quoteString(p.brief) + ', { agentType: "gsd-executor", isolation: "worktree" }),'); + } + // Replace trailing comma on the last agent line with nothing. + const lastIdx = lines.length - 1; + lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); + lines.push(')'); + } + } + if (wi < waves.length - 1) lines.push(''); + } + + lines.push('// Each agent writes SUMMARY.md on its worktree branch; commits land there'); + lines.push('// and are merged by the orchestrator exactly as in inline wave dispatch.'); + + const script = lines.join('\n'); + + return { + ok: true, + script, + summary: { + waves: waves.length, + plans: totalPlans, + stagesByWave, + resumeRunId: runId, + budgetTokens, + }, + }; +} + +// ─── Exports ────────────────────────────────────────────────────────────────── + +export = { + detectWorkflowBackend, + emitWorkflowScript, + compareSemver, + isValidSemver, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + WORKFLOW_RUNTIME, +}; diff --git a/tests/check-gap-analysis-plan-post-e2e.test.cjs b/tests/check-gap-analysis-plan-post-e2e.test.cjs index 4a4a7930a..62b4bf87a 100644 --- a/tests/check-gap-analysis-plan-post-e2e.test.cjs +++ b/tests/check-gap-analysis-plan-post-e2e.test.cjs @@ -493,7 +493,7 @@ describe('resolveLoopHooks plan:post — pure function against real registry', ( assert.strictEqual(result.activeHooks[0].capId, 'gap-analysis'); }); - test('[happy] real registry byLoopPoint plan:post has 1 step (mempalace), 1 contribution (external-job planner fragment), and 1 gate (gap-analysis)', () => { + test('[happy] real registry byLoopPoint plan:post has 1 step (mempalace), 2 contributions (external-job planner + claude-orchestration ultraplan ownership), and 1 gate (gap-analysis)', () => { const entry = realRegistry.byLoopPoint['plan:post']; assert.ok(entry, 'plan:post must exist in byLoopPoint'); assert.ok(Array.isArray(entry.steps), 'steps must be an array'); @@ -501,8 +501,12 @@ describe('resolveLoopHooks plan:post — pure function against real registry', ( assert.ok(Array.isArray(entry.gates), 'gates must be an array'); assert.strictEqual(entry.steps.length, 1, 'plan:post must have 1 step (mempalace capture)'); assert.strictEqual(entry.steps[0].capId, 'mempalace', 'plan:post step must be from mempalace'); - assert.strictEqual(entry.contributions.length, 1, 'plan:post must have 1 contribution (external-job planner runtime-budget fragment)'); - assert.strictEqual(entry.contributions[0].capId, 'external-job', 'plan:post contribution must be from external-job'); + // #1143: claude-orchestration registers a plan:post contribution declaring + // ultraplan plan-offload ownership under its runtime gate (default-off). + assert.strictEqual(entry.contributions.length, 2, 'plan:post must have 2 contributions (external-job planner + claude-orchestration ultraplan ownership)'); + const capIds = entry.contributions.map(c => c.capId).sort(); + assert.deepStrictEqual(capIds, ['claude-orchestration', 'external-job'], + `plan:post contributions must be external-job + claude-orchestration; got ${capIds.join(',')}`); assert.strictEqual(entry.gates.length, 1, 'plan:post must have exactly one gate'); assert.strictEqual(entry.gates[0].capId, 'gap-analysis'); }); diff --git a/tests/claude-orchestration-command-router.test.cjs b/tests/claude-orchestration-command-router.test.cjs new file mode 100644 index 000000000..e0a408f1b --- /dev/null +++ b/tests/claude-orchestration-command-router.test.cjs @@ -0,0 +1,169 @@ +'use strict'; + +/** + * claude-orchestration-command-router.test.cjs — end-to-end tests for the + * `gsd-tools claude-orchestration` command surface (#1143). + * + * Exercises the full dispatch path: gsd-tools → dispatchCapabilityCommand → + * routeClaudeOrchestrationCommand → the pure detect/emitter functions. + * + * NOTE: runGsdTools() returns { success, output, exitCode, error } — it does + * NOT throw on a non-zero exit (helpers.cjs). Tests assert `.success` and parse + * `.output`, and check `.error` on the failure path. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs'); + +// ─── Fixtures ───────────────────────────────────────────────────────────────── + +const WAVES_MANIFEST = { + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] }, + { id: 'p2', brief: 'Wire the bar seam', files_modified: ['src/bar.cts'] }, + ], + }, + ], +}; + +function writeManifest(tmpDir) { + const manifestPath = path.join(tmpDir, 'waves.json'); + fs.writeFileSync(manifestPath, JSON.stringify(WAVES_MANIFEST), 'utf8'); + return manifestPath; +} + +/** Run the command and assert it succeeded, returning the parsed JSON output. */ +function runAndParse(args, cwd) { + const res = runGsdTools(args, cwd); + assert.strictEqual(res.success, true, 'command should succeed; stderr: ' + (res.error || '')); + assert.ok(res.output.length > 0, 'command should emit output'); + return JSON.parse(res.output); +} + +// ─── emit-workflow ──────────────────────────────────────────────────────────── + +describe('claude-orchestration emit-workflow (CLI)', () => { + test('emits a Workflow script with parallel barriers, gsd-executor + worktree, and resumeFromRunId', () => { + const tmp = createTempProject('claw-emit-'); + try { + const manifestPath = writeManifest(tmp); + const parsed = runAndParse([ + 'claude-orchestration', 'emit-workflow', + '--waves', manifestPath, + '--run-id', 'run-cli-1143', + '--phase-dir', '.planning/phases/01-foo', + ], tmp); + assert.ok(typeof parsed.script === 'string' && parsed.script.length > 0); + assert.ok(parsed.script.includes('parallel('), 'parallel() barrier emitted'); + assert.ok(parsed.script.includes('gsd-executor'), 'gsd-executor agentType'); + assert.ok(parsed.script.includes('worktree'), 'worktree isolation'); + assert.ok(parsed.script.includes('resumeFromRunId'), 'resumeFromRunId wired'); + assert.ok(parsed.script.includes('run-cli-1143'), 'carries the run id'); + assert.strictEqual(parsed.summary.resumeRunId, 'run-cli-1143'); + assert.strictEqual(parsed.summary.waves, 1); + assert.strictEqual(parsed.summary.plans, 2); + } finally { + cleanup(tmp); + } + }); + + test('budget flag threads a shared token pool into the script', () => { + const tmp = createTempProject('claw-budget-'); + try { + const manifestPath = writeManifest(tmp); + const parsed = runAndParse([ + 'claude-orchestration', 'emit-workflow', + '--waves', manifestPath, + '--run-id', 'r', + '--budget', '750000', + ], tmp); + assert.ok(parsed.script.includes('budget('), 'budget() pool emitted'); + assert.ok(parsed.script.includes('750000')); + } finally { + cleanup(tmp); + } + }); + + test('missing --waves -> non-zero exit with a diagnostic', () => { + const tmp = createTempProject('claw-noargs-'); + try { + const res = runGsdTools([ + 'claude-orchestration', 'emit-workflow', '--run-id', 'r', + ], tmp); + assert.strictEqual(res.success, false, 'missing --waves must fail'); + assert.ok(res.exitCode !== 0, 'non-zero exit'); + assert.match(res.error || '', /--waves/); + } finally { + cleanup(tmp); + } + }); +}); + +// ─── detect-backend ─────────────────────────────────────────────────────────── + +describe('claude-orchestration detect-backend (CLI)', () => { + test('default config (capability off) -> inline, even on Claude', () => { + const tmp = createTempProject('claw-detect-'); + try { + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'claude', + '--agent-sdk-version', '1.0.0', + ], tmp); + assert.strictEqual(parsed.backend, 'inline'); + assert.strictEqual(parsed.available, false); + assert.match(parsed.reason, /disabled/); + } finally { + cleanup(tmp); + } + }); + + test('enabled + claude + capable + new-enough SDK -> workflow', () => { + const tmp = createTempProject('claw-detect-on-'); + try { + fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(tmp, '.planning', 'config.json'), + JSON.stringify({ claude_orchestration: { enabled: true, execution_backend: 'auto' } }), + 'utf8', + ); + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'claude', + '--agent-sdk-version', '1.2.0', + ], tmp); + assert.strictEqual(parsed.backend, 'workflow'); + assert.strictEqual(parsed.available, true); + } finally { + cleanup(tmp); + } + }); + + test('non-Claude runtime -> inline (criterion 6)', () => { + const tmp = createTempProject('claw-detect-nonclaude-'); + try { + fs.mkdirSync(path.join(tmp, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(tmp, '.planning', 'config.json'), + JSON.stringify({ claude_orchestration: { enabled: true } }), + 'utf8', + ); + const parsed = runAndParse([ + 'claude-orchestration', 'detect-backend', + '--runtime', 'codex', + '--agent-sdk-version', '1.0.0', + ], tmp); + assert.strictEqual(parsed.backend, 'inline'); + assert.match(parsed.reason, /claude/i); + } finally { + cleanup(tmp); + } + }); +}); diff --git a/tests/claude-orchestration.test.cjs b/tests/claude-orchestration.test.cjs new file mode 100644 index 000000000..808501e1e --- /dev/null +++ b/tests/claude-orchestration.test.cjs @@ -0,0 +1,615 @@ +'use strict'; + +/** + * claude-orchestration.test.cjs — Behavioral tests for the Claude orchestration + * capability (#1143): Workflow-tool backend detection, Workflow-script emission, + * capability-declaration validation, registry integration, and inline-fallback parity. + * + * The capability is default-off + BETA + claude-only. On any runtime lacking the + * Workflow tool it must be a byte-identical no-op. These tests encode that contract. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); + +const fc = require('fast-check'); + +const { + detectWorkflowBackend, + emitWorkflowScript, + WORKFLOW_TOOL_FLOOR_VERSION, + BACKEND_VALUES, + compareSemver, +} = require('../gsd-core/bin/lib/claude-orchestration.cjs'); + +const { + validateCapability, + validateAgainstContract, + loadAndValidate, + buildRegistry, + serializeRegistry, + normalizeLineEndings, + stripGeneratedComment, +} = require('../scripts/gen-capability-registry.cjs'); + +const ROOT = path.resolve(__dirname, '..'); +const CAP_PATH = path.join(ROOT, 'capabilities', 'claude-orchestration', 'capability.json'); +const REGISTRY_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs'); + +// ─── Fixtures ───────────────────────────────────────────────────────────────── + +/** A host-integration descriptor whose dispatch axis signals Workflow-tool capability. */ +const CAPABLE_HOST = { + dispatch: { namedDispatch: true, nested: true, background: true, backgroundDispatch: false }, +}; + +/** Read the real capability declaration (data file — not a source grep). */ +function loadCap() { + return JSON.parse(fs.readFileSync(CAP_PATH, 'utf8')); +} + +/** A minimal single-plan wave manifest. */ +function singleWaveManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Implement the foo module', files_modified: ['src/foo.cts'] }, + ], + }, + ], + }; +} + +/** Two plans in one wave that DO NOT overlap (parallel-safe in a single stage). */ +function nonOverlappingManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Plan A', files_modified: ['src/a.cts'] }, + { id: 'p2', brief: 'Plan B', files_modified: ['src/b.cts'] }, + ], + }, + ], + }; +} + +/** Two plans in one wave that DO overlap on files_modified (must split into stages). */ +function overlappingManifest() { + return { + phaseDir: '.planning/phases/01-foo', + runId: 'run-abc-1143', + waves: [ + { + id: 'w1', + plans: [ + { id: 'p1', brief: 'Plan A', files_modified: ['src/shared.cts', 'src/a.cts'] }, + { id: 'p2', brief: 'Plan B', files_modified: ['src/shared.cts', 'src/b.cts'] }, + ], + }, + ], + }; +} + +// ─── 1. detectWorkflowBackend ───────────────────────────────────────────────── + +describe('detectWorkflowBackend', () => { + + test('capability disabled (default-off) -> inline, even on Claude with the tool', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': false }, + }); + assert.strictEqual(r.available, false); + assert.strictEqual(r.backend, 'inline'); + assert.match(r.reason, /disabled/); + }); + + test('non-Claude runtime -> inline (criterion 6: no change to non-Claude loop)', () => { + for (const runtimeId of ['codex', 'cursor', 'opencode', 'copilot', ' Windsurf'.trim()]) { + const r = detectWorkflowBackend({ + runtimeId, + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'workflow' }, + }); + assert.strictEqual(r.backend, 'inline', runtimeId + ' should be inline'); + assert.strictEqual(r.available, false, runtimeId + ' should be unavailable'); + assert.match(r.reason, /claude/i, runtimeId + ' reason should mention claude'); + } + }); + + test('Claude + auto + capable host + new-enough SDK -> workflow', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.2.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'workflow'); + assert.strictEqual(r.available, true); + }); + + test('Claude + execution_backend:"workflow" forces workflow when tool is capable', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'workflow' }, + }); + assert.strictEqual(r.backend, 'workflow'); + assert.strictEqual(r.available, true); + }); + + test('Claude + execution_backend:"inline" -> inline even when tool is capable', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'inline' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.match(r.reason, /inline/); + }); + + test('Claude + auto + host lacking nested dispatch -> inline (fail-closed)', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: { dispatch: { nested: false, background: true } }, + agentSdkVersion: '1.0.0', + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.strictEqual(r.available, false); + }); + + test('Claude + unknown agentSdkVersion -> inline fail-closed (criterion 3 fallback)', () => { + const r = detectWorkflowBackend({ + runtimeId: 'claude', + hostIntegration: CAPABLE_HOST, + agentSdkVersion: undefined, + config: { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }, + }); + assert.strictEqual(r.backend, 'inline'); + assert.strictEqual(r.available, false); + assert.match(r.reason, /version|sdk|unknown/i); + }); + + test('agent SDK version boundary: floor-1 -> inline, floor -> workflow, floor+patch -> workflow', () => { + const floor = WORKFLOW_TOOL_FLOOR_VERSION; + const [maj, min, pat] = floor.split('.').map((n) => parseInt(n, 10)); + // Robust "below" derivation with full borrow chain (works for .0.0 floors too). + let below; + if (pat > 0) below = `${maj}.${min}.${pat - 1}`; + else if (min > 0) below = `${maj}.${min - 1}.999`; + else if (maj > 0) below = `${maj - 1}.999.999`; + else { assert.ok(false, 'cannot derive below for 0.0.0 floor'); return; } + const above = `${maj}.${min}.${pat + 1}`; + // Sanity: confirm below really is below per the comparator under test. + assert.ok(compareSemver(below, floor) < 0, below + ' must compare below ' + floor); + + const cfg = { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }; + + const rBelow = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: below, config: cfg }); + assert.strictEqual(rBelow.backend, 'inline', below + ' (floor-1) must be inline'); + + const rAt = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: floor, config: cfg }); + assert.strictEqual(rAt.backend, 'workflow', floor + ' (exact floor) must be workflow'); + + const rAbove = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: above, config: cfg }); + assert.strictEqual(rAbove.backend, 'workflow', above + ' (floor+patch) must be workflow'); + }); + + test('config-level min_agent_sdk_version override raises/lowers the floor', () => { + const cfg = { + 'claude_orchestration.enabled': true, + 'claude_orchestration.execution_backend': 'auto', + 'claude_orchestration.min_agent_sdk_version': '2.0.0', + }; + const r1 = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '1.9.9', config: cfg }); + assert.strictEqual(r1.backend, 'inline', 'below raised floor -> inline'); + const r2 = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '2.0.0', config: cfg }); + assert.strictEqual(r2.backend, 'workflow', 'at raised floor -> workflow'); + }); + + test('execution_backend:"workflow" + SDK below floor -> inline (M-1: floor applies in both modes)', () => { + const cfg = { + 'claude_orchestration.enabled': true, + 'claude_orchestration.execution_backend': 'workflow', + }; + const r = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '0.3.0', config: cfg }); + assert.strictEqual(r.backend, 'inline', 'workflow mode must still honor the SDK floor (fail-closed)'); + assert.strictEqual(r.available, false); + assert.match(r.reason, /floor|version/); + }); + + test('pre-release of the floor (0.3.149-rc.1) -> inline (pre-release < GA per SemVer)', () => { + const cfg = { 'claude_orchestration.enabled': true, 'claude_orchestration.execution_backend': 'auto' }; + // Explicitly assert the precedence rule: a pre-release tag is below the GA release. + assert.ok(compareSemver('0.3.149-rc.1', '0.3.149') < 0, 'pre-release must compare below GA'); + const r = detectWorkflowBackend({ runtimeId: 'claude', hostIntegration: CAPABLE_HOST, agentSdkVersion: '0.3.149-rc.1', config: cfg }); + assert.strictEqual(r.backend, 'inline', 'pre-release of the floor must not activate the BETA backend'); + assert.strictEqual(r.available, false); + }); + + test('two pre-releases of the same triple order by their identifiers (SemVer §11)', () => { + assert.ok(compareSemver('0.3.149-rc.0', '0.3.149-rc.1') < 0, 'rc.0 < rc.1'); + assert.ok(compareSemver('1.0.0-alpha.1', '1.0.0-alpha.2') < 0, 'alpha.1 < alpha.2'); + assert.ok(compareSemver('1.0.0-rc.1', '1.0.0-rc.2') < 0, 'rc.1 < rc.2'); + // numeric < alphanumeric at the same position + assert.ok(compareSemver('1.0.0-1', '1.0.0-alpha') < 0, 'numeric identifier < alphanumeric'); + }); + + test('missing/empty input -> inline, never throws (Postel: liberal-in-input)', () => { + assert.strictEqual(detectWorkflowBackend({}).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend(null).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend(undefined).backend, 'inline'); + assert.strictEqual(detectWorkflowBackend({ runtimeId: 'claude' }).backend, 'inline'); + }); + + test('BACKEND_VALUES exposes the closed enum', () => { + assert.deepStrictEqual([...BACKEND_VALUES].sort(), ['auto', 'inline', 'workflow']); + }); + + test('property: pure & deterministic (same input -> same output)', () => { + fc.assert(fc.property( + fc.record({ + runtimeId: fc.constantFrom('claude', 'codex', 'cursor', 'opencode'), + sdk: fc.option(fc.string({ minLength: 1, maxLength: 8 }).filter((s) => /^\d/.test(s)), { nil: undefined }), + backend: fc.constantFrom('auto', 'workflow', 'inline'), + enabled: fc.boolean(), + }), + (input) => { + const cfg = { + 'claude_orchestration.enabled': input.enabled, + 'claude_orchestration.execution_backend': input.backend, + }; + const a = detectWorkflowBackend({ runtimeId: input.runtimeId, hostIntegration: CAPABLE_HOST, agentSdkVersion: input.sdk, config: cfg }); + const b = detectWorkflowBackend({ runtimeId: input.runtimeId, hostIntegration: CAPABLE_HOST, agentSdkVersion: input.sdk, config: cfg }); + assert.deepStrictEqual(a, b); + assert.ok(['workflow', 'inline'].includes(a.backend)); + }, + )); + }); +}); + +// ─── 2. compareSemver helper ────────────────────────────────────────────────── + +describe('compareSemver', () => { + test('ordering', () => { + assert.ok(compareSemver('1.0.0', '0.9.9') > 0); + assert.ok(compareSemver('1.0.0', '1.0.0') === 0); + assert.ok(compareSemver('1.0.0', '1.0.1') < 0); + assert.ok(compareSemver('2.0.0', '1.9.9') > 0); + }); + test('garbage versions compare as -1 (fail-closed)', () => { + assert.strictEqual(compareSemver('garbage', '1.0.0'), -1); + assert.strictEqual(compareSemver('1.0.0', ''), -1); + }); +}); + +// ─── 3. emitWorkflowScript ──────────────────────────────────────────────────── + +describe('emitWorkflowScript', () => { + + test('single-wave single-plan -> one parallel barrier, one agent, executor+worktree', () => { + const { ok, script, summary } = emitWorkflowScript(singleWaveManifest()); + assert.strictEqual(ok, true); + assert.ok(typeof script === 'string' && script.length > 0); + + const parallelCount = (script.match(/parallel\s*\(/g) || []).length; + assert.ok(parallelCount >= 1, 'at least one parallel() barrier'); + assert.ok(script.includes('agent('), 'agent() call per plan'); + assert.ok(script.includes('gsd-executor'), 'uses gsd-executor agentType'); + assert.ok(script.includes('worktree'), 'uses worktree isolation'); + assert.ok(script.includes('SUMMARY.md'), 'produces SUMMARY.md (same artifact as inline path)'); + + assert.deepStrictEqual(summary.waves, 1); + assert.deepStrictEqual(summary.plans, 1); + }); + + test('multi-wave -> one parallel() barrier per wave (sequential barriers)', () => { + const r = emitWorkflowScript({ + phaseDir: '.planning/phases/01-foo', + runId: 'run-multi', + waves: [ + { id: 'w1', plans: [{ id: 'p1', brief: 'A', files_modified: ['src/a.cts'] }] }, + { id: 'w2', plans: [{ id: 'p2', brief: 'B', files_modified: ['src/b.cts'] }] }, + { id: 'w3', plans: [{ id: 'p3', brief: 'C', files_modified: ['src/c.cts'] }] }, + ], + }); + assert.strictEqual(r.ok, true); + const parallelCount = (r.script.match(/parallel\s*\(/g) || []).length; + assert.strictEqual(parallelCount, 3, 'one parallel() per wave'); + assert.strictEqual(r.summary.waves, 3); + assert.strictEqual(r.summary.plans, 3); + }); + + test('overlapping files_modified -> plans split into separate sequential stages (criterion 2)', () => { + const r = emitWorkflowScript(overlappingManifest()); + assert.strictEqual(r.ok, true); + // Two plans sharing src/shared.cts must NOT be in the same stage. + const stages = r.summary.stagesByWave[0]; // wave w1 + assert.ok(Array.isArray(stages), 'stagesByWave present'); + assert.strictEqual(stages.length, 2, 'overlapping plans split into 2 stages'); + const stagePlanSets = stages.map((s) => s.slice().sort()); + const allPlans = stagePlanSets.flat().sort(); + assert.deepStrictEqual(allPlans, ['p1', 'p2']); + // p1 and p2 must be in different stages + assert.ok(stages[0].length === 1 && stages[1].length === 1, 'one plan per stage when they overlap'); + }); + + test('non-overlapping plans -> coalesced into a single parallel stage', () => { + const r = emitWorkflowScript(nonOverlappingManifest()); + assert.strictEqual(r.ok, true); + const stages = r.summary.stagesByWave[0]; + assert.strictEqual(stages.length, 1, 'non-overlapping plans share one stage'); + assert.deepStrictEqual(stages[0].slice().sort(), ['p1', 'p2']); + }); + + test('resumeFromRunId wired to the provided runId (criterion 4)', () => { + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(r.script.includes('resumeFromRunId'), 'references resumeFromRunId'); + assert.ok(r.script.includes('run-abc-1143'), 'carries the run id'); + assert.strictEqual(r.summary.resumeRunId, 'run-abc-1143'); + }); + + test('shared budget pool emitted when budgetTokens provided', () => { + const r = emitWorkflowScript({ ...singleWaveManifest(), budgetTokens: 500000 }); + assert.ok(r.script.includes('budget('), 'emits budget() pool'); + assert.ok(r.script.includes('500000')); + }); + + test('no budget() emitted when budgetTokens omitted', () => { + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(!r.script.includes('budget('), 'no budget() when unset'); + }); + + test('invalid input -> ok:false with a reason, never throws', () => { + const empty = emitWorkflowScript({ phaseDir: '.p', runId: 'r', waves: [] }); + assert.strictEqual(empty.ok, false); + assert.ok(typeof empty.reason === 'string' && empty.reason.length > 0); + + const noRun = emitWorkflowScript({ phaseDir: '.p', runId: '', waves: singleWaveManifest().waves }); + assert.strictEqual(noRun.ok, false); + + const noPhase = emitWorkflowScript({ phaseDir: '', runId: 'r', waves: singleWaveManifest().waves }); + assert.strictEqual(noPhase.ok, false); + + const badWave = emitWorkflowScript({ phaseDir: '.p', runId: 'r', waves: [{ id: 'w1', plans: [] }] }); + assert.strictEqual(badWave.ok, false); + }); + + test('SECURITY: runId/phaseDir/wave.id/plan.id with injection chars -> ok:false (never reach the script)', () => { + // runId is interpolated inside resumeFromRunId("...") — a quote/backslash/newline + // could break out of the call. Identifier validation must reject it. + const injectRun = emitWorkflowScript({ phaseDir: '.p', runId: 'x");evil("y', waves: singleWaveManifest().waves }); + assert.strictEqual(injectRun.ok, false); + assert.match(injectRun.reason, /runId/i); + + const newlineRun = emitWorkflowScript({ phaseDir: '.p', runId: 'r\nbreakout', waves: singleWaveManifest().waves }); + assert.strictEqual(newlineRun.ok, false); + + const injectPhase = emitWorkflowScript({ phaseDir: '.p"; drop table', runId: 'r', waves: singleWaveManifest().waves }); + assert.strictEqual(injectPhase.ok, false); + + const injectWave = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1\nagent("evil")', plans: [{ id: 'p1', brief: 'b', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(injectWave.ok, false); + + const injectPlan = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1";x("y', brief: 'b', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(injectPlan.ok, false); + }); + + test('SECURITY: a brief containing quotes/backslash/newlines is neutralised (never breaks the string literal)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1', brief: 'he said "hi" \\ then \n newline', files_modified: ['a.cts'] }] }], + }); + assert.strictEqual(r.ok, true); + // The emitted script must not contain a raw unescaped quote that closes the + // agent() string literal, nor a raw newline inside the brief. + assert.ok(!r.script.includes('he said "hi" \\\\'), 'no unescaped breakout'); + // The full brief text never appears verbatim with its dangerous chars intact. + assert.ok(!r.script.includes('"hi"'), 'the inner quote must be JSON-escaped, not raw'); + }); + + test('duplicate plan id within a wave -> ok:false (L-5: no silent brief loss)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [ + { id: 'p1', brief: 'first', files_modified: ['a.cts'] }, + { id: 'p1', brief: 'second', files_modified: ['b.cts'] }, + ] }], + }); + assert.strictEqual(r.ok, false); + assert.match(r.reason, /duplicate/i); + }); + + test('non-string files_modified entries -> ok:false (L-7: strict element typing)', () => { + const r = emitWorkflowScript({ + phaseDir: '.p', runId: 'r', + waves: [{ id: 'w1', plans: [{ id: 'p1', brief: 'b', files_modified: ['ok.cts', 42, { path: 'x' }] }] }], + }); + assert.strictEqual(r.ok, false); + assert.match(r.reason, /files_modified/); + }); + + test('property: deterministic (same input -> identical script)', () => { + fc.assert(fc.property( + fc.record({ + runId: fc.string({ minLength: 1, maxLength: 12 }).filter((s) => /^[a-zA-Z0-9-]+$/.test(s)), + nPlans: fc.integer({ min: 1, max: 5 }), + }), + ({ runId, nPlans }) => { + const waves = [{ + id: 'w1', + plans: Array.from({ length: nPlans }, (_, i) => ({ + id: 'p' + i, + brief: 'brief ' + i, + files_modified: ['src/file' + i + '.cts'], + })), + }]; + const a = emitWorkflowScript({ phaseDir: '.planning/phases/01-x', runId, waves }); + const b = emitWorkflowScript({ phaseDir: '.planning/phases/01-x', runId, waves }); + assert.strictEqual(a.script, b.script); + assert.deepStrictEqual(a.summary, b.summary); + }, + )); + }); +}); + +// ─── 4. Capability declaration validation ───────────────────────────────────── + +describe('capability declaration (capabilities/claude-orchestration/capability.json)', () => { + + test('file exists and parses', () => { + const cap = loadCap(); + assert.strictEqual(cap.id, 'claude-orchestration'); + }); + + test('passes per-file validateCapability', () => { + const errors = validateCapability(loadCap(), 'claude-orchestration'); + assert.deepEqual(errors, [], 'Expected no validation errors: ' + JSON.stringify(errors)); + }); + + test('passes contract validation (contribution.into roles, when references)', () => { + const errors = validateAgainstContract(loadCap(), 'claude-orchestration'); + assert.deepEqual(errors, [], 'Expected no contract errors: ' + JSON.stringify(errors)); + }); + + test('default-off: activationKey default is false and points at the enabled key', () => { + const cap = loadCap(); + assert.strictEqual(cap.activationKey, 'claude_orchestration.enabled'); + assert.strictEqual(cap.config['claude_orchestration.enabled'].default, false); + assert.strictEqual(cap.config['claude_orchestration.enabled'].type, 'boolean'); + }); + + test('runtimeCompat is claude-only (criterion 6)', () => { + const cap = loadCap(); + assert.deepStrictEqual(cap.runtimeCompat.supported, ['claude']); + assert.deepStrictEqual(cap.runtimeCompat.unsupported, []); + }); + + test('BETA posture: tier full, role feature', () => { + const cap = loadCap(); + assert.strictEqual(cap.role, 'feature'); + assert.strictEqual(cap.tier, 'full'); + }); + + test('execution_backend is an enum with auto|workflow|inline defaulting to auto', () => { + const slice = loadCap().config['claude_orchestration.execution_backend']; + assert.strictEqual(slice.type, 'enum'); + assert.deepStrictEqual(slice.values, ['auto', 'workflow', 'inline']); + assert.strictEqual(slice.default, 'auto'); + }); + + test('registers at WIRED points only (execute:wave:post, plan:post)', () => { + const cap = loadCap(); + const points = cap.contributions.map((c) => c.point); + for (const p of points) { + assert.ok( + ['discuss:pre', 'discuss:post', 'plan:pre', 'plan:post', 'execute:post', 'execute:wave:post', 'verify:post', 'ship:pre', 'ship:post'].includes(p), + 'contribution point ' + p + ' must be a wired point', + ); + } + assert.ok(points.includes('execute:wave:post'), 'registers the execute wave hook'); + assert.ok(points.includes('plan:post'), 'declares plan:* ownership for ultraplan (criterion 5)'); + }); + + test('all contributions gated by the enabled key + onError:skip (default-resilient)', () => { + const cap = loadCap(); + for (const c of cap.contributions) { + assert.strictEqual(c.when, 'claude_orchestration.enabled', 'every contribution gated by enabled'); + assert.strictEqual(c.onError, 'skip', 'every contribution onError:skip'); + } + }); +}); + +// ─── 5. Registry integration ────────────────────────────────────────────────── + +describe('registry integration', () => { + + test('loadAndValidate includes claude-orchestration with no errors', () => { + const { capMap, errors } = loadAndValidate(new Set()); // empty central keys = no collision noise + // Filter errors to only those touching our capability. + const ours = errors.filter((e) => e.includes('claude-orchestration')); + assert.deepEqual(ours, [], 'our capability produced errors: ' + JSON.stringify(ours)); + assert.ok(capMap.has('claude-orchestration'), 'capMap includes claude-orchestration'); + }); + + test('buildRegistry surfaces the federated config keys in configSchema', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + assert.ok(registry.configSchema['claude_orchestration.enabled'], 'enabled key federated'); + assert.ok(registry.configSchema['claude_orchestration.execution_backend'], 'execution_backend key federated'); + assert.strictEqual(registry.configSchema['claude_orchestration.enabled'].owner, 'claude-orchestration'); + assert.strictEqual(registry.configSchema['claude_orchestration.execution_backend'].default, 'auto'); + }); + + test('byLoopPoint[execute:wave:post].contributions includes our capability', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + const contribs = registry.byLoopPoint['execute:wave:post'].contributions; + const ours = contribs.find((c) => c.capId === 'claude-orchestration'); + assert.ok(ours, 'our execute:wave:post contribution is registered'); + assert.strictEqual(ours.into, 'executor'); + }); + + test('committed registry is in sync (gen-capability-registry --check)', () => { + const { capMap } = loadAndValidate(new Set()); + const registry = buildRegistry(capMap); + const live = serializeRegistry(registry, capMap); + const committed = fs.readFileSync(REGISTRY_PATH, 'utf8'); + assert.strictEqual( + normalizeLineEndings(stripGeneratedComment(committed)), + normalizeLineEndings(stripGeneratedComment(live)), + 'registry is stale — run: node scripts/gen-capability-registry.cjs --write', + ); + }); +}); + +// ─── 6. Inline-fallback parity (criterion 3 + 6) ────────────────────────────── + +describe('inline-fallback parity', () => { + + test('default config (capability off) -> inline on every runtime, including Claude', () => { + // The capability ships default-off; with no user opt-in the backend is always inline. + const defaultCfg = {}; // nothing set + for (const runtimeId of ['claude', 'codex', 'cursor', 'opencode']) { + const r = detectWorkflowBackend({ + runtimeId, + hostIntegration: CAPABLE_HOST, + agentSdkVersion: '1.0.0', + config: defaultCfg, + }); + assert.strictEqual(r.backend, 'inline', runtimeId + ' default must be inline'); + assert.strictEqual(r.available, false, runtimeId + ' default must be unavailable'); + } + }); + + test('generated Workflow script preserves the inline-path contract (same agent + isolation + artifact)', () => { + // Criterion 2: the emitted Workflow composes the SAME gsd-executor agent and worktree + // isolation the inline path uses, and produces the same SUMMARY.md artifact. + const r = emitWorkflowScript(singleWaveManifest()); + assert.ok(r.script.includes('gsd-executor'), 'same executor agent as inline dispatch'); + assert.ok(r.script.includes('worktree'), 'same worktree isolation as inline dispatch'); + assert.ok(r.script.includes('SUMMARY.md'), 'same SUMMARY.md artifact as inline dispatch'); + }); +}); diff --git a/tests/execute-wave-post-gate-pipeline-e2e.test.cjs b/tests/execute-wave-post-gate-pipeline-e2e.test.cjs index 99fdbeec9..c01f1903d 100644 --- a/tests/execute-wave-post-gate-pipeline-e2e.test.cjs +++ b/tests/execute-wave-post-gate-pipeline-e2e.test.cjs @@ -629,15 +629,17 @@ describe('F. Real registry execute:wave:post shape — guard against accidental `ui.safety-gate onError must be 'halt'; got ${uiGate.onError}`); }); - test('[happy] real registry: execute:wave:post has no steps and 2 contributions (mempalace capture-problems + external-job executor fragment)', () => { + test('[happy] real registry: execute:wave:post has no steps and 3 contributions (claude-orchestration executor + external-job executor + mempalace capture-problems)', () => { const point = realRegistry.byLoopPoint['execute:wave:post']; assert.strictEqual(point.steps.length, 0, `execute:wave:post steps must be empty; got ${point.steps.length}`); - assert.strictEqual(point.contributions.length, 2, - `execute:wave:post must have 2 contributions (mempalace + external-job); got ${point.contributions.length}`); + // #1143: claude-orchestration registers an execute:wave:post contribution + // providing the Workflow-tool backend guidance (default-off, claude-only). + assert.strictEqual(point.contributions.length, 3, + `execute:wave:post must have 3 contributions (claude-orchestration + external-job + mempalace); got ${point.contributions.length}`); const capIds = point.contributions.map(c => c.capId).sort(); - assert.deepStrictEqual(capIds, ['external-job', 'mempalace'], - `execute:wave:post contributions must be mempalace + external-job; got ${capIds.join(',')}`); + assert.deepStrictEqual(capIds, ['claude-orchestration', 'external-job', 'mempalace'], + `execute:wave:post contributions must be claude-orchestration + external-job + mempalace; got ${capIds.join(',')}`); }); });