* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable Every emitted script was rejected. Four invalid constructs, the first fatal on its own, so the Workflow backend could never dispatch a wave: 1. no `export const meta = {…}` first statement -> whole script rejected 2. resumeFromRunId("<id>") -> "resumeFromRunId is not defined". It is a Workflow TOOL INPUT parameter, not a script function. The run id still reaches the caller via summary.resumeRunId, to pass as that input. 3. budget(<n>) -> "budget is not a function". `budget` is a read-only object { total, spent(), remaining() } fed by the caller's token directive; a script cannot set it. Recorded as intent in a comment. 4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions". Now parallel([() => agent(…), …]) — passing agent() results directly also started every agent eagerly, before parallel() could bound concurrency. The single-plan stage had its own branch with the same parallel() defect; both branches are now one array-emitting path. Waves also emit phase() calls whose titles match meta.phases exactly, so progress groups correctly. Two secondary defects kept the script from ever being REACHED — which is why this shipped undetected: 5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no scriptable way" to introspect it and told callers to omit the flag, so gate 5 returned agent_sdk_version_unknown on every automated run while `capability state` still reported active:true. True for bash, false for Node: the router now reads the installed @anthropic-ai/claude-agent-sdk version, walking node_modules up the tree and reading package.json directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the SDK's exports map does not expose ./package.json. Precedence: explicit flag > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an unresolvable version still declines to inline. A too-old SDK now reports the truthful agent_sdk_version_below_floor instead of unknown. 6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any invocation without --runtime reported runtime_not_claude on an ordinary Claude project. Now delegates to runtime-slash.resolveRuntime. The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` snippet is removed rather than repaired: it was also shell-dependent — zsh does not word-split unquoted parameter expansions, so it collapsed to a single argv element, argValue() never matched, and the run failed into the same agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto- resolution removes the need for the construct entirely. Verified with the issue's own repro: no flags now reaches the version gate; an SDK above the floor yields backend:"workflow" with a script that parses as a real ES module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids Findings from the isolated review, all fixed. HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was already RED because of it. The registry embeds the fragment text INLINE, so the shipped/installed copy still taught the exact broken contract this PR fixes: the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown" guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of checking its exit code, so I recorded a red chain as green — checking $? now.) HIGH — three existing tests asserted the OLD broken shape and would have failed CI; none was touched by the first commit: tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…") tests/claude-orchestration.test.cjs — .includes('budget(') tests/claude-orchestration-command-router.test.cjs — .includes('budget(') Each now asserts the corrected contract: the id/pool reaches the caller via summary, and neither construct is ever CALLED. Two sibling assertions had also gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the new explanatory COMMENT contains that substring, not because anything is wired. Rewritten to assert the real property. MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked within a wave, but nothing checked wave ids across waves. That was harmless before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a matching meta.phases entry and the tool matches titles by exact string — two waves sharing an id would collapse into one progress group and misattribute the second wave's agents to the first. Rejected at validation, with tests either side of the boundary. MEDIUM — the fragment contradicted itself (its "Manifest construction" header still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never told the orchestrator to pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to a tool-invocation input, an implementer following only the fragment would have silently regressed phase-resume to a no-op. Both fixed. MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and docs/explanation/claude-orchestration-capability.md documented `resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output — teaching the bug as the feature. Updated to the real contract, including the required meta block and the thunk-array parallel() form. (The changeset is `Fixed`, so the docs gate exempts this; it is corrected because it is wrong, not because a gate demanded it.) LOW — the router's top-of-file comment still described the divergent `--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the fix; and inserting resolveInstalledAgentSdkVersion had orphaned resolveDetectionArgs' JSDoc above the wrong function. Both repaired. lint:ci now exits 0 (verified by exit code, not by reading output). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2590): backfill changeset pr number (#2681) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
9.1 KiB
How to enable and use the Claude orchestration backend (BETA)
Run GSD's execute-phase waves through Claude Code's Workflow tool (/effort ultracode, Agent SDK ≥ v0.3.149) instead of the default one-agent-per-message dispatch, and fold the gsd-ultraplan-phase plan-offload under the same gate. On Claude Code this restores the wave parallelism that backgrounded-agent nesting (#853) otherwise forces inline.
BETA. This capability tracks a Claude Code preview surface. It is default-off, fail-closed, and Claude-only. Every detection miss degrades silently to today's inline behaviour — enabling it can never break the loop. See the explanation doc for the why, and ADR-1143 for the design.
What you need:
- GSD installed with the
fullprofile (the capability istier: full). - Claude Code with the Workflow tool available (Agent SDK ≥
0.3.149). On any other runtime the capability is an explicit no-op — you can flip the switch safely, nothing happens. - A GSD project with at least one planned phase (you need a wave/plan manifest to emit a script for).
Step 1 — Enable the capability
The capability ships disabled. Turn on the master switch inside your GSD project:
gsd-tools query config-set claude_orchestration.enabled true
That single key gates everything — both the Workflow-backend hook at execute:wave:post and the ultraplan ownership declaration at plan:post. All other claude_orchestration.* keys are optional refinements.
Verify it took:
gsd-tools query config-get claude_orchestration.enabled
# → true
Step 2 — Check whether your runtime qualifies
Detection is fail-closed: the Workflow backend activates only when every gate opens. Before relying on it, confirm your runtime reports as capable:
gsd-tools claude-orchestration detect-backend \
--runtime claude \
--agent-sdk-version 1.2.0
You will get one of two results:
backend |
available |
Meaning |
|---|---|---|
workflow |
true |
Every gate passed — the emitter will produce a Workflow script the orchestrator can run. |
inline |
false |
A gate failed. The reason field tells you which: capability_disabled, runtime_not_claude, backend_inline, workflow_tool_unavailable, agent_sdk_version_unknown, or agent_sdk_version_below_floor. |
The CLI is a simulation harness, not a probe.
detect-backendassumes a capable host descriptor unless you pass--no-nested-dispatch. It exists so you (and the orchestrator) can ask "given these facts, would the backend activate?" The real detection the loop uses is the puredetectWorkflowBackendfunction, called with the live host descriptor.
If detection returns inline
Work through the reason:
runtime_not_claude— you are on Codex / Cursor / opencode / etc. The Workflow tool is Claude-specific; there is nothing to enable here. Your loop is unchanged.agent_sdk_version_below_floor— upgrade Claude Code / the Agent SDK to at leastclaude_orchestration.min_agent_sdk_version(default0.3.149). A pre-release of the floor (e.g.0.3.149-rc.1) compares below the GA release and will not activate.workflow_tool_unavailable— your host descriptor does not advertise nested + background dispatch. This is unusual on Claude Code; if you see it, the Workflow tool is not present in this session.agent_sdk_version_unknown— the version could not be determined. Supply it explicitly via--agent-sdk-version.
Pin a higher floor (optional)
If you want to gate the BETA behind a newer Agent SDK than the default:
gsd-tools query config-set claude_orchestration.min_agent_sdk_version 1.0.0
Step 3 — Choose the execution backend
claude_orchestration.execution_backend controls how aggressively the backend is used once detection passes:
| Value | Behaviour |
|---|---|
auto (default) |
Use the Workflow backend if detection passes; otherwise inline. The safe, recommended value. |
workflow |
Force the Workflow backend when the tool is present (still fails closed to inline if the tool is absent or the SDK is too old — the floor applies in both modes). |
inline |
Force today's manual one-agent-per-message dispatch, even on a capable Claude Code runtime. Use this to A/B compare or to temporarily retire the BETA. |
Switch with:
gsd-tools query config-set claude_orchestration.execution_backend workflow
Step 4 — Emit a Workflow script for a phase
With the capability enabled and detection passing, generate the Workflow script for a phase's wave/plan manifest. The manifest is the wave/plan model execute-phase already builds:
{
"waves": [
{
"id": "w1",
"plans": [
{ "id": "p1", "brief": "Implement the foo module", "files_modified": ["src/foo.cts"] },
{ "id": "p2", "brief": "Wire the bar seam", "files_modified": ["src/bar.cts"] }
]
}
]
}
Emit the script:
gsd-tools claude-orchestration emit-workflow \
--waves .planning/phases/01-foo/waves.json \
--run-id phase-01-foo \
--phase-dir .planning/phases/01-foo \
--budget 500000
The output is a generated Workflow script that maps GSD's model 1:1 onto Workflow primitives:
- an
export const meta = { name, description, phases }block as the first statement (the Workflow tool rejects any script without it), - waves → a
phase("Wave <id>")group and a sequentialawait parallel([...])barrier (split into separate stages within a wave whenfiles_modifiedoverlap), - plans →
() => agent(brief, { agentType: "gsd-executor", isolation: "worktree" })— the same executor agent and worktree isolation the inline path uses, each wrapped in a thunk becauseparallel()takes an array of functions, - the phase run id in
summary.resumeRunId, and the intended token pool insummary.budgetTokens.
resumeFromRunId and budget are not emitted as calls (#2590). resumeFromRunId is a Workflow tool input, and budget is a read-only object ({ total, spent(), remaining() }) fed by the caller's token directive — calling either from a script throws and the whole script is rejected.
Because the script composes the same gsd-executor agent + worktree isolation + SUMMARY.md artifact as the inline path, the artifacts and commits it produces are identical — only the execution vehicle differs.
Run the emitted script
Feed the emitted script to Claude Code's Workflow tool (/effort ultracode, or an Agent SDK Workflow invocation), passing summary.resumeRunId as the tool's resumeFromRunId input. The orchestrator runs it; each agent() call spawns a gsd-executor in its own worktree, and waves barrier between each other. Omitting that input silently regresses phase-resume to a no-op — an interrupted phase re-runs completed plans.
Step 5 — Ultraplan plan-offload
Enabling the capability also folds gsd-ultraplan-phase under the same runtime gate. When the capability is on, the planner may offer the /gsd-ultraplan-phase path (offload plan-phase to Claude Code's ultraplan cloud) as an alternative to local /gsd-plan-phase. This is advisory — the stable local planner remains the default.
If the capability is off, or the runtime is not Claude Code, ultraplan offload is not surfaced and /gsd-plan-phase runs as normal.
Disabling
To turn the capability off and return to byte-identical inline behaviour:
gsd-tools query config-set claude_orchestration.enabled false
Or force inline dispatch while leaving the capability otherwise on:
gsd-tools query config-set claude_orchestration.execution_backend inline
Either step is sufficient — no uninstall or resurface needed. The federated config keys live only in the capability registry, so they vanish cleanly if the capability is ever removed.
What is and is not wired in BETA v1
Working today:
- Detection (
detectWorkflowBackend/gsd-tools claude-orchestration detect-backend) — fail-closed, tested across every gate. - Emission (
emitWorkflowScript/gsd-tools claude-orchestration emit-workflow) — waves→barriers, overlap→stages, resume, budget, anti-injection. - The contribution fragments at
execute:wave:postandplan:post(gated,onError: skip). - Inline fallback on every non-capable combination (regression-tested).
Not yet wired (follow-ups):
execute-phase.mddoes not yet auto-branch to emit-and-run the Workflow script. Today you emit the script explicitly (Step 4) and run it via the Workflow tool. Automatic dispatch inside the loop is the next milestone.- The plan-checker and verifier still run inline — this capability delivers the parallel-execution backend, not those gates.
- Full install-profile migration of the
gsd-ultraplan-phaseskill into the capability'sskills[](it is currently declared in the manifest; the skill's own runtime gate continues to no-op on non-Claude runtimes).
If a preview-API change breaks detection, the capability degrades to inline; it cannot destabilise the core loop.