Files
msd-core/docs/explanation/claude-orchestration-capability.md
Tom Boucher 0d08c32048 fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable

Every emitted script was rejected. Four invalid constructs, the first fatal on
its own, so the Workflow backend could never dispatch a wave:

  1. no `export const meta = {…}` first statement -> whole script rejected
  2. resumeFromRunId("<id>")  -> "resumeFromRunId is not defined". It is a
     Workflow TOOL INPUT parameter, not a script function. The run id still
     reaches the caller via summary.resumeRunId, to pass as that input.
  3. budget(<n>)              -> "budget is not a function". `budget` is a
     read-only object { total, spent(), remaining() } fed by the caller's token
     directive; a script cannot set it. Recorded as intent in a comment.
  4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions".
     Now parallel([() => agent(…), …]) — passing agent() results directly also
     started every agent eagerly, before parallel() could bound concurrency.

The single-plan stage had its own branch with the same parallel() defect; both
branches are now one array-emitting path. Waves also emit phase() calls whose
titles match meta.phases exactly, so progress groups correctly.

Two secondary defects kept the script from ever being REACHED — which is why
this shipped undetected:

  5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no
     scriptable way" to introspect it and told callers to omit the flag, so
     gate 5 returned agent_sdk_version_unknown on every automated run while
     `capability state` still reported active:true. True for bash, false for
     Node: the router now reads the installed @anthropic-ai/claude-agent-sdk
     version, walking node_modules up the tree and reading package.json
     directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the
     SDK's exports map does not expose ./package.json. Precedence: explicit flag
     > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an
     unresolvable version still declines to inline. A too-old SDK now reports
     the truthful agent_sdk_version_below_floor instead of unknown.
  6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
     from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
     invocation without --runtime reported runtime_not_claude on an ordinary
     Claude project. Now delegates to runtime-slash.resolveRuntime.

The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}`
snippet is removed rather than repaired: it was also shell-dependent — zsh does
not word-split unquoted parameter expansions, so it collapsed to a single argv
element, argValue() never matched, and the run failed into the same
agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto-
resolution removes the need for the construct entirely.

Verified with the issue's own repro: no flags now reaches the version gate; an
SDK above the floor yields backend:"workflow" with a script that parses as a
real ES module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids

Findings from the isolated review, all fixed.

HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was
already RED because of it. The registry embeds the fragment text INLINE, so the
shipped/installed copy still taught the exact broken contract this PR fixes:
the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown"
guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of
checking its exit code, so I recorded a red chain as green — checking $? now.)

HIGH — three existing tests asserted the OLD broken shape and would have failed
CI; none was touched by the first commit:
  tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…")
  tests/claude-orchestration.test.cjs                 — .includes('budget(')
  tests/claude-orchestration-command-router.test.cjs  — .includes('budget(')
Each now asserts the corrected contract: the id/pool reaches the caller via
summary, and neither construct is ever CALLED. Two sibling assertions had also
gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the
new explanatory COMMENT contains that substring, not because anything is wired.
Rewritten to assert the real property.

MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked
within a wave, but nothing checked wave ids across waves. That was harmless
before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a
matching meta.phases entry and the tool matches titles by exact string — two
waves sharing an id would collapse into one progress group and misattribute the
second wave's agents to the first. Rejected at validation, with tests either
side of the boundary.

MEDIUM — the fragment contradicted itself (its "Manifest construction" header
still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never
told the orchestrator to pass summary.resumeRunId as the Workflow tool's
resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to
a tool-invocation input, an implementer following only the fragment would have
silently regressed phase-resume to a no-op. Both fixed.

MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and
docs/explanation/claude-orchestration-capability.md documented
`resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output —
teaching the bug as the feature. Updated to the real contract, including the
required meta block and the thunk-array parallel() form. (The changeset is
`Fixed`, so the docs gate exempts this; it is corrected because it is wrong,
not because a gate demanded it.)

LOW — the router's top-of-file comment still described the divergent
`--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the
fix; and inserting resolveInstalledAgentSdkVersion had orphaned
resolveDetectionArgs' JSDoc above the wrong function. Both repaired.

lint:ci now exits 0 (verified by exit code, not by reading output).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2590): backfill changeset pr number (#2681)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:49:38 -04:00

109 lines
5.9 KiB
Markdown

# Claude orchestration capability (BETA)
> **Explanation** — *why this capability exists and how it fits the loop.* For the
> step-by-step, see the [capability reference](../reference/capability-matrix.md);
> for the design record, see [ADR-1143](../adr/1143-claude-orchestration-capability.md).
## The problem
GSD's `execute-phase` is wave-based: plans carry a wave number, waves run
sequentially, and plans *within* a wave run in parallel when their
`files_modified` sets don't overlap. On most runtimes GSD realizes that by
fanning out one backgrounded `gsd-executor` agent (in a worktree) per plan.
On **Claude Code** that fan-out degrades. Backgrounded agents on Claude Code have
no `Agent`/`Task` tool, so they cannot nest subagents ([#853]). The autonomous
loop therefore falls back to **inline sequential execution** — and with it
silently drops wave parallelism, the plan-checker, and the verifier — on the one
runtime most GSD users run.
Claude Code ships an orchestration primitive that sidesteps exactly this: the
**Workflow tool** (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149).
A Workflow script *is* the orchestrator — it runs from the main loop and spawns
subagents itself via `agent()`, `parallel()` (barrier), `pipeline()`, and
`phase()`, with `isolation: 'worktree'`. Two related capabilities are **tool
inputs rather than script functions**: the token `budget` is a read-only object
a script reads but cannot set, and `resumeFromRunId` is a parameter passed when
invoking the tool.
## The capability
`claude-orchestration` is a **default-off, BETA, claude-only** capability that
adopts the Workflow tool as an optional, runtime-gated parallel-execution
backend, and folds the existing `gsd-ultraplan-phase` plan-offload under the same
gate. It is blocked-on-nothing now that the ADR-857 capability system is released.
- **`role: feature`**, `runtimeCompat.supported: ["claude"]`, `tier: full`.
- **`activationKey: claude_orchestration.enabled`** — default `false`. Nothing
changes until you opt in.
- Registers at two **wired** loop points: `execute:wave:pre` (into the executor)
and `plan:post` (into the planner). Both are `onError: skip` and gated by the
`enabled` key. The dispatch-backend selector fires at `execute:wave:pre` — the
seam that runs immediately BEFORE a wave's agents are dispatched — because a
selector fired *after* a wave already dispatched inline (the original
`execute:wave:post` placement, [#2285]) is structurally too late to change how
dispatch happens.
## How it decides whether to activate
Detection is a pure, **fail-closed** function — `detectWorkflowBackend`. The
Workflow backend activates only when *every* gate passes; any miss degrades to
`inline` (today's behaviour):
1. `claude_orchestration.enabled` is true.
2. The runtime is Claude (the Workflow tool is Claude / Agent SDK-specific).
3. `claude_orchestration.execution_backend` is `auto` or `workflow` (not `inline`).
4. The host descriptor advertises `dispatch.nested` **and** `dispatch.background`
(the nesting-capable Claude-Code shape — a proxy for Workflow-tool presence,
meaningful only after gate 2).
5. The Agent SDK reports a valid semver version.
6. That version is `>= claude_orchestration.min_agent_sdk_version`
(default `0.3.149`). A pre-release of the floor (e.g. `0.3.149-rc.1`) compares
*below* the GA release per SemVer, so the preview backend stays off.
## What the executor runs when the backend is active
`emitWorkflowScript` maps the phase's wave/plan model onto Workflow primitives:
| GSD concept | Workflow primitive |
|---|---|
| Wave | `parallel()` stage barrier |
| Plan (`use_worktree` not `false`) | `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })` |
| Plan (`use_worktree: false`) | `agent(brief, { agentType: 'gsd-executor' })` (no isolation) |
| `files_modified` overlap | forces the plans into separate sequential stages |
| Wave | a `phase("Wave <id>")` group, matching a `meta.phases` entry |
| Phase run id | `summary.resumeRunId` → pass as the Workflow tool's `resumeFromRunId` **input** |
| Phase token cap | recorded in `summary.budgetTokens`; `budget` is read-only in a script |
Because the emitted script composes the **same** `gsd-executor` agent the
inline path uses, with worktree isolation applied **per plan** from the
manifest's `use_worktree` field, it produces the same `SUMMARY.md` artifacts
and commits — the only difference is the execution vehicle. `use_worktree`
mirrors execute-phase.md step 2.5's per-plan submodule safety gate exactly: a
plan that touches a submodule path is never forced into worktree isolation,
whichever backend dispatches it ([#2772]).
## The fallback contract
On any runtime lacking the Workflow tool — or when the capability is disabled,
the SDK is too old, or detection fails for any reason — execute-phase proceeds
with the standard inline wave dispatch. This is a release gate, not a nicety: a
regression test asserts the inline fallback on every non-capable combination, so
the capability is default-off and low-risk by construction.
## BETA scope (v1)
The first slice ships **detection + emission + declarative ultraplan ownership**.
The emitter is exercised at the contract level (structure, overlap splitting,
resume, budget, anti-injection). End-to-end execution through the Workflow tool
is verifiable only inside Claude Code with the tool present. Full install-profile
migration of the `gsd-ultraplan-phase` skill into the capability's `skills[]`
array is a follow-up (it touches the cluster/profile machinery); for v1 the
manifest *declares* ultraplan ownership at `plan:post` and the existing skill's
own runtime gate continues to no-op on non-Claude runtimes.
[#853]: https://github.com/open-gsd/gsd-core/issues/853
[#1143]: https://github.com/open-gsd/gsd-core/issues/1143
[#2772]: https://github.com/open-gsd/gsd-core/issues/2772
[#2285]: https://github.com/open-gsd/gsd-core/issues/2285