Files
msd-core/docs/explanation/claude-orchestration-capability.md
Tom Boucher 06845717fe feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition

Failing-first coverage for the Loop Host Contract role partition. At this
commit crossCheckRoleFamilies does not exist, so the rows throw
"crossCheckRoleFamilies is not a function" -- the RED proof they bind to
behavior rather than restating it.

ADR-894 section 3 assigns roles per step but parenthesises the assignment as
"(illustrative roles)", and nothing enforced it. The only thing standing in the
way was a single deepEqual in this same file, which is editable prose.

Rows cover: each step's own family accepted; a strict subset accepted; a
foreign role rejected at every step; an unknown role rejected; an unknown step
failing CLOSED; capitalization not silently matched; every offending role
reported rather than only the first; and purity, because buildContract puts the
same array into the generated contract.

Two rows exist because an earlier cut of this suite was vacuous. The purity
fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot
fail an in-place sort(), and the mutant was being killed by three unrelated
rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the
exact same role-name domain, both directions: they are parallel constants over
one domain, so divergence is the generative-fix class CLAUDE.md names.

Every negative row asserts the offending ROLE NAME and the STEP NAME appear in
the message. A count-only assertion survives a mutant that reports the wrong
role, which the 80% Stryker gate would surface only after a full CI round-trip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#4740): reject a cross-family agent-role declaration

Orchestration and execution are distinct functions of the loop and must not
drift into one another. That partition was real but unenforced: ADR-894
section 3 calls its own role assignment "illustrative", and the generator
accepted anything. Adding orchestrator to execute-phase.md's agent-roles line
compiled, --check passed once regenerated, and capability-validator.cjs then
began accepting into:"orchestrator" at every execute point.

ROLE_FAMILY maps every role to one of orchestration, planning or execution.
EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family.
crossCheckRoleFamilies rejects a cross-family role, a role outside the
vocabulary, and an unknown step. It reports every offender, not the first.

It fails CLOSED on an unknown step, deliberately diverging from
assertPointsCoverage's "unknown step -- caught elsewhere". For points that is
true: the canonical-set and duplicate checks catch it. For roles there is no
second net, so failing open would leave an unknown step as the one input that
bypasses the gate.

crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles
to agent FILES and the orchestrator is the host, owning none -- admissibility
and agent-file presence are separate concerns with separate checks.

Additive to section 3's existing rule that contribution.into must be a member
of the step's agentRoles, which is unchanged. That governs what a CAPABILITY
may target; this governs what a WORKFLOW may declare. No capability is
affected, and all five workflows already declare single-family sets, so the
gate is green on the commit that introduces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4740): make the ADR-894 role assignment normative

Section 3 parenthesises its per-step role assignment as "(illustrative roles)".
That word was accurate about the list's PURPOSE -- it illustrated the shape of
a generated contract entry -- and wrong about its STATUS, because the
assignment was load-bearing from the moment the generator consumed it. Read
literally it makes the partition an example rather than a rule.

Appended as a dated in-place section per docs/contributor-standards.md, which
records that an accepted ADR is never rewritten and names this the default
pattern. Section 3's original body is untouched.

The amendment states the three disjoint families, the one family each step
admits, that a step may declare a strict subset but never outside it, and why
this is a clarification rather than a new decision: the contract is generated
from the workflow markers "so it cannot drift into a lie", and all five
workflows have always declared single-family sets. What was absent was any
statement that it is required, and any check that it holds.

It also pins the distinction that is easy to re-merge: contribution.into being
a member of agentRoles governs what a CAPABILITY may target and is unchanged;
the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary
entry for the Loop Host Contract records the same, beside the agent-reference
drift guard it already documented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): backfill changeset pr number

Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with
GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both
go from invalid_pr(0) to ok. Without that env both report success without
evaluating the branch at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): stop injecting the orchestrator procedure into executors

claude-orchestration declared a contribution at execute:wave:pre with
into:"executor". loop-hook-dispatch.md defines a contribution as "inject
fragment.inline verbatim into the context for the role named in into", so its
267 lines were injected into EXECUTOR prompts whenever the capability was
enabled. Those lines are orchestration end to end -- construct a wave manifest,
resolve the dispatch backend, invoke the Workflow tool to spawn executors,
bridge per-agent results into the merge chain. An executor can act on none of
it.

Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT
carries no orchestrator entry by design: the orchestrator IS the host, and the
host's procedure lives in execute-phase.md. A step's agentRoles enumerates
agents a capability may inject context INTO, so adding orchestrator there would
model the host as an injectable agent -- the same category error pointed the
other way, and it would need an exception carved into the partition the same
issue just made normative.

So the defect is the mechanism, not the label. A contribution injects into an
agent's context; "replace step 3's inline dispatch loop" is a change to what
the HOST does. The contribution channel was serving as a host-behaviour
directive because it was the only channel available at an execute point.

The entry is removed. plan:post into:"planner" is correct and untouched. The
procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the
capability -- it is the only copy in the repo -- and is no longer injected
anywhere.

Consequence, not softened: the Workflow backend now has no loop wiring.
Detection, emission and config remain and the design is intact, but nothing
dispatches it. Under the separation ADR-1143 itself asserts it never had a
legitimate channel; ADR-1143's own audit already records the end-to-end path
has never been exercised. Wiring it properly needs a host-level mechanism that
does not exist today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): invert the stale execute:wave:pre registry assertions

Removing the contribution left four surfaces asserting or describing the old
state. Caught by an isolated review before a verification run was spent, which
is the point of reviewing first: the first of these was a guaranteed CI red.

execute-wave-post-gate-pipeline-e2e asserted against the REAL generated
registry that byLoopPoint['execute:wave:pre'] held exactly one contribution
with capId claude-orchestration. It now holds zero. Inverted to assert exactly
0 -- not a vague >= 0 -- and the #2285 comment above it now explains the
current state rather than the one it was written for.

CONTEXT.md's Claude Orchestration entry claimed two contributions at wired
points. It is now one, and the entry's execute:wave:post label was already
wrong before this change: the manifest said execute:wave:pre. Rewritten to one
plan:post contribution, why the execute-point one was removed, and where the
procedure now lives.

One assertion in claude-orchestration.test.cjs could not fail. It tested for
the prose "(into the executor)" while the doc says "(`into: executor`)", so no
plausible wording matched it and the paired plan:post assertion was carrying
the row. Replaced with a check on the structural claim, and proved RED by
restoring the two-contribution wording before reverting.

The moved procedure keeps section headings that speak as a live contribution --
"When this contribution is active", "Why execute:wave:pre". Preserving the body
verbatim was deliberate, so the headings stay and an editor's note under the
header explains why they read that way.

A sweep of all 17 files referencing byLoopPoint found no further siblings: the
remaining hits are a synthetic capability fixture and an empty-points test that
already expected no active hooks, both correct before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:30:47 -04:00

6.3 KiB

Claude orchestration capability (BETA)

Explanation — why this capability exists and how it fits the loop. For the step-by-step, see the capability reference; for the design record, see ADR-1143.

The problem

GSD's execute-phase is wave-based: plans carry a wave number, waves run sequentially, and plans within a wave run in parallel when their files_modified sets don't overlap. On most runtimes GSD realizes that by fanning out one backgrounded gsd-executor agent (in a worktree) per plan.

On Claude Code that fan-out degrades. Backgrounded agents on Claude Code have no Agent/Task tool, so they cannot nest subagents (#853). The autonomous loop therefore falls back to inline sequential execution — and with it silently drops wave parallelism, the plan-checker, and the verifier — on the one runtime most GSD users run.

Claude Code ships an orchestration primitive that sidesteps exactly this: the Workflow tool (the engine behind /effort ultracode, Agent SDK ≥ v0.3.149). A Workflow script is the orchestrator — it runs from the main loop and spawns subagents itself via agent(), parallel() (barrier), pipeline(), and phase(), with isolation: 'worktree'. Two related capabilities are tool inputs rather than script functions: the token budget is a read-only object a script reads but cannot set, and resumeFromRunId is a parameter passed when invoking the tool.

The capability

claude-orchestration is a default-off, BETA, claude-only capability that adopts the Workflow tool as an optional, runtime-gated parallel-execution backend, and folds the existing gsd-ultraplan-phase plan-offload under the same gate. It is blocked-on-nothing now that the ADR-857 capability system is released.

  • role: feature, runtimeCompat.supported: ["claude"], tier: full.
  • activationKey: claude_orchestration.enabled — default false. Nothing changes until you opt in.
  • Registers at one wired loop point: plan:post (into the planner), onError: skip and gated by the enabled key. The capability previously also contributed at execute:wave:pre (into: executor), but that fragment was pure orchestrator procedure — build a wave manifest, resolve the dispatch backend, spawn executor agents — with nothing an executor agent can act on. Injecting orchestrator instructions into executor prompts violates the loop's role partition (the orchestrator orchestrates, the executor executes), so the contribution was removed (#4740). The procedure text is preserved at capabilities/claude-orchestration/docs/workflow-backend-dispatch.md as reference for wiring the orchestrator side through a host-level mechanism; it is not injected anywhere.

How it decides whether to activate

Detection is a pure, fail-closed function — detectWorkflowBackend. The Workflow backend activates only when every gate passes; any miss degrades to inline (today's behaviour):

  1. claude_orchestration.enabled is true.
  2. The runtime is Claude (the Workflow tool is Claude / Agent SDK-specific).
  3. claude_orchestration.execution_backend is auto or workflow (not inline).
  4. The host descriptor advertises dispatch.nested and dispatch.background (the nesting-capable Claude-Code shape — a proxy for Workflow-tool presence, meaningful only after gate 2).
  5. The Agent SDK reports a valid semver version.
  6. That version is >= claude_orchestration.min_agent_sdk_version (default 0.3.149). A pre-release of the floor (e.g. 0.3.149-rc.1) compares below the GA release per SemVer, so the preview backend stays off.

What the executor runs when the backend is active

emitWorkflowScript maps the phase's wave/plan model onto Workflow primitives:

GSD concept Workflow primitive
Wave parallel() stage barrier
Plan (use_worktree not false) agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })
Plan (use_worktree: false) agent(brief, { agentType: 'gsd-executor' }) (no isolation)
files_modified overlap forces the plans into separate sequential stages
Wave a phase("Wave <id>") group, matching a meta.phases entry
Phase run id summary.resumeRunId → pass as the Workflow tool's resumeFromRunId input
Phase token cap recorded in summary.budgetTokens; budget is read-only in a script

Because the emitted script composes the same gsd-executor agent the inline path uses, with worktree isolation applied per plan from the manifest's use_worktree field, it produces the same SUMMARY.md artifacts and commits — the only difference is the execution vehicle. use_worktree mirrors execute-phase.md step 2.5's per-plan submodule safety gate exactly: a plan that touches a submodule path is never forced into worktree isolation, whichever backend dispatches it (#2772).

The fallback contract

On any runtime lacking the Workflow tool — or when the capability is disabled, the SDK is too old, or detection fails for any reason — execute-phase proceeds with the standard inline wave dispatch. This is a release gate, not a nicety: a regression test asserts the inline fallback on every non-capable combination, so the capability is default-off and low-risk by construction.

BETA scope (v1)

The first slice ships detection + emission + declarative ultraplan ownership. The emitter is exercised at the contract level (structure, overlap splitting, resume, budget, anti-injection). End-to-end execution through the Workflow tool is verifiable only inside Claude Code with the tool present. Full install-profile migration of the gsd-ultraplan-phase skill into the capability's skills[] array is a follow-up (it touches the cluster/profile machinery); for v1 the manifest declares ultraplan ownership at plan:post and the existing skill's own runtime gate continues to no-op on non-Claude runtimes.