* test(#4740): pin the per-step role-family partition Failing-first coverage for the Loop Host Contract role partition. At this commit crossCheckRoleFamilies does not exist, so the rows throw "crossCheckRoleFamilies is not a function" -- the RED proof they bind to behavior rather than restating it. ADR-894 section 3 assigns roles per step but parenthesises the assignment as "(illustrative roles)", and nothing enforced it. The only thing standing in the way was a single deepEqual in this same file, which is editable prose. Rows cover: each step's own family accepted; a strict subset accepted; a foreign role rejected at every step; an unknown role rejected; an unknown step failing CLOSED; capitalization not silently matched; every offending role reported rather than only the first; and purity, because buildContract puts the same array into the generated contract. Two rows exist because an earlier cut of this suite was vacuous. The purity fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot fail an in-place sort(), and the mutant was being killed by three unrelated rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the exact same role-name domain, both directions: they are parallel constants over one domain, so divergence is the generative-fix class CLAUDE.md names. Every negative row asserts the offending ROLE NAME and the STEP NAME appear in the message. A count-only assertion survives a mutant that reports the wrong role, which the 80% Stryker gate would surface only after a full CI round-trip. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#4740): reject a cross-family agent-role declaration Orchestration and execution are distinct functions of the loop and must not drift into one another. That partition was real but unenforced: ADR-894 section 3 calls its own role assignment "illustrative", and the generator accepted anything. Adding orchestrator to execute-phase.md's agent-roles line compiled, --check passed once regenerated, and capability-validator.cjs then began accepting into:"orchestrator" at every execute point. ROLE_FAMILY maps every role to one of orchestration, planning or execution. EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family. crossCheckRoleFamilies rejects a cross-family role, a role outside the vocabulary, and an unknown step. It reports every offender, not the first. It fails CLOSED on an unknown step, deliberately diverging from assertPointsCoverage's "unknown step -- caught elsewhere". For points that is true: the canonical-set and duplicate checks catch it. For roles there is no second net, so failing open would leave an unknown step as the one input that bypasses the gate. crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles to agent FILES and the orchestrator is the host, owning none -- admissibility and agent-file presence are separate concerns with separate checks. Additive to section 3's existing rule that contribution.into must be a member of the step's agentRoles, which is unchanged. That governs what a CAPABILITY may target; this governs what a WORKFLOW may declare. No capability is affected, and all five workflows already declare single-family sets, so the gate is green on the commit that introduces it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4740): make the ADR-894 role assignment normative Section 3 parenthesises its per-step role assignment as "(illustrative roles)". That word was accurate about the list's PURPOSE -- it illustrated the shape of a generated contract entry -- and wrong about its STATUS, because the assignment was load-bearing from the moment the generator consumed it. Read literally it makes the partition an example rather than a rule. Appended as a dated in-place section per docs/contributor-standards.md, which records that an accepted ADR is never rewritten and names this the default pattern. Section 3's original body is untouched. The amendment states the three disjoint families, the one family each step admits, that a step may declare a strict subset but never outside it, and why this is a clarification rather than a new decision: the contract is generated from the workflow markers "so it cannot drift into a lie", and all five workflows have always declared single-family sets. What was absent was any statement that it is required, and any check that it holds. It also pins the distinction that is easy to re-merge: contribution.into being a member of agentRoles governs what a CAPABILITY may target and is unchanged; the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary entry for the Loop Host Contract records the same, beside the agent-reference drift guard it already documented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4740): backfill changeset pr number Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both go from invalid_pr(0) to ok. Without that env both report success without evaluating the branch at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): stop injecting the orchestrator procedure into executors claude-orchestration declared a contribution at execute:wave:pre with into:"executor". loop-hook-dispatch.md defines a contribution as "inject fragment.inline verbatim into the context for the role named in into", so its 267 lines were injected into EXECUTOR prompts whenever the capability was enabled. Those lines are orchestration end to end -- construct a wave manifest, resolve the dispatch backend, invoke the Workflow tool to spawn executors, bridge per-agent results into the merge chain. An executor can act on none of it. Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT carries no orchestrator entry by design: the orchestrator IS the host, and the host's procedure lives in execute-phase.md. A step's agentRoles enumerates agents a capability may inject context INTO, so adding orchestrator there would model the host as an injectable agent -- the same category error pointed the other way, and it would need an exception carved into the partition the same issue just made normative. So the defect is the mechanism, not the label. A contribution injects into an agent's context; "replace step 3's inline dispatch loop" is a change to what the HOST does. The contribution channel was serving as a host-behaviour directive because it was the only channel available at an execute point. The entry is removed. plan:post into:"planner" is correct and untouched. The procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the capability -- it is the only copy in the repo -- and is no longer injected anywhere. Consequence, not softened: the Workflow backend now has no loop wiring. Detection, emission and config remain and the design is intact, but nothing dispatches it. Under the separation ADR-1143 itself asserts it never had a legitimate channel; ADR-1143's own audit already records the end-to-end path has never been exercised. Wiring it properly needs a host-level mechanism that does not exist today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4740): invert the stale execute:wave:pre registry assertions Removing the contribution left four surfaces asserting or describing the old state. Caught by an isolated review before a verification run was spent, which is the point of reviewing first: the first of these was a guaranteed CI red. execute-wave-post-gate-pipeline-e2e asserted against the REAL generated registry that byLoopPoint['execute:wave:pre'] held exactly one contribution with capId claude-orchestration. It now holds zero. Inverted to assert exactly 0 -- not a vague >= 0 -- and the #2285 comment above it now explains the current state rather than the one it was written for. CONTEXT.md's Claude Orchestration entry claimed two contributions at wired points. It is now one, and the entry's execute:wave:post label was already wrong before this change: the manifest said execute:wave:pre. Rewritten to one plan:post contribution, why the execute-point one was removed, and where the procedure now lives. One assertion in claude-orchestration.test.cjs could not fail. It tested for the prose "(into the executor)" while the doc says "(`into: executor`)", so no plausible wording matched it and the paired plan:post assertion was carrying the row. Replaced with a check on the structural claim, and proved RED by restoring the two-contribution wording before reverting. The moved procedure keeps section headings that speak as a live contribution -- "When this contribution is active", "Why execute:wave:pre". Preserving the body verbatim was deliberate, so the headings stay and an editor's note under the header explains why they read that way. A sweep of all 17 files referencing byLoopPoint found no further siblings: the remaining hits are a synthetic capability fixture and an empty-points test that already expected no active hooks, both correct before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
18 KiB
ADR-1143: Claude orchestration capability — Workflow tool (ultracode) as a runtime-gated loop execution backend [Proposed]
- Status: Proposed
- Date: 2026-06-12
- Issue: #1143
- Builds on: ADR-857 — host/core vs plug-in split, loop extension points,
runtimeCompat, federated config - Blocked by: #857 being released (Proposed → Accepted + capability infrastructure shipped). Not actionable until then.
- Relates to: #853 (Claude Code backgrounded agents cannot nest subagents), existing BETA skill
gsd-ultraplan-phase
Why this is still Proposed (audited 2026-07-17)
Confirmed shipped, on-tree: the capability is real and registered, not vaporware. capabilities/claude-orchestration/capability.json exists with detection + emission (detectWorkflowBackend / emitWorkflowScript) in src/claude-orchestration.cts (compiled to gsd-core/bin/lib/claude-orchestration.cjs), federated config (claude_orchestration.enabled / execution_backend / min_agent_sdk_version), and 1,552 lines of tests across tests/claude-orchestration.test.cjs (which now includes the #2285 wiring-fix regression coverage, folded in per #3334) and tests/claude-orchestration-command-router.test.cjs. The previously-fatal wiring bug, #2285 ("claude-orchestration capability (#1143) registered as active but never wired into execute-phase orchestrator prompt"), is closed COMPLETED (2026-07-15) — one day before this audit — and the owning feature issue #1143 is also closed COMPLETED.
The blocker. The ADR sets its own bar for ratification in its own Amendment (above): "flipping to Accepted follows maintainer sign-off on the E2E behaviour once exercised on Claude Code with the Workflow tool present." No such exercise is recorded anywhere in issues, PRs, or tests. Every test in the two files above operates at the contract or CLI-subprocess layer — asserting the shape of an emitted script or the return value of resolve-wave-dispatch — none constructs or executes an actual Workflow-tool run (grep -rn "Workflow(" tests/claude-orchestration*.test.cjs returns no hits). Two further gaps sit inside the ADR's own Decision section: (1) Decision §1's claimed net effect — "wave parallelism, the plan-checker, and the verifier are restored" — is narrower than what shipped: capability.json's own description says "the plan-checker and verifier remain inline until separately wired — this capability delivers the parallel-execution backend, not those gates"; (2) Decision §3's fold-in of the gsd-ultraplan-phase skill into the capability's skills[] has not happened — capability.json still shows "skills": [], and no follow-up issue for the migration the Amendment promises exists (searched via gh issue list --search, no result).
Unblock condition. Ratify once: (a) a real Claude Code session with the Workflow tool present and claude_orchestration.enabled=true drives an execute-phase wave through the Workflow backend, and the result is recorded (issue comment, PR, or a test that actually builds/executes a Workflow script rather than asserting emitted-script shape) — that is the maintainer sign-off the ADR itself asks for; and (b) the Decision section's "plan-checker and verifier restored" language is reconciled with the shipped scope (either corrected to match, or backed by a tracked issue for the deferred wiring capability.json already discloses). The skills[] migration (item 3) is lower priority since it is openly disclosed as deferred rather than silently dropped, but should carry a tracked issue number before ratification so it doesn't quietly vanish.
Context
Claude Code ships two orchestration primitives GSD does not yet treat as first-class:
-
ultraplan(/ultraplan <objective>,claude --teleport) — hands a planning task to a Claude Code web session running in plan mode; the plan is drafted in the cloud, reviewed/commented in the browser, then approved back to the terminal (which archives the web session). GSD already wraps this as the BETA skillgsd-ultraplan-phase(gsd-core/workflows/ultraplan-phase.md), which constructs a plan prompt, triggers/ultraplan, and relies on the user manually running/gsd:import --from <file>to bring the plan back. -
ultracode(/effort ultracode, or anultracode:prompt prefix) — setsxhighreasoning plus automatic Workflow orchestration. It is the trigger for the Workflow tool (Agent SDK ≥ v0.3.149): a deterministic JavaScript orchestrator that fans subagents out withagent(),parallel()(barrier),pipeline()(no barrier), andphase(), supportingisolation: 'worktree', a shared tokenbudget, customagentType, andresumeFromRunId. It can be activated per-session via/effort,--settings, or an Agent SDK control request ("ultracode": true). GSD does not use the Workflow tool at all.
Why this matters for the loop
GSD's execute-phase is wave-based: plans carry a wave number, waves run sequentially, and plans within a wave run in parallel when their files_modified sets do not overlap. Today GSD realizes that by dispatching, one Agent call per message, a backgrounded gsd-executor in a worktree.
On Claude Code this degrades. Backgrounded agents on Claude Code have no Agent/Task tool, so they cannot nest subagents (#853). The autonomous loop therefore falls back to inline sequential execution — and with it silently drops wave parallelism, the plan-checker, and the verifier — on the single runtime most GSD users run. The degradation is structural and outside GSD's control to fix at the agent layer.
The Workflow tool sidesteps #853 precisely because it is invoked from the main loop, not from a backgrounded agent: the script is the orchestrator and spawns the subagents itself. Its primitives map almost 1:1 onto GSD's existing wave model, and it adds three things GSD's hand-rolled fan-out lacks: determinism, a shared token budget, and resume.
Post ADR-857, GSD now has a home for exactly this kind of runtime-specific, preview-grade, toggleable feature: a Capability with runtimeCompat, a federated config slice, a tier, and loop-hook registration at the stable extension points. Before 857 there was nowhere clean to put runtime-gated loop behavior without forking core prose; after 857 it is a data manifest plus a registry regen.
Decision
Introduce a Claude orchestration Capability — capabilities/claude-orchestration/, role: feature, runtimeCompat: claude, default-off, BETA — that adopts Claude Code's orchestration primitives as optional, runtime-gated GSD surfaces. It is blocked on #857 being released and ships only after the capability infrastructure is Accepted.
1. Workflow tool as an optional loop execution backend (primary, NEW)
When the active runtime exposes the Workflow tool, execute-phase may emit a generated Workflow script in place of manual one-agent-per-message dispatch. The translation preserves GSD's existing model:
| GSD execute-phase concept | Workflow primitive |
|---|---|
| Wave (barrier between waves) | parallel() barrier; or pipeline() when plans flow through stages without a cross-plan barrier |
| Plan executor | agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' }) |
files_modified overlap forces sequential |
overlap forces the plan into a later pipeline stage (same rule, declaratively) |
| Cost discipline | the Workflow budget pool |
Resume / gsd-undo |
resumeFromRunId keyed to the phase manifest |
| Orchestrator-only STATE.md / ROADMAP.md writes | unchanged — only the orchestrator (the script's return) writes shared files; executors stay in their worktrees |
Net effect on Claude Code: wave parallelism, the plan-checker, and the verifier are restored — the capability path is not a backgrounded agent, so #853 does not apply — and GSD gains determinism, a shared budget, and resume it does not have today.
2. Registration at ADR-857 loop extension points
The capability registers hooks at execute:wave:pre, execute:wave:post, and execute:post, resolved through gsd_run loop render-hooks <point>. Activation is gated by:
- Runtime/tool detection — the active runtime must actually expose the Workflow tool (Claude Code / Agent SDK ≥ the pinned minimum). Detection miss → no-op.
- Federated config — a
claude_orchestration.*slice withexecution_backend: auto | workflow | inline(defaultauto, which selectsworkflowonly when the tool is present, elseinline).
When neither gate passes, the loop runs exactly as it does today. This is the central safety property: the inline/manual path remains the default and the only path on non-Claude runtimes and on Claude versions without the Workflow tool.
3. Fold ultraplan under the same capability
The existing gsd-ultraplan-phase BETA skill and commands/gsd/ultraplan-phase.md become capability-owned at plan:*. Plan-offload (ultraplan) and execute-orchestration (ultracode/Workflow) then share one runtime gate, one federated config slice, and one BETA boundary — instead of ultraplan living as a stray BETA skill while a second preview surface is added elsewhere. This also lets a future iteration close ultraplan's manual /gsd:import --from round-trip behind the same capability without touching core.
4. (Stretch) Other fan-out points opt into the same backend
map-codebase (parallel mappers), project/phase research (parallel researchers), and plan-review-convergence (parallel reviewers) are already manual parallel-shaped dispatches. Behind the same toggle and gate they can adopt the Workflow backend incrementally. Out of scope for the first cut; recorded so the capability is named for the general seam, not just execute-phase.
Resolved design details
Why a capability and not core
The Workflow tool is Claude Code / Agent SDK-specific and preview-grade. ADR-857 is explicit that the five-step loop plus shared-infrastructure skills are the privileged core and "every other feature is a Capability — a plug-in selectable at install and toggleable after restart." A runtime-specific, BETA execution backend is the textbook case for role: feature + runtimeCompat: claude + tier. Shipping it in core would re-introduce exactly the runtime-coupling 857 removed.
Fallback is the contract, not an afterthought
The capability must be byte-for-byte transparent when absent: identical artifacts, commits, and STATE/ROADMAP writes whether a wave ran via the Workflow backend or the inline path. A parity regression test asserting this on a runtime without the Workflow tool is a release gate, not a nicety — it is what lets the capability be default-off and low-risk.
Maintenance posture
The capability tracks two moving Claude-Code preview surfaces. Containment: one capability + BETA gate (mirroring today's gsd-ultraplan-phase BETA isolation), a pinned minimum Agent-SDK/Workflow-tool version in the runtime gate, and inline fallback on any detection miss. A preview-API change can degrade the capability to inline; it cannot break the core loop.
Relationship to cross-AI delegation and plan-review-convergence
These existing multi-model features (execute-phase cross_ai_delegation, the convergence reviewer loop) are model/runtime delegation, not intra-runtime fan-out orchestration. They are orthogonal: cross-AI routes a plan to a different CLI; the Workflow backend parallelizes plans within the current Claude Code runtime. They compose rather than conflict.
Consequences
- Positive: restores loop parallelism + plan-checker + verifier on Claude Code (undoing the #853 inline degradation); adds determinism, shared token budget, and resume to GSD execution; consolidates two Claude preview surfaces behind one gated capability; sets a reusable pattern for the other fan-out points.
- Negative / cost: another preview surface to track; a Workflow-script emitter and runtime/tool detection to maintain; BETA-grade until the upstream tool stabilizes.
- Neutral: no effect on non-Claude runtimes by construction; no behavior change until explicitly enabled.
Governance note: This ADR is a draft design accompanying feature request #1143. Per CONTRIBUTING, it is PR'd only after the issue receives
approved-feature, and the capability is implemented only after #857 is released.
Amendment (2026-07-06): BETA v1 implementation landed
#857 is released (CLOSED); the capability infrastructure is live. The BETA v1
of this capability has shipped as capabilities/claude-orchestration/ with the
scope agreed in the Decision, refined to the lowest-risk first slice:
- Detection + emission live as pure, fail-closed functions in
gsd-core/bin/lib/claude-orchestration.cjs(sourcesrc/claude-orchestration.cts):detectWorkflowBackend(gate ladder: enabled → Claude runtime → execution_backend ≠ inline → host dispatch nested+background → valid Agent SDK → SDK ≥claude_orchestration.min_agent_sdk_version, default0.3.149) andemitWorkflowScript(waves →parallel()stage barriers, plans →agent({ agentType: 'gsd-executor', isolation: 'worktree' }),files_modifiedoverlap → separate sequential stages,resumeFromRunIdwired to the phase run id, sharedbudget(tokens)pool). All interpolated values are validated as script-safe identifiers or JSON-quoted (review Finding 1). - Loop registration is at the two wired points the loop host contract
actually renders:
execute:wave:post into:executor(Workflow-backend guidance) andplan:post into:planner(ultraplan ownership declaration).execute:wave:preandexecute:preare declared in the contract but not wired today, so the capability registers atwave:post(the constraintexternal-jobalso documents). - Config is federated (
claude_orchestration.enableddefault false /activationKey,execution_backendenumauto|workflow|inlinedefaultauto,min_agent_sdk_version); the keys live only in the registry, so uninstall removes them cleanly. - ultraplan ownership is declared in the manifest (
plan:postcontribution); full install-profile migration of thegsd-ultraplan-phaseskill into the capability'sskills[]is deferred to a follow-up (it triggers the CLUSTERS / profile membership gate and is a heavier, install-machinery change).
Status remains Proposed — the BETA is default-off and the end-to-end Workflow execution path (actual orchestration via the Workflow tool inside Claude Code) is not verifiable outside that runtime. The capability is structurally complete and tested at the contract level; flipping to Accepted follows maintainer sign-off on the E2E behaviour once exercised on Claude Code with the Workflow tool present.
Amendment (2026-09-14): the execute:wave:* contribution is removed — orchestration is not an agent contribution
Issue #4740.
Decision §2 and the 2026-07-06 amendment above are historical record and are not rewritten; this section supersedes them on the single point of loop registration.
capabilities/claude-orchestration/capability.json declared a contribution at execute:wave:pre
with into: "executor". gsd-core/references/loop-hook-dispatch.md defines a contribution as
"Inject fragment.inline verbatim into the context for the role named in into", so that
fragment was injected into executor prompts whenever claude_orchestration.enabled.
Its 267 lines are orchestration end to end: construct a wave manifest, resolve the dispatch backend, invoke the Workflow tool to spawn executors, bridge per-agent results into the merge chain. An executor can act on none of it.
Retargeting it to into: "orchestrator" would not have been a fix. ROLE_TO_AGENT carries no
orchestrator entry by design — the orchestrator IS the host, not an agent, and the host's
procedure lives in gsd-core/workflows/execute-phase.md. A step's agentRoles enumerates agents a
capability may inject context INTO. Adding orchestrator there would model the host as an
injectable agent: the same category error pointed the other way, and it would have required
carving an exception into the role partition ADR-894 §3's 2026-09-14 amendment had just made
normative.
The defect is therefore the mechanism, not the label. A contribution injects into an agent's
context; "replace step 3's inline dispatch loop" (execute-phase.md:588) is a change to what the
host does. The contribution channel was being used as a host-behaviour directive because it was
the only channel available at an execute:* point.
What changed: the execute:wave:pre contribution is removed. The plan:post /
into: "planner" contribution is correct and is untouched. The orchestrator-side procedure is
preserved verbatim at capabilities/claude-orchestration/docs/workflow-backend-dispatch.md — it is
the only copy in the repo — and is no longer injected anywhere.
Consequence, stated plainly: the Workflow execution backend now has no loop wiring. Its
detection and emission code (detectWorkflowBackend, emitWorkflowScript) and its federated
config remain, and its design is intact in the preserved document, but nothing dispatches it. Under
the orchestration/execution separation this ADR itself asserts — "only the orchestrator (the
script's return) writes shared files; executors stay in their worktrees" — it never had a
legitimate channel. This makes that visible rather than changing it, and the "Why this is still
Proposed" audit above already records that the end-to-end path has never been exercised.
Wiring it properly needs a host-level mechanism for a capability to alter the orchestrator's own dispatch procedure. That does not exist today and is not proposed here.