Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted.
15 KiB
Claude orchestration — Workflow execution backend (BETA)
This is the orchestrator-side procedure for the Workflow execution backend. It is NOT a loop contribution and is not injected anywhere. It was previously declared as an
execute:wave:precontribution withinto: "executor", which routed orchestrator instructions into executor prompts — a role-partition violation (the orchestrator only orchestrates, the executor only executes). It is retained here as the reference for whoever wires the orchestrator side through a host-level mechanism. See issue #4740.
Editor's note: the body below is preserved byte-for-byte from when this file WAS the
execute:wave:precontribution fragment, so its section headings and prose ("this contribution", "injected", etc.) still speak in those terms. That is intentional — it is not being rewritten to match its new status — and it is retained purely as the orchestrator-side reference described above.
When this contribution is active
The Claude orchestration capability is default-off and BETA. It activates only when ALL of the following hold:
claude_orchestration.enabledistruein.planning/config.json, AND- the active runtime is Claude Code (the Workflow tool is Claude / Agent SDK-specific), AND
claude_orchestration.execution_backendresolves toworkflow— either explicitly, or viaauto— and the Agent SDK version is>= claude_orchestration.min_agent_sdk_version(default0.3.149). The SDK floor applies in bothautoandworkflowmodes (fail-closed: a pre-release or older SDK never activates the preview backend).
Detection is fail-closed: any miss degrades to inline, manual, one-agent-per- message dispatch — exactly today's behaviour. On a non-Claude runtime this contribution is a no-op.
Why execute:wave:pre (not execute:wave:post)
This is a dispatch-backend selector — it decides HOW a wave's executor agents
are spawned. That decision has to be made BEFORE the wave's Agent() calls in
execute-phase.md step 3, not after the wave has already finished (#2285). The
capability previously registered at execute:wave:post, which fires only after
worktree merge/post-merge tests/tracking updates — by then the wave was already
dispatched inline, so the contribution was structurally unable to change how
dispatch happened. This fragment is injected at the point that actually precedes
dispatch.
What the orchestrator does when the Workflow backend is active
Before spawning executor agents for the current wave (execute-phase.md step 3), resolve the dispatch backend through the single composed CLI seam:
msd-tools claude-orchestration resolve-wave-dispatch \
--waves "$WAVE_MANIFEST_PATH" --run-id "$PHASE_RUN_ID" \
--runtime "$RUNTIME" \
--phase-dir "$PHASE_DIR" --raw
--agent-sdk-version is no longer passed here (#2590). The router resolves the
installed Agent SDK version itself; see Agent SDK version below. The former
${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"} line was also
shell-dependent: zsh does not word-split unquoted parameter expansions, so it
collapsed to a SINGLE argv element there, argValue() never matched, and the run
failed into agent_sdk_version_unknown — indistinguishable from genuinely
unknown. Pass --agent-sdk-version <ver> explicitly only to pin a version.
This composes detectWorkflowBackend (the gate ladder above) with
emitWorkflowScript (the wave→plan mapping below) in ONE call — the pure
function backing it is resolveWaveDispatch in
msd-core/bin/lib/claude-orchestration.cjs. Response shape:
{ backend: 'inline'|'workflow', reason, script?, summary? }.
Manifest construction ($WAVE_MANIFEST_PATH, $PHASE_RUN_ID, $PHASE_DIR)
These are NOT pre-existing execute-phase.md variables — the orchestrator builds
them at this step, from data it already has in-context from discover_and_group_plans
(the PLAN_INDEX JSON) and step 2.5 (the per-plan USE_WORKTREES_FOR_PLAN decision):
-
$PHASE_DIR— reuse{phase_dir}from theINITbundle (already loaded in theinitializestep). No new value needed. -
$PHASE_RUN_ID— a stable identifier for THIS phase-execution attempt, soresumeFromRunIdcan resume an interrupted run without re-dispatching plans the Workflow tool already completed. Construct it deterministically —execute-{phase_number}-{phase_slug}— fromINIT'sphase_number/phase_slug(both are already validated identifiers used elsewhere in this workflow, so they satisfyemitWorkflowScript'sisScriptableIdentifiercheck). Do NOT mint a new random id per wave — the SAME$PHASE_RUN_IDis reused for every wave in the phase so the Workflow tool can correctly track cross-wave resume state. -
$WAVE_MANIFEST_PATH— a fresh temp file for THIS wave's manifest (one wave = onewavesarray with a single entry, matching the wave-by-wave dispatch loop; do not batch multiple waves into one manifest — waves are dispatched in wave order, not all at once):WAVE_MANIFEST_PATH=$(mktemp "${TMPDIR:-/tmp}/msd-wave-dispatch-XXXXXX") && mv "$WAVE_MANIFEST_PATH" "$WAVE_MANIFEST_PATH.json" && WAVE_MANIFEST_PATH="$WAVE_MANIFEST_PATH.json"Then use the Write tool (not a bash/jq pipeline — the orchestrator already has every field parsed in-context) to write the manifest JSON to
$WAVE_MANIFEST_PATH:{ "waves": [ { "id": "wave-{N}", "plans": [ { "id": "{plan_id}", "brief": "{the SAME <objective>...<success_criteria> prompt block step 3 builds for this plan's inline Agent() call}", "files_modified": ["{from PLAN_INDEX.plans[].files_modified for this plan}"], "use_worktree": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan} } ] } ] }id— the plan id fromPLAN_INDEX, e.g."01-01".brief— MUST carry the same task content as step 3's inlineAgent()prompt (the<objective>/<execution_context>/<required_reading>/<success_criteria>block, with{plan_number}/{phase_number}/{phase_name}substituted) — a short summary here would NOT reproduce step 3's behavior and would violate the "identical artifacts" contract.files_modified— copy verbatim from the plan'sPLAN_INDEXentry.use_worktree—truefor every plan UNLESS step 2.5's per-plan worktree gate (execute-phase/steps/per-plan-worktree-gate.md) setUSE_WORKTREES_FOR_PLAN=falsefor that plan (submodule-touching plan, or project-levelUSE_WORKTREES=false) — in which case passfalsehere soemitWorkflowScriptomitsisolation: "worktree"for that plan (#2772 / #2285 finding 1). Never hardcodetrue— that would force worktree isolation on a plan the inline path explicitly keeps out of worktrees.
-
$AGENT_SDK_VERSION— no longer built here; the router resolves it.
Agent SDK version: the orchestrator has no bash-computable way to introspect the live Agent SDK version — but the router runs in Node, so it resolves the version itself (#2590), in this order:
- an explicit
--agent-sdk-version <ver>(pin a version), MSD_AGENT_SDK_VERSION,- the installed
@anthropic-ai/claude-agent-sdkpackage version, read from itspackage.jsonon disk by walkingnode_modulesup the tree. (Read directly rather than viarequire.resolve: the SDK'sexportsmap does not expose./package.json, sorequire.resolvethrowsERR_PACKAGE_PATH_NOT_EXPORTED.)
Previously nothing computed this at all, so gate 5 returned
agent_sdk_version_unknown on every automated run and the Workflow backend
could never activate — while msd-tools capability state still reported the
capability active: true. Fail-closed is preserved: when no version can be
resolved, gate 5 still declines to inline. What changed is that a resolvable
version is now actually found, so a genuinely-too-old SDK reports
agent_sdk_version_below_floor — the truthful reason — instead of unknown.
If backend == "workflow": run the emitted script via the Workflow tool
for THIS wave instead of the per-message Agent() loop in step 3. The script
composes the SAME msd-executor agent type the inline path uses, with
worktree isolation applied PER PLAN from the manifest's use_worktree field
(see emitWorkflowScript):
- waves → one or more sequential
parallel()barriers — each wave is a barrier group; when plans within a wave sharefiles_modified, they are split into separate sequential stages within that wave's barrier. - plans →
agent(brief, { agentType: 'msd-executor', isolation: 'worktree' })whenuse_worktreeis notfalse, oragent(brief, { agentType: 'msd-executor' })(no isolation) when it is — so the producedSUMMARY.mdand commits are identical to inline dispatch, INCLUDING the inline path's submodule safety gate (#2772 / #2285 finding 1). files_modifiedoverlap → separate sequential stages — the same overlap rule execute-phase already applies inline (step 1 of the wave loop).resumeFromRunId— passsummary.resumeRunIdas the Workflow tool'sresumeFromRunIdINPUT when you invoke the tool. It is a tool parameter, not a script function; the script deliberately does not call it (#2590 — doing so threw "resumeFromRunId is not defined" and rejected the entire script). Omitting it from the tool invocation silently regresses phase-resume to a no-op: an interrupted phase re-runs completed plans.
After the run: manifest bridge into the merge chain (#3302)
The single Workflow tool call replaces step 3's per-plan Agent() loop — which also
means step 3's manifest bookkeeping (creation + per-agent recording) does NOT happen on
this path. The orchestrator MUST bridge the run's per-agent results into the SAME
manifest-scoped merge chain inline dispatch uses, before steps 4–5.8, which then run
unchanged:
-
Create the manifest BEFORE invoking the tool (this is step 3's creation block, which this path skips). When ANY plan in the wave has
use_worktreenotfalse:if [ -z "${WAVE_WORKTREE_MANIFEST:-}" ]; then M=$(mktemp "${TMPDIR:-/tmp}/msd-worktree-wave-XXXXXX") && mv "$M" "$M.json" && WAVE_WORKTREE_MANIFEST="$M.json" || exit 1 # XXXXXX must be path-final on BSD/macOS (#1520) # Persist the dispatch-time orchestrator worktree root so wave-cleanup pins back # to the orchestrator's OWN worktree (#630), exactly as inline dispatch does. ORCH_ROOT=$(git rev-parse --show-toplevel) ORCH_ROOT="$ORCH_ROOT" MANIFEST="$WAVE_WORKTREE_MANIFEST" node -e 'const fs=require("fs");fs.writeFileSync(process.env.MANIFEST,JSON.stringify({orchestrator_root:process.env.ORCH_ROOT||null,worktrees:[]})+"\n")' export WAVE_WORKTREE_MANIFEST fi -
Invoke the Workflow tool with the emitted script and
resumeFromRunId: summary.resumeRunId. The script top-levelreturns one entry per dispatched plan:{ plan, expects_worktree, metadata }.metadatais that plan's executor<worktree_metadata>JSON ({agent_id, worktree_path, branch, expected_base}— captured by the executor itself peragents/msd-executor.md), ornullwhen the agent's result carried none (interrupted agent, resumed-from-cache plan, or a non-worktree plan). -
Record every worktree plan exactly as inline dispatch does at step 3's "After each
Agent()returns" — oneworktree.record-agentper returned entry withexpects_worktree: trueand complete metadata:msd_run query worktree.record-agent --manifest "$WAVE_WORKTREE_MANIFEST" \ --agent-id "<metadata.agent_id>" --path "<metadata.worktree_path>" \ --branch "<metadata.branch>" --base "<metadata.expected_base>" \ --files "<plan files_modified, space-separated>" \ --deletions "<plan files_deleted, space-separated>"--deletions(#3003) carries the plan's declaredfiles_deletedso a plan that scoped a file removal merges throughcleanup-waveinstead of being blocked. Unlike--filesit is not advisory: omitting it leaves the deletions guard blocking on any deletion at all, so this dispatch path must pass it or plans declaring a removal fail to merge here while succeeding on the inline path.The verb's write-strict validation applies as inline: on a non-zero exit or any missing field, stop and ask for recovery — do not append an under-populated entry.
-
HALT on uncapturable metadata — never a silently-empty manifest (#3302). After recording, the manifest must hold one entry per
expects_worktree: trueoutcome (summary.worktreePlansfromresolve-wave-dispatchis the expected count). Any shortfall — anullmetadata, a missing/empty field, or a count mismatch — means commits are stranded on theirworktree-wf_*branches andworktree.cleanup-wavewould merge nothing while the phase looks green. STOP the phase with the failing plan id and the recovery hint below; do NOT runworktree.cleanup-waveand do NOT proceed to step 4.Recovery hint: the unmerged
worktree-wf_*branch still holds the work. Recover the missing metadata from the run's per-agent result journal (journal.jsonl— one{"type":"result",…}line per agent — in the Workflow run's transcript dir), re-runworktree.record-agentby hand, then re-run cleanup. If the journal cannot be recovered either, merge the branch manually after review — never discard it. -
Resume (
resumeFromRunId). Cached/resumed agents do not re-emit their final messages, so a previously-completed plan can return withmetadata: null. Recover that plan's metadata from the ORIGINAL run's journal (same hint as above). If it cannot be recovered, fail loudly per rule 4 — a resumed run must never report success over silently-dropped agent work. -
Non-worktree plans (
expects_worktree: false—use_worktree: falsein the manifest): they ran without isolation; their commits are already on the main working tree. No record-agent entry, no manifest write.
With the manifest populated, steps 4–5.8 (wait/completion bookkeeping, step 5.5's
manifest-scoped worktree.cleanup-wave, post-merge gate, tracking update) run
UNCHANGED — the Workflow backend replaces HOW agents are spawned and returns their
metadata; the merge chain itself is the inline path's own, now with real input.
If backend == "inline" (any gate miss, or resolve-wave-dispatch itself
unavailable/erroring): proceed to step 3's standard per-message Agent()
dispatch — the default, byte-identical-to-today path. onError: skip on this
contribution means a resolve-wave-dispatch command failure is treated exactly
like an inline result, never as a fatal wave error.
Fallback contract
Detection is fail-closed end-to-end: capability disabled, non-Claude runtime,
execution_backend:"inline", missing/incapable host descriptor, unknown or
below-floor Agent SDK version, or an emitWorkflowScript failure on a malformed
wave manifest — ANY of these degrades to backend:"inline" and execute-phase's
standard inline dispatch (step 3) runs unmodified. The Workflow backend never
partially activates; the executor MUST NOT assume parallelism, a shared budget,
or resume-from-run-id semantics when backend == "inline".