Files
msd-core/capabilities/claude-orchestration/docs/workflow-backend-dispatch.md
Tom Boucher 06845717fe feat(#4740): make the Loop Host Contract role partition normative and enforced (#4742)
* test(#4740): pin the per-step role-family partition

Failing-first coverage for the Loop Host Contract role partition. At this
commit crossCheckRoleFamilies does not exist, so the rows throw
"crossCheckRoleFamilies is not a function" -- the RED proof they bind to
behavior rather than restating it.

ADR-894 section 3 assigns roles per step but parenthesises the assignment as
"(illustrative roles)", and nothing enforced it. The only thing standing in the
way was a single deepEqual in this same file, which is editable prose.

Rows cover: each step's own family accepted; a strict subset accepted; a
foreign role rejected at every step; an unknown role rejected; an unknown step
failing CLOSED; capitalization not silently matched; every offending role
reported rather than only the first; and purity, because buildContract puts the
same array into the generated contract.

Two rows exist because an earlier cut of this suite was vacuous. The purity
fixture is deliberately UNSORTED -- an alphabetically-sorted fixture cannot
fail an in-place sort(), and the mutant was being killed by three unrelated
rows instead. A parity row asserts ROLE_FAMILY and ROLE_TO_AGENT cover the
exact same role-name domain, both directions: they are parallel constants over
one domain, so divergence is the generative-fix class CLAUDE.md names.

Every negative row asserts the offending ROLE NAME and the STEP NAME appear in
the message. A count-only assertion survives a mutant that reports the wrong
role, which the 80% Stryker gate would surface only after a full CI round-trip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#4740): reject a cross-family agent-role declaration

Orchestration and execution are distinct functions of the loop and must not
drift into one another. That partition was real but unenforced: ADR-894
section 3 calls its own role assignment "illustrative", and the generator
accepted anything. Adding orchestrator to execute-phase.md's agent-roles line
compiled, --check passed once regenerated, and capability-validator.cjs then
began accepting into:"orchestrator" at every execute point.

ROLE_FAMILY maps every role to one of orchestration, planning or execution.
EXPECTED_FAMILY_BY_STEP gives each of the five steps exactly one family.
crossCheckRoleFamilies rejects a cross-family role, a role outside the
vocabulary, and an unknown step. It reports every offender, not the first.

It fails CLOSED on an unknown step, deliberately diverging from
assertPointsCoverage's "unknown step -- caught elsewhere". For points that is
true: the canonical-set and duplicate checks catch it. For roles there is no
second net, so failing open would leave an unknown step as the one input that
bypasses the gate.

crossCheckRoles' orchestrator exemption is untouched. ROLE_TO_AGENT maps roles
to agent FILES and the orchestrator is the host, owning none -- admissibility
and agent-file presence are separate concerns with separate checks.

Additive to section 3's existing rule that contribution.into must be a member
of the step's agentRoles, which is unchanged. That governs what a CAPABILITY
may target; this governs what a WORKFLOW may declare. No capability is
affected, and all five workflows already declare single-family sets, so the
gate is green on the commit that introduces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4740): make the ADR-894 role assignment normative

Section 3 parenthesises its per-step role assignment as "(illustrative roles)".
That word was accurate about the list's PURPOSE -- it illustrated the shape of
a generated contract entry -- and wrong about its STATUS, because the
assignment was load-bearing from the moment the generator consumed it. Read
literally it makes the partition an example rather than a rule.

Appended as a dated in-place section per docs/contributor-standards.md, which
records that an accepted ADR is never rewritten and names this the default
pattern. Section 3's original body is untouched.

The amendment states the three disjoint families, the one family each step
admits, that a step may declare a strict subset but never outside it, and why
this is a clarification rather than a new decision: the contract is generated
from the workflow markers "so it cannot drift into a lie", and all five
workflows have always declared single-family sets. What was absent was any
statement that it is required, and any check that it holds.

It also pins the distinction that is easy to re-merge: contribution.into being
a member of agentRoles governs what a CAPABILITY may target and is unchanged;
the family rule governs what a WORKFLOW may declare. The CONTEXT.md glossary
entry for the Loop Host Contract records the same, beside the agent-reference
drift guard it already documented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4740): backfill changeset pr number

Replaces the pr:0 placeholder with 4742 now that the PR exists. Verified with
GITHUB_BASE_REF=next, the way CI runs them: changeset lint and lint:docs both
go from invalid_pr(0) to ok. Without that env both report success without
evaluating the branch at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): stop injecting the orchestrator procedure into executors

claude-orchestration declared a contribution at execute:wave:pre with
into:"executor". loop-hook-dispatch.md defines a contribution as "inject
fragment.inline verbatim into the context for the role named in into", so its
267 lines were injected into EXECUTOR prompts whenever the capability was
enabled. Those lines are orchestration end to end -- construct a wave manifest,
resolve the dispatch backend, invoke the Workflow tool to spawn executors,
bridge per-agent results into the merge chain. An executor can act on none of
it.

Retargeting to into:"orchestrator" would not have been a fix. ROLE_TO_AGENT
carries no orchestrator entry by design: the orchestrator IS the host, and the
host's procedure lives in execute-phase.md. A step's agentRoles enumerates
agents a capability may inject context INTO, so adding orchestrator there would
model the host as an injectable agent -- the same category error pointed the
other way, and it would need an exception carved into the partition the same
issue just made normative.

So the defect is the mechanism, not the label. A contribution injects into an
agent's context; "replace step 3's inline dispatch loop" is a change to what
the HOST does. The contribution channel was serving as a host-behaviour
directive because it was the only channel available at an execute point.

The entry is removed. plan:post into:"planner" is correct and untouched. The
procedure is preserved verbatim at docs/workflow-backend-dispatch.md inside the
capability -- it is the only copy in the repo -- and is no longer injected
anywhere.

Consequence, not softened: the Workflow backend now has no loop wiring.
Detection, emission and config remain and the design is intact, but nothing
dispatches it. Under the separation ADR-1143 itself asserts it never had a
legitimate channel; ADR-1143's own audit already records the end-to-end path
has never been exercised. Wiring it properly needs a host-level mechanism that
does not exist today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4740): invert the stale execute:wave:pre registry assertions

Removing the contribution left four surfaces asserting or describing the old
state. Caught by an isolated review before a verification run was spent, which
is the point of reviewing first: the first of these was a guaranteed CI red.

execute-wave-post-gate-pipeline-e2e asserted against the REAL generated
registry that byLoopPoint['execute:wave:pre'] held exactly one contribution
with capId claude-orchestration. It now holds zero. Inverted to assert exactly
0 -- not a vague >= 0 -- and the #2285 comment above it now explains the
current state rather than the one it was written for.

CONTEXT.md's Claude Orchestration entry claimed two contributions at wired
points. It is now one, and the entry's execute:wave:post label was already
wrong before this change: the manifest said execute:wave:pre. Rewritten to one
plan:post contribution, why the execute-point one was removed, and where the
procedure now lives.

One assertion in claude-orchestration.test.cjs could not fail. It tested for
the prose "(into the executor)" while the doc says "(`into: executor`)", so no
plausible wording matched it and the paired plan:post assertion was carrying
the row. Replaced with a check on the structural claim, and proved RED by
restoring the two-contribution wording before reverting.

The moved procedure keeps section headings that speak as a live contribution --
"When this contribution is active", "Why execute:wave:pre". Preserving the body
verbatim was deliberate, so the headings stay and an editor's note under the
header explains why they read that way.

A sweep of all 17 files referencing byLoopPoint found no further siblings: the
remaining hits are a synthetic capability fixture and an empty-points test that
already expected no active hooks, both correct before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:30:47 -04:00

15 KiB
Raw Blame History

Claude orchestration — Workflow execution backend (BETA)

This is the orchestrator-side procedure for the Workflow execution backend. It is NOT a loop contribution and is not injected anywhere. It was previously declared as an execute:wave:pre contribution with into: "executor", which routed orchestrator instructions into executor prompts — a role-partition violation (the orchestrator only orchestrates, the executor only executes). It is retained here as the reference for whoever wires the orchestrator side through a host-level mechanism. See issue #4740.

Editor's note: the body below is preserved byte-for-byte from when this file WAS the execute:wave:pre contribution fragment, so its section headings and prose ("this contribution", "injected", etc.) still speak in those terms. That is intentional — it is not being rewritten to match its new status — and it is retained purely as the orchestrator-side reference described above.

When this contribution is active

The Claude orchestration capability is default-off and BETA. It activates only when ALL of the following hold:

  1. claude_orchestration.enabled is true in .planning/config.json, AND
  2. the active runtime is Claude Code (the Workflow tool is Claude / Agent SDK-specific), AND
  3. claude_orchestration.execution_backend resolves to workflow — either explicitly, or via auto — and the Agent SDK version is >= claude_orchestration.min_agent_sdk_version (default 0.3.149). The SDK floor applies in both auto and workflow modes (fail-closed: a pre-release or older SDK never activates the preview backend).

Detection is fail-closed: any miss degrades to inline, manual, one-agent-per- message dispatch — exactly today's behaviour. On a non-Claude runtime this contribution is a no-op.

Why execute:wave:pre (not execute:wave:post)

This is a dispatch-backend selector — it decides HOW a wave's executor agents are spawned. That decision has to be made BEFORE the wave's Agent() calls in execute-phase.md step 3, not after the wave has already finished (#2285). The capability previously registered at execute:wave:post, which fires only after worktree merge/post-merge tests/tracking updates — by then the wave was already dispatched inline, so the contribution was structurally unable to change how dispatch happened. This fragment is injected at the point that actually precedes dispatch.

What the orchestrator does when the Workflow backend is active

Before spawning executor agents for the current wave (execute-phase.md step 3), resolve the dispatch backend through the single composed CLI seam:

gsd-tools claude-orchestration resolve-wave-dispatch \
  --waves "$WAVE_MANIFEST_PATH" --run-id "$PHASE_RUN_ID" \
  --runtime "$RUNTIME" \
  --phase-dir "$PHASE_DIR" --raw

--agent-sdk-version is no longer passed here (#2590). The router resolves the installed Agent SDK version itself; see Agent SDK version below. The former ${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"} line was also shell-dependent: zsh does not word-split unquoted parameter expansions, so it collapsed to a SINGLE argv element there, argValue() never matched, and the run failed into agent_sdk_version_unknown — indistinguishable from genuinely unknown. Pass --agent-sdk-version <ver> explicitly only to pin a version.

This composes detectWorkflowBackend (the gate ladder above) with emitWorkflowScript (the wave→plan mapping below) in ONE call — the pure function backing it is resolveWaveDispatch in gsd-core/bin/lib/claude-orchestration.cjs. Response shape: { backend: 'inline'|'workflow', reason, script?, summary? }.

Manifest construction ($WAVE_MANIFEST_PATH, $PHASE_RUN_ID, $PHASE_DIR)

These are NOT pre-existing execute-phase.md variables — the orchestrator builds them at this step, from data it already has in-context from discover_and_group_plans (the PLAN_INDEX JSON) and step 2.5 (the per-plan USE_WORKTREES_FOR_PLAN decision):

  1. $PHASE_DIR — reuse {phase_dir} from the INIT bundle (already loaded in the initialize step). No new value needed.

  2. $PHASE_RUN_ID — a stable identifier for THIS phase-execution attempt, so resumeFromRunId can resume an interrupted run without re-dispatching plans the Workflow tool already completed. Construct it deterministically — execute-{phase_number}-{phase_slug} — from INIT's phase_number/phase_slug (both are already validated identifiers used elsewhere in this workflow, so they satisfy emitWorkflowScript's isScriptableIdentifier check). Do NOT mint a new random id per wave — the SAME $PHASE_RUN_ID is reused for every wave in the phase so the Workflow tool can correctly track cross-wave resume state.

  3. $WAVE_MANIFEST_PATH — a fresh temp file for THIS wave's manifest (one wave = one waves array with a single entry, matching the wave-by-wave dispatch loop; do not batch multiple waves into one manifest — waves are dispatched in wave order, not all at once):

    WAVE_MANIFEST_PATH=$(mktemp "${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX") && mv "$WAVE_MANIFEST_PATH" "$WAVE_MANIFEST_PATH.json" && WAVE_MANIFEST_PATH="$WAVE_MANIFEST_PATH.json"
    

    Then use the Write tool (not a bash/jq pipeline — the orchestrator already has every field parsed in-context) to write the manifest JSON to $WAVE_MANIFEST_PATH:

    {
      "waves": [
        {
          "id": "wave-{N}",
          "plans": [
            {
              "id": "{plan_id}",
              "brief": "{the SAME <objective>...<success_criteria> prompt block step 3 builds for this plan's inline Agent() call}",
              "files_modified": ["{from PLAN_INDEX.plans[].files_modified for this plan}"],
              "use_worktree": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}
            }
          ]
        }
      ]
    }
    
    • id — the plan id from PLAN_INDEX, e.g. "01-01".
    • brief — MUST carry the same task content as step 3's inline Agent() prompt (the <objective>/<execution_context>/<required_reading>/ <success_criteria> block, with {plan_number}/{phase_number}/ {phase_name} substituted) — a short summary here would NOT reproduce step 3's behavior and would violate the "identical artifacts" contract.
    • files_modified — copy verbatim from the plan's PLAN_INDEX entry.
    • use_worktree — true for every plan UNLESS step 2.5's per-plan worktree gate (execute-phase/steps/per-plan-worktree-gate.md) set USE_WORKTREES_FOR_PLAN=false for that plan (submodule-touching plan, or project-level USE_WORKTREES=false) — in which case pass false here so emitWorkflowScript omits isolation: "worktree" for that plan (#2772 / #2285 finding 1). Never hardcode true — that would force worktree isolation on a plan the inline path explicitly keeps out of worktrees.
  4. $AGENT_SDK_VERSION — no longer built here; the router resolves it.

Agent SDK version: the orchestrator has no bash-computable way to introspect the live Agent SDK version — but the router runs in Node, so it resolves the version itself (#2590), in this order:

  1. an explicit --agent-sdk-version <ver> (pin a version),
  2. GSD_AGENT_SDK_VERSION,
  3. the installed @anthropic-ai/claude-agent-sdk package version, read from its package.json on disk by walking node_modules up the tree. (Read directly rather than via require.resolve: the SDK's exports map does not expose ./package.json, so require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED.)

Previously nothing computed this at all, so gate 5 returned agent_sdk_version_unknown on every automated run and the Workflow backend could never activate — while gsd-tools capability state still reported the capability active: true. Fail-closed is preserved: when no version can be resolved, gate 5 still declines to inline. What changed is that a resolvable version is now actually found, so a genuinely-too-old SDK reports agent_sdk_version_below_floor — the truthful reason — instead of unknown.

If backend == "workflow": run the emitted script via the Workflow tool for THIS wave instead of the per-message Agent() loop in step 3. The script composes the SAME gsd-executor agent type the inline path uses, with worktree isolation applied PER PLAN from the manifest's use_worktree field (see emitWorkflowScript):

  • waves → one or more sequential parallel() barriers — each wave is a barrier group; when plans within a wave share files_modified, they are split into separate sequential stages within that wave's barrier.
  • plans → agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' }) when use_worktree is not false, or agent(brief, { agentType: 'gsd-executor' }) (no isolation) when it is — so the produced SUMMARY.md and commits are identical to inline dispatch, INCLUDING the inline path's submodule safety gate (#2772 / #2285 finding 1).
  • files_modified overlap → separate sequential stages — the same overlap rule execute-phase already applies inline (step 1 of the wave loop).
  • resumeFromRunId — pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT when you invoke the tool. It is a tool parameter, not a script function; the script deliberately does not call it (#2590 — doing so threw "resumeFromRunId is not defined" and rejected the entire script). Omitting it from the tool invocation silently regresses phase-resume to a no-op: an interrupted phase re-runs completed plans.

After the run: manifest bridge into the merge chain (#3302)

The single Workflow tool call replaces step 3's per-plan Agent() loop — which also means step 3's manifest bookkeeping (creation + per-agent recording) does NOT happen on this path. The orchestrator MUST bridge the run's per-agent results into the SAME manifest-scoped merge chain inline dispatch uses, before steps 4–5.8, which then run unchanged:

  1. Create the manifest BEFORE invoking the tool (this is step 3's creation block, which this path skips). When ANY plan in the wave has use_worktree not false:

    if [ -z "${WAVE_WORKTREE_MANIFEST:-}" ]; then
      M=$(mktemp "${TMPDIR:-/tmp}/gsd-worktree-wave-XXXXXX") && mv "$M" "$M.json" && WAVE_WORKTREE_MANIFEST="$M.json" || exit 1  # XXXXXX must be path-final on BSD/macOS (#1520)
      # Persist the dispatch-time orchestrator worktree root so wave-cleanup pins back
      # to the orchestrator's OWN worktree (#630), exactly as inline dispatch does.
      ORCH_ROOT=$(git rev-parse --show-toplevel)
      ORCH_ROOT="$ORCH_ROOT" MANIFEST="$WAVE_WORKTREE_MANIFEST" node -e 'const fs=require("fs");fs.writeFileSync(process.env.MANIFEST,JSON.stringify({orchestrator_root:process.env.ORCH_ROOT||null,worktrees:[]})+"\n")'
      export WAVE_WORKTREE_MANIFEST
    fi
    
  2. Invoke the Workflow tool with the emitted script and resumeFromRunId: summary.resumeRunId. The script top-level returns one entry per dispatched plan: { plan, expects_worktree, metadata }. metadata is that plan's executor <worktree_metadata> JSON ({agent_id, worktree_path, branch, expected_base} — captured by the executor itself per agents/gsd-executor.md), or null when the agent's result carried none (interrupted agent, resumed-from-cache plan, or a non-worktree plan).

  3. Record every worktree plan exactly as inline dispatch does at step 3's "After each Agent() returns" — one worktree.record-agent per returned entry with expects_worktree: true and complete metadata:

    gsd_run query worktree.record-agent --manifest "$WAVE_WORKTREE_MANIFEST" \
      --agent-id "<metadata.agent_id>" --path "<metadata.worktree_path>" \
      --branch "<metadata.branch>" --base "<metadata.expected_base>" \
      --files "<plan files_modified, space-separated>" \
      --deletions "<plan files_deleted, space-separated>"
    

    --deletions (#3003) carries the plan's declared files_deleted so a plan that scoped a file removal merges through cleanup-wave instead of being blocked. Unlike --files it is not advisory: omitting it leaves the deletions guard blocking on any deletion at all, so this dispatch path must pass it or plans declaring a removal fail to merge here while succeeding on the inline path.

    The verb's write-strict validation applies as inline: on a non-zero exit or any missing field, stop and ask for recovery — do not append an under-populated entry.

  4. HALT on uncapturable metadata — never a silently-empty manifest (#3302). After recording, the manifest must hold one entry per expects_worktree: true outcome (summary.worktreePlans from resolve-wave-dispatch is the expected count). Any shortfall — a null metadata, a missing/empty field, or a count mismatch — means commits are stranded on their worktree-wf_* branches and worktree.cleanup-wave would merge nothing while the phase looks green. STOP the phase with the failing plan id and the recovery hint below; do NOT run worktree.cleanup-wave and do NOT proceed to step 4.

    Recovery hint: the unmerged worktree-wf_* branch still holds the work. Recover the missing metadata from the run's per-agent result journal (journal.jsonl — one {"type":"result",…} line per agent — in the Workflow run's transcript dir), re-run worktree.record-agent by hand, then re-run cleanup. If the journal cannot be recovered either, merge the branch manually after review — never discard it.

  5. Resume (resumeFromRunId). Cached/resumed agents do not re-emit their final messages, so a previously-completed plan can return with metadata: null. Recover that plan's metadata from the ORIGINAL run's journal (same hint as above). If it cannot be recovered, fail loudly per rule 4 — a resumed run must never report success over silently-dropped agent work.

  6. Non-worktree plans (expects_worktree: false — use_worktree: false in the manifest): they ran without isolation; their commits are already on the main working tree. No record-agent entry, no manifest write.

With the manifest populated, steps 4–5.8 (wait/completion bookkeeping, step 5.5's manifest-scoped worktree.cleanup-wave, post-merge gate, tracking update) run UNCHANGED — the Workflow backend replaces HOW agents are spawned and returns their metadata; the merge chain itself is the inline path's own, now with real input.

If backend == "inline" (any gate miss, or resolve-wave-dispatch itself unavailable/erroring): proceed to step 3's standard per-message Agent() dispatch — the default, byte-identical-to-today path. onError: skip on this contribution means a resolve-wave-dispatch command failure is treated exactly like an inline result, never as a fatal wave error.

Fallback contract

Detection is fail-closed end-to-end: capability disabled, non-Claude runtime, execution_backend:"inline", missing/incapable host descriptor, unknown or below-floor Agent SDK version, or an emitWorkflowScript failure on a malformed wave manifest — ANY of these degrades to backend:"inline" and execute-phase's standard inline dispatch (step 3) runs unmodified. The Workflow backend never partially activates; the executor MUST NOT assume parallelism, a shared budget, or resume-from-run-id semantics when backend == "inline".