diff --git a/.changeset/curious-lemurs-tumble.md b/.changeset/curious-lemurs-tumble.md new file mode 100644 index 000000000..9385370d3 --- /dev/null +++ b/.changeset/curious-lemurs-tumble.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 2681 +--- +**The Claude-orchestration Workflow backend can now actually dispatch a wave** — every script `emitWorkflowScript` generated was rejected by the Workflow tool. It omitted the required `export const meta = {…}` first statement (fatal on its own), called `resumeFromRunId()` and `budget()` which are a tool input parameter and a read-only object rather than script functions, and passed `parallel(agent(…), agent(…))` where an array of thunks is required. Two further defects meant the script was never even reached: nothing resolved the Agent SDK version, so the gate ladder returned `agent_sdk_version_unknown` on every automated run while `capability state` still reported the capability active; and the runtime fallback diverged from the canonical `GSD_RUNTIME > config.runtime > 'claude'` chain, so any invocation without `--runtime` reported `runtime_not_claude`. The router now resolves the installed SDK version itself and defers to the canonical runtime resolver, and the emitted script is valid ES module syntax with `phase()` titles matching `meta.phases`. (#2590) diff --git a/capabilities/claude-orchestration/fragments/execute-wave-pre.md b/capabilities/claude-orchestration/fragments/execute-wave-pre.md index 1af4f39dc..d87900da3 100644 --- a/capabilities/claude-orchestration/fragments/execute-wave-pre.md +++ b/capabilities/claude-orchestration/fragments/execute-wave-pre.md @@ -41,17 +41,24 @@ resolve the dispatch backend through the single composed CLI seam: gsd-tools claude-orchestration resolve-wave-dispatch \ --waves "$WAVE_MANIFEST_PATH" --run-id "$PHASE_RUN_ID" \ --runtime "$RUNTIME" \ - ${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"} \ --phase-dir "$PHASE_DIR" --raw ``` +`--agent-sdk-version` is no longer passed here (#2590). The router resolves the +installed Agent SDK version itself; see **Agent SDK version** below. The former +`${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` line was also +**shell-dependent**: zsh does not word-split unquoted parameter expansions, so it +collapsed to a SINGLE argv element there, `argValue()` never matched, and the run +failed into `agent_sdk_version_unknown` — indistinguishable from genuinely +unknown. Pass `--agent-sdk-version ` explicitly only to pin a version. + This composes `detectWorkflowBackend` (the gate ladder above) with `emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure function backing it is `resolveWaveDispatch` in `gsd-core/bin/lib/claude-orchestration.cjs`. Response shape: `{ backend: 'inline'|'workflow', reason, script?, summary? }`. -### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`, `$AGENT_SDK_VERSION`) +### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`) These are NOT pre-existing execute-phase.md variables — the orchestrator builds them at this step, from data it already has in-context from `discover_and_group_plans` @@ -116,15 +123,27 @@ them at this step, from data it already has in-context from `discover_and_group_ #2285 finding 1). **Never** hardcode `true` — that would force worktree isolation on a plan the inline path explicitly keeps out of worktrees. -4. **`$AGENT_SDK_VERSION`** — see below; OMIT when unknown (fails closed). +4. **`$AGENT_SDK_VERSION`** — no longer built here; the router resolves it. -**Agent SDK version:** the orchestrator has no scriptable (bash-computable) way -to introspect the live Agent SDK version. When it can determine the version -(e.g. from a host-exposed value it can read directly), pass -`--agent-sdk-version`. When it cannot, OMIT the flag — `resolveWaveDispatch`'s -gate 5 (`agent_sdk_version_unknown`) then fails closed to `inline` by design; -this is not a bug, it is the same fail-closed posture documented above applied -to a real absence of information. +**Agent SDK version:** the orchestrator has no *bash-computable* way to +introspect the live Agent SDK version — but the router runs in Node, so it +resolves the version itself (#2590), in this order: + +1. an explicit `--agent-sdk-version ` (pin a version), +2. `GSD_AGENT_SDK_VERSION`, +3. the **installed** `@anthropic-ai/claude-agent-sdk` package version, read from + its `package.json` on disk by walking `node_modules` up the tree. (Read + directly rather than via `require.resolve`: the SDK's `exports` map does not + expose `./package.json`, so `require.resolve` throws + `ERR_PACKAGE_PATH_NOT_EXPORTED`.) + +Previously nothing computed this at all, so gate 5 returned +`agent_sdk_version_unknown` on **every** automated run and the Workflow backend +could never activate — while `gsd-tools capability state` still reported the +capability `active: true`. Fail-closed is preserved: when no version can be +resolved, gate 5 still declines to `inline`. What changed is that a resolvable +version is now actually found, so a genuinely-too-old SDK reports +`agent_sdk_version_below_floor` — the truthful reason — instead of `unknown`. **If `backend == "workflow"`:** run the emitted `script` via the Workflow tool for THIS wave instead of the per-message `Agent()` loop in step 3. The script @@ -142,8 +161,12 @@ worktree isolation applied PER PLAN from the manifest's `use_worktree` field gate (#2772 / #2285 finding 1). - **`files_modified` overlap → separate sequential stages** — the same overlap rule execute-phase already applies inline (step 1 of the wave loop). -- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase - resumes without re-running completed plans. +- **`resumeFromRunId`** — **pass `summary.resumeRunId` as the Workflow tool's + `resumeFromRunId` INPUT when you invoke the tool.** It is a tool parameter, + not a script function; the script deliberately does not call it (#2590 — doing + so threw "resumeFromRunId is not defined" and rejected the entire script). + Omitting it from the tool invocation silently regresses phase-resume to a + no-op: an interrupted phase re-runs completed plans. The orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup, post-merge gate, tracking update) exactly as it does for inline dispatch — the diff --git a/docs/explanation/claude-orchestration-capability.md b/docs/explanation/claude-orchestration-capability.md index f7c7d0517..c14c15dc5 100644 --- a/docs/explanation/claude-orchestration-capability.md +++ b/docs/explanation/claude-orchestration-capability.md @@ -21,8 +21,10 @@ Claude Code ships an orchestration primitive that sidesteps exactly this: the **Workflow tool** (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149). A Workflow script *is* the orchestrator — it runs from the main loop and spawns subagents itself via `agent()`, `parallel()` (barrier), `pipeline()`, and -`phase()`, with `isolation: 'worktree'`, a shared token `budget`, and -`resumeFromRunId`. +`phase()`, with `isolation: 'worktree'`. Two related capabilities are **tool +inputs rather than script functions**: the token `budget` is a read-only object +a script reads but cannot set, and `resumeFromRunId` is a parameter passed when +invoking the tool. ## The capability @@ -69,8 +71,9 @@ Workflow backend activates only when *every* gate passes; any miss degrades to | Plan (`use_worktree` not `false`) | `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })` | | Plan (`use_worktree: false`) | `agent(brief, { agentType: 'gsd-executor' })` (no isolation) | | `files_modified` overlap | forces the plans into separate sequential stages | -| Phase run id | `resumeFromRunId("")` | -| Phase token cap | `budget()` | +| Wave | a `phase("Wave ")` group, matching a `meta.phases` entry | +| Phase run id | `summary.resumeRunId` → pass as the Workflow tool's `resumeFromRunId` **input** | +| Phase token cap | recorded in `summary.budgetTokens`; `budget` is read-only in a script | Because the emitted script composes the **same** `gsd-executor` agent the inline path uses, with worktree isolation applied **per plan** from the diff --git a/docs/how-to/enable-claude-orchestration-workflow-backend.md b/docs/how-to/enable-claude-orchestration-workflow-backend.md index a8fea05e1..2c0cc4342 100644 --- a/docs/how-to/enable-claude-orchestration-workflow-backend.md +++ b/docs/how-to/enable-claude-orchestration-workflow-backend.md @@ -116,16 +116,18 @@ gsd-tools claude-orchestration emit-workflow \ The output is a generated Workflow script that maps GSD's model 1:1 onto Workflow primitives: -- **waves → sequential `parallel()` barriers** (split into separate stages within a wave when `files_modified` overlap), -- **plans → `agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses, -- **`resumeFromRunId("")`** wired to the phase run id, -- **`budget()`** — a shared token pool across the whole phase (omit `--budget` to skip). +- an `export const meta = { name, description, phases }` block as the **first statement** (the Workflow tool rejects any script without it), +- **waves → a `phase("Wave ")` group and a sequential `await parallel([...])` barrier** (split into separate stages within a wave when `files_modified` overlap), +- **plans → `() => agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses, each wrapped in a thunk because `parallel()` takes an **array of functions**, +- the phase run id in **`summary.resumeRunId`**, and the intended token pool in **`summary.budgetTokens`**. + +`resumeFromRunId` and `budget` are **not** emitted as calls (#2590). `resumeFromRunId` is a Workflow **tool input**, and `budget` is a read-only object (`{ total, spent(), remaining() }`) fed by the caller's token directive — calling either from a script throws and the whole script is rejected. Because the script composes the same `gsd-executor` agent + worktree isolation + `SUMMARY.md` artifact as the inline path, the artifacts and commits it produces are identical — only the execution vehicle differs. ### Run the emitted script -Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation). The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, waves barrier between each other, and `resumeFromRunId` lets an interrupted phase resume without re-running completed plans. +Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation), **passing `summary.resumeRunId` as the tool's `resumeFromRunId` input**. The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, and waves barrier between each other. Omitting that input silently regresses phase-resume to a no-op — an interrupted phase re-runs completed plans. --- diff --git a/gsd-core/bin/lib/capability-registry.cjs b/gsd-core/bin/lib/capability-registry.cjs index 65b7cfc8d..076142d30 100644 --- a/gsd-core/bin/lib/capability-registry.cjs +++ b/gsd-core/bin/lib/capability-registry.cjs @@ -603,7 +603,7 @@ const capabilities = { "into": "executor", "fragment": { "path": "fragments/execute-wave-pre.md", - "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:pre` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## Why `execute:wave:pre` (not `execute:wave:post`)\n\nThis is a **dispatch-backend selector** — it decides HOW a wave's executor agents\nare spawned. That decision has to be made BEFORE the wave's `Agent()` calls in\n`execute-phase.md` step 3, not after the wave has already finished (#2285). The\ncapability previously registered at `execute:wave:post`, which fires only after\nworktree merge/post-merge tests/tracking updates — by then the wave was already\ndispatched inline, so the contribution was structurally unable to change how\ndispatch happened. This fragment is injected at the point that actually precedes\ndispatch.\n\n## What the orchestrator does when the Workflow backend is active\n\nBefore spawning executor agents for the current wave (execute-phase.md step 3),\nresolve the dispatch backend through the single composed CLI seam:\n\n```bash\ngsd-tools claude-orchestration resolve-wave-dispatch \\\n --waves \"$WAVE_MANIFEST_PATH\" --run-id \"$PHASE_RUN_ID\" \\\n --runtime \"$RUNTIME\" \\\n ${AGENT_SDK_VERSION:+--agent-sdk-version \"$AGENT_SDK_VERSION\"} \\\n --phase-dir \"$PHASE_DIR\" --raw\n```\n\nThis composes `detectWorkflowBackend` (the gate ladder above) with\n`emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure\nfunction backing it is `resolveWaveDispatch` in\n`gsd-core/bin/lib/claude-orchestration.cjs`. Response shape:\n`{ backend: 'inline'|'workflow', reason, script?, summary? }`.\n\n### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`, `$AGENT_SDK_VERSION`)\n\nThese are NOT pre-existing execute-phase.md variables — the orchestrator builds\nthem at this step, from data it already has in-context from `discover_and_group_plans`\n(the `PLAN_INDEX` JSON) and step 2.5 (the per-plan `USE_WORKTREES_FOR_PLAN` decision):\n\n1. **`$PHASE_DIR`** — reuse `{phase_dir}` from the `INIT` bundle (already loaded\n in the `initialize` step). No new value needed.\n\n2. **`$PHASE_RUN_ID`** — a stable identifier for THIS phase-execution attempt, so\n `resumeFromRunId` can resume an interrupted run without re-dispatching plans\n the Workflow tool already completed. Construct it deterministically —\n `execute-{phase_number}-{phase_slug}` — from `INIT`'s `phase_number`/`phase_slug`\n (both are already validated identifiers used elsewhere in this workflow, so\n they satisfy `emitWorkflowScript`'s `isScriptableIdentifier` check). Do NOT\n mint a new random id per wave — the SAME `$PHASE_RUN_ID` is reused for every\n wave in the phase so the Workflow tool can correctly track cross-wave resume\n state.\n\n3. **`$WAVE_MANIFEST_PATH`** — a fresh temp file for THIS wave's manifest (one\n wave = one `waves` array with a single entry, matching the wave-by-wave\n dispatch loop; do not batch multiple waves into one manifest — waves are\n dispatched in wave order, not all at once):\n\n ```bash\n WAVE_MANIFEST_PATH=$(mktemp \"${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX\") && mv \"$WAVE_MANIFEST_PATH\" \"$WAVE_MANIFEST_PATH.json\" && WAVE_MANIFEST_PATH=\"$WAVE_MANIFEST_PATH.json\"\n ```\n\n Then **use the Write tool** (not a bash/jq pipeline — the orchestrator already\n has every field parsed in-context) to write the manifest JSON to\n `$WAVE_MANIFEST_PATH`:\n\n ```json\n {\n \"waves\": [\n {\n \"id\": \"wave-{N}\",\n \"plans\": [\n {\n \"id\": \"{plan_id}\",\n \"brief\": \"{the SAME ... prompt block step 3 builds for this plan's inline Agent() call}\",\n \"files_modified\": [\"{from PLAN_INDEX.plans[].files_modified for this plan}\"],\n \"use_worktree\": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}\n }\n ]\n }\n ]\n }\n ```\n\n - **`id`** — the plan id from `PLAN_INDEX`, e.g. `\"01-01\"`.\n - **`brief`** — MUST carry the same task content as step 3's inline `Agent()`\n prompt (the ``/``/``/\n `` block, with `{plan_number}`/`{phase_number}`/\n `{phase_name}` substituted) — a short summary here would NOT reproduce\n step 3's behavior and would violate the \"identical artifacts\" contract.\n - **`files_modified`** — copy verbatim from the plan's `PLAN_INDEX` entry.\n - **`use_worktree`** — `true` for every plan UNLESS step 2.5's per-plan\n worktree gate (`execute-phase/steps/per-plan-worktree-gate.md`) set\n `USE_WORKTREES_FOR_PLAN=false` for that plan (submodule-touching plan, or\n project-level `USE_WORKTREES=false`) — in which case pass `false` here so\n `emitWorkflowScript` omits `isolation: \"worktree\"` for that plan (#2772 /\n #2285 finding 1). **Never** hardcode `true` — that would force worktree\n isolation on a plan the inline path explicitly keeps out of worktrees.\n\n4. **`$AGENT_SDK_VERSION`** — see below; OMIT when unknown (fails closed).\n\n**Agent SDK version:** the orchestrator has no scriptable (bash-computable) way\nto introspect the live Agent SDK version. When it can determine the version\n(e.g. from a host-exposed value it can read directly), pass\n`--agent-sdk-version`. When it cannot, OMIT the flag — `resolveWaveDispatch`'s\ngate 5 (`agent_sdk_version_unknown`) then fails closed to `inline` by design;\nthis is not a bug, it is the same fail-closed posture documented above applied\nto a real absence of information.\n\n**If `backend == \"workflow\"`:** run the emitted `script` via the Workflow tool\nfor THIS wave instead of the per-message `Agent()` loop in step 3. The script\ncomposes the SAME `gsd-executor` agent type the inline path uses, with\nworktree isolation applied PER PLAN from the manifest's `use_worktree` field\n(see `emitWorkflowScript`):\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier.\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n when `use_worktree` is not `false`, or `agent(brief, { agentType: 'gsd-executor' })`\n (no isolation) when it is — so the produced `SUMMARY.md` and commits are\n identical to inline dispatch, INCLUDING the inline path's submodule safety\n gate (#2772 / #2285 finding 1).\n- **`files_modified` overlap → separate sequential stages** — the same overlap\n rule execute-phase already applies inline (step 1 of the wave loop).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n\nThe orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup,\npost-merge gate, tracking update) exactly as it does for inline dispatch — the\nWorkflow backend only replaces HOW agents are spawned for this wave, not what\nhappens after they return.\n\n**If `backend == \"inline\"`** (any gate miss, or `resolve-wave-dispatch` itself\nunavailable/erroring): proceed to step 3's standard per-message `Agent()`\ndispatch — the default, byte-identical-to-today path. `onError: skip` on this\ncontribution means a `resolve-wave-dispatch` command failure is treated exactly\nlike an `inline` result, never as a fatal wave error.\n\n## Fallback contract\n\nDetection is fail-closed end-to-end: capability disabled, non-Claude runtime,\n`execution_backend:\"inline\"`, missing/incapable host descriptor, unknown or\nbelow-floor Agent SDK version, or an `emitWorkflowScript` failure on a malformed\nwave manifest — ANY of these degrades to `backend:\"inline\"` and execute-phase's\nstandard inline dispatch (step 3) runs unmodified. The Workflow backend never\npartially activates; the executor MUST NOT assume parallelism, a shared budget,\nor resume-from-run-id semantics when `backend == \"inline\"`.\n" + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:pre` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## Why `execute:wave:pre` (not `execute:wave:post`)\n\nThis is a **dispatch-backend selector** — it decides HOW a wave's executor agents\nare spawned. That decision has to be made BEFORE the wave's `Agent()` calls in\n`execute-phase.md` step 3, not after the wave has already finished (#2285). The\ncapability previously registered at `execute:wave:post`, which fires only after\nworktree merge/post-merge tests/tracking updates — by then the wave was already\ndispatched inline, so the contribution was structurally unable to change how\ndispatch happened. This fragment is injected at the point that actually precedes\ndispatch.\n\n## What the orchestrator does when the Workflow backend is active\n\nBefore spawning executor agents for the current wave (execute-phase.md step 3),\nresolve the dispatch backend through the single composed CLI seam:\n\n```bash\ngsd-tools claude-orchestration resolve-wave-dispatch \\\n --waves \"$WAVE_MANIFEST_PATH\" --run-id \"$PHASE_RUN_ID\" \\\n --runtime \"$RUNTIME\" \\\n --phase-dir \"$PHASE_DIR\" --raw\n```\n\n`--agent-sdk-version` is no longer passed here (#2590). The router resolves the\ninstalled Agent SDK version itself; see **Agent SDK version** below. The former\n`${AGENT_SDK_VERSION:+--agent-sdk-version \"$AGENT_SDK_VERSION\"}` line was also\n**shell-dependent**: zsh does not word-split unquoted parameter expansions, so it\ncollapsed to a SINGLE argv element there, `argValue()` never matched, and the run\nfailed into `agent_sdk_version_unknown` — indistinguishable from genuinely\nunknown. Pass `--agent-sdk-version ` explicitly only to pin a version.\n\nThis composes `detectWorkflowBackend` (the gate ladder above) with\n`emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure\nfunction backing it is `resolveWaveDispatch` in\n`gsd-core/bin/lib/claude-orchestration.cjs`. Response shape:\n`{ backend: 'inline'|'workflow', reason, script?, summary? }`.\n\n### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`)\n\nThese are NOT pre-existing execute-phase.md variables — the orchestrator builds\nthem at this step, from data it already has in-context from `discover_and_group_plans`\n(the `PLAN_INDEX` JSON) and step 2.5 (the per-plan `USE_WORKTREES_FOR_PLAN` decision):\n\n1. **`$PHASE_DIR`** — reuse `{phase_dir}` from the `INIT` bundle (already loaded\n in the `initialize` step). No new value needed.\n\n2. **`$PHASE_RUN_ID`** — a stable identifier for THIS phase-execution attempt, so\n `resumeFromRunId` can resume an interrupted run without re-dispatching plans\n the Workflow tool already completed. Construct it deterministically —\n `execute-{phase_number}-{phase_slug}` — from `INIT`'s `phase_number`/`phase_slug`\n (both are already validated identifiers used elsewhere in this workflow, so\n they satisfy `emitWorkflowScript`'s `isScriptableIdentifier` check). Do NOT\n mint a new random id per wave — the SAME `$PHASE_RUN_ID` is reused for every\n wave in the phase so the Workflow tool can correctly track cross-wave resume\n state.\n\n3. **`$WAVE_MANIFEST_PATH`** — a fresh temp file for THIS wave's manifest (one\n wave = one `waves` array with a single entry, matching the wave-by-wave\n dispatch loop; do not batch multiple waves into one manifest — waves are\n dispatched in wave order, not all at once):\n\n ```bash\n WAVE_MANIFEST_PATH=$(mktemp \"${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX\") && mv \"$WAVE_MANIFEST_PATH\" \"$WAVE_MANIFEST_PATH.json\" && WAVE_MANIFEST_PATH=\"$WAVE_MANIFEST_PATH.json\"\n ```\n\n Then **use the Write tool** (not a bash/jq pipeline — the orchestrator already\n has every field parsed in-context) to write the manifest JSON to\n `$WAVE_MANIFEST_PATH`:\n\n ```json\n {\n \"waves\": [\n {\n \"id\": \"wave-{N}\",\n \"plans\": [\n {\n \"id\": \"{plan_id}\",\n \"brief\": \"{the SAME ... prompt block step 3 builds for this plan's inline Agent() call}\",\n \"files_modified\": [\"{from PLAN_INDEX.plans[].files_modified for this plan}\"],\n \"use_worktree\": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}\n }\n ]\n }\n ]\n }\n ```\n\n - **`id`** — the plan id from `PLAN_INDEX`, e.g. `\"01-01\"`.\n - **`brief`** — MUST carry the same task content as step 3's inline `Agent()`\n prompt (the ``/``/``/\n `` block, with `{plan_number}`/`{phase_number}`/\n `{phase_name}` substituted) — a short summary here would NOT reproduce\n step 3's behavior and would violate the \"identical artifacts\" contract.\n - **`files_modified`** — copy verbatim from the plan's `PLAN_INDEX` entry.\n - **`use_worktree`** — `true` for every plan UNLESS step 2.5's per-plan\n worktree gate (`execute-phase/steps/per-plan-worktree-gate.md`) set\n `USE_WORKTREES_FOR_PLAN=false` for that plan (submodule-touching plan, or\n project-level `USE_WORKTREES=false`) — in which case pass `false` here so\n `emitWorkflowScript` omits `isolation: \"worktree\"` for that plan (#2772 /\n #2285 finding 1). **Never** hardcode `true` — that would force worktree\n isolation on a plan the inline path explicitly keeps out of worktrees.\n\n4. **`$AGENT_SDK_VERSION`** — no longer built here; the router resolves it.\n\n**Agent SDK version:** the orchestrator has no *bash-computable* way to\nintrospect the live Agent SDK version — but the router runs in Node, so it\nresolves the version itself (#2590), in this order:\n\n1. an explicit `--agent-sdk-version ` (pin a version),\n2. `GSD_AGENT_SDK_VERSION`,\n3. the **installed** `@anthropic-ai/claude-agent-sdk` package version, read from\n its `package.json` on disk by walking `node_modules` up the tree. (Read\n directly rather than via `require.resolve`: the SDK's `exports` map does not\n expose `./package.json`, so `require.resolve` throws\n `ERR_PACKAGE_PATH_NOT_EXPORTED`.)\n\nPreviously nothing computed this at all, so gate 5 returned\n`agent_sdk_version_unknown` on **every** automated run and the Workflow backend\ncould never activate — while `gsd-tools capability state` still reported the\ncapability `active: true`. Fail-closed is preserved: when no version can be\nresolved, gate 5 still declines to `inline`. What changed is that a resolvable\nversion is now actually found, so a genuinely-too-old SDK reports\n`agent_sdk_version_below_floor` — the truthful reason — instead of `unknown`.\n\n**If `backend == \"workflow\"`:** run the emitted `script` via the Workflow tool\nfor THIS wave instead of the per-message `Agent()` loop in step 3. The script\ncomposes the SAME `gsd-executor` agent type the inline path uses, with\nworktree isolation applied PER PLAN from the manifest's `use_worktree` field\n(see `emitWorkflowScript`):\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier.\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n when `use_worktree` is not `false`, or `agent(brief, { agentType: 'gsd-executor' })`\n (no isolation) when it is — so the produced `SUMMARY.md` and commits are\n identical to inline dispatch, INCLUDING the inline path's submodule safety\n gate (#2772 / #2285 finding 1).\n- **`files_modified` overlap → separate sequential stages** — the same overlap\n rule execute-phase already applies inline (step 1 of the wave loop).\n- **`resumeFromRunId`** — **pass `summary.resumeRunId` as the Workflow tool's\n `resumeFromRunId` INPUT when you invoke the tool.** It is a tool parameter,\n not a script function; the script deliberately does not call it (#2590 — doing\n so threw \"resumeFromRunId is not defined\" and rejected the entire script).\n Omitting it from the tool invocation silently regresses phase-resume to a\n no-op: an interrupted phase re-runs completed plans.\n\nThe orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup,\npost-merge gate, tracking update) exactly as it does for inline dispatch — the\nWorkflow backend only replaces HOW agents are spawned for this wave, not what\nhappens after they return.\n\n**If `backend == \"inline\"`** (any gate miss, or `resolve-wave-dispatch` itself\nunavailable/erroring): proceed to step 3's standard per-message `Agent()`\ndispatch — the default, byte-identical-to-today path. `onError: skip` on this\ncontribution means a `resolve-wave-dispatch` command failure is treated exactly\nlike an `inline` result, never as a fatal wave error.\n\n## Fallback contract\n\nDetection is fail-closed end-to-end: capability disabled, non-Claude runtime,\n`execution_backend:\"inline\"`, missing/incapable host descriptor, unknown or\nbelow-floor Agent SDK version, or an `emitWorkflowScript` failure on a malformed\nwave manifest — ANY of these degrades to `backend:\"inline\"` and execute-phase's\nstandard inline dispatch (step 3) runs unmodified. The Workflow backend never\npartially activates; the executor MUST NOT assume parallelism, a shared budget,\nor resume-from-run-id semantics when `backend == \"inline\"`.\n" }, "produces": [], "consumes": [ @@ -3541,7 +3541,7 @@ const byLoopPoint = { "into": "executor", "fragment": { "path": "fragments/execute-wave-pre.md", - "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:pre` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## Why `execute:wave:pre` (not `execute:wave:post`)\n\nThis is a **dispatch-backend selector** — it decides HOW a wave's executor agents\nare spawned. That decision has to be made BEFORE the wave's `Agent()` calls in\n`execute-phase.md` step 3, not after the wave has already finished (#2285). The\ncapability previously registered at `execute:wave:post`, which fires only after\nworktree merge/post-merge tests/tracking updates — by then the wave was already\ndispatched inline, so the contribution was structurally unable to change how\ndispatch happened. This fragment is injected at the point that actually precedes\ndispatch.\n\n## What the orchestrator does when the Workflow backend is active\n\nBefore spawning executor agents for the current wave (execute-phase.md step 3),\nresolve the dispatch backend through the single composed CLI seam:\n\n```bash\ngsd-tools claude-orchestration resolve-wave-dispatch \\\n --waves \"$WAVE_MANIFEST_PATH\" --run-id \"$PHASE_RUN_ID\" \\\n --runtime \"$RUNTIME\" \\\n ${AGENT_SDK_VERSION:+--agent-sdk-version \"$AGENT_SDK_VERSION\"} \\\n --phase-dir \"$PHASE_DIR\" --raw\n```\n\nThis composes `detectWorkflowBackend` (the gate ladder above) with\n`emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure\nfunction backing it is `resolveWaveDispatch` in\n`gsd-core/bin/lib/claude-orchestration.cjs`. Response shape:\n`{ backend: 'inline'|'workflow', reason, script?, summary? }`.\n\n### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`, `$AGENT_SDK_VERSION`)\n\nThese are NOT pre-existing execute-phase.md variables — the orchestrator builds\nthem at this step, from data it already has in-context from `discover_and_group_plans`\n(the `PLAN_INDEX` JSON) and step 2.5 (the per-plan `USE_WORKTREES_FOR_PLAN` decision):\n\n1. **`$PHASE_DIR`** — reuse `{phase_dir}` from the `INIT` bundle (already loaded\n in the `initialize` step). No new value needed.\n\n2. **`$PHASE_RUN_ID`** — a stable identifier for THIS phase-execution attempt, so\n `resumeFromRunId` can resume an interrupted run without re-dispatching plans\n the Workflow tool already completed. Construct it deterministically —\n `execute-{phase_number}-{phase_slug}` — from `INIT`'s `phase_number`/`phase_slug`\n (both are already validated identifiers used elsewhere in this workflow, so\n they satisfy `emitWorkflowScript`'s `isScriptableIdentifier` check). Do NOT\n mint a new random id per wave — the SAME `$PHASE_RUN_ID` is reused for every\n wave in the phase so the Workflow tool can correctly track cross-wave resume\n state.\n\n3. **`$WAVE_MANIFEST_PATH`** — a fresh temp file for THIS wave's manifest (one\n wave = one `waves` array with a single entry, matching the wave-by-wave\n dispatch loop; do not batch multiple waves into one manifest — waves are\n dispatched in wave order, not all at once):\n\n ```bash\n WAVE_MANIFEST_PATH=$(mktemp \"${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX\") && mv \"$WAVE_MANIFEST_PATH\" \"$WAVE_MANIFEST_PATH.json\" && WAVE_MANIFEST_PATH=\"$WAVE_MANIFEST_PATH.json\"\n ```\n\n Then **use the Write tool** (not a bash/jq pipeline — the orchestrator already\n has every field parsed in-context) to write the manifest JSON to\n `$WAVE_MANIFEST_PATH`:\n\n ```json\n {\n \"waves\": [\n {\n \"id\": \"wave-{N}\",\n \"plans\": [\n {\n \"id\": \"{plan_id}\",\n \"brief\": \"{the SAME ... prompt block step 3 builds for this plan's inline Agent() call}\",\n \"files_modified\": [\"{from PLAN_INDEX.plans[].files_modified for this plan}\"],\n \"use_worktree\": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}\n }\n ]\n }\n ]\n }\n ```\n\n - **`id`** — the plan id from `PLAN_INDEX`, e.g. `\"01-01\"`.\n - **`brief`** — MUST carry the same task content as step 3's inline `Agent()`\n prompt (the ``/``/``/\n `` block, with `{plan_number}`/`{phase_number}`/\n `{phase_name}` substituted) — a short summary here would NOT reproduce\n step 3's behavior and would violate the \"identical artifacts\" contract.\n - **`files_modified`** — copy verbatim from the plan's `PLAN_INDEX` entry.\n - **`use_worktree`** — `true` for every plan UNLESS step 2.5's per-plan\n worktree gate (`execute-phase/steps/per-plan-worktree-gate.md`) set\n `USE_WORKTREES_FOR_PLAN=false` for that plan (submodule-touching plan, or\n project-level `USE_WORKTREES=false`) — in which case pass `false` here so\n `emitWorkflowScript` omits `isolation: \"worktree\"` for that plan (#2772 /\n #2285 finding 1). **Never** hardcode `true` — that would force worktree\n isolation on a plan the inline path explicitly keeps out of worktrees.\n\n4. **`$AGENT_SDK_VERSION`** — see below; OMIT when unknown (fails closed).\n\n**Agent SDK version:** the orchestrator has no scriptable (bash-computable) way\nto introspect the live Agent SDK version. When it can determine the version\n(e.g. from a host-exposed value it can read directly), pass\n`--agent-sdk-version`. When it cannot, OMIT the flag — `resolveWaveDispatch`'s\ngate 5 (`agent_sdk_version_unknown`) then fails closed to `inline` by design;\nthis is not a bug, it is the same fail-closed posture documented above applied\nto a real absence of information.\n\n**If `backend == \"workflow\"`:** run the emitted `script` via the Workflow tool\nfor THIS wave instead of the per-message `Agent()` loop in step 3. The script\ncomposes the SAME `gsd-executor` agent type the inline path uses, with\nworktree isolation applied PER PLAN from the manifest's `use_worktree` field\n(see `emitWorkflowScript`):\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier.\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n when `use_worktree` is not `false`, or `agent(brief, { agentType: 'gsd-executor' })`\n (no isolation) when it is — so the produced `SUMMARY.md` and commits are\n identical to inline dispatch, INCLUDING the inline path's submodule safety\n gate (#2772 / #2285 finding 1).\n- **`files_modified` overlap → separate sequential stages** — the same overlap\n rule execute-phase already applies inline (step 1 of the wave loop).\n- **`resumeFromRunId`** — wired to the phase run id, so an interrupted phase\n resumes without re-running completed plans.\n\nThe orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup,\npost-merge gate, tracking update) exactly as it does for inline dispatch — the\nWorkflow backend only replaces HOW agents are spawned for this wave, not what\nhappens after they return.\n\n**If `backend == \"inline\"`** (any gate miss, or `resolve-wave-dispatch` itself\nunavailable/erroring): proceed to step 3's standard per-message `Agent()`\ndispatch — the default, byte-identical-to-today path. `onError: skip` on this\ncontribution means a `resolve-wave-dispatch` command failure is treated exactly\nlike an `inline` result, never as a fatal wave error.\n\n## Fallback contract\n\nDetection is fail-closed end-to-end: capability disabled, non-Claude runtime,\n`execution_backend:\"inline\"`, missing/incapable host descriptor, unknown or\nbelow-floor Agent SDK version, or an `emitWorkflowScript` failure on a malformed\nwave manifest — ANY of these degrades to `backend:\"inline\"` and execute-phase's\nstandard inline dispatch (step 3) runs unmodified. The Workflow backend never\npartially activates; the executor MUST NOT assume parallelism, a shared budget,\nor resume-from-run-id semantics when `backend == \"inline\"`.\n" + "inline": "# Claude orchestration — Workflow execution backend (BETA)\n\n> Injected at `execute:wave:pre` `into: executor` only when\n> `claude_orchestration.enabled` is true. Default-off; `onError: skip`.\n\n## When this contribution is active\n\nThe Claude orchestration capability is **default-off and BETA**. It activates only\nwhen ALL of the following hold:\n\n1. `claude_orchestration.enabled` is `true` in `.planning/config.json`, AND\n2. the active runtime is **Claude Code** (the Workflow tool is Claude / Agent\n SDK-specific), AND\n3. `claude_orchestration.execution_backend` resolves to `workflow` — either\n explicitly, or via `auto` — **and** the Agent SDK version is\n `>= claude_orchestration.min_agent_sdk_version` (default `0.3.149`). The SDK\n floor applies in both `auto` and `workflow` modes (fail-closed: a pre-release\n or older SDK never activates the preview backend).\n\nDetection is fail-closed: any miss degrades to **inline, manual, one-agent-per-\nmessage dispatch** — exactly today's behaviour. On a non-Claude runtime this\ncontribution is a no-op.\n\n## Why `execute:wave:pre` (not `execute:wave:post`)\n\nThis is a **dispatch-backend selector** — it decides HOW a wave's executor agents\nare spawned. That decision has to be made BEFORE the wave's `Agent()` calls in\n`execute-phase.md` step 3, not after the wave has already finished (#2285). The\ncapability previously registered at `execute:wave:post`, which fires only after\nworktree merge/post-merge tests/tracking updates — by then the wave was already\ndispatched inline, so the contribution was structurally unable to change how\ndispatch happened. This fragment is injected at the point that actually precedes\ndispatch.\n\n## What the orchestrator does when the Workflow backend is active\n\nBefore spawning executor agents for the current wave (execute-phase.md step 3),\nresolve the dispatch backend through the single composed CLI seam:\n\n```bash\ngsd-tools claude-orchestration resolve-wave-dispatch \\\n --waves \"$WAVE_MANIFEST_PATH\" --run-id \"$PHASE_RUN_ID\" \\\n --runtime \"$RUNTIME\" \\\n --phase-dir \"$PHASE_DIR\" --raw\n```\n\n`--agent-sdk-version` is no longer passed here (#2590). The router resolves the\ninstalled Agent SDK version itself; see **Agent SDK version** below. The former\n`${AGENT_SDK_VERSION:+--agent-sdk-version \"$AGENT_SDK_VERSION\"}` line was also\n**shell-dependent**: zsh does not word-split unquoted parameter expansions, so it\ncollapsed to a SINGLE argv element there, `argValue()` never matched, and the run\nfailed into `agent_sdk_version_unknown` — indistinguishable from genuinely\nunknown. Pass `--agent-sdk-version ` explicitly only to pin a version.\n\nThis composes `detectWorkflowBackend` (the gate ladder above) with\n`emitWorkflowScript` (the wave→plan mapping below) in ONE call — the pure\nfunction backing it is `resolveWaveDispatch` in\n`gsd-core/bin/lib/claude-orchestration.cjs`. Response shape:\n`{ backend: 'inline'|'workflow', reason, script?, summary? }`.\n\n### Manifest construction (`$WAVE_MANIFEST_PATH`, `$PHASE_RUN_ID`, `$PHASE_DIR`)\n\nThese are NOT pre-existing execute-phase.md variables — the orchestrator builds\nthem at this step, from data it already has in-context from `discover_and_group_plans`\n(the `PLAN_INDEX` JSON) and step 2.5 (the per-plan `USE_WORKTREES_FOR_PLAN` decision):\n\n1. **`$PHASE_DIR`** — reuse `{phase_dir}` from the `INIT` bundle (already loaded\n in the `initialize` step). No new value needed.\n\n2. **`$PHASE_RUN_ID`** — a stable identifier for THIS phase-execution attempt, so\n `resumeFromRunId` can resume an interrupted run without re-dispatching plans\n the Workflow tool already completed. Construct it deterministically —\n `execute-{phase_number}-{phase_slug}` — from `INIT`'s `phase_number`/`phase_slug`\n (both are already validated identifiers used elsewhere in this workflow, so\n they satisfy `emitWorkflowScript`'s `isScriptableIdentifier` check). Do NOT\n mint a new random id per wave — the SAME `$PHASE_RUN_ID` is reused for every\n wave in the phase so the Workflow tool can correctly track cross-wave resume\n state.\n\n3. **`$WAVE_MANIFEST_PATH`** — a fresh temp file for THIS wave's manifest (one\n wave = one `waves` array with a single entry, matching the wave-by-wave\n dispatch loop; do not batch multiple waves into one manifest — waves are\n dispatched in wave order, not all at once):\n\n ```bash\n WAVE_MANIFEST_PATH=$(mktemp \"${TMPDIR:-/tmp}/gsd-wave-dispatch-XXXXXX\") && mv \"$WAVE_MANIFEST_PATH\" \"$WAVE_MANIFEST_PATH.json\" && WAVE_MANIFEST_PATH=\"$WAVE_MANIFEST_PATH.json\"\n ```\n\n Then **use the Write tool** (not a bash/jq pipeline — the orchestrator already\n has every field parsed in-context) to write the manifest JSON to\n `$WAVE_MANIFEST_PATH`:\n\n ```json\n {\n \"waves\": [\n {\n \"id\": \"wave-{N}\",\n \"plans\": [\n {\n \"id\": \"{plan_id}\",\n \"brief\": \"{the SAME ... prompt block step 3 builds for this plan's inline Agent() call}\",\n \"files_modified\": [\"{from PLAN_INDEX.plans[].files_modified for this plan}\"],\n \"use_worktree\": {true unless step 2.5 set USE_WORKTREES_FOR_PLAN=false for this plan}\n }\n ]\n }\n ]\n }\n ```\n\n - **`id`** — the plan id from `PLAN_INDEX`, e.g. `\"01-01\"`.\n - **`brief`** — MUST carry the same task content as step 3's inline `Agent()`\n prompt (the ``/``/``/\n `` block, with `{plan_number}`/`{phase_number}`/\n `{phase_name}` substituted) — a short summary here would NOT reproduce\n step 3's behavior and would violate the \"identical artifacts\" contract.\n - **`files_modified`** — copy verbatim from the plan's `PLAN_INDEX` entry.\n - **`use_worktree`** — `true` for every plan UNLESS step 2.5's per-plan\n worktree gate (`execute-phase/steps/per-plan-worktree-gate.md`) set\n `USE_WORKTREES_FOR_PLAN=false` for that plan (submodule-touching plan, or\n project-level `USE_WORKTREES=false`) — in which case pass `false` here so\n `emitWorkflowScript` omits `isolation: \"worktree\"` for that plan (#2772 /\n #2285 finding 1). **Never** hardcode `true` — that would force worktree\n isolation on a plan the inline path explicitly keeps out of worktrees.\n\n4. **`$AGENT_SDK_VERSION`** — no longer built here; the router resolves it.\n\n**Agent SDK version:** the orchestrator has no *bash-computable* way to\nintrospect the live Agent SDK version — but the router runs in Node, so it\nresolves the version itself (#2590), in this order:\n\n1. an explicit `--agent-sdk-version ` (pin a version),\n2. `GSD_AGENT_SDK_VERSION`,\n3. the **installed** `@anthropic-ai/claude-agent-sdk` package version, read from\n its `package.json` on disk by walking `node_modules` up the tree. (Read\n directly rather than via `require.resolve`: the SDK's `exports` map does not\n expose `./package.json`, so `require.resolve` throws\n `ERR_PACKAGE_PATH_NOT_EXPORTED`.)\n\nPreviously nothing computed this at all, so gate 5 returned\n`agent_sdk_version_unknown` on **every** automated run and the Workflow backend\ncould never activate — while `gsd-tools capability state` still reported the\ncapability `active: true`. Fail-closed is preserved: when no version can be\nresolved, gate 5 still declines to `inline`. What changed is that a resolvable\nversion is now actually found, so a genuinely-too-old SDK reports\n`agent_sdk_version_below_floor` — the truthful reason — instead of `unknown`.\n\n**If `backend == \"workflow\"`:** run the emitted `script` via the Workflow tool\nfor THIS wave instead of the per-message `Agent()` loop in step 3. The script\ncomposes the SAME `gsd-executor` agent type the inline path uses, with\nworktree isolation applied PER PLAN from the manifest's `use_worktree` field\n(see `emitWorkflowScript`):\n\n- **waves → one or more sequential `parallel()` barriers** — each wave is a\n barrier group; when plans within a wave share `files_modified`, they are split\n into separate sequential stages within that wave's barrier.\n- **plans → `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })`**\n when `use_worktree` is not `false`, or `agent(brief, { agentType: 'gsd-executor' })`\n (no isolation) when it is — so the produced `SUMMARY.md` and commits are\n identical to inline dispatch, INCLUDING the inline path's submodule safety\n gate (#2772 / #2285 finding 1).\n- **`files_modified` overlap → separate sequential stages** — the same overlap\n rule execute-phase already applies inline (step 1 of the wave loop).\n- **`resumeFromRunId`** — **pass `summary.resumeRunId` as the Workflow tool's\n `resumeFromRunId` INPUT when you invoke the tool.** It is a tool parameter,\n not a script function; the script deliberately does not call it (#2590 — doing\n so threw \"resumeFromRunId is not defined\" and rejected the entire script).\n Omitting it from the tool invocation silently regresses phase-resume to a\n no-op: an interrupted phase re-runs completed plans.\n\nThe orchestrator still runs steps 4–5.8 (wait for completion, worktree cleanup,\npost-merge gate, tracking update) exactly as it does for inline dispatch — the\nWorkflow backend only replaces HOW agents are spawned for this wave, not what\nhappens after they return.\n\n**If `backend == \"inline\"`** (any gate miss, or `resolve-wave-dispatch` itself\nunavailable/erroring): proceed to step 3's standard per-message `Agent()`\ndispatch — the default, byte-identical-to-today path. `onError: skip` on this\ncontribution means a `resolve-wave-dispatch` command failure is treated exactly\nlike an `inline` result, never as a fatal wave error.\n\n## Fallback contract\n\nDetection is fail-closed end-to-end: capability disabled, non-Claude runtime,\n`execution_backend:\"inline\"`, missing/incapable host descriptor, unknown or\nbelow-floor Agent SDK version, or an `emitWorkflowScript` failure on a malformed\nwave manifest — ANY of these degrades to `backend:\"inline\"` and execute-phase's\nstandard inline dispatch (step 3) runs unmodified. The Workflow backend never\npartially activates; the executor MUST NOT assume parallelism, a shared budget,\nor resume-from-run-id semantics when `backend == \"inline\"`.\n" }, "produces": [], "consumes": [ diff --git a/gsd-core/bin/lib/claude-orchestration-command-router.cjs b/gsd-core/bin/lib/claude-orchestration-command-router.cjs index cf54eedae..f6ff36fb4 100644 --- a/gsd-core/bin/lib/claude-orchestration-command-router.cjs +++ b/gsd-core/bin/lib/claude-orchestration-command-router.cjs @@ -14,8 +14,11 @@ * * Subcommands: * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] - * Resolves whether the Workflow backend should activate. `--runtime` - * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * Resolves whether the Workflow backend should activate. Both flags are + * OPTIONAL (#2590): `--runtime` falls back to the canonical + * `GSD_RUNTIME > config.runtime > 'claude'` chain, and + * `--agent-sdk-version` to `GSD_AGENT_SDK_VERSION` then the installed + * @anthropic-ai/claude-agent-sdk version. Reads the * `claude_orchestration.*` keys from .planning/config.json. Emits * { available, backend, reason }. * @@ -48,6 +51,8 @@ const io = require("./io.cjs"); const core = require("./claude-orchestration.cjs"); // eslint-disable-next-line @typescript-eslint/no-require-imports const configLoader = require("./config-loader.cjs"); +// eslint-disable-next-line @typescript-eslint/no-require-imports +const runtimeSlash = require("./runtime-slash.cjs"); const { output } = io; const { detectWorkflowBackend, emitWorkflowScript, resolveWaveDispatch } = core; const CAPABLE_HOST = { dispatch: { nested: true, background: true } }; @@ -86,14 +91,69 @@ function resolveFlatClaudeOrchestrationConfig(cwd) { } return flatConfig; } +/** + * Resolve the installed Agent SDK version (#2590). + * + * The `execute:wave:pre` fragment claimed the orchestrator "has no scriptable + * way to introspect the live Agent SDK version" and told callers to omit the + * flag — so gate 5 returned `agent_sdk_version_unknown` on every automated run + * and the Workflow backend never activated, while `capability state` still + * reported it `active: true`. That claim is true for BASH, but this router runs + * in Node: the installed package's own package.json is authoritative and + * requires no flag at all. + * + * Resolution is side-effect-free and fails closed to undefined (gate 5 then + * declines, exactly as before) rather than guessing a version. + */ +const AGENT_SDK_PKG = node_path_1.default.join('@anthropic-ai', 'claude-agent-sdk', 'package.json'); +function resolveInstalledAgentSdkVersion(cwd) { + // Walk node_modules up the tree by hand rather than require.resolve: the SDK's + // `exports` map does not expose './package.json', so require.resolve throws + // ERR_PACKAGE_PATH_NOT_EXPORTED. Reading the file directly is exports-map + // independent and cannot execute package code. + for (const start of [cwd, __dirname]) { + let dir; + try { + dir = node_path_1.default.resolve(start); + } + catch { + continue; + } + for (;;) { + try { + const pkgPath = node_path_1.default.join(dir, 'node_modules', AGENT_SDK_PKG); + if (node_fs_1.default.existsSync(pkgPath)) { + const parsed = JSON.parse(node_fs_1.default.readFileSync(pkgPath, 'utf8')); + if (typeof parsed.version === 'string' && parsed.version.length > 0) + return parsed.version; + } + } + catch { /* unreadable/malformed — keep walking */ } + const parent = node_path_1.default.dirname(dir); + if (parent === dir) + break; + dir = parent; + } + } + return undefined; +} /** * Resolve `--runtime`/`--agent-sdk-version`/`--no-nested-dispatch` into the * `{ runtimeId, hostIntegration, agentSdkVersion }` triple both `detect-backend` * and `resolve-wave-dispatch` pass to the pure detection seam. */ -function resolveDetectionArgs(args) { - const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; - const agentSdkVersion = argValue(args, '--agent-sdk-version'); +function resolveDetectionArgs(args, cwd) { + // #2590: the old fallback chain was `--runtime > GSD_RUNTIME > 'unknown'`, + // diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'` used + // by runtime-slash.resolveRuntime — so ANY manual invocation without + // --runtime reported `runtime_not_claude` on a perfectly ordinary Claude + // project. Delegate to the canonical resolver instead of re-deriving it. + const runtimeId = argValue(args, '--runtime') || runtimeSlash.resolveRuntime(cwd || null); + // Explicit flag wins (lets a caller pin a version); then the environment; + // then the actually-installed SDK. + const agentSdkVersion = argValue(args, '--agent-sdk-version') + || process.env['GSD_AGENT_SDK_VERSION'] + || resolveInstalledAgentSdkVersion(cwd || process.cwd()); const noNested = args.includes('--no-nested-dispatch'); const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; return { runtimeId, hostIntegration, agentSdkVersion }; @@ -128,7 +188,7 @@ function readWavesManifest(wavesPath, error) { * SDK version come from flags (the orchestrator already knows these) or env. */ function cmdDetectBackend(args, cwd, raw) { - const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args); + const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args, cwd); const flatConfig = resolveFlatClaudeOrchestrationConfig(cwd); const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); output(result, raw); @@ -188,7 +248,7 @@ function cmdResolveWaveDispatch(args, cwd, raw, error) { const read = readWavesManifest(wavesPath, (msg) => error('resolve-wave-dispatch: ' + msg)); if (!read.ok) return; // read/parse failure — error() already surfaced it loudly above - const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args); + const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args, cwd); const flatConfig = resolveFlatClaudeOrchestrationConfig(cwd); const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; const budget = (typeof budgetTokens === 'number' && !Number.isNaN(budgetTokens)) ? budgetTokens : undefined; diff --git a/gsd-core/bin/lib/claude-orchestration.cjs b/gsd-core/bin/lib/claude-orchestration.cjs index 95264bdfd..d2f730722 100644 --- a/gsd-core/bin/lib/claude-orchestration.cjs +++ b/gsd-core/bin/lib/claude-orchestration.cjs @@ -314,6 +314,12 @@ function emitWorkflowScript(input) { if (!Array.isArray(waves) || waves.length === 0) { return { ok: false, reason: 'waves must be a non-empty array' }; } + // Wave ids must be unique ACROSS waves, not just plan ids within one (#2590). + // Each wave emits a `phase("Wave ")` call plus a matching meta.phases + // entry, and the Workflow tool matches phase titles by exact string — two + // waves sharing an id would collapse into one progress group and misattribute + // every agent in the second wave to the first. + const seenWaveIds = new Set(); for (let i = 0; i < waves.length; i++) { const w = waves[i]; if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { @@ -322,6 +328,10 @@ function emitWorkflowScript(input) { if (!isScriptableIdentifier(w.id)) { return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; } + if (seenWaveIds.has(w.id)) { + return { ok: false, reason: 'duplicate wave id "' + w.id + '" — wave ids must be unique (phase titles must map 1:1)' }; + } + seenWaveIds.add(w.id); if (!Array.isArray(w.plans) || w.plans.length === 0) { return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; } @@ -352,15 +362,43 @@ function emitWorkflowScript(input) { ? Math.floor(input.budgetTokens) : null; const lines = []; + // `export const meta = {…}` MUST be the first statement in the script — the + // Workflow tool rejects the whole script otherwise (#2590). Leading comments + // are not statements, but the meta block is emitted first regardless so the + // contract holds under the strictest reading of "first statement". + // + // meta.phases must be a PURE LITERAL (no variables, calls, spreads, or + // template interpolation), and its titles are matched EXACTLY against the + // phase() calls emitted below. + lines.push('export const meta = {'); + lines.push(' name: ' + quoteString('gsd-execute-' + runId) + ','); + lines.push(' description: ' + quoteString('GSD wave dispatch for ' + phaseDir) + ','); + lines.push(' phases: ['); + for (const w of waves) { + lines.push(' { title: ' + quoteString('Wave ' + w.id) + ', detail: ' + + quoteString(w.plans.length + ' plan(s)') + ' },'); + } + lines.push(' ],'); + lines.push('}'); + lines.push(''); lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); lines.push('// phase: ' + phaseDir); lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); lines.push('// Composes the SAME gsd-executor agent as the inline path, so artifacts (SUMMARY.md)'); lines.push('// and commits are produced identically. Worktree isolation is per-plan (use_worktree)'); lines.push('// and mirrors execute-phase.md step 2.5\'s submodule gate exactly (#2772 / #2285).'); - lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + lines.push('//'); + // resumeFromRunId is a Workflow TOOL INPUT parameter, not a script function — + // calling it threw "resumeFromRunId is not defined" (#2590). The run id is + // carried in summary.resumeRunId for the caller to pass as that input. + lines.push('// resume: pass ' + quoteString(runId) + ' as the Workflow tool\'s resumeFromRunId input'); + lines.push('// (it is a tool parameter, NOT a script function).'); if (budgetTokens !== null) { - lines.push('budget(' + budgetTokens + ')'); + // `budget` is a read-only object ({ total, spent(), remaining() }) supplied + // by the caller's token directive — a script cannot SET it, and `budget(n)` + // threw "budget is not a function" (#2590). Recorded as intent only. + lines.push('// budget: ' + budgetTokens + ' output tokens intended for this run; `budget` is'); + lines.push('// read-only in a Workflow script — set it via the caller\'s token directive.'); } lines.push(''); const stagesByWave = []; @@ -371,6 +409,8 @@ function emitWorkflowScript(input) { stagesByWave.push(stages); totalPlans += wave.plans.length; lines.push('// Wave ' + wave.id); + // Title must match this wave's meta.phases entry EXACTLY. + lines.push('phase(' + quoteString('Wave ' + wave.id) + ')'); for (let si = 0; si < stages.length; si++) { const stagePlanIds = stages[si]; // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). @@ -378,22 +418,15 @@ function emitWorkflowScript(input) { if (stages.length > 1) { lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); } - if (stagePlans.length === 1) { - const p = stagePlans[0]; - lines.push('parallel('); - lines.push(' agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + ')'); - lines.push(')'); - } - else { - lines.push('parallel('); - for (const p of stagePlans) { - lines.push(' agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + '),'); - } - // Replace trailing comma on the last agent line with nothing. - const lastIdx = lines.length - 1; - lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); - lines.push(')'); + // parallel() takes an ARRAY OF THUNKS — `parallel(agent(…), agent(…))` + // threw "parallel() expects an array of functions" (#2590). Passing + // agent() results directly would also start every agent eagerly, before + // parallel() could bound concurrency. + lines.push('await parallel(['); + for (const p of stagePlans) { + lines.push(' () => agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + '),'); } + lines.push('])'); } if (wi < waves.length - 1) lines.push(''); diff --git a/src/claude-orchestration-command-router.cts b/src/claude-orchestration-command-router.cts index 02a0e576d..bd7964efe 100644 --- a/src/claude-orchestration-command-router.cts +++ b/src/claude-orchestration-command-router.cts @@ -13,8 +13,11 @@ * * Subcommands: * detect-backend [--runtime ] [--agent-sdk-version ] [--no-nested-dispatch] - * Resolves whether the Workflow backend should activate. `--runtime` - * defaults to the GSD_RUNTIME env var (or 'unknown'). Reads the + * Resolves whether the Workflow backend should activate. Both flags are + * OPTIONAL (#2590): `--runtime` falls back to the canonical + * `GSD_RUNTIME > config.runtime > 'claude'` chain, and + * `--agent-sdk-version` to `GSD_AGENT_SDK_VERSION` then the installed + * @anthropic-ai/claude-agent-sdk version. Reads the * `claude_orchestration.*` keys from .planning/config.json. Emits * { available, backend, reason }. * @@ -45,6 +48,8 @@ import io = require('./io.cjs'); import core = require('./claude-orchestration.cjs'); // eslint-disable-next-line @typescript-eslint/no-require-imports import configLoader = require('./config-loader.cjs'); +// eslint-disable-next-line @typescript-eslint/no-require-imports +import runtimeSlash = require('./runtime-slash.cjs'); const { output } = io; const { detectWorkflowBackend, emitWorkflowScript, resolveWaveDispatch } = core; @@ -98,14 +103,62 @@ function resolveFlatClaudeOrchestrationConfig(cwd: string): Record 0) return parsed.version; + } + } catch { /* unreadable/malformed — keep walking */ } + const parent = path.dirname(dir); + if (parent === dir) break; + dir = parent; + } + } + return undefined; +} + /** * Resolve `--runtime`/`--agent-sdk-version`/`--no-nested-dispatch` into the * `{ runtimeId, hostIntegration, agentSdkVersion }` triple both `detect-backend` * and `resolve-wave-dispatch` pass to the pure detection seam. */ -function resolveDetectionArgs(args: string[]): { runtimeId: string; hostIntegration: { dispatch: { nested: boolean; background: boolean } }; agentSdkVersion: string | undefined } { - const runtimeId = argValue(args, '--runtime') || process.env['GSD_RUNTIME'] || 'unknown'; - const agentSdkVersion = argValue(args, '--agent-sdk-version'); +function resolveDetectionArgs(args: string[], cwd?: string): { runtimeId: string; hostIntegration: { dispatch: { nested: boolean; background: boolean } }; agentSdkVersion: string | undefined } { + // #2590: the old fallback chain was `--runtime > GSD_RUNTIME > 'unknown'`, + // diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'` used + // by runtime-slash.resolveRuntime — so ANY manual invocation without + // --runtime reported `runtime_not_claude` on a perfectly ordinary Claude + // project. Delegate to the canonical resolver instead of re-deriving it. + const runtimeId = argValue(args, '--runtime') || runtimeSlash.resolveRuntime(cwd || null); + // Explicit flag wins (lets a caller pin a version); then the environment; + // then the actually-installed SDK. + const agentSdkVersion = argValue(args, '--agent-sdk-version') + || process.env['GSD_AGENT_SDK_VERSION'] + || resolveInstalledAgentSdkVersion(cwd || process.cwd()); const noNested = args.includes('--no-nested-dispatch'); const hostIntegration = noNested ? { dispatch: { nested: false, background: true } } : CAPABLE_HOST; return { runtimeId, hostIntegration, agentSdkVersion }; @@ -146,7 +199,7 @@ function readWavesManifest(wavesPath: string, error: (msg: string, reason?: stri * SDK version come from flags (the orchestrator already knows these) or env. */ function cmdDetectBackend(args: string[], cwd: string, raw: boolean): void { - const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args); + const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args, cwd); const flatConfig = resolveFlatClaudeOrchestrationConfig(cwd); const result = detectWorkflowBackend({ runtimeId, hostIntegration, config: flatConfig, agentSdkVersion }); output(result, raw); @@ -214,7 +267,7 @@ function cmdResolveWaveDispatch(args: string[], cwd: string, raw: boolean, error const read = readWavesManifest(wavesPath, (msg) => error('resolve-wave-dispatch: ' + msg)); if (!read.ok) return; // read/parse failure — error() already surfaced it loudly above - const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args); + const { runtimeId, hostIntegration, agentSdkVersion } = resolveDetectionArgs(args, cwd); const flatConfig = resolveFlatClaudeOrchestrationConfig(cwd); const budgetTokens = budgetRaw !== undefined ? parseInt(budgetRaw, 10) : undefined; diff --git a/src/claude-orchestration.cts b/src/claude-orchestration.cts index a8a0cc914..ef0713e0b 100644 --- a/src/claude-orchestration.cts +++ b/src/claude-orchestration.cts @@ -402,6 +402,12 @@ function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitE if (!Array.isArray(waves) || waves.length === 0) { return { ok: false, reason: 'waves must be a non-empty array' }; } + // Wave ids must be unique ACROSS waves, not just plan ids within one (#2590). + // Each wave emits a `phase("Wave ")` call plus a matching meta.phases + // entry, and the Workflow tool matches phase titles by exact string — two + // waves sharing an id would collapse into one progress group and misattribute + // every agent in the second wave to the first. + const seenWaveIds = new Set(); for (let i = 0; i < waves.length; i++) { const w = waves[i]; if (w === null || typeof w !== 'object' || typeof w.id !== 'string') { @@ -410,6 +416,10 @@ function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitE if (!isScriptableIdentifier(w.id)) { return { ok: false, reason: 'waves[' + i + '].id must not contain newlines/quotes/backslash/control chars' }; } + if (seenWaveIds.has(w.id)) { + return { ok: false, reason: 'duplicate wave id "' + w.id + '" — wave ids must be unique (phase titles must map 1:1)' }; + } + seenWaveIds.add(w.id); if (!Array.isArray(w.plans) || w.plans.length === 0) { return { ok: false, reason: 'waves[' + i + '] must have a non-empty plans array' }; } @@ -442,15 +452,43 @@ function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitE : null; const lines: string[] = []; + // `export const meta = {…}` MUST be the first statement in the script — the + // Workflow tool rejects the whole script otherwise (#2590). Leading comments + // are not statements, but the meta block is emitted first regardless so the + // contract holds under the strictest reading of "first statement". + // + // meta.phases must be a PURE LITERAL (no variables, calls, spreads, or + // template interpolation), and its titles are matched EXACTLY against the + // phase() calls emitted below. + lines.push('export const meta = {'); + lines.push(' name: ' + quoteString('gsd-execute-' + runId) + ','); + lines.push(' description: ' + quoteString('GSD wave dispatch for ' + phaseDir) + ','); + lines.push(' phases: ['); + for (const w of waves) { + lines.push(' { title: ' + quoteString('Wave ' + w.id) + ', detail: ' + + quoteString(w.plans.length + ' plan(s)') + ' },'); + } + lines.push(' ],'); + lines.push('}'); + lines.push(''); lines.push('// GSD Workflow script — generated by the claude-orchestration capability (#1143)'); lines.push('// phase: ' + phaseDir); lines.push('// BETA: preview-grade; on any failure the orchestrator falls back to inline dispatch.'); lines.push('// Composes the SAME gsd-executor agent as the inline path, so artifacts (SUMMARY.md)'); lines.push('// and commits are produced identically. Worktree isolation is per-plan (use_worktree)'); lines.push('// and mirrors execute-phase.md step 2.5\'s submodule gate exactly (#2772 / #2285).'); - lines.push('resumeFromRunId(' + quoteString(runId) + ')'); + lines.push('//'); + // resumeFromRunId is a Workflow TOOL INPUT parameter, not a script function — + // calling it threw "resumeFromRunId is not defined" (#2590). The run id is + // carried in summary.resumeRunId for the caller to pass as that input. + lines.push('// resume: pass ' + quoteString(runId) + ' as the Workflow tool\'s resumeFromRunId input'); + lines.push('// (it is a tool parameter, NOT a script function).'); if (budgetTokens !== null) { - lines.push('budget(' + budgetTokens + ')'); + // `budget` is a read-only object ({ total, spent(), remaining() }) supplied + // by the caller's token directive — a script cannot SET it, and `budget(n)` + // threw "budget is not a function" (#2590). Recorded as intent only. + lines.push('// budget: ' + budgetTokens + ' output tokens intended for this run; `budget` is'); + lines.push('// read-only in a Workflow script — set it via the caller\'s token directive.'); } lines.push(''); @@ -464,6 +502,8 @@ function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitE totalPlans += wave.plans.length; lines.push('// Wave ' + wave.id); + // Title must match this wave's meta.phases entry EXACTLY. + lines.push('phase(' + quoteString('Wave ' + wave.id) + ')'); for (let si = 0; si < stages.length; si++) { const stagePlanIds = stages[si]; // Resolve back to plan objects for briefs (ids are unique within a wave — validated above). @@ -471,21 +511,15 @@ function emitWorkflowScript(input: EmitInput | null | undefined): EmitOk | EmitE if (stages.length > 1) { lines.push('// Stage ' + si + (si > 0 ? ' (sequential — files_modified overlap)' : '')); } - if (stagePlans.length === 1) { - const p = stagePlans[0]; - lines.push('parallel('); - lines.push(' agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + ')'); - lines.push(')'); - } else { - lines.push('parallel('); - for (const p of stagePlans) { - lines.push(' agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + '),'); - } - // Replace trailing comma on the last agent line with nothing. - const lastIdx = lines.length - 1; - lines[lastIdx] = lines[lastIdx].replace(/,$/, ''); - lines.push(')'); + // parallel() takes an ARRAY OF THUNKS — `parallel(agent(…), agent(…))` + // threw "parallel() expects an array of functions" (#2590). Passing + // agent() results directly would also start every agent eagerly, before + // parallel() could bound concurrency. + lines.push('await parallel(['); + for (const p of stagePlans) { + lines.push(' () => agent(' + quoteString(p.brief) + ', ' + agentOptions(p) + '),'); } + lines.push('])'); } if (wi < waves.length - 1) lines.push(''); } diff --git a/tests/claude-orchestration-command-router.test.cjs b/tests/claude-orchestration-command-router.test.cjs index e0a408f1b..89dd105a7 100644 --- a/tests/claude-orchestration-command-router.test.cjs +++ b/tests/claude-orchestration-command-router.test.cjs @@ -64,7 +64,8 @@ describe('claude-orchestration emit-workflow (CLI)', () => { assert.ok(parsed.script.includes('parallel('), 'parallel() barrier emitted'); assert.ok(parsed.script.includes('gsd-executor'), 'gsd-executor agentType'); assert.ok(parsed.script.includes('worktree'), 'worktree isolation'); - assert.ok(parsed.script.includes('resumeFromRunId'), 'resumeFromRunId wired'); + // #2590: never CALLED — it is a Workflow tool input, not a script function. + assert.ok(!/^\s*resumeFromRunId\s*\(/m.test(parsed.script), 'must not CALL resumeFromRunId'); assert.ok(parsed.script.includes('run-cli-1143'), 'carries the run id'); assert.strictEqual(parsed.summary.resumeRunId, 'run-cli-1143'); assert.strictEqual(parsed.summary.waves, 1); @@ -84,8 +85,11 @@ describe('claude-orchestration emit-workflow (CLI)', () => { '--run-id', 'r', '--budget', '750000', ], tmp); - assert.ok(parsed.script.includes('budget('), 'budget() pool emitted'); + // #2590: `budget` is read-only; budget(750000) threw "budget is not a + // function". The intended pool is recorded, not called. + assert.ok(!/^\s*budget\s*\(/m.test(parsed.script), 'must not CALL budget()'); assert.ok(parsed.script.includes('750000')); + assert.strictEqual(parsed.summary.budgetTokens, 750000); } finally { cleanup(tmp); } diff --git a/tests/claude-orchestration.test.cjs b/tests/claude-orchestration.test.cjs index 2eaa89e6f..d274ee874 100644 --- a/tests/claude-orchestration.test.cjs +++ b/tests/claude-orchestration.test.cjs @@ -358,17 +358,24 @@ describe('emitWorkflowScript', () => { assert.deepStrictEqual(stages[0].slice().sort(), ['p1', 'p2']); }); - test('resumeFromRunId wired to the provided runId (criterion 4)', () => { + test('runId carried for the caller, never CALLED as resumeFromRunId (criterion 4, #2590)', () => { const r = emitWorkflowScript(singleWaveManifest()); - assert.ok(r.script.includes('resumeFromRunId'), 'references resumeFromRunId'); - assert.ok(r.script.includes('run-abc-1143'), 'carries the run id'); + // resumeFromRunId is a Workflow TOOL INPUT parameter, not a script function; + // emitting a call threw "resumeFromRunId is not defined" and rejected the + // whole script. The id reaches the caller via summary.resumeRunId. + assert.ok(!/^\s*resumeFromRunId\s*\(/m.test(r.script), 'must not CALL resumeFromRunId'); + assert.ok(r.script.includes('run-abc-1143'), 'carries the run id for the caller'); assert.strictEqual(r.summary.resumeRunId, 'run-abc-1143'); }); - test('shared budget pool emitted when budgetTokens provided', () => { + test('budgetTokens recorded as intent, never CALLED as budget() (#2590)', () => { const r = emitWorkflowScript({ ...singleWaveManifest(), budgetTokens: 500000 }); - assert.ok(r.script.includes('budget('), 'emits budget() pool'); - assert.ok(r.script.includes('500000')); + // `budget` is a read-only object { total, spent(), remaining() } supplied by + // the caller's token directive; `budget(500000)` threw "budget is not a + // function". The intent is recorded in a comment and in the summary. + assert.ok(!/^\s*budget\s*\(/m.test(r.script), 'must not CALL budget()'); + assert.ok(r.script.includes('500000'), 'records the intended budget'); + assert.strictEqual(r.summary.budgetTokens, 500000); }); test('no budget() emitted when budgetTokens omitted', () => { diff --git a/tests/fix-2285-claude-orchestration-wiring.test.cjs b/tests/fix-2285-claude-orchestration-wiring.test.cjs index c99a8a84e..ec0d82ab2 100644 --- a/tests/fix-2285-claude-orchestration-wiring.test.cjs +++ b/tests/fix-2285-claude-orchestration-wiring.test.cjs @@ -107,7 +107,11 @@ describe('A. resolveWaveDispatch — enabled + all gates satisfied → workflow assert.strictEqual(result.backend, 'workflow'); assert.strictEqual(result.reason, 'workflow_backend_active'); assert.ok(typeof result.script === 'string' && result.script.length > 0, 'script must be a non-empty string'); - assert.match(result.script, /resumeFromRunId\("run-2285-1"\)/); + // #2590: resumeFromRunId is a Workflow TOOL INPUT, not a script function — + // calling it threw "resumeFromRunId is not defined". The run id must still + // reach the caller (which passes it as that input), but never as a call. + assert.ok(!/^\s*resumeFromRunId\s*\(/m.test(result.script), 'must not CALL resumeFromRunId'); + assert.strictEqual(result.summary.resumeRunId, 'run-2285-1'); assert.match(result.script, /agentType: "gsd-executor", isolation: "worktree"/); assert.ok(result.summary && result.summary.plans === 1, 'summary.plans must reflect the manifest'); }); diff --git a/tests/fix-2590-workflow-script-contract.test.cjs b/tests/fix-2590-workflow-script-contract.test.cjs new file mode 100644 index 000000000..38721627f --- /dev/null +++ b/tests/fix-2590-workflow-script-contract.test.cjs @@ -0,0 +1,277 @@ +/** + * #2590 — every emitted Workflow script was rejected by the Workflow tool. + * + * `emitWorkflowScript` generated four constructs the tool does not accept. The + * first was fatal on its own, so the backend could never dispatch a wave: + * + * 1. no `export const meta = {…}` first statement -> whole script rejected + * 2. `resumeFromRunId("")` -> "resumeFromRunId is not defined" + * (it is a Workflow TOOL INPUT parameter, not a script function) + * 3. `budget()` -> "budget is not a function" + * (`budget` is a read-only object { total, spent(), remaining() }) + * 4. `parallel(agent(…), agent(…))` -> "parallel() expects an array of functions" + * + * Plus two secondary defects that kept the emitted script from ever being + * REACHED, which is why this shipped undetected: + * + * 5. nothing resolved the Agent SDK version, so gate 5 returned + * `agent_sdk_version_unknown` on every automated run + * 6. the runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging + * from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any + * invocation without --runtime reported `runtime_not_claude` + * + * The script assertions parse the emitted text as a real ES module rather than + * pattern-matching it, so a syntactically invalid script fails outright. + */ + +'use strict'; + +process.env.GSD_TEST_MODE = '1'; + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const fs = require('node:fs'); +const path = require('node:path'); +const { execFileSync } = require('node:child_process'); +const { createTempDir, cleanup } = require('./helpers.cjs'); + +const core = require('../gsd-core/bin/lib/claude-orchestration.cjs'); +const TOOLS = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs'); + +function emit(overrides) { + const input = Object.assign({ + phaseDir: '.planning/phases/01', + runId: 'execute-1', + waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }], + }, overrides || {}); + const r = core.emitWorkflowScript(input); + assert.ok(r.ok, `emit failed: ${JSON.stringify(r)}`); + return r; +} + +/** First non-comment, non-blank line — the script's first actual statement. */ +function firstStatement(script) { + return script.split('\n').map((l) => l.trim()) + .find((l) => l.length > 0 && !l.startsWith('//')) || ''; +} + +describe('#2590: emitted Workflow scripts satisfy the Workflow tool contract', () => { + test('the emitted script is syntactically valid as an ES module', () => { + // `export const meta` + top-level `await` only parse in module context — + // which is exactly the context the Workflow tool runs the script in. + const { script } = emit({ + waves: [ + { id: 'w1', plans: [ + { id: 'a', brief: 'one', files_modified: ['a.ts'] }, + { id: 'b', brief: 'two', files_modified: ['b.ts'] }, + ] }, + { id: 'w2', plans: [{ id: 'c', brief: 'three', files_modified: ['c.ts'] }] }, + ], + }); + // .mjs so node parses it in module context, inside a helper temp dir so + // cleanup() carries the Windows-EBUSY retry budget. + const dir = createTempDir('gsd-2590-parse-'); + const f = path.join(dir, 'emitted.mjs'); + fs.writeFileSync(f, script); + try { + execFileSync(process.execPath, ['--check', f], { stdio: 'pipe' }); + } catch (e) { + assert.fail(`emitted script does not parse: ${e.stderr ? e.stderr.toString() : e.message}`); + } finally { + cleanup(dir); + } + }); + + test('1. `export const meta` is the first statement', () => { + const { script } = emit(); + assert.match( + firstStatement(script), + /^export const meta = \{/, + 'the Workflow tool rejects any script whose first statement is not the meta block', + ); + }); + + test('meta.phases titles match the emitted phase() calls exactly', () => { + // The tool matches phase titles by exact string; a mismatch silently splits + // progress into an unnamed group. + const { script } = emit({ + waves: [ + { id: 'alpha', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] }, + { id: 'beta', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] }, + ], + }); + const metaTitles = [...script.matchAll(/\{ title: "([^"]+)"/g)].map((m) => m[1]); + const phaseTitles = [...script.matchAll(/^phase\("([^"]+)"\)/gm)].map((m) => m[1]); + assert.deepEqual(metaTitles, ['Wave alpha', 'Wave beta']); + assert.deepEqual(phaseTitles, metaTitles, 'phase() titles must match meta.phases exactly'); + }); + + test('duplicate wave ids are rejected (phase titles must map 1:1)', () => { + // Two waves sharing an id emit two identical `phase("Wave x")` calls and two + // identical meta.phases entries; the tool matches titles by exact string, so + // the second wave's agents would be attributed to the first's progress group. + const r = core.emitWorkflowScript({ + phaseDir: '.planning/phases/01', + runId: 'execute-1', + waves: [ + { id: 'dup', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] }, + { id: 'dup', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] }, + ], + }); + assert.equal(r.ok, false, 'duplicate wave ids must be rejected, not silently merged'); + assert.match(String(r.reason), /duplicate wave id/); + }); + + test('distinct wave ids are still accepted (the boundary either side)', () => { + const r = core.emitWorkflowScript({ + phaseDir: '.planning/phases/01', + runId: 'execute-1', + waves: [ + { id: 'w1', plans: [{ id: 'a', brief: 'one', files_modified: ['a.ts'] }] }, + { id: 'w2', plans: [{ id: 'b', brief: 'two', files_modified: ['b.ts'] }] }, + ], + }); + assert.equal(r.ok, true, `distinct wave ids must pass: ${JSON.stringify(r)}`); + }); + + test('2. resumeFromRunId is never CALLED (it is a tool input, not a function)', () => { + const { script, summary } = emit({ runId: 'execute-7' }); + assert.ok( + !/^\s*resumeFromRunId\s*\(/m.test(script), + 'calling resumeFromRunId() throws "resumeFromRunId is not defined"', + ); + // The run id must still reach the caller, which passes it as the tool input. + assert.equal(summary.resumeRunId, 'execute-7'); + }); + + test('3. budget is never CALLED, at and around the boundary', () => { + // budgetTokens is floored at > 0; check 0 (rejected), 1 (accepted), and a + // large value — none may produce a budget(...) call. + for (const tokens of [0, 1, 500000]) { + const { script, summary } = emit({ budgetTokens: tokens }); + assert.ok( + !/^\s*budget\s*\(/m.test(script), + `budgetTokens=${tokens}: calling budget() throws "budget is not a function"`, + ); + assert.equal(summary.budgetTokens, tokens > 0 ? tokens : null); + } + }); + + test('4. parallel() receives an array of thunks, not agent() results', () => { + const { script } = emit({ + waves: [{ id: 'w', plans: [ + { id: 'a', brief: 'one', files_modified: ['a.ts'] }, + { id: 'b', brief: 'two', files_modified: ['b.ts'] }, + ] }], + }); + assert.ok(/parallel\(\[/.test(script), 'parallel() expects an array of functions'); + assert.ok( + !/parallel\(\s*agent\(/.test(script), + 'passing agent() results directly both throws and starts every agent eagerly', + ); + // Each agent must be wrapped in a thunk so parallel() can bound concurrency. + const agents = [...script.matchAll(/agent\("/g)].length; + const thunks = [...script.matchAll(/\(\) => agent\("/g)].length; + assert.equal(thunks, agents, 'every agent() must be wrapped in a () => thunk'); + }); + + test('single-plan stages also emit an array (regression: the 1-plan branch)', () => { + // The pre-fix code had a SEPARATE single-plan branch that emitted + // `parallel(\n agent(...)\n)` — valid-looking but the same defect. + const { script } = emit(); + assert.ok(/parallel\(\[/.test(script)); + assert.equal([...script.matchAll(/\(\) => agent\("/g)].length, 1); + }); + + test('per-plan worktree isolation still mirrors use_worktree', () => { + const { script } = emit({ + waves: [{ id: 'w', plans: [ + { id: 'a', brief: 'iso', files_modified: ['a.ts'] }, + { id: 'b', brief: 'noiso', files_modified: ['b.ts'], use_worktree: false }, + ] }], + }); + assert.match(script, /agent\("iso", \{ agentType: "gsd-executor", isolation: "worktree" \}\)/); + assert.match(script, /agent\("noiso", \{ agentType: "gsd-executor" \}\)/); + }); +}); + +describe('#2590: the backend is reachable without hand-passed flags', () => { + function repro() { + const dir = createTempDir('gsd-2590-repro-'); + fs.mkdirSync(path.join(dir, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(dir, '.planning', 'config.json'), + JSON.stringify({ claude_orchestration: { enabled: true } }), + ); + fs.writeFileSync( + path.join(dir, 'waves.json'), + JSON.stringify({ waves: [{ id: 'wave-1', plans: [{ id: '01-01', brief: 'noop', files_modified: ['a.ts'] }] }] }), + ); + return dir; + } + + function resolve(dir, extraArgs) { + const out = execFileSync(process.execPath, [ + TOOLS, 'claude-orchestration', 'resolve-wave-dispatch', + '--waves', 'waves.json', '--run-id', 'execute-1', + '--phase-dir', '.planning/phases/01', '--raw', + ...(extraArgs || []), + ], { cwd: dir, encoding: 'utf8' }); + return JSON.parse(out); + } + + test('5+6. no --runtime and no --agent-sdk-version still reaches the version gate', () => { + const dir = repro(); + try { + const r = resolve(dir); + // Pre-fix this was `agent_sdk_version_unknown` (nothing resolved a + // version) or `runtime_not_claude` (the divergent fallback). Either is a + // regression; the version gate must now be reached and answer truthfully. + assert.notEqual(r.reason, 'agent_sdk_version_unknown', + 'the router must resolve the installed SDK version itself'); + assert.notEqual(r.reason, 'runtime_not_claude', + 'runtime must fall back to the canonical config.runtime > claude chain'); + } finally { + cleanup(dir); + } + }); + + test('an SDK version above the floor activates the workflow backend end to end', () => { + const dir = repro(); + try { + const r = resolve(dir, ['--agent-sdk-version', '0.3.149']); + assert.equal(r.backend, 'workflow', `expected workflow backend, got ${JSON.stringify(r)}`); + assert.ok(typeof r.script === 'string' && r.script.length > 0); + assert.match(firstStatement(r.script), /^export const meta = \{/); + } finally { + cleanup(dir); + } + }); + + test('an explicit --agent-sdk-version still wins over the installed one', () => { + const dir = repro(); + try { + // A deliberately ancient pin must be honored (and decline), proving the + // flag is not ignored now that a fallback exists. + const r = resolve(dir, ['--agent-sdk-version', '0.0.1']); + assert.equal(r.backend, 'inline'); + assert.equal(r.reason, 'agent_sdk_version_below_floor'); + } finally { + cleanup(dir); + } + }); + + test('GSD_AGENT_SDK_VERSION is honored between the flag and the installed version', () => { + const dir = repro(); + try { + const out = execFileSync(process.execPath, [ + TOOLS, 'claude-orchestration', 'resolve-wave-dispatch', + '--waves', 'waves.json', '--run-id', 'execute-1', + '--phase-dir', '.planning/phases/01', '--raw', + ], { cwd: dir, encoding: 'utf8', env: { ...process.env, GSD_AGENT_SDK_VERSION: '0.3.149' } }); + assert.equal(JSON.parse(out).backend, 'workflow'); + } finally { + cleanup(dir); + } + }); +});