fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)

* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable

Every emitted script was rejected. Four invalid constructs, the first fatal on
its own, so the Workflow backend could never dispatch a wave:

  1. no `export const meta = {…}` first statement -> whole script rejected
  2. resumeFromRunId("<id>")  -> "resumeFromRunId is not defined". It is a
     Workflow TOOL INPUT parameter, not a script function. The run id still
     reaches the caller via summary.resumeRunId, to pass as that input.
  3. budget(<n>)              -> "budget is not a function". `budget` is a
     read-only object { total, spent(), remaining() } fed by the caller's token
     directive; a script cannot set it. Recorded as intent in a comment.
  4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions".
     Now parallel([() => agent(…), …]) — passing agent() results directly also
     started every agent eagerly, before parallel() could bound concurrency.

The single-plan stage had its own branch with the same parallel() defect; both
branches are now one array-emitting path. Waves also emit phase() calls whose
titles match meta.phases exactly, so progress groups correctly.

Two secondary defects kept the script from ever being REACHED — which is why
this shipped undetected:

  5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no
     scriptable way" to introspect it and told callers to omit the flag, so
     gate 5 returned agent_sdk_version_unknown on every automated run while
     `capability state` still reported active:true. True for bash, false for
     Node: the router now reads the installed @anthropic-ai/claude-agent-sdk
     version, walking node_modules up the tree and reading package.json
     directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the
     SDK's exports map does not expose ./package.json. Precedence: explicit flag
     > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an
     unresolvable version still declines to inline. A too-old SDK now reports
     the truthful agent_sdk_version_below_floor instead of unknown.
  6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
     from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
     invocation without --runtime reported runtime_not_claude on an ordinary
     Claude project. Now delegates to runtime-slash.resolveRuntime.

The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}`
snippet is removed rather than repaired: it was also shell-dependent — zsh does
not word-split unquoted parameter expansions, so it collapsed to a single argv
element, argValue() never matched, and the run failed into the same
agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto-
resolution removes the need for the construct entirely.

Verified with the issue's own repro: no flags now reaches the version gate; an
SDK above the floor yields backend:"workflow" with a script that parses as a
real ES module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids

Findings from the isolated review, all fixed.

HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was
already RED because of it. The registry embeds the fragment text INLINE, so the
shipped/installed copy still taught the exact broken contract this PR fixes:
the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown"
guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of
checking its exit code, so I recorded a red chain as green — checking $? now.)

HIGH — three existing tests asserted the OLD broken shape and would have failed
CI; none was touched by the first commit:
  tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…")
  tests/claude-orchestration.test.cjs                 — .includes('budget(')
  tests/claude-orchestration-command-router.test.cjs  — .includes('budget(')
Each now asserts the corrected contract: the id/pool reaches the caller via
summary, and neither construct is ever CALLED. Two sibling assertions had also
gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the
new explanatory COMMENT contains that substring, not because anything is wired.
Rewritten to assert the real property.

MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked
within a wave, but nothing checked wave ids across waves. That was harmless
before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a
matching meta.phases entry and the tool matches titles by exact string — two
waves sharing an id would collapse into one progress group and misattribute the
second wave's agents to the first. Rejected at validation, with tests either
side of the boundary.

MEDIUM — the fragment contradicted itself (its "Manifest construction" header
still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never
told the orchestrator to pass summary.resumeRunId as the Workflow tool's
resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to
a tool-invocation input, an implementer following only the fragment would have
silently regressed phase-resume to a no-op. Both fixed.

MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and
docs/explanation/claude-orchestration-capability.md documented
`resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output —
teaching the bug as the feature. Updated to the real contract, including the
required meta block and the thunk-array parallel() form. (The changeset is
`Fixed`, so the docs gate exempts this; it is corrected because it is wrong,
not because a gate demanded it.)

LOW — the router's top-of-file comment still described the divergent
`--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the
fix; and inserting resolveInstalledAgentSdkVersion had orphaned
resolveDetectionArgs' JSDoc above the wrong function. Both repaired.

lint:ci now exits 0 (verified by exit code, not by reading output).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2590): backfill changeset pr number (#2681)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-07-26 20:49:38 -04:00
committed by GitHub
parent c3958018dd
commit 0d08c32048
13 changed files with 584 additions and 79 deletions

View File

@@ -21,8 +21,10 @@ Claude Code ships an orchestration primitive that sidesteps exactly this: the
**Workflow tool** (the engine behind `/effort ultracode`, Agent SDK ≥ v0.3.149).
A Workflow script *is* the orchestrator — it runs from the main loop and spawns
subagents itself via `agent()`, `parallel()` (barrier), `pipeline()`, and
`phase()`, with `isolation: 'worktree'`, a shared token `budget`, and
`resumeFromRunId`.
`phase()`, with `isolation: 'worktree'`. Two related capabilities are **tool
inputs rather than script functions**: the token `budget` is a read-only object
a script reads but cannot set, and `resumeFromRunId` is a parameter passed when
invoking the tool.
## The capability
@@ -69,8 +71,9 @@ Workflow backend activates only when *every* gate passes; any miss degrades to
| Plan (`use_worktree` not `false`) | `agent(brief, { agentType: 'gsd-executor', isolation: 'worktree' })` |
| Plan (`use_worktree: false`) | `agent(brief, { agentType: 'gsd-executor' })` (no isolation) |
| `files_modified` overlap | forces the plans into separate sequential stages |
| Phase run id | `resumeFromRunId("<id>")` |
| Phase token cap | `budget(<tokens>)` |
| Wave | a `phase("Wave <id>")` group, matching a `meta.phases` entry |
| Phase run id | `summary.resumeRunId` → pass as the Workflow tool's `resumeFromRunId` **input** |
| Phase token cap | recorded in `summary.budgetTokens`; `budget` is read-only in a script |
Because the emitted script composes the **same** `gsd-executor` agent the
inline path uses, with worktree isolation applied **per plan** from the

View File

@@ -116,16 +116,18 @@ gsd-tools claude-orchestration emit-workflow \
The output is a generated Workflow script that maps GSD's model 1:1 onto Workflow primitives:
- **waves → sequential `parallel()` barriers** (split into separate stages within a wave when `files_modified` overlap),
- **plans → `agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses,
- **`resumeFromRunId("<run-id>")`** wired to the phase run id,
- **`budget(<tokens>)`** — a shared token pool across the whole phase (omit `--budget` to skip).
- an `export const meta = { name, description, phases }` block as the **first statement** (the Workflow tool rejects any script without it),
- **waves → a `phase("Wave <id>")` group and a sequential `await parallel([...])` barrier** (split into separate stages within a wave when `files_modified` overlap),
- **plans → `() => agent(brief, { agentType: "gsd-executor", isolation: "worktree" })`** — the **same** executor agent and worktree isolation the inline path uses, each wrapped in a thunk because `parallel()` takes an **array of functions**,
- the phase run id in **`summary.resumeRunId`**, and the intended token pool in **`summary.budgetTokens`**.
`resumeFromRunId` and `budget` are **not** emitted as calls (#2590). `resumeFromRunId` is a Workflow **tool input**, and `budget` is a read-only object (`{ total, spent(), remaining() }`) fed by the caller's token directive — calling either from a script throws and the whole script is rejected.
Because the script composes the same `gsd-executor` agent + worktree isolation + `SUMMARY.md` artifact as the inline path, the artifacts and commits it produces are identical — only the execution vehicle differs.
### Run the emitted script
Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation). The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, waves barrier between each other, and `resumeFromRunId` lets an interrupted phase resume without re-running completed plans.
Feed the emitted script to Claude Code's Workflow tool (`/effort ultracode`, or an Agent SDK `Workflow` invocation), **passing `summary.resumeRunId` as the tool's `resumeFromRunId` input**. The orchestrator runs it; each `agent()` call spawns a `gsd-executor` in its own worktree, and waves barrier between each other. Omitting that input silently regresses phase-resume to a no-op — an interrupted phase re-runs completed plans.
---