* feat(execute-phase): classify quota/rate-limit failures across runtimes (#3095)
Dispatched executor subagents that die from provider quota or rate-limit
errors currently look identical to a crashed agent to the orchestrator —
so step 7's recovery prompt offers "retry now" when the right action is
"wait for reset and resume". This adds a runtime-agnostic classifier and
wires execute-phase step 7 to it.
- `agent.classify-failure` SDK query returns
`{class: 'quota-exceeded' | 'classify-handoff-bug' | 'unknown-failure',
sentinel?, retryAfterSeconds?}`. Sentinels cover Claude Code
(`usage limit`, `429`), Copilot CLI (`rate_limit`,
`user_weekly_rate_limited`), Codex (`usage_limit_reached`,
`too many requests`), and Gemini (`RESOURCE_EXHAUSTED`,
`exceeded your`).
- `execute-phase.md` step 7 now branches on the class. Quota-exceeded
presents a wait-for-reset prompt and points at the safe-resume gate
landing in #3212 instead of re-dispatching a fresh executor.
- `docs/research/provider-rate-limit-signals.md` records the proactive
(header / SDK event) signals each provider exposes and the upstream
Claude Code / Copilot / Codex issues blocking hook-side detection —
the forward path once host runtimes surface them.
Resume-from-partial-worktree and context-load metrics from the original
report are deliberately out of scope; they overlap #3212's
`state.verify-against-disk` work already in flight.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(execute): render quota retry hint and refresh alias artifacts
* fix(workflow): restore slash namespace and execute-phase size budget
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>