* feat(execute-phase): classify quota/rate-limit failures across runtimes (#3095) Dispatched executor subagents that die from provider quota or rate-limit errors currently look identical to a crashed agent to the orchestrator — so step 7's recovery prompt offers "retry now" when the right action is "wait for reset and resume". This adds a runtime-agnostic classifier and wires execute-phase step 7 to it. - `agent.classify-failure` SDK query returns `{class: 'quota-exceeded' | 'classify-handoff-bug' | 'unknown-failure', sentinel?, retryAfterSeconds?}`. Sentinels cover Claude Code (`usage limit`, `429`), Copilot CLI (`rate_limit`, `user_weekly_rate_limited`), Codex (`usage_limit_reached`, `too many requests`), and Gemini (`RESOURCE_EXHAUSTED`, `exceeded your`). - `execute-phase.md` step 7 now branches on the class. Quota-exceeded presents a wait-for-reset prompt and points at the safe-resume gate landing in #3212 instead of re-dispatching a fresh executor. - `docs/research/provider-rate-limit-signals.md` records the proactive (header / SDK event) signals each provider exposes and the upstream Claude Code / Copilot / Codex issues blocking hook-side detection — the forward path once host runtimes surface them. Resume-from-partial-worktree and context-load metrics from the original report are deliberately out of scope; they overlap #3212's `state.verify-against-disk` work already in flight. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(execute): render quota retry hint and refresh alias artifacts * fix(workflow): restore slash namespace and execute-phase size budget --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
15 KiB
15 KiB