You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When the SUB-003 subscription auto-switch fires on a 429 Too Many Requests or auth-class failure mid-execution, the triggering execution row is marked FAILED (see task_execution_service.py:600-660). #606 closed as "by design" for the paths that exist today, but one path is genuinely uncovered: one-shot scheduled executions (ad-hoc manual triggers, not a recurring cron). Neither of the recovery patterns below applies, and the user sees the FAILED row without any retry.
What already exists (do not redo)
Interactive chat — already retries.src/backend/routers/chat.py:466-493 returns retry_after: 15 + an auto_switch payload to the frontend when a subscription-mode agent hits 429/503 and a successful switch occurred. The Vue chat UI handles the client-driven retry. Interactive callers do not need this issue.
Recurring scheduled executions — recover at next tick. A schedule with a real cron expression fires again on the next tick using the newly assigned subscription. Transient FAILED rows are recovered by the schedule's natural recurrence (not by the platform replaying the failed execution).
Gap this issue addresses
One-shot triggers — POST /api/agents/{name}/schedules/{id}/trigger, webhook fires on schedules with cron_expression that doesn't recur in any meaningful window, manual trigger_agent_schedule MCP calls — produce a single execution. If the very first attempt hits the 429 that fires the auto-switch, the execution is marked FAILED and the user has to manually retrigger after the (typically successful) switch.
Design surface
Each of these is a non-trivial decision; reviewers should pick a path explicitly before implementation.
Idempotency model. Is the agent's executed request idempotent under retry? If the original message kicked off external side effects before hitting 429 (e.g. a Slack post, a file write, an MCP tool call), the retry must not double-fire. How is this guaranteed — agent-level dedup keys, platform-level execution-id reuse, or scoping the retry to executions that didn't reach a tool-call boundary?
Retry budget. One retry per execution, or N? Exponential backoff? Hard cap to prevent infinite cascades if the new subscription is also exhausted?
Capacity-slot lifecycle. The original execution holds a CapacityManager slot. Release-and-reacquire (race-prone against other queued tasks) or hold-through (occupies capacity longer than necessary)?
Subscription usage attribution. Cost / token usage for the retried execution — attribute to the new subscription or the old one that triggered the 429? Affects fairness for shared subscription pools.
Failure attribution. Does the original 429 still get recorded as the cause of the FAILED-then-recovered execution, or is the audit trail rewritten?
Interaction with the 2h skip-list. If the new subscription is also rate-limited mid-retry (cascade auth/quota failure), do we cascade-switch again? Hard-fail after N? The 2h skip-list in select_best_alternative_subscription already guards against ping-pong but only at the selection layer — retry adds a new dimension.
src/backend/services/task_execution_service.py:600-660 — the 429/auth exception block that fire-and-forgets auto-switch and marks the execution FAILED.
Context
When the SUB-003 subscription auto-switch fires on a
429 Too Many Requestsor auth-class failure mid-execution, the triggering execution row is marked FAILED (seetask_execution_service.py:600-660). #606 closed as "by design" for the paths that exist today, but one path is genuinely uncovered: one-shot scheduled executions (ad-hoc manual triggers, not a recurring cron). Neither of the recovery patterns below applies, and the user sees the FAILED row without any retry.What already exists (do not redo)
src/backend/routers/chat.py:466-493returnsretry_after: 15+ anauto_switchpayload to the frontend when a subscription-mode agent hits 429/503 and a successful switch occurred. The Vue chat UI handles the client-driven retry. Interactive callers do not need this issue.Gap this issue addresses
One-shot triggers —
POST /api/agents/{name}/schedules/{id}/trigger, webhook fires on schedules withcron_expressionthat doesn't recur in any meaningful window, manualtrigger_agent_scheduleMCP calls — produce a single execution. If the very first attempt hits the 429 that fires the auto-switch, the execution is marked FAILED and the user has to manually retrigger after the (typically successful) switch.Design surface
Each of these is a non-trivial decision; reviewers should pick a path explicitly before implementation.
auto_retrying), or reusepending_retryfrom RETRY-001 / bug(SUB-003): rate-limit events never age out due to SQLite string-compare bug; retries amplify outages #476? How does the retried execution appear inagent_activitiesandchat_messages— one row or two? Linked via FK?CapacityManagerslot. Release-and-reacquire (race-prone against other queued tasks) or hold-through (occupies capacity longer than necessary)?select_best_alternative_subscriptionalready guards against ping-pong but only at the selection layer — retry adds a new dimension.assign_subscription_to_agent+container_stop+start_agent_internal(raised by the bug: auto-switch (SUB-003) marks triggering execution failed due to spurious .credentials.enc injection after container restart #606 Codex eng review). If two 429s for the same agent fire near-simultaneously, retries could trample each other. Lock as part of this issue, or punt to a separate concurrency-hardening issue?Out of scope
routers/chat.py:466-493)..credentials.encimport noise (covered by bug: agent start response shows credentials_injection: failed for subscription-mode agents #612 + the regression test pinning the_restart_agentchain in the bug: auto-switch (SUB-003) marks triggering execution failed due to spurious .credentials.enc injection after container restart #606 close PR).References
lifecycle.py:155).src/backend/routers/chat.py:466-493— existing interactive retry hint.src/backend/services/task_execution_service.py:600-660— the 429/auth exception block that fire-and-forgets auto-switch and marks the execution FAILED.src/backend/services/subscription_auto_switch.py:145-259—_perform_auto_switch/_restart_agent.