Repository navigation
Conversation
…RC-1) Issue #904 RC-1 — UI freeze on slow agent. Backend's `task_execution_service.agent_post_with_retry` had no fan-out bound, so a single misbehaving agent's 11.5-min HTTP call could leave many backend coroutines `await`ing on `httpx.post` while each emitted periodic synchronous `sqlite3` calls (`db/connection.py:18` — `sqlite3.connect(timeout=30.0)`). Sync DB inside async coroutines stalls the event loop momentarily; with enough concurrent long-runners + writes to the same SQLite file, the writer-lock contention drove the Docker healthcheck past its 10s ceiling and the operator dashboard's parallel API fan-out queued until the offending agent was restarted by hand. This PR adds two layered semaphores around outbound agent calls, keeping the call shape (await on httpx) untouched: 1. **Per-agent semaphore**, sized to the agent's `max_parallel_tasks` (default 3 on lookup miss). Bounds fan-out per agent so one bad citizen can't dominate. 2. **Global semaphore** sized to `BACKEND_AGENT_CALL_LIMIT` env (default 8). Caps total concurrent outbound calls — the backend keeps spare async capacity for dashboard / healthcheck even when every agent is mid-call. Acquire-with-timeout (`BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S`, default 30s) raises `BackendAgentCallBudgetExhausted`; both `task_execution_service.execute_task` and `routers/chat.py` get dedicated except-branches that mark the execution FAILED and (in the chat router) return HTTP 503 with the budget message. SUB-003 auto-switch does NOT fire on this path — the rejection is local to the backend, the subscription is unrelated. Files: - `src/backend/services/agent_call_limiter.py` (new) — primitives: `acquire_agent_call_slot`, `BackendAgentCallBudgetExhausted`, `_reset_for_testing` test hook - `src/backend/services/task_execution_service.py` — wrap each connect-retry attempt in `agent_post_with_retry` with the slot context manager; add dedicated except in `execute_task` - `src/backend/routers/chat.py` — same dedicated except, 503 to the caller - `docker-compose.yml` — pass both env vars to backend (commented defaults: 8 / 30) - `docs/memory/requirements.md` — §10.4.2 explaining behavior + the explicit out-of-scope note (sync→async DB is a separate refactor) Verified live on local instance with `BACKEND_AGENT_CALL_LIMIT=2` and a sleeper-shim agent (replaces `/usr/bin/claude` with a `sleep 300`): 3 concurrent /api/chat calls; 2 acquired immediately, the 3rd returned HTTP 503 in 2176ms with detail "Backend call budget exhausted for rc1-test after 2001ms (agent_cap=3, global_cap=2)". Dashboard /health stayed responsive. Out of scope (separate follow-up): - True sync→async-DB migration (`run_in_executor`-wrapped sqlite3 or `aiosqlite`). The semaphore reduces contention but doesn't eliminate it. - RC-4 cgroup OOM observability. Related to #904 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes RC-1 of #904 — the UI-freeze symptom. Stacks on top of #907 (RC-2 + RC-3) but is branched directly off
devso it can land in either order. Together with #907, both layers (false signal cascade + worker saturation) of the bug are fixed; RC-4 (cgroup OOM observability) remains open as a follow-up.What the fix is
Two layered semaphores around outbound agent HTTP calls in
services.task_execution_service.agent_post_with_retry:max_parallel_tasks(default 3 on DB miss). One bad citizen can't dominate.BACKEND_AGENT_CALL_LIMITenv (default 8). Backend always has spare async capacity for dashboard / healthcheck.Acquire with
BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S(default 30s) → raisesBackendAgentCallBudgetExhausted. Bothexecute_task(TaskExecutionResult FAILED) androuters/chat.py(HTTP 503) have dedicated except-branches. SUB-003 auto-switch is explicitly bypassed on this path — the rejection is local to the backend, the subscription is unrelated.Why this shape (vs. alternatives)
--workerscount: mitigation only — a misbehaving agent can still hold every worker.The semaphore reduces sync-DB contention by bounding fan-out. It does NOT fix the underlying sync
sqlite3calls indb/connection.py— that's a separate refactor (run_in_executor-wrapped DB oraiosqlite).Test plan
tests/unit/test_904_agent_call_limiter.py— happy path, per-agent cap, global cap, timeout, release-on-exception, unknown-agent fallbacktask_execution,capacity,subscription,chat_router,backlog)/api/chatto a sleeper-shim agent withBACKEND_AGENT_CALL_LIMIT=2and 2s queue timeout. Chat 1 returned HTTP 503 in 2176ms with"Backend call budget exhausted for rc1-test after 2001ms (agent_cap=3, global_cap=2)". Chats 2 & 3 held the slots until the shim was killed. Dashboard /health stayed responsive.BACKEND_AGENT_CALL_LIMIT=8(matches typical fleet); monitor for unexpected 503s insubscription_rate_limit_events(must stay zero — SUB-003 skip is enforced).Related to #904
🤖 Generated with Claude Code