Skip to content

fix(#904): per-agent + global semaphore on backend agent HTTP calls (RC-1) - #908

Closed
dolho wants to merge 1 commit into
devfrom
feature/904-rc1-agent-call-semaphore
Closed

dolho wants to merge 1 commit into
devfrom
feature/904-rc1-agent-call-semaphore

Conversation

@dolho

@dolho dolho commented May 21, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes RC-1 of #904 — the UI-freeze symptom. Stacks on top of #907 (RC-2 + RC-3) but is branched directly off dev so it can land in either order. Together with #907, both layers (false signal cascade + worker saturation) of the bug are fixed; RC-4 (cgroup OOM observability) remains open as a follow-up.

What the fix is

Two layered semaphores around outbound agent HTTP calls in services.task_execution_service.agent_post_with_retry:

  • Per-agent semaphore — sized to the agent's max_parallel_tasks (default 3 on DB miss). One bad citizen can't dominate.
  • Global semaphore — sized to BACKEND_AGENT_CALL_LIMIT env (default 8). Backend always has spare async capacity for dashboard / healthcheck.

Acquire with BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S (default 30s) → raises BackendAgentCallBudgetExhausted. Both execute_task (TaskExecutionResult FAILED) and routers/chat.py (HTTP 503) have dedicated except-branches. SUB-003 auto-switch is explicitly bypassed on this path — the rejection is local to the backend, the subscription is unrelated.

Why this shape (vs. alternatives)

  • Detach long calls to background tasks: cleaner architecturally but ~400 LoC and touches result-delivery (chat router, scheduler poll, frontend WS). Higher blast radius. Left as future work.
  • Raise --workers count: mitigation only — a misbehaving agent can still hold every worker.
  • Semaphore (this PR) — minimal scope, no architecture change, deterministic local repro.

The semaphore reduces sync-DB contention by bounding fan-out. It does NOT fix the underlying sync sqlite3 calls in db/connection.py — that's a separate refactor (run_in_executor-wrapped DB or aiosqlite).

Test plan

  • 7 unit tests in tests/unit/test_904_agent_call_limiter.py — happy path, per-agent cap, global cap, timeout, release-on-exception, unknown-agent fallback
  • 107 touched-area unit tests still pass (task_execution, capacity, subscription, chat_router, backlog)
  • Live local repro: 3 concurrent /api/chat to a sleeper-shim agent with BACKEND_AGENT_CALL_LIMIT=2 and 2s queue timeout. Chat 1 returned HTTP 503 in 2176ms with "Backend call budget exhausted for rc1-test after 2001ms (agent_cap=3, global_cap=2)". Chats 2 & 3 held the slots until the shim was killed. Dashboard /health stayed responsive.
  • Stage rollout: deploy with default BACKEND_AGENT_CALL_LIMIT=8 (matches typical fleet); monitor for unexpected 503s in subscription_rate_limit_events (must stay zero — SUB-003 skip is enforced).

Related to #904

🤖 Generated with Claude Code

…RC-1)

Issue #904 RC-1 — UI freeze on slow agent. Backend's
`task_execution_service.agent_post_with_retry` had no fan-out bound,
so a single misbehaving agent's 11.5-min HTTP call could leave many
backend coroutines `await`ing on `httpx.post` while each emitted
periodic synchronous `sqlite3` calls (`db/connection.py:18` —
`sqlite3.connect(timeout=30.0)`). Sync DB inside async coroutines
stalls the event loop momentarily; with enough concurrent
long-runners + writes to the same SQLite file, the writer-lock
contention drove the Docker healthcheck past its 10s ceiling and the
operator dashboard's parallel API fan-out queued until the offending
agent was restarted by hand.

This PR adds two layered semaphores around outbound agent calls,
keeping the call shape (await on httpx) untouched:

  1. **Per-agent semaphore**, sized to the agent's
     `max_parallel_tasks` (default 3 on lookup miss). Bounds
     fan-out per agent so one bad citizen can't dominate.
  2. **Global semaphore** sized to `BACKEND_AGENT_CALL_LIMIT` env
     (default 8). Caps total concurrent outbound calls — the backend
     keeps spare async capacity for dashboard / healthcheck even
     when every agent is mid-call.

Acquire-with-timeout (`BACKEND_AGENT_CALL_QUEUE_TIMEOUT_S`, default
30s) raises `BackendAgentCallBudgetExhausted`; both
`task_execution_service.execute_task` and `routers/chat.py` get
dedicated except-branches that mark the execution FAILED and (in
the chat router) return HTTP 503 with the budget message. SUB-003
auto-switch does NOT fire on this path — the rejection is local
to the backend, the subscription is unrelated.

Files:
- `src/backend/services/agent_call_limiter.py` (new) — primitives:
  `acquire_agent_call_slot`, `BackendAgentCallBudgetExhausted`,
  `_reset_for_testing` test hook
- `src/backend/services/task_execution_service.py` — wrap each
  connect-retry attempt in `agent_post_with_retry` with the slot
  context manager; add dedicated except in `execute_task`
- `src/backend/routers/chat.py` — same dedicated except, 503 to
  the caller
- `docker-compose.yml` — pass both env vars to backend (commented
  defaults: 8 / 30)
- `docs/memory/requirements.md` — §10.4.2 explaining behavior + the
  explicit out-of-scope note (sync→async DB is a separate refactor)

Verified live on local instance with `BACKEND_AGENT_CALL_LIMIT=2`
and a sleeper-shim agent (replaces `/usr/bin/claude` with a `sleep
300`): 3 concurrent /api/chat calls; 2 acquired immediately, the
3rd returned HTTP 503 in 2176ms with detail "Backend call budget
exhausted for rc1-test after 2001ms (agent_cap=3, global_cap=2)".
Dashboard /health stayed responsive.

Out of scope (separate follow-up):
- True sync→async-DB migration (`run_in_executor`-wrapped sqlite3
  or `aiosqlite`). The semaphore reduces contention but doesn't
  eliminate it.
- RC-4 cgroup OOM observability.

Related to #904

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@dolho

dolho commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

Consolidating #904 fixes into a single PR — RC-1 commit cherry-picked into #907 (the SIGKILL classification PR). Closing this one; #907 now covers RC-2 + RC-3 + RC-1.

@dolho dolho closed this May 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant