Summary
When an MCP synchronous call to an agent is dropped mid-execution (e.g., the MCP client times out or disconnects), the broken HTTP connection propagates as [Errno 32] Broken pipe inside the agent's Claude Code runtime. This single event triggers the circuit breaker, which stays open and accumulates failures from subsequent retries/polling on the still-running concurrent executions — eventually opening after 7+ failures and blocking all new task dispatches to the agent, even though the agent container is healthy and the other executions are still running successfully.
Component
Backend / Agent Client / Circuit Breaker / MCP Server
Priority
P2
Error
# Agent container log:
ERROR:agent_server.services.claude_code:[Headless Task] Execution error: [Errno 32] Broken pipe
# Backend log sequence:
ERROR services.task_execution_service: [TaskExecService] Failed to execute task on [AGENT]: Task execution error: [Errno 32] Broken pipe
WARNING services.agent_client: Circuit OPENED for agent [AGENT] after 3 failures
WARNING services.agent_client: Circuit OPENED for agent [AGENT] after 4 failures
WARNING services.agent_client: Circuit OPENED for agent [AGENT] after 5 failures
WARNING services.agent_client: Circuit OPENED for agent [AGENT] after 6 failures
WARNING services.agent_client: Circuit OPENED for agent [AGENT] after 7 failures
Timeline
- T+0: Agent is running 2 concurrent executions (slots 1/3 and 2/3)
- T+0: MCP sync call dispatches a 3rd task (slot 3/3)
- T+10s: MCP client connection drops; agent gets EPIPE writing response → HTTP 500
- T+60s onward: Concurrent running executions trigger probe/retry calls → circuit opens repeatedly
- Result: Agent healthy but circuit blocked, no new tasks accepted
Root Cause
Two interacting issues:
-
Broken pipe propagation: When the backend closes its HTTP connection to the agent (due to MCP client disconnect), the agent server gets EPIPE trying to write back. This is surfaced as a task execution error rather than a connection-level error, causing it to count against the circuit breaker failure threshold.
-
Circuit breaker over-sensitivity: A single EPIPE from a dropped client connection (not an agent health issue) increments the circuit breaker failure counter. The subsequent concurrent executions' completion callbacks are then misinterpreted as additional failures, opening the circuit after only 3 real errors.
Reproduction Steps
- Start an agent with 2+ parallel slot capacity
- Start 2 long-running executions on the agent (filling 2 slots)
- Make an MCP synchronous call that dispatches a 3rd task to the same agent
- Drop/cancel the MCP client connection before the task completes
- Observe
[Errno 32] Broken pipe error followed by circuit breaker opening
- Observe the circuit continues to open on subsequent attempts even though the agent container is running fine
Suggested Fix
Two-pronged fix:
-
Distinguish connection errors from execution errors: In the agent client HTTP layer, catch BrokenPipeError / ConnectionResetError and classify them as connection-level failures rather than execution failures. These should not count toward the circuit breaker threshold, or should count much less aggressively.
-
Circuit breaker should not re-open on concurrent execution completions: After the initial EPIPE trip, the circuit breaker should not continue accumulating failures from the completion callbacks of the other still-running (healthy) executions. Consider resetting the circuit to half-open after the first recovery attempt rather than letting concurrent callbacks pile on.
Environment
- Trinity version:
615a266
- OS: Ubuntu 22.04
Related
Summary
When an MCP synchronous call to an agent is dropped mid-execution (e.g., the MCP client times out or disconnects), the broken HTTP connection propagates as
[Errno 32] Broken pipeinside the agent's Claude Code runtime. This single event triggers the circuit breaker, which stays open and accumulates failures from subsequent retries/polling on the still-running concurrent executions — eventually opening after 7+ failures and blocking all new task dispatches to the agent, even though the agent container is healthy and the other executions are still running successfully.Component
Backend / Agent Client / Circuit Breaker / MCP Server
Priority
P2
Error
Timeline
Root Cause
Two interacting issues:
Broken pipe propagation: When the backend closes its HTTP connection to the agent (due to MCP client disconnect), the agent server gets
EPIPEtrying to write back. This is surfaced as a task execution error rather than a connection-level error, causing it to count against the circuit breaker failure threshold.Circuit breaker over-sensitivity: A single EPIPE from a dropped client connection (not an agent health issue) increments the circuit breaker failure counter. The subsequent concurrent executions' completion callbacks are then misinterpreted as additional failures, opening the circuit after only 3 real errors.
Reproduction Steps
[Errno 32] Broken pipeerror followed by circuit breaker openingSuggested Fix
Two-pronged fix:
Distinguish connection errors from execution errors: In the agent client HTTP layer, catch
BrokenPipeError/ConnectionResetErrorand classify them as connection-level failures rather than execution failures. These should not count toward the circuit breaker threshold, or should count much less aggressively.Circuit breaker should not re-open on concurrent execution completions: After the initial EPIPE trip, the circuit breaker should not continue accumulating failures from the completion callbacks of the other still-running (healthy) executions. Consider resetting the circuit to half-open after the first recovery attempt rather than letting concurrent callbacks pile on.
Environment
615a266Related
services/agent_client.py