You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
After a Claude Code subprocess exits, _kill_orphan_pipe_writers occasionally blocks for ~15 minutes because its /proc scan stalls on a D-state (uninterruptible sleep) kernel process. During this window, agent health probes time out and the circuit breaker opens, causing all scheduled executions to fail until the CB is manually reset or times out. PRs #730 and #747 (the two #728 fixes) are deployed but do not eliminate the D-state blocking — they cap the CPU spin and make the scan D-state-safer, but the block still occurs.
P1 — scheduled executions fail for the duration of the CB open window (observed 4+ hour outages)
Error
WARNING:agent_server.utils.subprocess_pgroup:[Subprocess] Reader thread(s) still busy after process exit (pid=XXXX, stuck_count=1) — killing process group, then waiting 30s for natural drain
WARNING:agent_server.services.subprocess_lifecycle:[Subprocess] Drain budget (90s) exceeded — safe_close_pipes may have deadlocked with reader thread's TextIOWrapper lock; reader threads are leaked daemon threads. Issue #728.
WARNING:agent_server.utils.subprocess_pgroup:[Subprocess] _kill_orphan_pipe_writers still running after 10s (pid=XXXX) — /proc scan may be blocked on a D-state process
ERROR:agent_server.utils.subprocess_pgroup:[Subprocess] Reader thread(s) still stuck after 917.6s post-kill grace — force-closing pipes; some buffered data may be lost (pid=XXXX, stuck_count=1)
_kill_orphan_pipe_writers hangs for ~15 minutes until the D-state process eventually unblocks or the grace timer force-closes
During steps 3–5, the agent's HTTP interface is degraded (event loop partially blocked) → health probes fail → circuit breaker opens
The 917-second stuck duration (observed in production) is the grace period before pipes are force-closed. The CB opens well before that and does not auto-recover once the CB goes dormant (~40 min of failed probes).
Reproduction Steps
Deploy an agent with a scheduled task that spawns a subprocess (e.g. a Claude Code execution that runs a tool with a long-lived child process)
Allow the scheduled execution to complete normally
Observe _kill_orphan_pipe_writers log warning in agent container logs shortly after execution ends
Monitor circuit breaker state — CB opens within minutes of the warning
CB remains open until manually reset or backend restart
Note: Trigger is non-deterministic; occurs when a subprocess cleanup coincides with a D-state kernel process in the container's PID namespace.
Suggested Fix
Two complementary approaches:
1. Time-box _kill_orphan_pipe_writers unconditionally — the 10s warning exists but the scan itself has no hard kill. Wrap the entire /proc iteration in a concurrent.futures.ThreadPoolExecutor with a strict timeout so a stalled scan cannot block the event loop indefinitely regardless of D-state.
2. Move pipe cleanup entirely off the event loop — run safe_close_pipes and _kill_orphan_pipe_writers in a dedicated daemon thread with a hard wall-clock deadline, so a stuck cleanup cannot degrade health probe responsiveness.
Environment
Trinity version: 3a3e9a7 (dev branch, 2026-05-11)
Agent runtime: Claude Code subprocess via subprocess_pgroup
Summary
After a Claude Code subprocess exits,
_kill_orphan_pipe_writersoccasionally blocks for ~15 minutes because its/procscan stalls on a D-state (uninterruptible sleep) kernel process. During this window, agent health probes time out and the circuit breaker opens, causing all scheduled executions to fail until the CB is manually reset or times out. PRs #730 and #747 (the two #728 fixes) are deployed but do not eliminate the D-state blocking — they cap the CPU spin and make the scan D-state-safer, but the block still occurs.Component
Agent Runtime —
subprocess_pgroup/subprocess_lifecyclePriority
P1 — scheduled executions fail for the duration of the CB open window (observed 4+ hour outages)
Error
Location
src/agent_server/utils/subprocess_pgroup.py(_kill_orphan_pipe_writers)src/agent_server/services/subprocess_lifecycle.py(safe_close_pipes)Root Cause
The sequence:
TextIOWrapperlock on the stdout/stderr pipesafe_close_pipesdeadlocks waiting for the reader thread — drain budget (90s cap from PR fix(agent-server): cap drain executor thread at 90 s to fix CPU spin on auth failure (#728) #730) expires_kill_orphan_pipe_writersis invoked to kill processes holding the pipe open via/procscan/procscan itself stalls because one of the scanned PIDs is in D-state (uninterruptible I/O sleep) — the os.readlink improvement from PR fix(subprocess): replace os.stat() with os.readlink() in orphan pipe scan (#728) #747 reduces the chance but does not eliminate it_kill_orphan_pipe_writershangs for ~15 minutes until the D-state process eventually unblocks or the grace timer force-closesThe 917-second stuck duration (observed in production) is the grace period before pipes are force-closed. The CB opens well before that and does not auto-recover once the CB goes dormant (~40 min of failed probes).
Reproduction Steps
_kill_orphan_pipe_writerslog warning in agent container logs shortly after execution endsNote: Trigger is non-deterministic; occurs when a subprocess cleanup coincides with a D-state kernel process in the container's PID namespace.
Suggested Fix
Two complementary approaches:
1. Time-box
_kill_orphan_pipe_writersunconditionally — the 10s warning exists but the scan itself has no hard kill. Wrap the entire/prociteration in aconcurrent.futures.ThreadPoolExecutorwith a strict timeout so a stalled scan cannot block the event loop indefinitely regardless of D-state.2. Move pipe cleanup entirely off the event loop — run
safe_close_pipesand_kill_orphan_pipe_writersin a dedicated daemon thread with a hard wall-clock deadline, so a stuck cleanup cannot degrade health probe responsiveness.Environment
3a3e9a7(dev branch, 2026-05-11)subprocess_pgroupRelated
cap drain executor thread at 90s— fixes CPU spin, not D-state blockingreplace os.stat() with os.readlink() in orphan pipe scan— reduces D-state risk, does not eliminate