Skip to content

bug: D-state subprocess blocks _kill_orphan_pipe_writers for ~15 min, opening circuit breaker #808

Description

@vybe

Summary

After a Claude Code subprocess exits, _kill_orphan_pipe_writers occasionally blocks for ~15 minutes because its /proc scan stalls on a D-state (uninterruptible sleep) kernel process. During this window, agent health probes time out and the circuit breaker opens, causing all scheduled executions to fail until the CB is manually reset or times out. PRs #730 and #747 (the two #728 fixes) are deployed but do not eliminate the D-state blocking — they cap the CPU spin and make the scan D-state-safer, but the block still occurs.

Component

Agent Runtime — subprocess_pgroup / subprocess_lifecycle

Priority

P1 — scheduled executions fail for the duration of the CB open window (observed 4+ hour outages)

Error

WARNING:agent_server.utils.subprocess_pgroup:[Subprocess] Reader thread(s) still busy after process exit (pid=XXXX, stuck_count=1) — killing process group, then waiting 30s for natural drain
WARNING:agent_server.services.subprocess_lifecycle:[Subprocess] Drain budget (90s) exceeded — safe_close_pipes may have deadlocked with reader thread's TextIOWrapper lock; reader threads are leaked daemon threads. Issue #728.
WARNING:agent_server.utils.subprocess_pgroup:[Subprocess] _kill_orphan_pipe_writers still running after 10s (pid=XXXX) — /proc scan may be blocked on a D-state process
ERROR:agent_server.utils.subprocess_pgroup:[Subprocess] Reader thread(s) still stuck after 917.6s post-kill grace — force-closing pipes; some buffered data may be lost (pid=XXXX, stuck_count=1)

Location

  • File: src/agent_server/utils/subprocess_pgroup.py (_kill_orphan_pipe_writers)
  • File: src/agent_server/services/subprocess_lifecycle.py (safe_close_pipes)

Root Cause

The sequence:

  1. Claude Code subprocess exits; reader thread is stuck holding the TextIOWrapper lock on the stdout/stderr pipe
  2. safe_close_pipes deadlocks waiting for the reader thread — drain budget (90s cap from PR fix(agent-server): cap drain executor thread at 90 s to fix CPU spin on auth failure (#728) #730) expires
  3. _kill_orphan_pipe_writers is invoked to kill processes holding the pipe open via /proc scan
  4. The /proc scan itself stalls because one of the scanned PIDs is in D-state (uninterruptible I/O sleep) — the os.readlink improvement from PR fix(subprocess): replace os.stat() with os.readlink() in orphan pipe scan (#728) #747 reduces the chance but does not eliminate it
  5. _kill_orphan_pipe_writers hangs for ~15 minutes until the D-state process eventually unblocks or the grace timer force-closes
  6. During steps 3–5, the agent's HTTP interface is degraded (event loop partially blocked) → health probes fail → circuit breaker opens

The 917-second stuck duration (observed in production) is the grace period before pipes are force-closed. The CB opens well before that and does not auto-recover once the CB goes dormant (~40 min of failed probes).

Reproduction Steps

  1. Deploy an agent with a scheduled task that spawns a subprocess (e.g. a Claude Code execution that runs a tool with a long-lived child process)
  2. Allow the scheduled execution to complete normally
  3. Observe _kill_orphan_pipe_writers log warning in agent container logs shortly after execution ends
  4. Monitor circuit breaker state — CB opens within minutes of the warning
  5. CB remains open until manually reset or backend restart

Note: Trigger is non-deterministic; occurs when a subprocess cleanup coincides with a D-state kernel process in the container's PID namespace.

Suggested Fix

Two complementary approaches:

1. Time-box _kill_orphan_pipe_writers unconditionally — the 10s warning exists but the scan itself has no hard kill. Wrap the entire /proc iteration in a concurrent.futures.ThreadPoolExecutor with a strict timeout so a stalled scan cannot block the event loop indefinitely regardless of D-state.

2. Move pipe cleanup entirely off the event loop — run safe_close_pipes and _kill_orphan_pipe_writers in a dedicated daemon thread with a hard wall-clock deadline, so a stuck cleanup cannot degrade health probe responsiveness.

Environment

  • Trinity version: 3a3e9a7 (dev branch, 2026-05-11)
  • Agent runtime: Claude Code subprocess via subprocess_pgroup
  • OS: Ubuntu 22.04 LTS / Linux kernel (container)

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions