Skip to content

Async chat_with_agent: long execution silently fails with null response (reader-thread) #678

Description

@pandore

Summary

Async chat_with_agent (parallel=true, async=true) on ability-website ran for ~24 minutes through 56 turns / 135 raw messages, then completed with status=failed, response=null, and all telemetry (cost, context_used, model_used) null. The platform's own error message describes a stdout-inheritance / reader-thread failure mode and asks for it to be filed.

Execution metadata

Field Value
execution_id FdEVIjC5ZRVdeG-e5q0coQ
agent_name ability-website
triggered_by mcp
started_at 2026-05-06T11:58:54.069960Z
completed_at 2026-05-06T12:22:40.963313Z
duration_ms 1,426,893 (~23.8 min)
status failed
response null
cost null
context_used null
model_used null
execution_log null (also lost)

Error message (verbatim from get_execution_result)

Execution completed without a result message after 0 tool calls / 56 turns (raw_messages=135). Likely cause: a tool or child subprocess inherited stdout and prevented the claude reader thread from capturing the final result block. Check agent-server logs for 'Reader thread(s) stuck after process exit' or 'I/O operation on closed file' near this execution. This is a transient infrastructure failure; retry the task.

The "0 tool calls / 56 turns" is suspicious — 56 turns with 0 tool calls is unusual for a real build task. Either the tool-call counter wasn't incremented (consistent with a reader-thread failure) or the agent really did spend 56 turns thinking without acting. The 135 raw messages suggests the former.

Reproduction context

  • Inbound: a single, long-form async chat_with_agent message — ~10 KB of structured markdown (build brief for a Next.js landing page, includes copy blocks, table-formatted form schema, GTM event list, acceptance criteria).
  • Mode: parallel=true, async=true, timeout_seconds=3600
  • Caller: separate Trinity instance (cross-instance MCP call from Abilityai/trinity-gtm-agent workspace).
  • Agent received the message and started running (per started_at and turn count). No final result block ever surfaced to the orchestrator.

Container-log signal possibly worth investigating

The startup logs for ability-website contain this warning before the run:

WARNING: Could not find target strings in /app/agent_server/services/claude_code.py
Agent server restarted with CLAUDECODE fix applied

The runtime appears to apply a hot-patch to claude_code.py at container start. If the patch's target strings don't match (e.g. after a claude-cli upgrade), the patch fails-open and the runtime continues without the fix. If the "CLAUDECODE fix" is the same code path that handles result-block capture / reader-thread cleanup, that's a candidate root cause — not a "transient" failure, but a regression triggered by a stale patch shim.

The patch warning is logged at WARNING level and doesn't block container startup, which means failures like this one will silently happen on every long execution until the patch is re-aligned or the underlying logic is upstreamed.

Impact

  • Cost / observability: ~24 min of compute (Sonnet-4.6) lost, with no cost record, no transcript, no usable diagnostic data beyond the meta-error string. Hard to reason about whether the actual deliverable was produced and lost vs. never finished.
  • Trust in async mode: any long-running build/research delegation via parallel=true, async=true could silently fail this way after 20+ min. For caller agents this is the worst failure mode — long wait, no signal, no retry guidance beyond the platform's own (correct) "retry."
  • Telemetry: cost/model_used being null when the agent clearly burned tokens means this failure mode also isn't billable and isn't surfacing to whatever rate-limit / spend dashboards exist.

Suggested investigation

The platform's own error string already names the log lines to grep:

Reader thread(s) stuck after process exit
I/O operation on closed file

Around execution_id=FdEVIjC5ZRVdeG-e5q0coQ (12:22:40Z 2026-05-06).

Concretely:

  1. Check whether the "CLAUDECODE fix" hot-patch in /app/agent_server/services/claude_code.py succeeded or fell through to the warning path on this container's last build.
  2. If the patch fell through: identify what claude-cli / claude_code.py change caused the target strings to drift, and either upstream the patch or re-align the strings.
  3. If the patch succeeded: a real subprocess-stdout-inheritance bug. Likely candidates: an MCP server (this agent has gscServer, trinity connected — both shown as connected, but earlier sessions on other agents showed seranking and fibery-mcp-server as failed, suggesting MCP-init flakiness in the broader fleet).
  4. Add a hard timeout on the orchestrator-side reader thread distinct from the agent execution timeout, so a stuck reader doesn't hold the execution open until the 24-min mark and lose all telemetry.
  5. Surface non-null cost and model_used even on failed executions, so async failures don't disappear from spend tracking.

Workaround

Per the platform's own message, retry. But for long deliverables that's a 24-minute coin flip; sequential mode (parallel=false) may be more reliable for content-heavy briefs until this is resolved.

Filed by

GTM Agent on behalf of Oleksii Nikitin (o.nikitin@ability.ai) from Abilityai/trinity-gtm-agent workspace. Not retrying the original task pending resolution — Oleksii will run the brief manually via the website agent UI in the meantime.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions