Summary
Async chat_with_agent (parallel=true, async=true) on ability-website ran for ~24 minutes through 56 turns / 135 raw messages, then completed with status=failed, response=null, and all telemetry (cost, context_used, model_used) null. The platform's own error message describes a stdout-inheritance / reader-thread failure mode and asks for it to be filed.
Execution metadata
| Field |
Value |
execution_id |
FdEVIjC5ZRVdeG-e5q0coQ |
agent_name |
ability-website |
triggered_by |
mcp |
started_at |
2026-05-06T11:58:54.069960Z |
completed_at |
2026-05-06T12:22:40.963313Z |
duration_ms |
1,426,893 (~23.8 min) |
status |
failed |
response |
null |
cost |
null |
context_used |
null |
model_used |
null |
execution_log |
null (also lost) |
Error message (verbatim from get_execution_result)
Execution completed without a result message after 0 tool calls / 56 turns (raw_messages=135). Likely cause: a tool or child subprocess inherited stdout and prevented the claude reader thread from capturing the final result block. Check agent-server logs for 'Reader thread(s) stuck after process exit' or 'I/O operation on closed file' near this execution. This is a transient infrastructure failure; retry the task.
The "0 tool calls / 56 turns" is suspicious — 56 turns with 0 tool calls is unusual for a real build task. Either the tool-call counter wasn't incremented (consistent with a reader-thread failure) or the agent really did spend 56 turns thinking without acting. The 135 raw messages suggests the former.
Reproduction context
- Inbound: a single, long-form async
chat_with_agent message — ~10 KB of structured markdown (build brief for a Next.js landing page, includes copy blocks, table-formatted form schema, GTM event list, acceptance criteria).
- Mode:
parallel=true, async=true, timeout_seconds=3600
- Caller: separate Trinity instance (cross-instance MCP call from
Abilityai/trinity-gtm-agent workspace).
- Agent received the message and started running (per
started_at and turn count). No final result block ever surfaced to the orchestrator.
Container-log signal possibly worth investigating
The startup logs for ability-website contain this warning before the run:
WARNING: Could not find target strings in /app/agent_server/services/claude_code.py
Agent server restarted with CLAUDECODE fix applied
The runtime appears to apply a hot-patch to claude_code.py at container start. If the patch's target strings don't match (e.g. after a claude-cli upgrade), the patch fails-open and the runtime continues without the fix. If the "CLAUDECODE fix" is the same code path that handles result-block capture / reader-thread cleanup, that's a candidate root cause — not a "transient" failure, but a regression triggered by a stale patch shim.
The patch warning is logged at WARNING level and doesn't block container startup, which means failures like this one will silently happen on every long execution until the patch is re-aligned or the underlying logic is upstreamed.
Impact
- Cost / observability: ~24 min of compute (Sonnet-4.6) lost, with no cost record, no transcript, no usable diagnostic data beyond the meta-error string. Hard to reason about whether the actual deliverable was produced and lost vs. never finished.
- Trust in async mode: any long-running build/research delegation via
parallel=true, async=true could silently fail this way after 20+ min. For caller agents this is the worst failure mode — long wait, no signal, no retry guidance beyond the platform's own (correct) "retry."
- Telemetry:
cost/model_used being null when the agent clearly burned tokens means this failure mode also isn't billable and isn't surfacing to whatever rate-limit / spend dashboards exist.
Suggested investigation
The platform's own error string already names the log lines to grep:
Reader thread(s) stuck after process exit
I/O operation on closed file
Around execution_id=FdEVIjC5ZRVdeG-e5q0coQ (12:22:40Z 2026-05-06).
Concretely:
- Check whether the "CLAUDECODE fix" hot-patch in
/app/agent_server/services/claude_code.py succeeded or fell through to the warning path on this container's last build.
- If the patch fell through: identify what claude-cli / claude_code.py change caused the target strings to drift, and either upstream the patch or re-align the strings.
- If the patch succeeded: a real subprocess-stdout-inheritance bug. Likely candidates: an MCP server (this agent has
gscServer, trinity connected — both shown as connected, but earlier sessions on other agents showed seranking and fibery-mcp-server as failed, suggesting MCP-init flakiness in the broader fleet).
- Add a hard timeout on the orchestrator-side reader thread distinct from the agent execution timeout, so a stuck reader doesn't hold the execution open until the 24-min mark and lose all telemetry.
- Surface non-null
cost and model_used even on failed executions, so async failures don't disappear from spend tracking.
Workaround
Per the platform's own message, retry. But for long deliverables that's a 24-minute coin flip; sequential mode (parallel=false) may be more reliable for content-heavy briefs until this is resolved.
Filed by
GTM Agent on behalf of Oleksii Nikitin (o.nikitin@ability.ai) from Abilityai/trinity-gtm-agent workspace. Not retrying the original task pending resolution — Oleksii will run the brief manually via the website agent UI in the meantime.
Summary
Async
chat_with_agent(parallel=true, async=true) onability-websiteran for ~24 minutes through 56 turns / 135 raw messages, then completed withstatus=failed,response=null, and all telemetry (cost, context_used, model_used) null. The platform's own error message describes a stdout-inheritance / reader-thread failure mode and asks for it to be filed.Execution metadata
execution_idFdEVIjC5ZRVdeG-e5q0coQagent_nameability-websitetriggered_bymcpstarted_at2026-05-06T11:58:54.069960Zcompleted_at2026-05-06T12:22:40.963313Zduration_ms1,426,893(~23.8 min)statusfailedresponsenullcostnullcontext_usednullmodel_usednullexecution_lognull(also lost)Error message (verbatim from
get_execution_result)The "0 tool calls / 56 turns" is suspicious — 56 turns with 0 tool calls is unusual for a real build task. Either the tool-call counter wasn't incremented (consistent with a reader-thread failure) or the agent really did spend 56 turns thinking without acting. The 135 raw messages suggests the former.
Reproduction context
chat_with_agentmessage — ~10 KB of structured markdown (build brief for a Next.js landing page, includes copy blocks, table-formatted form schema, GTM event list, acceptance criteria).parallel=true, async=true, timeout_seconds=3600Abilityai/trinity-gtm-agentworkspace).started_atand turn count). No final result block ever surfaced to the orchestrator.Container-log signal possibly worth investigating
The startup logs for
ability-websitecontain this warning before the run:The runtime appears to apply a hot-patch to
claude_code.pyat container start. If the patch's target strings don't match (e.g. after aclaude-cliupgrade), the patch fails-open and the runtime continues without the fix. If the "CLAUDECODE fix" is the same code path that handles result-block capture / reader-thread cleanup, that's a candidate root cause — not a "transient" failure, but a regression triggered by a stale patch shim.The patch warning is logged at WARNING level and doesn't block container startup, which means failures like this one will silently happen on every long execution until the patch is re-aligned or the underlying logic is upstreamed.
Impact
parallel=true, async=truecould silently fail this way after 20+ min. For caller agents this is the worst failure mode — long wait, no signal, no retry guidance beyond the platform's own (correct) "retry."cost/model_usedbeing null when the agent clearly burned tokens means this failure mode also isn't billable and isn't surfacing to whatever rate-limit / spend dashboards exist.Suggested investigation
The platform's own error string already names the log lines to grep:
Around
execution_id=FdEVIjC5ZRVdeG-e5q0coQ(12:22:40Z 2026-05-06).Concretely:
/app/agent_server/services/claude_code.pysucceeded or fell through to the warning path on this container's last build.gscServer,trinityconnected — both shown asconnected, but earlier sessions on other agents showedserankingandfibery-mcp-serverasfailed, suggesting MCP-init flakiness in the broader fleet).costandmodel_usedeven on failed executions, so async failures don't disappear from spend tracking.Workaround
Per the platform's own message, retry. But for long deliverables that's a 24-minute coin flip; sequential mode (
parallel=false) may be more reliable for content-heavy briefs until this is resolved.Filed by
GTM Agent on behalf of Oleksii Nikitin (
o.nikitin@ability.ai) fromAbilityai/trinity-gtm-agentworkspace. Not retrying the original task pending resolution — Oleksii will run the brief manually via the website agent UI in the meantime.