Summary
When trinity-backend starts up, the Slack Socket Mode adapter makes one attempt (10 s timeout) per configured connection. If every initial attempt times out, the adapter logs Backend continues without Slack and never retries. The watchdog (which would normally bring a dead socket back up) only starts after a successful initial connect, so it can't help here. Result: a fully-running backend with Slack permanently offline until someone manually restarts the container.
Impact
- Severity: P1 (Slack-bound agents stop receiving messages with no operator-visible signal beyond a single ERROR line at startup)
- Reproduced on
[INSTANCE] 2026-05-07: backend restarted at 10:21:12 UTC, both initial Slack connect attempts timed out at 10 s, adapter gave up. Slack remained down for ~4 h 34 m until an operator noticed messages weren't reaching agents and restarted the backend manually. Slack edge was healthy the whole time — the very next attempt (after restart) succeeded in <1 s.
Reproduction
- Cause a transient network blip (or DNS slow-resolve, or Slack edge throttle) that exceeds the 10 s connect timeout during backend startup.
- Observe in
trinity-backend logs:
ERROR adapters.transports.slack_socket Slack Socket Mode [c=0]: connection timed out (10s). Check app token and network.
ERROR adapters.transports.slack_socket Slack Socket Mode: all connection attempts failed. Backend continues without Slack.
- Wait any amount of time. Slack never reconnects. No further log lines.
- Verify Slack edge is reachable from the backend container:
curl -m 5 https://wss-primary.slack.com returns 200 in ~80 ms.
- Restart
trinity-backend. Connection succeeds on first attempt (<1 s).
Location
docker/base-image/agent_server/adapters/transports/slack_socket.py — startup path in start() (~lines 150–172):
# Connect all clients in parallel — total startup time stays at
# ~10s ceiling rather than N * 10s sequential.
results = await asyncio.gather(
*(self._start_one_client(i) for i in range(n)),
return_exceptions=True,
)
for i, result in enumerate(results):
if isinstance(result, _ClientCtx):
self.contexts.append(result)
else:
logger.warning(f"Slack Socket Mode [c={i}]: did not connect")
if not self.contexts:
logger.error(
"Slack Socket Mode: all connection attempts failed. "
"Backend continues without Slack."
)
return # <-- silent, permanent giveup; nothing else tries again
The return here means no watchdog is spawned (watchdogs are created per-context just below this block) and no other code path in the adapter ever reattempts start().
Root cause
The watchdog model assumes the connection has been established at least once. There is no equivalent "connection-not-established" recovery loop. A 10 s startup timeout is treated as a permanent failure rather than a transient one.
Suggested fix
Two options, in increasing order of invasiveness:
Option A — bounded retry inside start() before giving up
Wrap the asyncio.gather(...) block in a retry loop, e.g. 3 attempts with 5 s / 15 s / 30 s sleep between them. If all 3 attempts fail, then log the error and return. Doesn't add startup latency in the happy path.
Option B — promote the watchdog to also handle "not yet connected"
Even on total initial failure, spawn a single "reconnect-on-startup" task that retries every 60 s (matching WATCHDOG_INTERVAL_SECONDS) with the same backoff schedule the watchdog already uses (60 → 120 → 240 → 300 s cap). Once any client connects, switch to the per-client watchdog model.
Option B is the more general fix — it eliminates the special case of "never connected" entirely. Option A is the smaller change and likely sufficient since the empirical failure mode is a single 10 s blip, not a sustained outage.
Defensive: the Task was destroyed but it is pending! log
After the timeout, the SDK leaks two process_messages tasks (asyncio garbage collector logs them ~5 minutes later). The _start_one_client cleanup path on timeout doesn't await/cancel the SDK's internal tasks. Worth tightening as a separate small PR.
Workaround
docker restart trinity-backend brings Slack back. ~10–30 s API downtime; agent containers unaffected. No state loss other than the in-flight HTTP requests.
Detection / monitoring (operator-side)
Currently the only signal is a single ERROR log line at startup. Suggested:
- Emit a
slack_socket_connected gauge (0/1) so dashboards / alerts can fire when it stays at 0.
/api/observability/status already exists — add Slack adapter state to it so /fleet-health surfaces the issue.
Related
- Builds on the watchdog work that handles silent socket death post-connect (existing).
- Companion forensic report (post-connect disconnect cadence): kept private in ops repo.
Summary
When
trinity-backendstarts up, the Slack Socket Mode adapter makes one attempt (10 s timeout) per configured connection. If every initial attempt times out, the adapter logsBackend continues without Slackand never retries. The watchdog (which would normally bring a dead socket back up) only starts after a successful initial connect, so it can't help here. Result: a fully-running backend with Slack permanently offline until someone manually restarts the container.Impact
[INSTANCE]2026-05-07: backend restarted at 10:21:12 UTC, both initial Slack connect attempts timed out at 10 s, adapter gave up. Slack remained down for ~4 h 34 m until an operator noticed messages weren't reaching agents and restarted the backend manually. Slack edge was healthy the whole time — the very next attempt (after restart) succeeded in <1 s.Reproduction
trinity-backendlogs:curl -m 5 https://wss-primary.slack.comreturns 200 in ~80 ms.trinity-backend. Connection succeeds on first attempt (<1 s).Location
docker/base-image/agent_server/adapters/transports/slack_socket.py— startup path instart()(~lines 150–172):The
returnhere means no watchdog is spawned (watchdogs are created per-context just below this block) and no other code path in the adapter ever reattemptsstart().Root cause
The watchdog model assumes the connection has been established at least once. There is no equivalent "connection-not-established" recovery loop. A 10 s startup timeout is treated as a permanent failure rather than a transient one.
Suggested fix
Two options, in increasing order of invasiveness:
Option A — bounded retry inside
start()before giving upWrap the
asyncio.gather(...)block in a retry loop, e.g. 3 attempts with 5 s / 15 s / 30 s sleep between them. If all 3 attempts fail, then log the error and return. Doesn't add startup latency in the happy path.Option B — promote the watchdog to also handle "not yet connected"
Even on total initial failure, spawn a single "reconnect-on-startup" task that retries every 60 s (matching
WATCHDOG_INTERVAL_SECONDS) with the same backoff schedule the watchdog already uses (60 → 120 → 240 → 300 s cap). Once any client connects, switch to the per-client watchdog model.Option B is the more general fix — it eliminates the special case of "never connected" entirely. Option A is the smaller change and likely sufficient since the empirical failure mode is a single 10 s blip, not a sustained outage.
Defensive: the
Task was destroyed but it is pending!logAfter the timeout, the SDK leaks two
process_messagestasks (asyncio garbage collector logs them ~5 minutes later). The_start_one_clientcleanup path on timeout doesn't await/cancel the SDK's internal tasks. Worth tightening as a separate small PR.Workaround
docker restart trinity-backendbrings Slack back. ~10–30 s API downtime; agent containers unaffected. No state loss other than the in-flight HTTP requests.Detection / monitoring (operator-side)
Currently the only signal is a single ERROR log line at startup. Suggested:
slack_socket_connectedgauge (0/1) so dashboards / alerts can fire when it stays at 0./api/observability/statusalready exists — add Slack adapter state to it so/fleet-healthsurfaces the issue.Related