Skip to content

bug(slack): startup connect failure leaves Slack permanently offline — no retry, watchdog never spawns #708

Description

@pavshulin

Summary

When trinity-backend starts up, the Slack Socket Mode adapter makes one attempt (10 s timeout) per configured connection. If every initial attempt times out, the adapter logs Backend continues without Slack and never retries. The watchdog (which would normally bring a dead socket back up) only starts after a successful initial connect, so it can't help here. Result: a fully-running backend with Slack permanently offline until someone manually restarts the container.

Impact

  • Severity: P1 (Slack-bound agents stop receiving messages with no operator-visible signal beyond a single ERROR line at startup)
  • Reproduced on [INSTANCE] 2026-05-07: backend restarted at 10:21:12 UTC, both initial Slack connect attempts timed out at 10 s, adapter gave up. Slack remained down for ~4 h 34 m until an operator noticed messages weren't reaching agents and restarted the backend manually. Slack edge was healthy the whole time — the very next attempt (after restart) succeeded in <1 s.

Reproduction

  1. Cause a transient network blip (or DNS slow-resolve, or Slack edge throttle) that exceeds the 10 s connect timeout during backend startup.
  2. Observe in trinity-backend logs:
    ERROR  adapters.transports.slack_socket  Slack Socket Mode [c=0]: connection timed out (10s). Check app token and network.
    ERROR  adapters.transports.slack_socket  Slack Socket Mode: all connection attempts failed. Backend continues without Slack.
    
  3. Wait any amount of time. Slack never reconnects. No further log lines.
  4. Verify Slack edge is reachable from the backend container: curl -m 5 https://wss-primary.slack.com returns 200 in ~80 ms.
  5. Restart trinity-backend. Connection succeeds on first attempt (<1 s).

Location

docker/base-image/agent_server/adapters/transports/slack_socket.py — startup path in start() (~lines 150–172):

# Connect all clients in parallel — total startup time stays at
# ~10s ceiling rather than N * 10s sequential.
results = await asyncio.gather(
    *(self._start_one_client(i) for i in range(n)),
    return_exceptions=True,
)

for i, result in enumerate(results):
    if isinstance(result, _ClientCtx):
        self.contexts.append(result)
    else:
        logger.warning(f"Slack Socket Mode [c={i}]: did not connect")

if not self.contexts:
    logger.error(
        "Slack Socket Mode: all connection attempts failed. "
        "Backend continues without Slack."
    )
    return    # <-- silent, permanent giveup; nothing else tries again

The return here means no watchdog is spawned (watchdogs are created per-context just below this block) and no other code path in the adapter ever reattempts start().

Root cause

The watchdog model assumes the connection has been established at least once. There is no equivalent "connection-not-established" recovery loop. A 10 s startup timeout is treated as a permanent failure rather than a transient one.

Suggested fix

Two options, in increasing order of invasiveness:

Option A — bounded retry inside start() before giving up

Wrap the asyncio.gather(...) block in a retry loop, e.g. 3 attempts with 5 s / 15 s / 30 s sleep between them. If all 3 attempts fail, then log the error and return. Doesn't add startup latency in the happy path.

Option B — promote the watchdog to also handle "not yet connected"

Even on total initial failure, spawn a single "reconnect-on-startup" task that retries every 60 s (matching WATCHDOG_INTERVAL_SECONDS) with the same backoff schedule the watchdog already uses (60 → 120 → 240 → 300 s cap). Once any client connects, switch to the per-client watchdog model.

Option B is the more general fix — it eliminates the special case of "never connected" entirely. Option A is the smaller change and likely sufficient since the empirical failure mode is a single 10 s blip, not a sustained outage.

Defensive: the Task was destroyed but it is pending! log

After the timeout, the SDK leaks two process_messages tasks (asyncio garbage collector logs them ~5 minutes later). The _start_one_client cleanup path on timeout doesn't await/cancel the SDK's internal tasks. Worth tightening as a separate small PR.

Workaround

docker restart trinity-backend brings Slack back. ~10–30 s API downtime; agent containers unaffected. No state loss other than the in-flight HTTP requests.

Detection / monitoring (operator-side)

Currently the only signal is a single ERROR log line at startup. Suggested:

  • Emit a slack_socket_connected gauge (0/1) so dashboards / alerts can fire when it stays at 0.
  • /api/observability/status already exists — add Slack adapter state to it so /fleet-health surfaces the issue.

Related

  • Builds on the watchdog work that handles silent socket death post-connect (existing).
  • Companion forensic report (post-connect disconnect cadence): kept private in ops repo.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions