Skip to content

hub: server wedges (100% CPU, silent HTTP death) under direct fleet reconnect — front-door proxy deployed as mitigation; root cause open #775

Description

@aarontrowbridge

Incident

2026-09-03, erlich (canonical hub). Every direct server boot from 11:13 wedged: process ~100% CPU, TCP accepts but HTTP never answers, logs stop, every fleet client panel stuck at boot. Downtime ~11:25–13:29.

Two boot signatures, both fatal:

  • 4971e93a (local dev build, 0.0.0--202609031509): 12 MaxListenersExceededWarning dumps in minute 1 (Effect runLoop/addListener stacks), wedge ~T+12 min, RSS 5.1 GB.
  • .18 release (8b1b1aad, from v1.18.10-amicode.18): silent, wedge ~T+60–90 s.

The 4-day survivor (c99d6eab) predates the era and was destroyed by the 11:12 hub-restart.sh swap (no archive — separate hardening issue).

Reproduction matrix (all NEGATIVE — see vault sessions/session-20260903-hub-wedge-forensics.md)

On the .18 release binary with the faithful hub env (spliced config, staging cwd, live DB):

  • idle poll 15 min; SSE reconnect churn (hundreds of connect/abandons); /config, /config/providers, /agent?directory=…, /command?directory=…; full 30 MB message-stream fetch of the wedging-era session; 60-socket true-simultaneous SSE storms; 96 authed (real fleet token) SSE connect-abandons; 360 interleaved threaded SSE lifecycles — healthy through all.
  • Real fleet traffic, direct → wedges ≤ 90 s.
  • Real fleet traffic through a connection-serializing Python proxy (4096→4095) → healthy for hours, including across a proxy restart (reconnect storm survived).

So the trigger needs the real clients' arrival pattern (macbook panel + the mini's webviews: two VS Code versions, /global/event ×2, fleet-token authed), not just their requests.

Current mitigation (deployed, supervised)

  • co.harmoniqs.amicode-server.service pinned to 4095 (drop-in port-pin.conf), running the .18 release binary.
  • co.harmoniqs.amicode-frontdoor.service: ~/.amico/server/hub-frontdoor.py, 0.0.0.0:4096 → 127.0.0.1:4095, Restart=always, request log ~/.amico/server/frontdoor.log.
  • Removal of the front door is gated on this issue closing.

Artifacts (erlich ~/.amico/server/incidents/20260903-wedge/)

server.log (warning stacks), wedged process/journal snapshots, the preserved bad binary opencode.bad-4971e93a, capture-window.log (full proxy capture of real fleet traffic), harness scripts (storm.py, interleave.py, authtest.py, hub-proxy.py).

Suspects

  • SSE accept path under multi-connection arrival (interleaved accepts + requests).
  • GlobalBus listener/queue lifecycle on abandoned authed streams — the 5.1 GB RSS + the warning stacks point at event-queue buildup on abandoned subscribers (packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts — eager events.listen + finalizer-based cleanup, upstream lineage cb35493242).

Acceptance criteria

  • Deterministic offline repro (harness scripts + capture-window.log provided) or a named code path with a test.
  • Fix lands with a regression test exercising the real arrival pattern.
  • Front door removed; hub serves direct through a soak (idle + churn + storm) and a live fleet week.

Activity

  1. aarontrowbridge commented on Sep 3, 2026

    @aarontrowbridge
    MemberAuthor

    Update (2026-09-03 ~14:00) — mechanism confirmed as an SSE-churn leak; mitigation upgraded to an SSE splitter.

    New evidence since filing:

    • The v1 mitigation (connection-serializing proxy) re-wedged under load: RSS 718 MB → 2.35 GB → pegged. My own stress tests (~1,500 SSE lifecycles in 4 min) accelerated it — the positive signal was RSS trajectory, not HTTP responsiveness (+26 MB / 360 abandoned SSE connections in the harness, visible pre-wedge).
    • A restart under the aligned retry herd hits 1.6 GB RSS at 21 s through the plain proxy — serialization alone is insufficient at boot.

    Mechanism (for the root-cause work): each /event SSE connection allocates Queue.unbounded (per packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts); abandoned connections do not appear to release it (finalizer-based cleanup), so every leaked queue grows forever on heartbeats + events. Churn rate ∝ death speed: direct-aligned ~90 s, direct-idle ~12 min, proxied-stressed ~10 min. The original instance's 5.1 GB RSS matches.

    Current mitigation (deployed, supervised): front door v2.1 — SSE splitter. One permanent upstream each for /event and /global/event, fan-out to all clients (joins get the captured preamble: headers + server.connected). Everything else per-connection passthrough. The server now sees exactly two SSE connections for its lifetime; panel reconnect storms are absorbed as group joins. Fleet (macbook + mini webviews) serving through it, RSS at the normal attached baseline.

    Artifacts added to ~/.amico/server/incidents/20260903-wedge/: the v1 proxy + updated splitter (hub-frontdoor.py on the hub), RSS trajectories in the ledger (vault-aaron/sessions/session-20260903-hub-wedge-forensics.md).

    Root-cause AC unchanged; note the fix target is now precise: the abandoned-SSE queue cleanup path.

  2. aarontrowbridge commented on Sep 4, 2026

    @aarontrowbridge
    MemberAuthor

    Recurrence 2026-09-04 (~13:30 EDT) — the wedge hit WITH the front door up. Evidence for the root-cause analysis:

    • Server RSS 3.70–3.77 GB after ~21 h uptime, ~19.8% CPU, then wedged: TCP accepts, HTTP never answers, journal silent for the final 2 h (snapshot: erlich ~/.amico/server/incidents/20260904-wedge2/ — ps snapshot, service status, journal tail).
    • Wedged for all external clients (MacBook AND Mac mini both time out on 4095/4096); localhost on erlich also hung for HTTP; sshd unaffected (separate process).
    • Key data point: the front-door SSE splitter does NOT prevent this wedge. The splitter eliminated connection churn as the feed, yet the queue leak still grew to wedge over ~21 h — meaning the leak has feed sources beyond abandoned-connection churn (heartbeat/event accumulation in leaked queues, possibly hub-hosted agent sessions' event volume — the 2026-09-03 forensics noted 947 K event rows / 1.55 GB).
    • Remediation: systemctl --user restart co.harmoniqs.amicode-server.service → immediate recovery (200s on all paths in <20 ms, fresh process at ~414 MB).

    Implication: the front door buys stability against reconnect storms but is not a fix — the unbounded per-connection event queue (or equivalent) still wedges the server on a ~daily cadence under real load. Priority on finding the queue lifecycle leak itself.

  3. aarontrowbridge commented on Sep 4, 2026

    @aarontrowbridge
    MemberAuthor

    Watchdog-caught feed event (2026-09-04 14:35–14:40): attach+boot against a 1,449-session store. After the store merge, a client reloaded into fleet mode: RSS went 408 MB → 2.23 GB in ~5 minutes during the panel attach + app boot (watchdog trajectory log, erlich fleet-watchdog/rss-trajectory.log), wedged the server, and the 5-min RSS watchdog auto-restarted it — fleet recovered without human touch. New data: the leak feeds violently on attach/boot queries against a large store (home/session-list + message hydration scale with session count now). This is the strongest argument yet for (a) the #775 root-cause fix and (b) boot-path query budgets (the app should not hydrate unbounded history on attach).

  4. added 5 commits that reference this issue on Sep 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions