Repository navigation
hub: server wedges (100% CPU, silent HTTP death) under direct fleet reconnect — front-door proxy deployed as mitigation; root cause open #775
Description
Activity
Update (2026-09-03 ~14:00) — mechanism confirmed as an SSE-churn leak; mitigation upgraded to an SSE splitter.
New evidence since filing:
- The v1 mitigation (connection-serializing proxy) re-wedged under load: RSS 718 MB → 2.35 GB → pegged. My own stress tests (~1,500 SSE lifecycles in 4 min) accelerated it — the positive signal was RSS trajectory, not HTTP responsiveness (+26 MB / 360 abandoned SSE connections in the harness, visible pre-wedge).
- A restart under the aligned retry herd hits 1.6 GB RSS at 21 s through the plain proxy — serialization alone is insufficient at boot.
Mechanism (for the root-cause work): each
/eventSSE connection allocatesQueue.unbounded(perpackages/opencode/src/server/routes/instance/httpapi/handlers/event.ts); abandoned connections do not appear to release it (finalizer-based cleanup), so every leaked queue grows forever on heartbeats + events. Churn rate ∝ death speed: direct-aligned ~90 s, direct-idle ~12 min, proxied-stressed ~10 min. The original instance's 5.1 GB RSS matches.Current mitigation (deployed, supervised): front door v2.1 — SSE splitter. One permanent upstream each for
/eventand/global/event, fan-out to all clients (joins get the captured preamble: headers +server.connected). Everything else per-connection passthrough. The server now sees exactly two SSE connections for its lifetime; panel reconnect storms are absorbed as group joins. Fleet (macbook + mini webviews) serving through it, RSS at the normal attached baseline.Artifacts added to
~/.amico/server/incidents/20260903-wedge/: the v1 proxy + updated splitter (hub-frontdoor.pyon the hub), RSS trajectories in the ledger (vault-aaron/sessions/session-20260903-hub-wedge-forensics.md).Root-cause AC unchanged; note the fix target is now precise: the abandoned-SSE queue cleanup path.
Recurrence 2026-09-04 (~13:30 EDT) — the wedge hit WITH the front door up. Evidence for the root-cause analysis:
- Server RSS 3.70–3.77 GB after ~21 h uptime, ~19.8% CPU, then wedged: TCP accepts, HTTP never answers, journal silent for the final 2 h (snapshot: erlich
~/.amico/server/incidents/20260904-wedge2/— ps snapshot, service status, journal tail). - Wedged for all external clients (MacBook AND Mac mini both time out on 4095/4096); localhost on erlich also hung for HTTP; sshd unaffected (separate process).
- Key data point: the front-door SSE splitter does NOT prevent this wedge. The splitter eliminated connection churn as the feed, yet the queue leak still grew to wedge over ~21 h — meaning the leak has feed sources beyond abandoned-connection churn (heartbeat/event accumulation in leaked queues, possibly hub-hosted agent sessions' event volume — the 2026-09-03 forensics noted 947 K event rows / 1.55 GB).
- Remediation:
systemctl --user restart co.harmoniqs.amicode-server.service→ immediate recovery (200s on all paths in <20 ms, fresh process at ~414 MB).
Implication: the front door buys stability against reconnect storms but is not a fix — the unbounded per-connection event queue (or equivalent) still wedges the server on a ~daily cadence under real load. Priority on finding the queue lifecycle leak itself.
- Server RSS 3.70–3.77 GB after ~21 h uptime, ~19.8% CPU, then wedged: TCP accepts, HTTP never answers, journal silent for the final 2 h (snapshot: erlich
Watchdog-caught feed event (2026-09-04 14:35–14:40): attach+boot against a 1,449-session store. After the store merge, a client reloaded into fleet mode: RSS went 408 MB → 2.23 GB in ~5 minutes during the panel attach + app boot (watchdog trajectory log, erlich
fleet-watchdog/rss-trajectory.log), wedged the server, and the 5-min RSS watchdog auto-restarted it — fleet recovered without human touch. New data: the leak feeds violently on attach/boot queries against a large store (home/session-list + message hydration scale with session count now). This is the strongest argument yet for (a) the #775 root-cause fix and (b) boot-path query budgets (the app should not hydrate unbounded history on attach).- added 5 commits that reference this issue
on Sep 19, 2026 - added 17 commits that reference this issue
on Oct 4, 2026
Incident
2026-09-03, erlich (canonical hub). Every direct server boot from 11:13 wedged: process ~100% CPU, TCP accepts but HTTP never answers, logs stop, every fleet client panel stuck at boot. Downtime ~11:25–13:29.
Two boot signatures, both fatal:
4971e93a(local dev build,0.0.0--202609031509): 12MaxListenersExceededWarningdumps in minute 1 (EffectrunLoop/addListenerstacks), wedge ~T+12 min, RSS 5.1 GB..18release (8b1b1aad, from v1.18.10-amicode.18): silent, wedge ~T+60–90 s.The 4-day survivor (
c99d6eab) predates the era and was destroyed by the 11:12hub-restart.sh swap(no archive — separate hardening issue).Reproduction matrix (all NEGATIVE — see vault
sessions/session-20260903-hub-wedge-forensics.md)On the
.18release binary with the faithful hub env (spliced config, staging cwd, live DB):/config,/config/providers,/agent?directory=…,/command?directory=…; full 30 MB message-stream fetch of the wedging-era session; 60-socket true-simultaneous SSE storms; 96 authed (real fleet token) SSE connect-abandons; 360 interleaved threaded SSE lifecycles — healthy through all.So the trigger needs the real clients' arrival pattern (macbook panel + the mini's webviews: two VS Code versions,
/global/event×2, fleet-token authed), not just their requests.Current mitigation (deployed, supervised)
co.harmoniqs.amicode-server.servicepinned to 4095 (drop-inport-pin.conf), running the .18 release binary.co.harmoniqs.amicode-frontdoor.service:~/.amico/server/hub-frontdoor.py, 0.0.0.0:4096 → 127.0.0.1:4095,Restart=always, request log~/.amico/server/frontdoor.log.Artifacts (erlich
~/.amico/server/incidents/20260903-wedge/)server.log(warning stacks), wedged process/journal snapshots, the preserved bad binaryopencode.bad-4971e93a,capture-window.log(full proxy capture of real fleet traffic), harness scripts (storm.py,interleave.py,authtest.py,hub-proxy.py).Suspects
GlobalBuslistener/queue lifecycle on abandoned authed streams — the 5.1 GB RSS + the warning stacks point at event-queue buildup on abandoned subscribers (packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts— eagerevents.listen+ finalizer-based cleanup, upstream lineagecb35493242).Acceptance criteria