Skip to content

fix(ops): the RSS watchdog's blind pgrep pattern — match by binary path, not port - #1704

Open
aarontrowbridge wants to merge 6 commits into
mainfrom
ops/rss-watchdog-blind-pattern
Open

aarontrowbridge wants to merge 6 commits into
mainfrom
ops/rss-watchdog-blind-pattern

Conversation

@aarontrowbridge

Copy link
Copy Markdown
Member

What broke

The #775 wedge re-fired on erlich at ~2026-10-03T21:00Z and ran ~7 hours undetected: engine RSS 7.4 GB (threshold 1.5 GB), TCP accepts / HTTP silent, every client panel hanging at "booting". The canonical server restarted clean and all sessions are intact.

The watchdog was blind twice over:

  1. Wrong port — pattern opencode serve --port=4095; the M3 cutover (The hub cutover: main's architecture onto the production hub (spec-20260910-080000) #955) moved the engine to 4094.
  2. Wrong separator — the pattern says --port=4095 (equals), the engine cmdline is --port 4094 (space). pgrep -f could never have matched it on any port, before or after M3.

Every 5-minute sample since Oct 2 logged pid=none status=server-missing — and the watchdog exits 0 on that path, so nothing ever escalated. The deployed script also wasn't versioned anywhere, which is how it drifted.

The fix

  • Default WATCHDOG_PATTERN → binary path ($HOME/.amico/server/bin/opencode serve) — immune to port moves and separator style. Env override unchanged.
  • Version the script + rss-watchdog.{service,timer} under ops/ so the repo is the source of truth, with deploy + verify snippets in ops/fleet-watchdog/README.md.
  • ops/README.md gains the table row + a note that the on-erlich hotfix (applied live during the wedge, before this PR) is the one sanctioned exception to the never-edit-deployed-copies law.

Verify (already live on erlich)

$ tail -1 ~/.amico/server/fleet-watchdog/rss-trajectory.log
2026-10-04T04:14:50 pid=2063706 rss_kb=672656 threshold_kb=1536000

First real sample in two days. Incident capture (ps/threads/proc-status/log-tail/sha lines): ~/.amico/server/incidents/20261003-wedge/ on erlich.

Follow-up (not in this PR)

…th, not port

The #775 wedge re-fired 2026-10-03 and ran ~7 h undetected: the deployed
rss-watchdog (never versioned in ops/) matched 'opencode serve --port=4095'
— wrong port since the M3 cutover #955, and an equals-separator the engine's
space-separated cmdline ('--port 4094') could never match on ANY port.
Every 5-min sample logged pid=none status=server-missing since Oct 2.

Version the script + systemd pair here with the fixed default pattern
(binary path, immune to port moves and separator style) and document the
deploy + verify flow. Incident capture on erlich:
~/.amico/server/incidents/20261003-wedge/
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 76e9f882-8a15-4bc4-8a76-206e1d120ba8
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…ever restarted

The single-sample 1.5 GB threshold restart-stormed the hub 4x on
2026-10-04 (05:05/05:15/08:20/10:00): a healthy engine under active
session load rides 2.3-2.9 GB routinely. Each kill murdered in-flight
sessions mid-run (the very AFK-continuation behavior the fleet exists
to provide) and every stop risked the slow teardown outage window.

The #775 wedge is 'TCP accepts, HTTP silent' — RSS alone is not it.
RSS over threshold now triggers the /session corroboration probe first;
restart fires only on rss-over-threshold AND http-silent. Serving
engines log status=rss-high-but-serving and are left alone.

Same fix applied live to the deployed copy on erlich (the timer's
10:05 sample already ran the gated code path clean).
@aarontrowbridge

Copy link
Copy Markdown
Member Author

Follow-up (same day, ~5h later): the re-armed watchdog restart-stormed the hub 4x — 05:05, 05:15, 08:20, 10:00, each action=restart reason=rss-over-threshold at 2.3–2.9 GB. That's normal RSS under active session load, not the #775 leak. Each kill terminated in-flight sessions mid-run (the AFK-continuation behavior the fleet exists to provide) and the slow teardown made it read as 'stuck booting'.

Fix in d3bc4f8: the wedge is 'TCP accepts, HTTP silent' — RSS alone is not it. RSS over threshold now triggers a /session corroboration probe first; restart fires only on rss-over-threshold AND http-silent. A serving engine logs status=rss-high-but-serving and is never restarted. Knobs: WATCHDOG_HEALTH_PORT (4094), WATCHDOG_HEALTH_TIMEOUT (10s).

Deployed copy on erlich patched live (backup: rss-watchdog.sh.bak-20261004-single-sample-storm); the 10:05 timer sample already ran the gated path clean.

…atchdog wedge-capture companion)

Bun's JS inspector can pause a spinning main thread and report the exact
wedging JS frame — the one thing native gdb cannot read on the JIT
binary. Dependency-free CDP client (node >= 21 built-in WebSocket):
/json/list discovery -> Debugger.enable -> Debugger.pause -> print +
persist callFrames. Wired into the watchdog's wedge capture on erlich;
invocable by hand: node cdp-stack.mjs [port] [out-file].

Note: bun 1.4.2's inspector opens its port via BUN_INSPECT but registers
no discoverable page and prints no banner (no ws id) — /json/list stays
empty even under --inspect. The capture therefore currently degrades to
'no webSocketDebuggerUrl' on this bun; it is armed for the version where
discovery works, and the smaps/gdb legs of the wedge capture carry the
evidence load meanwhile.
The probe used /session — the heaviest route on the 5.4GB DB. Under the
user's legitimate 8-parallel-session load it can exceed the 10s probe
timeout while the engine is alive and working, and the corroboration
gate then killed healthy-but-saturated engines (the 04:45-05:25 series:
RSS climbing ~5MB/s under active work, strace showing a live task relay
— timers, pool threads, file reads — not a dead loop). The engine's
liveness is now measured by its cheapest substantive route (cached
per-directory config), and the single-sample RSS threshold rises to
2.5GB — the observed working set of this workload sits well above the
old 1.5GB rail.

The live-watch rig now records CHEAP vs HEAVY endpoint latency at
trigger (60s timeouts) — the next event definitively separates
saturation from a true wedge.
@aarontrowbridge

Copy link
Copy Markdown
Member Author

Another same-night lesson: the corroboration probe itself was the false-positive source. It used /session — the heaviest route on the 5.4 GB DB — with a 10 s timeout. Under legitimate 8-parallel-session load that route can exceed the timeout while the engine is alive and working (live strace of a "wedged" engine showed a functioning task relay: timers, pool threads, file reads — not a dead loop), so the corroboration gate killed healthy-but-saturated engines. The probe now hits the cheapest substantive route (cached per-directory config, ~30 ms idle) and the RSS rail rises to 2.5 GB — the workload's legitimate working set. A live-watch rig records cheap-vs-heavy endpoint latency (60 s timeouts) at the next event to definitively separate saturation from a true wedge.

… at 401MB

The #775-class wedge is not always a memory wedge: the 2026-10-07 07:55
event had the main thread parked in a timed futex (self-recovered after
~12 minutes when the wait expired), HTTP silent, RSS 401MB — far below
the 2.5GB rail, so the RSS gate would have left it wedged forever. The
watchdog now captures forensics and restarts on TWO consecutive silent
liveness probes regardless of RSS; a single silent tick only arms (one
blip must not kill in-flight sessions).
… markers + patterns

Two bugs bit together in the 2026-10-07 08:50 shard-1 wedge (5GB, 90% CPU,
HTTP dead ~40 min): (1) the http-silent marker path was SHARED across all
shard watchdog instances, so the healthy shards' serving ticks wiped the
wedged shard's marker every 5 minutes — it armed forever and never fired;
(2) the no-index default pgrep pattern matched ANY shard's engine, so the
shard-1 watchdog tracked shard 3's pid while probing shard 1's port. The
marker is now scoped per AMICODE_SHARD_INDEX and the default pattern pins
port 4094; all three instances verified ticking with their own pids.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant