You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
ops(hub): the engine-refresh campaign — bring the engine stack current through the pool's rolling deploy #1722
Problem — The hub's engine stack is three legs at three ages: engine binary from the Sep-10 M3 cutover (overlay last promoted Sep 13; the amicode delta itself July 15; upstream merged only through Aug 1), service dist Sep 15, app dist Oct 4. Since that stack froze, 241 commits landed on main, 49 of them touching the engine overlay — including fixes for exactly the failure class that has been killing the hub: bound and isolate background session warming (#1501, RSS-relevant), recover session reliability regressions (#1472), the never-blank stack (#1287), and the overlay sync onto upstream v1.18.30 (#1506). The watchdog has spent two consecutive nights restarting an old engine for bugs whose fixes are already merged and unshipped. The visible casualty is the harness picker (PR #1550's app leg deployed Oct 4; its service and engine legs never shipped — GET /amicode/harness answers the SPA fallback, so the chat-bar control silently renders nothing). Approach — One engine-refresh campaign: promote the overlay onto current upstream through the app-bundle refresh/materialize flow with its engine test, drift, and typecheck gates green, build the vendored binary via the release flow, ship the service dist from the same lineage, and deploy by rolling restart through the 3-shard pool (shards 2/3 first — empty, zero cost — then shard 1). The pool turns an engine upgrade from a fleet-wide kill event into a rolling maneuver; this is the first time a mid-week engine refresh is survivable on this hub. PR #1550's service + engine legs ride in the same lineage instead of becoming a fourth half-deploy. Scope — in: overlay promotion onto current upstream, the engine build, the service dist, PR #1550's missing legs, the rolling deploy, post-refresh verification. out: the pool itself (#1717 Phase 1 — PR #1721), per-shard DBs (#1717 Phase 2), adopting upstream features beyond the sync. Assumptions — The pool soaks clean tonight (per-shard trajectories show single-shard blast radius); PR #1721 merges before the refresh (the hub runs branch code); the app-bundle gates are the correctness bar for the overlay promotion; the release flow (amicode-release) is the build path.
Acceptance Criteria
The overlay manifest on main shows a fresh promotion: current upstream base, refreshed file SHAs, new promoted_at — and the engine test, drift, and overlay typecheck gates all pass.
A new vendored engine binary builds from that promotion and reports a version derived from the current lineage (not the v1.18.30-era string).
The service dist from the same lineage carries the harness route forward (commit a22007ee's seam or its merged successor); GET /amicode/harness on the hub answers JSON, not the SPA HTML fallback.
The harness picker renders in the VS Code chat bar on a fleet client, and a harness switch request round-trips (the picker's control + status chip both appear).
Rolling deploy order is 2 → 3 → 1: shards 2/3 restart to the new binary with zero session impact; shard 1's restart blips only its own pinned sessions; each shard returns to serving before the next restarts.
Post-refresh verification: per-shard watchdog trajectories clean for a soak night, and none of the kill events carry the old engine's now-fixed signatures (unbounded session warming, publisher fragment floods).
Testing Decisions
Reuse-first: the app-bundle gates (engine test, drift, overlay typecheck) for the promotion; the release flow's own smoke for the build; hub-upgrade-smoke for the hub deploy; the frontdoor suites from #1721 stay green (no router changes in this campaign). No new suites — the surfaces are all covered by existing gates.
Rolling order 2 → 3 → 1 — empty shards first, the shard holding all 1,274 pinned legacy sessions last, one shard verified live before the next restarts.
One lineage for all three legs — binary + service + app dist from the same promotion, deployed together; the freshness table in Notes becomes a deploy-record check, not a postmortem.
Constraints & Invariants
Deploys only through the ritual: backups, diff-first, hub-restart.sh/systemd-run — never an inline restart from an agent hosted on the hub.
Rollback = the .bak-20261006-pre-pool stack + a single-shard routing table; it must stay intact until the post-refresh soak passes.
The frontdoor router config (routing table, session map, Jev flag state) is untouched by the engine refresh — shard topology and placement survive the restarts.
The wedge-era watchdog stays armed throughout — per-shard probes must keep answering through the rollout.
Prior Art
The app-bundle scripts and gates (refresh manifest, materialize, engine/drift/typecheck gates) — the promotion tooling this campaign drives.
The release flow (amicode-release) — the vendored-binary build path.
Related to #1717 (the pool this rolls out through). Sequenced after: the #1721 merge and one clean pool soak night. Evidence recorded in the campaign ledger (session-20261005-jev-routed-pool).
Notes
The freshness table (measured 2026-10-06):
Leg
Hub state
Age
Engine binary
Sep-10 M3 cutover build, v1.18.30-era
overlay promoted Sep 13; delta branch July 15; upstream merged through Aug 1
Watchdog context: 9 kills overnight Oct 4/05, more Oct 5, another Oct 6 04:15 — all rss-over-threshold-and-http-silent on the old binary, several plausibly in classes the overlay fixes already address.
Important
Problem — The hub's engine stack is three legs at three ages: engine binary from the Sep-10 M3 cutover (overlay last promoted Sep 13; the amicode delta itself July 15; upstream merged only through Aug 1), service dist Sep 15, app dist Oct 4. Since that stack froze, 241 commits landed on main, 49 of them touching the engine overlay — including fixes for exactly the failure class that has been killing the hub: bound and isolate background session warming (#1501, RSS-relevant), recover session reliability regressions (#1472), the never-blank stack (#1287), and the overlay sync onto upstream v1.18.30 (#1506). The watchdog has spent two consecutive nights restarting an old engine for bugs whose fixes are already merged and unshipped. The visible casualty is the harness picker (PR #1550's app leg deployed Oct 4; its service and engine legs never shipped —
GET /amicode/harnessanswers the SPA fallback, so the chat-bar control silently renders nothing).Approach — One engine-refresh campaign: promote the overlay onto current upstream through the app-bundle refresh/materialize flow with its engine test, drift, and typecheck gates green, build the vendored binary via the release flow, ship the service dist from the same lineage, and deploy by rolling restart through the 3-shard pool (shards 2/3 first — empty, zero cost — then shard 1). The pool turns an engine upgrade from a fleet-wide kill event into a rolling maneuver; this is the first time a mid-week engine refresh is survivable on this hub. PR #1550's service + engine legs ride in the same lineage instead of becoming a fourth half-deploy.
Scope — in: overlay promotion onto current upstream, the engine build, the service dist, PR #1550's missing legs, the rolling deploy, post-refresh verification. out: the pool itself (#1717 Phase 1 — PR #1721), per-shard DBs (#1717 Phase 2), adopting upstream features beyond the sync.
Assumptions — The pool soaks clean tonight (per-shard trajectories show single-shard blast radius); PR #1721 merges before the refresh (the hub runs branch code); the app-bundle gates are the correctness bar for the overlay promotion; the release flow (amicode-release) is the build path.
Acceptance Criteria
a22007ee's seam or its merged successor);GET /amicode/harnesson the hub answers JSON, not the SPA HTML fallback.Testing Decisions
Reuse-first: the app-bundle gates (engine test, drift, overlay typecheck) for the promotion; the release flow's own smoke for the build;
hub-upgrade-smokefor the hub deploy; the frontdoor suites from #1721 stay green (no router changes in this campaign). No new suites — the surfaces are all covered by existing gates.Key Decisions
Constraints & Invariants
hub-restart.sh/systemd-run— never an inline restart from an agent hosted on the hub..bak-20261006-pre-poolstack + a single-shard routing table; it must stay intact until the post-refresh soak passes.Prior Art
Source
Related to #1717 (the pool this rolls out through). Sequenced after: the #1721 merge and one clean pool soak night. Evidence recorded in the campaign ledger (session-20261005-jev-routed-pool).
Notes
The freshness table (measured 2026-10-06):
live-test/harness-mainWatchdog context: 9 kills overnight Oct 4/05, more Oct 5, another Oct 6 04:15 — all
rss-over-threshold-and-http-silenton the old binary, several plausibly in classes the overlay fixes already address.