Skip to content

PRD: Fleet thin client — local shell, remote data #792

Description

@aarontrowbridge

PRD — Fleet thin client: local shell, remote data

Important

Problem

In fleet mode the panel is an iframe served by the hub: every UI asset, API call, and SSE handshake crosses the WAN link. On high-latency links (plane wifi, 750 ms RTT) the UI is structurally unusable — ~2.5 s per request against a 1.5 s attach budget (#777) — and the hub's request surface absorbs every client's render traffic, feeding the reconnect-storm class of incident (2026-09-03 wedge, §3.1).

Approach

Local shell, remote data. The client's vendored binary already ships the complete web app; in fleet mode it serves the UI locally and proxies the data plane (REST API + the two event streams) to the canonical hub, injecting fleet auth. The hub degrades to what it should be: a session store plus streams. The Sessions dropdown lists hub sessions natively (supersedes #779, which is held).

Approaches Considered

  • Relay mode in the local vendored server (chosen) — the server already owns HTTP, static assets, and SSE; one upstream connection per event endpoint (the front-door splitter pattern, mirrored client-side); extension-host reloads become relay-internal reattaches, never hub connection churn.
  • Extension-host proxy — zero fork work, but wedges a long-lived SSE multiplexer into the extension host process, which restarts on every window reload (churn = the wedge's favorite food).
  • Hub-side-only fixes (front door, budgets) — necessary but cannot fix render-over-WAN; the UI itself must move local.

Scope

In: a relay mode in the local server (serve app + path-preserving data proxy + auth injection); extension changes to spawn the local server in relay mode on fleet clients and point the panel at localhost; posture reporting (ties into #780's attach-state file); honest degraded state when the hub is unreachable.
Out: any session-data migration or mirroring (stores stay where they are); offline standalone continuity (deferred — "sessions-to-go" export is the future candidate); changing the hub front door; changes for standalone (non-fleet) machines; the hub-hosted agent-session containment question (§3.5).

Assumptions / Open Qs

  • Assumed the app's data fetches are same-origin relative paths (SDK baseUrl is configurable — a same-origin relay needs zero app changes; to be verified in the first slice).
  • Open: does relay mode fall back to the local store when the hub is unreachable (a "degraded but attached" posture), or fail honestly? Recommendation: fail honestly in v1 — dual-store fallback creates the confusion Read-only Fleet Sessions view: list hub sessions from a standalone client #779 was filed to avoid; revisit with sessions-to-go.
  • Open: SSE chain length (app → relay → tunnel → front door → server) — acceptable? The relay holds one upstream per endpoint, so it composes with the splitter rather than multiplying connections.

User Stories

  1. Plane wifi (the origin incident): Aaron opens Amicode on the MacBook over 750 ms wifi; the UI renders instantly from localhost; session data loads through the relay; SSE events arrive ~750 ms late but the studio is fully usable.
  2. Back at the hub: at home (26 ms), the same posture works with imperceptible latency; nothing changes because the link got faster.
  3. Hub goes down: the panel renders locally, then shows the honest posture (from Surface machine posture (hostname, hub reachability) in agent context #780's attach-state file): "canonical hub unreachable — data plane down." No fake local session list, no parked state machine.
  4. Reload the window: the extension host restarts; the relay reattaches its upstreams; the hub sees group joins, not a connection storm.

Modules & Interfaces

  • Fork (harmoniqs/opencode): relay mode on serve — flag takes the hub base URL; serves the app's static assets from the local binary; proxies data-plane paths (API + /event + /global/event) to the hub; injects fleet auth on proxied requests; SSE fan-out holds one upstream per event endpoint. Path-preserving; no app changes.
  • Extension (this repo): fleet-client attach flow spawns the local binary in relay mode (reusing ServerManager lifecycle + the guard), points the panel iframe at the local server instead of the tunnel URL; the tunnel remains the transport under the relay. Posture (Surface machine posture (hostname, hub reachability) in agent context #780) gains a relay mode state.
  • Auth: the relay injects the fleet token / per-boot password on the data plane; the app and the panel never handle credentials (Credential portability: take/push provider creds between client and hub over SSH #782's key-handling rule applies).

Testing Decisions

  • Relay integration test: boot a real local server in relay mode against a stub hub; assert UI served locally, API proxied path-preservingly, auth injected, SSE joined/reattached once.
  • Churn test: N extension-host restarts against a counting stub hub → exactly 2 SSE upstreams total (one per endpoint), not 2N.
  • Version-skew gate: relay refuses to start when client and hub pins disagree beyond the drift-gate tolerance (the brainstorm's pin-parity question, answered mechanically).

Risks

Source

Activity

  1. jeonghun-jj-lee commented on Sep 18, 2026

    @jeonghun-jj-lee
    Contributor

    Updated 2026-09-18 — reconciled after the durable-hub ADR 0024 was withdrawn (#1259 closed as superseded by #792/#1260; ADR 0024 is now 0024-fleet-thin-client-pluggable-transport). Corrected two engine-side file refs; added the mutation-over-tunnel (F14) and /global/event points.

    Gap analysis from a thin-client + transport investigation (with @jeonghun-jj-lee). #792 is the thin-client design of record — recording where the current build has drifted from it since this PRD was written (2026-09-04), plus the transport it deferred. Nothing here supersedes the PRD; these are amendments to fold in at decomposition.

    What changed under the PRD. The interface/server decoupling (#451/#822/#391/#955) moved /amicode/* out of the fork and into the extension-host amicode_service, which now owns those routes locally and reverse-proxies only the engine routes. So the relay this PRD describes is largely realized by amicode_service — but as an extension-host proxy, the approach this PRD's "Approaches Considered" explicitly rejected over SSE-multiplexer churn on window reload. That reconciliation is amendment B.

    A. /amicode/* must proxy to the host (host owns all state) — decided.
    Today amicode_service owns /amicode/* locally — registered routes served locally, the rest 404 at server.ts:247-250 — and never proxies them, in any mode. That is split-state: the client shows its own problems/catalog/runs while chat + solves live on the hub — for a research tool the Run Inspector desyncs from where compute happens. Decision (with @jeonghun-jj-lee): in fleet mode the client proxies /amicode/* to the host's authoritative amicode_service (running over the host's ~/.amico), so the host truly owns all state — this supersedes the PRD's "the hub degrades to … a session store plus streams." Changes the server.ts:247 precedence for fleet mode.
    Mutation safety (the reason this is sound over a tunnel): the credential/solver mutation routes (POST /amicode/connections, solver-mode) then travel client→host, but the host's amicode_service binds 127.0.0.1 (server.ts:319), so the loopback mutation guard (connections.ts:1236) is satisfied. This is the same inject-the-credential, keep-loopback model the #1259 review confirmed as ADR-0002/0005-consistent (anonymous-loopback was rejected): the relay injects the credential, the bind never leaves loopback.

    B. Re-base onto amicode_service; answer the churn objection.
    The PRD's "relay in the fork's serve" is superseded by the extension-host amicode_service — the approach the PRD rejected for reload churn. Either document how the churn is handled (does the service survive a window reload, or does #775's front-door splitter absorb it?), or move the relay to survive reload the way the engine does (ADR 0020). This decision picks the relay's locus, and A/C/D all attach to whatever process holds the relay — so it should be settled before decomposition.

    C. WebSocket / PTY upgrade proxying is missing.
    The data plane in scope is "API + /event + /global/event" — no WebSocket. amicode_service has no upgrade handler (server.ts:295) and both proxies are body-pipe-only, so the integrated terminal (a WS PTY route) is dead on a thin client. Add upgrade proxying to the relay's scope.

    D. SSE resumption (event loss, distinct from churn).
    The PRD covers connection churn (one upstream per endpoint, reattach). It does not cover event loss: the client's /event stream sends no Last-Event-ID (sse_client.ts:123-126) and the engine emits id: undefined (packages/app-bundle/.materialized/packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts:16 — stock opencode in the materialized base tree, not the overlay), so a blip drops every event in the gap. The engine already has a resumable per-session stream (/api/session/{id}/event?after=) that is unused. This applies to both streams the PRD proxies — /event and /global/event. Overlaps #775's root-cause surface (the GlobalBus queue on abandoned subscribers) and #415 (SSE resilience) — scope this to cursor emit + client resume, deferring subscriber-lifecycle cleanup to #775.

    E. Transport — the part #792 deferred → ADR 0024 + #1260.
    This PRD keeps "the tunnel remains the transport." That single SSH -L forward is the shared cause of #777 (unattachable at high RTT) and composes with #775. ADR 0024-fleet-thin-client-pluggable-transport makes the transport a pluggable provider — ssh (default), tailscale (opt-in; serve onto the loopback app so the bind guard stays intact), direct — tracked in #1260. Tailscale is never mandatory; SSH stays the floor.

    Related (link, don't duplicate): #1227 (silent local-shard spawn) — the attach-not-spawn model is exactly what prevents it; #781 (Go-Standalone hardening) — the honest-degraded posture is its seam.

    A–D are #792's to fold in when it is decomposed; E is tracked in #1260 and linked here.

  2. added 9 commits that reference this issue on Sep 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions