Skip to content

fix(windows): serve over loopback ZMQ and a selector event loop - #580

Draft
YevheniiKotyrlo wants to merge 6 commits into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-windows-runtime
Draft

YevheniiKotyrlo wants to merge 6 commits into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-windows-runtime

Conversation

@YevheniiKotyrlo

@YevheniiKotyrlo YevheniiKotyrlo commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Part of #384

Problem

With the build fixed (#575), ft serve on Windows starts and then answers nothing:

  1. Every link between the frontend, tokenizer and scheduler is ipc:///tmp/freetoken_N..., and the libzmq pyzmq ships for Windows has no ipc transport: zmq.has("ipc") is False and a bind fails with Protocol not supported (pyzmq 27.2.0, libzmq 4.3.5).
  2. uvicorn runs on the default Proactor event loop, and zmq.asyncio needs add_reader, which that loop does not implement: RuntimeError: Proactor event loop does not implement add_reader family of methods required for zmq.
  3. That error kills the frontend's listener task, which _create_listener_once creates and drops, so nothing observes it: the server stays up, accepts every request and answers none of them.
  4. When the backend dies, the server stops itself with os.kill(os.getpid(), signal.SIGTERM). On Windows that is TerminateProcess: no handler runs, so uvicorn's lifespan never terminates the remaining workers and they outlive the server. A Flash-Next serve whose scheduler died on a CUDA error left its detokenizer worker running with 1.9 GB.

Separately, every load asks torch for expandable segments, which torch compiles only off Windows (if(NOT WIN32) in c10/cuda/CMakeLists.txt): the engine logs Enabled expandable_segments, and the first allocation then warns expandable_segments not supported on this platform and runs without them. The request goes through torch.cuda.memory._set_allocator_settings, which torch 2.11 deprecates, so every start prints a FutureWarning, on Linux too.

Solution

  • SchedulerConfig chooses its five links once, in the parent (_choose_zmq_links): the same ipc paths as today where libzmq has ipc, loopback TCP ports where it does not. The links travel with the config to every spawned worker.
  • uvicorn runs on asyncio:SelectorEventLoop on Windows (UVICORN_LOOP), in both the serve and the shell path.
  • The listener task is held and given a done-callback: if it ends while the server is not shutting down, the server records why (fatal_error, maintenance_state = "failed") and exits the way it already does when the backend dies. This part is not Windows-specific - it is how the failure in (2) surfaced as silence.
  • That exit raises SIGTERM in-process (signal.raise_signal) instead of os.kill-ing our own pid, so uvicorn's handler runs the lifespan shutdown that terminates the workers. On Linux the two are the same; on Windows only the first runs a handler.
  • _ensure_expandable_segments leaves the allocator alone on Windows instead of requesting what torch does not build there, and elsewhere calls torch._C._accelerator_setAllocatorSettings, the binding the deprecated wrapper forwards to and the warning names.

pyzmq's own message suggests installing tornado to keep the Proactor loop; that adds a dependency and a selector thread per loop, where the selector loop needs neither.

Tests

tests/scheduler/test_zmq_links.py: without ipc every link is a distinct loopback port; with ipc the links are the per-process socket paths they were; a pickled config keeps its links on either transport (what a spawned worker receives).

tests/server/test_supervisor.py: a backend death stops the server through its own SIGTERM handler - a child process installs one and waits for it on a wakeup socket.

tests/engine/test_cache_budget.py: expandable segments are requested on Linux and not on Windows, without a deprecation warning; without this change the Windows case fails (['expandable_segments:True'] == []), and so does the warning check, on torch.cuda._set_allocator_settings is deprecated.

Verification

Windows 11, Python 3.13, pyzmq 27.2.0, RTX 3090 Ti: with #575 and #581 on main, ft serve --model RadixArk/Qwen3.8-Flash-Next-NVFP4 --moe-strategy offload --text-model-only --max-running-requests 1 --memory-ratio 0.85 reports ok on /health 127 s after it starts and answers a chat, an Anthropic and a Responses request. Its log has none of Enabled expandable_segments, torch's expandable_segments not supported on this platform and the deprecation warning, all three of which it printed without this change; a Qwen3-0.6B serve on Linux still logs Enabled expandable_segments, without the warning. tests/scheduler/test_zmq_links.py: 4 passed on Windows, where the module cannot import its link helpers without this change, and 4 passed on Linux (WSL2 Ubuntu 22.04, Python 3.10), where the links are the same ipc paths as before. tests/server/test_supervisor.py: 13 passed on both; without this change the new test fails on Windows (the child exits 15, killed before its handler runs) and passes on Linux, where os.kill reaches the handler. tests/engine/test_cache_budget.py: 30 passed on Windows; on Linux 27 passed and 1 skipped (WSL's pinning cap), with the two fi fixtures failing there as they do on main without flashinfer (#575 moves them to triton).

End to end with Qwen3-0.6B, killing the scheduler of a serving ft serve: without this change the server exited 10 s later with code 15 and one worker kept running (1.8 GB); with it uvicorn ran its shutdown (Application shutdown complete.) and no worker was left.

Known limits

A loopback port is chosen by binding it and releasing it, so another process can take it before zmq binds; that fails the start with a bind error rather than misrouting anything.

A hard kill of the server process on Windows (TerminateProcess, which is what Popen.terminate() does there) runs no handler, so its two worker processes outlive it: tests/e2e/test_cache_rebuild.py leaves them behind on every run here, about 1.4 GB and 0.7 GB. A job object with kill-on-close would tie them to the server; that is a separate change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant