Repository navigation
fix(windows): serve over loopback ZMQ and a selector event loop - #580
Draft
YevheniiKotyrlo wants to merge 6 commits into
Draft
YevheniiKotyrlo wants to merge 6 commits into
YevheniiKotyrlo wants to merge 6 commits into
Conversation
This was referenced Sep 30, 2026
The two transport tests already assert five distinct links each, and the platform test restated the branch it ran through.
YevheniiKotyrlo
force-pushed
the
fix-windows-runtime
branch
from
October 6, 2026 10:54
0e50062 to
1b0ee44
Compare
This was referenced Oct 6, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #384
Problem
With the build fixed (#575),
ft serveon Windows starts and then answers nothing:ipc:///tmp/freetoken_N..., and the libzmq pyzmq ships for Windows has no ipc transport:zmq.has("ipc")isFalseand a bind fails withProtocol not supported(pyzmq 27.2.0, libzmq 4.3.5).zmq.asyncioneedsadd_reader, which that loop does not implement:RuntimeError: Proactor event loop does not implement add_reader family of methods required for zmq._create_listener_oncecreates and drops, so nothing observes it: the server stays up, accepts every request and answers none of them.os.kill(os.getpid(), signal.SIGTERM). On Windows that isTerminateProcess: no handler runs, so uvicorn's lifespan never terminates the remaining workers and they outlive the server. A Flash-Next serve whose scheduler died on a CUDA error left its detokenizer worker running with 1.9 GB.Separately, every load asks torch for expandable segments, which torch compiles only off Windows (
if(NOT WIN32)inc10/cuda/CMakeLists.txt): the engine logsEnabled expandable_segments, and the first allocation then warnsexpandable_segments not supported on this platformand runs without them. The request goes throughtorch.cuda.memory._set_allocator_settings, which torch 2.11 deprecates, so every start prints a FutureWarning, on Linux too.Solution
SchedulerConfigchooses its five links once, in the parent (_choose_zmq_links): the sameipcpaths as today where libzmq has ipc, loopback TCP ports where it does not. The links travel with the config to every spawned worker.asyncio:SelectorEventLoopon Windows (UVICORN_LOOP), in both the serve and the shell path.fatal_error,maintenance_state = "failed") and exits the way it already does when the backend dies. This part is not Windows-specific - it is how the failure in (2) surfaced as silence.signal.raise_signal) instead ofos.kill-ing our own pid, so uvicorn's handler runs the lifespan shutdown that terminates the workers. On Linux the two are the same; on Windows only the first runs a handler._ensure_expandable_segmentsleaves the allocator alone on Windows instead of requesting what torch does not build there, and elsewhere callstorch._C._accelerator_setAllocatorSettings, the binding the deprecated wrapper forwards to and the warning names.pyzmq's own message suggests installing tornado to keep the Proactor loop; that adds a dependency and a selector thread per loop, where the selector loop needs neither.
Tests
tests/scheduler/test_zmq_links.py: without ipc every link is a distinct loopback port; with ipc the links are the per-process socket paths they were; a pickled config keeps its links on either transport (what a spawned worker receives).tests/server/test_supervisor.py: a backend death stops the server through its own SIGTERM handler - a child process installs one and waits for it on a wakeup socket.tests/engine/test_cache_budget.py: expandable segments are requested on Linux and not on Windows, without a deprecation warning; without this change the Windows case fails (['expandable_segments:True'] == []), and so does the warning check, ontorch.cuda._set_allocator_settings is deprecated.Verification
Windows 11, Python 3.13, pyzmq 27.2.0, RTX 3090 Ti: with #575 and #581 on
main,ft serve --model RadixArk/Qwen3.8-Flash-Next-NVFP4 --moe-strategy offload --text-model-only --max-running-requests 1 --memory-ratio 0.85reportsokon/health127 s after it starts and answers a chat, an Anthropic and a Responses request. Its log has none ofEnabled expandable_segments, torch'sexpandable_segments not supported on this platformand the deprecation warning, all three of which it printed without this change; a Qwen3-0.6B serve on Linux still logsEnabled expandable_segments, without the warning.tests/scheduler/test_zmq_links.py: 4 passed on Windows, where the module cannot import its link helpers without this change, and 4 passed on Linux (WSL2 Ubuntu 22.04, Python 3.10), where the links are the sameipcpaths as before.tests/server/test_supervisor.py: 13 passed on both; without this change the new test fails on Windows (the child exits 15, killed before its handler runs) and passes on Linux, whereos.killreaches the handler.tests/engine/test_cache_budget.py: 30 passed on Windows; on Linux 27 passed and 1 skipped (WSL's pinning cap), with the twofifixtures failing there as they do onmainwithout flashinfer (#575 moves them totriton).End to end with Qwen3-0.6B, killing the scheduler of a serving
ft serve: without this change the server exited 10 s later with code 15 and one worker kept running (1.8 GB); with it uvicorn ran its shutdown (Application shutdown complete.) and no worker was left.Known limits
A loopback port is chosen by binding it and releasing it, so another process can take it before zmq binds; that fails the start with a bind error rather than misrouting anything.
A hard kill of the server process on Windows (
TerminateProcess, which is whatPopen.terminate()does there) runs no handler, so its two worker processes outlive it:tests/e2e/test_cache_rebuild.pyleaves them behind on every run here, about 1.4 GB and 0.7 GB. A job object with kill-on-close would tie them to the server; that is a separate change.