Repository navigation
fix(engine): size memory pools against the VRAM other processes leave free on Windows - #583
Draft
YevheniiKotyrlo wants to merge 4 commits into
Draft
YevheniiKotyrlo wants to merge 4 commits into
YevheniiKotyrlo wants to merge 4 commits into
Conversation
_load_nvml exists, so the free-memory tests patch it without raising=False; a renamed loader now fails them at setup.
gpu_identity, _visible_of_physical and get_free_memory each spelled the nvidia-smi UUID of a visible device from its torch properties; gpu_uuid is that derivation once.
…form the test fakes
YevheniiKotyrlo
force-pushed
the
fix-windows-free-memory
branch
from
October 6, 2026 10:55
4f502c8 to
6d575bc
Compare
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 6, 2026
FlashML-org#618 deferred with FlashML-org#605 (pre-sm70 only); drafts FlashML-org#582, FlashML-org#583, FlashML-org#586, FlashML-org#613 deferred. Assisted-by: Claude Opus 5.5
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 6, 2026
FlashML-org#586 A stale local pr/N ref (fetched without + before a rebase) gave their pre-rebase SHAs; the rebased heads carry the same changes. Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #570
Problem
Under WDDM,
cudaMemGetInfoleaves out what other processes hold, soget_free_memoryreads a card the desktop uses as free: 23,313 MiB where NVML andnvidia-smiread 21,286 on mine. Every pool and the--vram-reserve-mbfit are sized against it, so a start beside a desktop overcommits the card, and WDDM pages the excess into shared memory rather than failing.Solution
gpu_select.nvml_free_bytes(uuid): the device's free memory across every process, asnvidia-smireports it, orNonewhen NVML cannot say. It uses the loader_nvml_uuidsalready had, split out as_load_nvml.get_free_memorytakes the smaller ofcudaMemGetInfo's figure and that one on Windows, finding the device by its UUID. Linux keepscudaMemGetInfoalone: it already counts every process there, and NVML is not loaded.Every pool and the reserve fit read through
get_free_memory, so nothing else changes.Tests
tests/engine/test_free_memory.pyfakes CUDA's reading, the device's UUID and the NVML library the engine loads, in four cases: the smaller figure on Windows; CUDA's where it is the smaller; CUDA's when NVML cannot be loaded; and CUDA's alone on Linux, with NVML never loaded. Without the fix the first fails on the figure itself,assert 24696061952 == (21 * 1073741824).Verification
Windows 11, RTX 3090 Ti,
RadixArk/Qwen3.8-Flash-Next-NVFP4at a 1M window with--memory-ratio 0.85 --vram-reserve-mb 2000, with a VRAM hog standing in for a busier desktop. Before the fix the engine read 22.75 GiB free at every one of these points (#570); with it:nvidia-smi)The engine's reading follows the card, 247-259 MiB under
nvidia-smi's before the start, and every fit's first prefill ran on the card, where before it ran 45-228 s in shared memory. The load times vary with the desktop beside it: the server ran at idle priority next to a working session.The engine, MoE, server and daemon suites over the three Windows branches, on a clear card: 803 passed, 13 skipped, and
test_batch_memcpy_roundtripfailed on the probe race #414 fixes. On Linux (WSL2 Ubuntu 22.04, Python 3.10, CUDA 13.0), the branch overmain:tests/engine114 passed and 3 skipped, and the twofitests oftest_cache_budget.pyfail there as they do onmain, without flashinfer.Related
#529 covers other Windows single-GPU serving limits (the WDDM pin budget, split residency, tiered KV), not this figure.