Skip to content

fix(engine): size memory pools against the VRAM other processes leave free on Windows - #583

Draft
YevheniiKotyrlo wants to merge 4 commits into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-windows-free-memory
Draft

YevheniiKotyrlo wants to merge 4 commits into
FlashML-org:mainfrom
YevheniiKotyrlo:fix-windows-free-memory

Conversation

@YevheniiKotyrlo

@YevheniiKotyrlo YevheniiKotyrlo commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Closes #570

Problem

Under WDDM, cudaMemGetInfo leaves out what other processes hold, so get_free_memory reads a card the desktop uses as free: 23,313 MiB where NVML and nvidia-smi read 21,286 on mine. Every pool and the --vram-reserve-mb fit are sized against it, so a start beside a desktop overcommits the card, and WDDM pages the excess into shared memory rather than failing.

Solution

  • gpu_select.nvml_free_bytes(uuid): the device's free memory across every process, as nvidia-smi reports it, or None when NVML cannot say. It uses the loader _nvml_uuids already had, split out as _load_nvml.
  • get_free_memory takes the smaller of cudaMemGetInfo's figure and that one on Windows, finding the device by its UUID. Linux keeps cudaMemGetInfo alone: it already counts every process there, and NVML is not loaded.

Every pool and the reserve fit read through get_free_memory, so nothing else changes.

Tests

tests/engine/test_free_memory.py fakes CUDA's reading, the device's UUID and the NVML library the engine loads, in four cases: the smaller figure on Windows; CUDA's where it is the smaller; CUDA's when NVML cannot be loaded; and CUDA's alone on Linux, with NVML never loaded. Without the fix the first fails on the figure itself, assert 24696061952 == (21 * 1073741824).

Verification

Windows 11, RTX 3090 Ti, RadixArk/Qwen3.8-Flash-Next-NVFP4 at a 1M window with --memory-ratio 0.85 --vram-reserve-mb 2000, with a VRAM hog standing in for a busier desktop. Before the fix the engine read 22.75 GiB free at every one of these points (#570); with it:

free before the start (nvidia-smi) read by the engine slots sized, kept the fit's first prefill load
21,547 MiB 20.79 GiB 2,253, 1,815 6.9 s, 0.95 GiB left 117 s
21,271 MiB 20.52 GiB 2,040, 1,401 9.6 s, 0.62 GiB left 243 s
20,778 MiB 20.05 GiB 2,010, 1,518 7.8 s, 0.81 GiB left 278 s
20,370 MiB 19.64 GiB 1,892, 1,328 7.8 s, 0.62 GiB left 250 s
20,329 MiB 19.61 GiB 1,854, 1,357 7.8 s, 0.80 GiB left 106 s
20,071 MiB 19.35 GiB 1,782, 1,247 7.3 s, 0.70 GiB left 199 s

The engine's reading follows the card, 247-259 MiB under nvidia-smi's before the start, and every fit's first prefill ran on the card, where before it ran 45-228 s in shared memory. The load times vary with the desktop beside it: the server ran at idle priority next to a working session.

The engine, MoE, server and daemon suites over the three Windows branches, on a clear card: 803 passed, 13 skipped, and test_batch_memcpy_roundtrip failed on the probe race #414 fixes. On Linux (WSL2 Ubuntu 22.04, Python 3.10, CUDA 13.0), the branch over main: tests/engine 114 passed and 3 skipped, and the two fi tests of test_cache_budget.py fail there as they do on main, without flashinfer.

Related

#529 covers other Windows single-GPU serving limits (the WDDM pin budget, split residency, tiered KV), not this figure.

_load_nvml exists, so the free-memory tests patch it without raising=False; a renamed loader now fails them at setup.
gpu_identity, _visible_of_physical and get_free_memory each spelled the
nvidia-smi UUID of a visible device from its torch properties; gpu_uuid is that
derivation once.
@YevheniiKotyrlo
YevheniiKotyrlo force-pushed the fix-windows-free-memory branch from 4f502c8 to 6d575bc Compare October 6, 2026 10:55
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 6, 2026
FlashML-org#618 deferred with FlashML-org#605 (pre-sm70 only); drafts FlashML-org#582, FlashML-org#583, FlashML-org#586, FlashML-org#613 deferred.

Assisted-by: Claude Opus 5.5
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 6, 2026
FlashML-org#586

A stale local pr/N ref (fetched without + before a rebase) gave their pre-rebase SHAs;
the rebased heads carry the same changes.

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

On Windows the engine sizes its pools as if no other process held any of the card

1 participant