You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
auto-fit produces worse plans than --offload-to-cpu on mid-range GPUs: resident weights are non-evictable and starve VAE decode #2118
On mid-range GPUs the default auto-fit placement can be significantly slower than the documented --offload-to-cpu recommendation — up to ~3x end-to-end on Wan2.2 TI2V-5B — because it keeps the diffusion model's weights resident on the GPU, and resident weights are non-evictable, which starves the VAE decode phase.
Root cause
Auto-fit assigns the DiT as resident on the main GPU whenever the weights fit in VRAM. Resident params are their own canonical copy (compute_backend == params_backend), so ModelManager::ensure_compute_backend_capacity can never evict them (reloadable == false). When VAE decode then needs a workspace larger than free VRAM - resident params, capacity checks fail (model manager cannot make enough memory available in -v logs) and decode falls back to tiled processing — currently jumping straight to 256px tiles (#2108).
Measured (RTX 5070 Ti Laptop, 12 GB, ~2-3.5 GB used by other apps, master f89d9b1, CUDA backend)
Model
Output
auto-fit (default) VAE decode
--offload-to-cpu VAE decode
End-to-end
Qwen Image 2.1, Q6_K DiT (5.6 GB)
1024x1024
15.9 s — 49 tiles at 256px
4.37 s — full frame (7.78 GB workspace)
~81 s vs ~68 s
Wan2.2 TI2V-5B, Q8_0 DiT (5.15 GB)
480x832x33 frames
366.9 s — 84 tiles at 112x128px
104.2 s — 15 tiles at 240x256px
402.9 s vs 138.9 s
Qwen Image 2.1
832x480
2.06 s — full frame (not triggered)
—
—
Qwen Image 2.1
1536x1536
32.6 s — 121 tiles at 256px
—
—
At 832x480 the same model decodes untiled in 2.06 s, so the trigger is resident params + decode peak > free VRAM, not the model itself. With --offload-to-cpu the manager evicts the staged DiT blocks at decode time (observable in -v logs) — exactly the behavior that resident placement forecloses.
The trigger band is common: quantized large models (DiT ~4-8 GB) on ~8-20 GB cards, i.e. the primary GGUF use case. Full-precision models do not hit this because their params never become resident in the first place. The same mechanism should affect any pipeline where an early stage leaves resident params while a later stage needs a large workspace (hires-fix second passes, upscalers, video decode) — measured above for two models; the generalization beyond them is inference from the shared code path, not direct measurement.
Why tiling does not save it
The VAE decode workspace for Wan-family 3D VAEs is ~7.8-22 GB at these sizes (f32 activations + conv3d intermediates), so a single resident DiT is enough to force the fallback. At 1536x1536 the fallback jumps straight to 256px tiles even though 768px tiles would fit — so even the tiling path itself is degraded.
Possible directions
Happy to implement whichever you prefer:
Let ensure_compute_backend_capacity demote file-reloadable resident states to ResidencyMode::Disk as a last-resort tier before failing. Demotion only unlocks evictability; release stays demand-driven via the existing release path, and reload goes through the existing disk-resident load path. Non-protected, unpinned, non-required states only.
Have auto-fit prefer evictable residency (params on CPU) when a component's resident params plus downstream peaks cannot fit — needs a better VAE reserve estimate (currently a fixed 1 GB vs ~8-22 GB actual for Wan-family 3D VAEs).
Workaround for users today: use --offload-to-cpu as shown in the model docs. Note that a partial --params-backend "diffusion=cpu" makes things worse: it disables auto-fit planning entirely and the remaining components (TE, VAE) become GPU-resident instead.
On mid-range GPUs the default auto-fit placement can be significantly slower than the documented
--offload-to-cpurecommendation — up to ~3x end-to-end on Wan2.2 TI2V-5B — because it keeps the diffusion model's weights resident on the GPU, and resident weights are non-evictable, which starves the VAE decode phase.Root cause
Auto-fit assigns the DiT as resident on the main GPU whenever the weights fit in VRAM. Resident params are their own canonical copy (
compute_backend == params_backend), soModelManager::ensure_compute_backend_capacitycan never evict them (reloadable == false). When VAE decode then needs a workspace larger thanfree VRAM - resident params, capacity checks fail (model manager cannot make enough memory availablein-vlogs) and decode falls back to tiled processing — currently jumping straight to 256px tiles (#2108).Measured (RTX 5070 Ti Laptop, 12 GB, ~2-3.5 GB used by other apps, master
f89d9b1, CUDA backend)--offload-to-cpuVAE decodeAt 832x480 the same model decodes untiled in 2.06 s, so the trigger is
resident params + decode peak > free VRAM, not the model itself. With--offload-to-cputhe manager evicts the staged DiT blocks at decode time (observable in-vlogs) — exactly the behavior that resident placement forecloses.The trigger band is common: quantized large models (DiT ~4-8 GB) on ~8-20 GB cards, i.e. the primary GGUF use case. Full-precision models do not hit this because their params never become resident in the first place. The same mechanism should affect any pipeline where an early stage leaves resident params while a later stage needs a large workspace (hires-fix second passes, upscalers, video decode) — measured above for two models; the generalization beyond them is inference from the shared code path, not direct measurement.
Why tiling does not save it
The VAE decode workspace for Wan-family 3D VAEs is ~7.8-22 GB at these sizes (f32 activations + conv3d intermediates), so a single resident DiT is enough to force the fallback. At 1536x1536 the fallback jumps straight to 256px tiles even though 768px tiles would fit — so even the tiling path itself is degraded.
Possible directions
Happy to implement whichever you prefer:
ensure_compute_backend_capacitydemote file-reloadable resident states toResidencyMode::Diskas a last-resort tier before failing. Demotion only unlocks evictability; release stays demand-driven via the existing release path, and reload goes through the existing disk-resident load path. Non-protected, unpinned, non-required states only.Workaround for users today: use
--offload-to-cpuas shown in the model docs. Note that a partial--params-backend "diffusion=cpu"makes things worse: it disables auto-fit planning entirely and the remaining components (TE, VAE) become GPU-resident instead.Reproduce