Skip to content

auto-fit produces worse plans than --offload-to-cpu on mid-range GPUs: resident weights are non-evictable and starve VAE decode #2118

Description

@losewayy

On mid-range GPUs the default auto-fit placement can be significantly slower than the documented --offload-to-cpu recommendation — up to ~3x end-to-end on Wan2.2 TI2V-5B — because it keeps the diffusion model's weights resident on the GPU, and resident weights are non-evictable, which starves the VAE decode phase.

Root cause

Auto-fit assigns the DiT as resident on the main GPU whenever the weights fit in VRAM. Resident params are their own canonical copy (compute_backend == params_backend), so ModelManager::ensure_compute_backend_capacity can never evict them (reloadable == false). When VAE decode then needs a workspace larger than free VRAM - resident params, capacity checks fail (model manager cannot make enough memory available in -v logs) and decode falls back to tiled processing — currently jumping straight to 256px tiles (#2108).

Measured (RTX 5070 Ti Laptop, 12 GB, ~2-3.5 GB used by other apps, master f89d9b1, CUDA backend)

Model Output auto-fit (default) VAE decode --offload-to-cpu VAE decode End-to-end
Qwen Image 2.1, Q6_K DiT (5.6 GB) 1024x1024 15.9 s — 49 tiles at 256px 4.37 s — full frame (7.78 GB workspace) ~81 s vs ~68 s
Wan2.2 TI2V-5B, Q8_0 DiT (5.15 GB) 480x832x33 frames 366.9 s — 84 tiles at 112x128px 104.2 s — 15 tiles at 240x256px 402.9 s vs 138.9 s
Qwen Image 2.1 832x480 2.06 s — full frame (not triggered) — —
Qwen Image 2.1 1536x1536 32.6 s — 121 tiles at 256px — —

At 832x480 the same model decodes untiled in 2.06 s, so the trigger is resident params + decode peak > free VRAM, not the model itself. With --offload-to-cpu the manager evicts the staged DiT blocks at decode time (observable in -v logs) — exactly the behavior that resident placement forecloses.

The trigger band is common: quantized large models (DiT ~4-8 GB) on ~8-20 GB cards, i.e. the primary GGUF use case. Full-precision models do not hit this because their params never become resident in the first place. The same mechanism should affect any pipeline where an early stage leaves resident params while a later stage needs a large workspace (hires-fix second passes, upscalers, video decode) — measured above for two models; the generalization beyond them is inference from the shared code path, not direct measurement.

Why tiling does not save it

The VAE decode workspace for Wan-family 3D VAEs is ~7.8-22 GB at these sizes (f32 activations + conv3d intermediates), so a single resident DiT is enough to force the fallback. At 1536x1536 the fallback jumps straight to 256px tiles even though 768px tiles would fit — so even the tiling path itself is degraded.

Possible directions

Happy to implement whichever you prefer:

  1. Let ensure_compute_backend_capacity demote file-reloadable resident states to ResidencyMode::Disk as a last-resort tier before failing. Demotion only unlocks evictability; release stays demand-driven via the existing release path, and reload goes through the existing disk-resident load path. Non-protected, unpinned, non-required states only.
  2. Have auto-fit prefer evictable residency (params on CPU) when a component's resident params plus downstream peaks cannot fit — needs a better VAE reserve estimate (currently a fixed 1 GB vs ~8-22 GB actual for Wan-family 3D VAEs).
  3. Fix the tile fallback granularity (VAE decode auto-fit jumps from 1024 to 256 px tiles, skipping 512 (2-4x slower decode on 16 GB GPUs) #2108) so intermediate sizes are tried — helps the unavoidably-tiled cases like 1536x1536 or video, but does not address this one.

Workaround for users today: use --offload-to-cpu as shown in the model docs. Note that a partial --params-backend "diffusion=cpu" makes things worse: it disables auto-fit planning entirely and the remaining components (TE, VAE) become GPU-resident instead.

Reproduce

# default auto-fit: watch for "Reducing VAE decode tiles ... to 256x256"
sd-cli --diffusion-model qwen_image_2.1_Q6_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \
  --llm Qwen3VL-8B-Instruct-Q4_K_M.gguf -p "<prompt>" -W 1024 -H 1024 \
  --steps 4 --cfg-scale 1 --sampling-method euler --fa -v

# same command plus --offload-to-cpu: full-frame decode, ~3.6x faster

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions