Repository navigation
Video regressions on 16 GB Vulkan: LTX-2.5 + MiniMax-H3 broken since master-864, Wan 2.2 broken by master-866 (graph segmentation / weight-budget behavior) — --auto-fit off only rescues Wan #1976
Description
Activity
- changed the title
[-]Regression master-864 → master-866 (16 GB Vulkan): Wan 2.2 / LTX-2.5 / MiniMax-H3 video generation all fail — memory manager segmentation/budget behavior changed (--auto-fit off only rescues Wan)[/-][+]Video regressions on 16 GB Vulkan: LTX-2.5 + MiniMax-H3 broken since master-864, Wan 2.2 broken by master-866 (graph segmentation / weight-budget behavior) — --auto-fit off only rescues Wan[/+]on Sep 15, 2026 Follow-up: we tested master-899
28b454b(2026-09-22 release, includes #2019 and #2020) on the same machine, same models, same commands. Full matrix below — net result: the Wan 2.2 regression is gone, LTX-2.5 now passes but is VRAM-margin sensitive, MiniMax-H3 is still broken with the identical signature.Results on master-899 (freshly booted, desktop VRAM usage 0.4/16 GB)
Path master-841 6b3edaamaster-864/866 master-899 28b454bFlux 1024² (dev, 4 steps) ✅ ✅ ✅ Wan 2.2 T2V 17f (exact recipe above) ❌ (#1946) ✅ (864) / ❌ (866) ✅ — sampling 2023 s vs 2116 s on 864 (−4 %), no --auto-fit offneededLTX-2.5 T2V 33f benchmark (2 steps, 768×512) ✅ ❌ ✅ on fresh boot — 53.75 s sampling, webm written — but ❌ on the same build before a reboot, see below MiniMax H3 Ref2VA turbo (exact recipe above) ✅ (2 segments) ❌ ❌ unchanged: minimax_h3 using 51 segments+vk::Queue::submit: ErrorOutOfDeviceMemory, sampling failed after 276.68 s (864: 262.86 s, 866: 358.87 s) — deterministic, reproduced on a freshly rebooted machine1. LTX-2.5 on master-899: passes, but the capacity check is desktop-VRAM-margin sensitive
Same command, same build, same machine, ~40 minutes apart:
Before reboot (desktop/apps holding ~0.8 GB VRAM) — fails at 21.5 s:
[WARN ] model_manager.cpp:1809 - model manager cannot make enough memory available on Vulkan0: need 908.75 MB device / 396.75 MB budget, available 787.26 MB device / 668.90 MB budget [ERROR ] ggml_runner.cpp:880 - ltxav segment 1/1 (graph) failed during workspace capacity check [ERROR ] video.cpp:1715 - sampling failed after 21.51sAfter reboot (desktop VRAM usage 0.8 → 0.4 GB) — passes end-to-end (53.75 s sampling, webm written, one non-fatal
Failed to allocate pinned memoryWARN).So the failure margin here was ~121 MB of "available device". Two observations worth flagging:
- The available-device under-reporting seems greatly reduced by fix: handle GPU memory reports and LLM encoding failures #2020. In our master-864/866 runs (original report above) the manager claimed
available 319.37 MB devicewhile ~15 GB was physically free and 0 MB of weights were resident in VRAM. On master-899 the same staging reportsavailable 787.26 MB device / total 16368.00 MB, tracked weights 14388.75 MB— consistent with the actual free VRAM minus non-tracked overhead. Whatever ghost reservation produced the 319 MB figure appears gone (matches fix: handle GPU memory reports and LLM encoding failures #2020 "handle GPU memory reports"). - The remaining ~120 MB shortfall was real desktop VRAM held by other processes. Practical consequence: on a 16 GB card the LTX-2.5 33-frame capacity check now passes but with little headroom — anything else holding a few hundred MB of VRAM flips it back to a refusal. We keep our master-841 parallel install for LTX/H3 production for that reason (our LTX I2V 65-frame production chains are also not re-validated on 899 yet).
2. MiniMax-H3: still the blocker
The 841→864 regression is not fixed by #2019/#2020: the DiT graph is still cut into 51 segments (control build: 2) and dies at submit with
ErrorOutOfDeviceMemoryseveral minutes into sampling, deterministically, on a fresh boot. H3 production on our side still runs on the master-841 parallel install. Happy to test any candidate fix or bisect the 841→864 window for this path — we have byte-stable A/B repro (same command/seed/assets, control webm from the 841 binary).- The available-device under-reporting seems greatly reduced by fix: handle GPU memory reports and LLM encoding failures #2020. In our master-864/866 runs (original report above) the manager claimed
Follow-up: we tested master-908
88411ef(2026-09-23 release, includes #1900 "add graph cuts for MiniMax-H3 text conditioning") on the same machine, same models, same commands as the matrix above, freshly rebooted (desktop VRAM usage 0.4/16 GB). Net result: Flux/LTX/Wan are stable or slightly faster, but MiniMax-H3 is still broken — the graph is now cut into 54 segments (was 51 on 864/866/899, control = 2 on 841) and still dies at submit withErrorOutOfDeviceMemory.Results on master-908 vs previous builds
Path master-841 6b3edaamaster-864→899 master-908 88411efFlux 1024² (dev, 4 steps) ✅ ✅ 26.98 s (899) ✅ 27.00 s LTX-2.5 T2V 33f benchmark (2 steps, 768×512) ✅ ✅ on fresh boot — 53.75 s (899) ✅ 54.21 s, 1 segment, no capacity-check refusal, same non-fatal Failed to allocate pinned memoryWARNWan 2.2 T2V 17f (exact recipe above) ✅ ✅ 2023 s (899, −4 % vs 864) ✅ 1999.5 s (−1.2 % vs 899), 1 resident segment (9211.74 MB weights, identical to 864/899), VAE CPU decode 209.5 s MiniMax H3 Ref2VA turbo (exact recipe above) ✅ (2 segments) ❌ using 51 segments+ OOM at submit (899: 276.68 s)❌ using 54 segments+ OOM at submit, sampling failed after 257.37 s#1900 changes the H3 graph cut (51 → 54 segments) but does not restore the healthy plan
Full failure sequence on master-908 (same command/seed/assets as the control webm produced by the 841 binary):
[VERBOSE] vae.hpp:262 - computing vae encode graph completed, taking 855.82s [VERBOSE] conditioner.hpp:3139 - computing condition graph completed, taking 92884 ms [VERBOSE] ggml_runner.cpp:864 - minimax_h3 using 54 segments [ERROR ] ggml_runner.cpp:654 - minimax_h3 graph execution failed on Vulkan0: vk::Queue::submit: ErrorOutOfDeviceMemory [ERROR ] diffusion_engine.cpp:2662 - Diffusion model sampling failed [ERROR ] video.cpp:1709 - sampling failed after 257.37sTwo observations:
- The segment count changed across the range that includes fix: reduce MiniMax-H3 VRAM spikes and redundant token refinemen #1900 (51 on 864/866/899 → 54 on 908), so the commit does affect how the H3 DiT graph is partitioned — but the result is still a fully fragmented plan (54 segments) versus the healthy 2-segment cut on master-841 (
segment 1/2: blocks.0..37+segment 2/2), and the deterministicErrorOutOfDeviceMemoryat submit remains. The 841→864 regression reported here is still open on master-908. - Everything else in the matrix is unchanged or marginally better (Wan −1.2 % sampling vs 899, LTX passes on a fresh boot with the same ~120 MB desktop-VRAM sensitivity noted above), so we kept 908 pinned — H3/LTX production on our side still runs on the master-841 parallel install.
Happy to test any candidate fix — we have a byte-stable A/B repro (same command/seed/assets, control webm + logs from the 841 binary).
- The segment count changed across the range that includes fix: reduce MiniMax-H3 VRAM spikes and redundant token refinemen #1900 (51 on 864/866/899 → 54 on 908), so the commit does affect how the H3 DiT graph is partitioned — but the result is still a fully fragmented plan (54 segments) versus the healthy 2-segment cut on master-841 (
Follow-up: we tested master-929
3f8527aon the same machine (RX 6950 XT 16 GB,
RDNA2, no matrix cores), same models, same commands as the matrix above, machine idle
(gate-checked). Net result: Wan 2.2 stays fixed and faster; the Wan I2V full path now
also passes; LTX-2.5 33f now fails EARLY and EXPLICITLY (the new memory manager reports
the numbers instead of crashing); MiniMax-H3 is still broken with the identical
54-segment OOM signature.Results on master-929
Path master-841 6b3edaamaster-864→908 master-929 3f8527aFlux 1024² (dev, 4 steps) ✅ ✅ ✅ Wan 2.2 T2V 17f ✅ ✅ (fixed by 899) ✅ (+8.4% vs 841) Wan I2V full path (25 steps, --vae-on-cpu)❌ 0xC0000409 (Sep 25) ? ✅ PASS (52 min, healthy) LTX-2.5 33f distilled (Q4_K_M, 768×512) ❌ (machine-state) ❌ ❌ explicit refusal (below) MiniMax-H3 turbo Ref2VA (validated recipe, 22f + LoRA) ✅ 2 segments ❌ 51→54 segments, OOM ❌ still 54 segments, OOM at submit LTX-2.5 33f on master-929 — no more silent crash, the memory manager now refuses
cleanly with the full picture (this looks like a 16 GB fit problem, not a crash bug):
stages 14,388.75 MB (4,349 tensors, 15 blocks) onto Vulkan0, leaving 776.09 MB free,
then:model manager cannot make enough memory available on Vulkan0: need 908.75 MB device / 396.75 MB budget, available 776.09 MB device / 668.90 MB budget→
ltxav segment 1/1 (graph) failed during workspace capacity checkat 21 s.
Question: should--offload-to-cpustream the LTX DiT weights for vid_gen (it does
rescue Wan 14B for us), or is 33f @ Q4_K_M simply above the 16 GB envelope?MiniMax-H3 turbo on master-929 — retested today with the full validated recipe
(--turboLoRA 8 steps, 864×480, 22 frames @ 24 fps,--offload-to-cpu --max-vram 10 --rng cpu, te+vae on CPU, ref tail + LoRA):minimax_h3 using 54 segments→
graph execution failed on Vulkan0: vk::Queue::submit: ErrorOutOfDeviceMemory(exit
0xC0000409). Control master-841 = 2 segments, runs fine (re-validated on the
parallel build). The #1900 graph cuts (51→54 segments) did not close the gap.
The question from the matrix still stands: is 16 GB below the supported envelope for
H3 even with offload + max-vram, or should the graph segmentation eventually make it
fit? Happy to test any candidate build — this machine is our standard 16 GB Vulkan
datapoint.I have similar problems, spent hours testing various configurations, various sd.cpp builds. Minimax H3 support is hopelessly broken and memory inefficient with stable-diffusion.cpp. Can only generate short clips on low res even with the 3090. Comfy can do 15 sec 1344x768 (H3's specified max) clips with H3 even on my 12 gb laptop 4080 without additional memory optimization nodes.
Confirming @nebkmb's findings with our own matrix (RX 6950 XT 16 GB, RDNA2, no matrix cores, machine gate-checked idle):
- MiniMax-H3: full validated recipe on master-929
3f8527a->minimax_h3 using 54 segments->vk::Queue::submit: ErrorOutOfDeviceMemory, exit 0xC0000409. Same signature on master-841. Since it also fails on your 24 GB 3090, this is not a 16 GB envelope issue — the segmentation path itself looks broken/memory-inefficient in sd.cpp (ComfyUI doing 15 s @ 1344x768 on a 12 GB 4080 corroborates that). - LTX-2.5: on master-929 the memory manager stages 14,388.75 MB (4,349 tensors) onto Vulkan0, leaves 838.65 MB free, then refuses cleanly:
cannot make enough memory available on Vulkan0: need 908.56 MB device / 396.56 MB budget->ltxav segment 1/1 (graph) failed during workspace capacity check(33 frames @ 768x512, 2 steps). LTX-2.5 33f does not fit on 16 GB with 929, while it runs fine on master-841. - Tested today:
--params-backend diskdoes not change the LTX outcome on 929 — same staging onto Vulkan0, same workspace refusal.
Net on RDNA2 16 GB: Wan 2.1/2.2 (incl. I2V + VACE) fixed and faster on 929; LTX-2.5 fits only on 841; H3 broken on both. Happy to test candidate patches/builds — the H3 envelope and the LTX offload behavior remain the two open questions from our previous comment.
- MiniMax-H3: full validated recipe on master-929
Follow-up: we tested master-945
a1ded76(2026-10-07 release, first build carrying the
H3-path change #2103 "keep MiniMax-H3 VAE weights resident across temporal chunks") on the
same machine (RX 6950 XT 16 GB, RDNA2, no matrix cores), same models, same commands as the
matrix above, machine gate-checked idle. Net result: MiniMax-H3 is still broken with the
identical signature, and LTX-2.5 still fails the same workspace-capacity refusal.Results on master-945
Path master-841 6b3edaamaster-864→929 master-945 a1ded76MiniMax-H3 Ref2VA 22 f turbo (factory full path, --turbo)✅ (2 segments) ❌ 51→54 segments + ErrorOutOfDeviceMemory❌ identical: minimax_h3 using 54 segments→vk::Queue::submit: ErrorOutOfDeviceMemoryat step ~10/28, exit 0xC0000409LTX-2.5 T2V 33 f (2-step distilled benchmark) ✅ ❌ early explicit refusal (908 MB / 776 free on 929) ❌ same class: ltxav segment 1/1 failed during workspace capacity check— tracked weights 14,388.75 MB on Vulkan0, need 893.75 MB device, reported free 694.57 MBFlux 1024² ✅ ✅ ✅ (real-path smoke, 160 s) Notes:
- fix: keep MiniMax-H3 VAE weights resident across temporal chunks #2103 did not change the H3 graph segmentation on our envelope: still 54 segments
(control = 2 on 841). The VAE-residency change may help other configurations, but the
DiT graph cut is unchanged and the submit OOM is byte-for-byte the 929 signature. - LTX numbers on 945 are the same failure mode as reported for 929 in the 10-01 comment;
--params-backend diskwas already tested negative on 929 (10-02). - Production on our side remains the pinned master-841 parallel install for H3 + LTX.
Not closing — the issue stays accurate for both regressions on 16 GB Vulkan. Happy to run
any experiment on 945+ (bisect offers from the original comment still stand).- fix: keep MiniMax-H3 VAE weights resident across temporal chunks #2103 did not change the H3 graph segmentation on our envelope: still 54 segments
- added a commit that references this issue
on Oct 8, 2026 Follow-up 2: we tested master-956
1b0ba10(2026-10-10 release, carrying #2119 "demote resident params to disk residency when memory reclamation fails") on the same machine (RX 6950 XT 16 GB, RDNA2, no matrix cores), same models, same commands as the matrix above (exact validated recipe:h3_ref2va --turbo, 864x480, 22 frames, seed 42, ref = 12 frames + 0.5 s WAV), machine gate-checked idle.MiniMax-H3 is still broken with the identical signature:
minimax_h3 using 54 segments->sampling failed after 325.43 s(submit OOM) -> exit 3221226505 (0xC0000409). Neither #2103 (945) nor #2119 (956) changes the DiT graph cut — the healthy control stays master-8416b3edaa(2 segments), still our production binary for H3/LTX.Updated row:
Path master-841 6b3edaamaster-864→945 master-956 1b0ba10MiniMax-H3 Ref2VA 22 f turbo (factory full path, --turbo)✅ (2 segments) ❌ 51→54 segments + ErrorOutOfDeviceMemory❌ identical: 54 segments → submit OOM at ~325 s, exit 0xC0000409 The issue stays accurate for the H3/LTX regressions. Separately, we hit what looks like a new master-956-specific regression on the Qwen-Image path (not #1976-related) — filing it as its own issue and linking it here if it turns out to share the memory-management root cause.
Addendum to the follow-up above: the Qwen-Image slowdown I mentioned (and filed as #2127, now closed) turned out to be environmental on our side — the models sit on an external USB 3.0 spinning drive and the post-graph-cut tensor re-read is disk/cache-bound, not build-related. No master-956 qwen regression; apologies for the cross-noise. The H3 datapoint in the table above (54 segments + submit OOM on 956) is unaffected by this and stands.
TL;DR: Two distinct video regressions on 16 GB Vulkan, both with deterministic A/B repro (same command/seed, control run on old build passes):
vk::Queue::submit: ErrorOutOfDeviceMemorydespite ~15 GB free VRAM — master-841 control passes with 2 segments.--auto-fit offrescues it.Image paths (Flux/SDXL/upscale) unaffected. We're pinning master-864 plus a parallel master-841 install meanwhile. Happy to bisect the 841→864 window or run any experiment.
Environment
sd-master-<sha>-bin-win-vulkan-x64.zip); the master-841 reference run used a parallel install from our archived binaries (--versionverified6b3edaa)Wan2.2-T2V-A14B-LowNoise-Q4_K_M.gguf+umt5-xxl-encoder-Q4_K_M.gguf+wan_2.1_vae.safetensors,LTX-2.5-Distilled-Q4_K_M.gguf,minimax_h3_ref2va_pruned-Q4_K_M.ggufSummary
Three video paths break differently across recent builds on a 16 GB Vulkan setup — with two distinct regression windows:
6b3edaaca37fad42d6c0a--auto-fit off)6b3edaaand these paths were simply never re-tested after upgrading; the failures below are deterministic (identical computed refusals across builds and days, on freshly booted machines).1. Wan 2.2 T2V A14B — fails at step 5/8 on master-866, works on master-864
Command (identical on both versions, seed fixed):
master-864 (working): DiT weights resident (
model manager prepared params backend buffers (9211.74 MB, 1095 tensors, 10 blocks, VRAM)); compute graph runs in 1 segment (Wan2.x-T2V-14B compute buffer size: 941.58 MB(VRAM) (peak across 1 segment)); sampling completes 8/8 (~264 s/it); VAE decode on CPU (--vae-on-cpu); webm written (2332 s total).master-866 (failing): identical weight placement (same 9211.74 MB / 1095 tensors resident in VRAM), but the graph is now cut into 42 segments (
compute buffer size: 787.47 MB(VRAM) (peak across 42 segments)), and generation dies at step 5/8 after 1390 s of sampling:The shape is reminiscent of #1946 (a few dozen MB short during weight preparation), but this is the new segmented path on a graph that master-864 kept fully resident.
Workaround (verified end-to-end on master-866): adding
--auto-fit offrestores the master-864 behavior — back to 1-segment resident compute (941.58 MB), sampling completes 8/8 (~272 s/it vs 264), webm written (2399 s total). Output is the same scene/trajectory with visible numerical divergence vs the master-864 run (middle-frame mean abs diff 14.8/255), as expected across different graph layouts.2. LTX-2.5 — broken since master-864, with AND without
--auto-fit off(no workaround found)Recipe: explicit
--backend diffusion=vulkan0,te=cpu,vae=cpu(all weights streamed from RAM: text encoder ~9.0 GB, DiT ~14.4 GB, VAE ~1.7 GB — 0 MB of weights in VRAM), 768×512, 33 frames, 8 steps with Lightricks distilled sigmas, euler_a, cfg 1.0. Last verified working on master-841; on both master-864 and master-866 it dies ~20 s in at the first workspace capacity check, with byte-identical numbers:With
--auto-fit offthe budget becomes "unlimited" but available device stays ~317 MB → same failure at 19 s. Note the paradox: with ~15 GB of VRAM physically free and 0 MB of weights resident in VRAM, the manager still reports ~320 MB "available device".I2V (65 frames, 832×480, init image) also fails on master-864 with
vk::Queue::submit: ErrorOutOfDeviceMemory~115 s into sampling; interestingly that run did not segment the graph at all, while the same command in 480×832 portrait did segment (50 segments) and still died at submit.3. MiniMax-H3 Ref2VA — broken since master-864 (same signature as master-866)
Recipe:
--backend diffusion=vulkan0,te=cpu,vae=cpu --offload-to-cpu --rng cpu --max-vram 10, turbo LoRA 8 steps, 22 frames, 864×480. On master-841 the graph-cut produced 2 segments and was stable. On both master-864 and master-866, after the CPU-side reference/audio conditioning completes (~15–20 min), the DiT reports 51 segments and dies at submit:Control run (master-841, parallel install, same machine 24 h apart, exact same command/seed/assets): PASSES —
minimax_h3 graph cut executing segment 1/2: minimax_h3.blocks.0..minimax_h3.blocks.37+segment 2/2, 8 steps × 2 segments, webm 864×480 + PCM 32 kHz stereo written in 2279 s (historical timing). This isolates the regression to the build, not the environment.What still works on master-866
Current mitigation on our workstation (until fixes land)
Two side-by-side installs, same GPU/driver, same models directory (
C:\Modeles_LLM\), only the binaries differ (sd-cli.exe --versionverified before each run):C:\SD\= master-864ca37fad(main install) — used for:-M upscale);--vae-on-cpumandatory on 16 GB (resident GPU VAE decode of 17 frames requests 19.4 GB; 33 frames gets a clean 2.2 GB workspace refusal). Validated end-to-end at 17 frames.C:\SD-6b3edaa\= master-8416b3edaa(parallel install, binaries restored from our archive into a separate directory, never mixed with the main install) — used for:Coverage: every validated recipe on this workstation has a working binary again. What we lose meanwhile: the 841→864 memory-manager improvements (fix #1946 for Wan/Flux, streaming perf) on the LTX/H3 paths — which is why a fix in the 841→864 window would let us collapse back to a single install.
Suspects
44dd137, preserve explicit backend assignments during auto-fit) is the only memory/auto-fit-related commit inca37fad...42d6c0a(the others: attention regex fix: bound plain-text runs in parse_prompt_attention regex #1919, Wan2.2 S2V feat: add Wan2.2 S2V (audio+img-to-video) support #1925, video metadata feat: Add generation parameters into video metadata #1901, MSVC warnings fix: resolve MSVC narrowing conversion warnings #1969).Happy to run any experiment that helps (bisect builds for the 841→864 window, extra logs, other flag combinations).