Skip to content

Video regressions on 16 GB Vulkan: LTX-2.5 + MiniMax-H3 broken since master-864, Wan 2.2 broken by master-866 (graph segmentation / weight-budget behavior) — --auto-fit off only rescues Wan #1976

Description

@laurentvv

TL;DR: Two distinct video regressions on 16 GB Vulkan, both with deterministic A/B repro (same command/seed, control run on old build passes):

  1. LTX-2.5 + MiniMax-H3 break in the 841→864 window (memory-manager rework refactor: unify runner lifecycles and weight residency #1940/refactor: unify model source and weight lifecycle management #1956/refactor: split generation pipeline out of stable-diffusion.cpp #1957): graph now cut into 50+ segments, dies with vk::Queue::submit: ErrorOutOfDeviceMemory despite ~15 GB free VRAM — master-841 control passes with 2 segments.
  2. Wan 2.2 T2V breaks in 864→866 (suspect feat: preserve explicit backend assignments during auto-fit #1967): DiT graph switches 1 segment → 42 segments, dies at step 5/8; --auto-fit off rescues it.

Image paths (Flux/SDXL/upscale) unaffected. We're pinning master-864 plus a parallel master-841 install meanwhile. Happy to bisect the 841→864 window or run any experiment.


Environment

  • GPU: AMD Radeon RX 6950 XT 16 GB (RDNA2, Vulkan, proprietary driver, no matrix cores)
  • OS: Windows x64
  • Binaries: official release assets (sd-master-<sha>-bin-win-vulkan-x64.zip); the master-841 reference run used a parallel install from our archived binaries (--version verified 6b3edaa)
  • Machine state for all runs: freshly booted or verified idle, 1.1–1.3/16 GB VRAM in use
  • Models: Wan2.2-T2V-A14B-LowNoise-Q4_K_M.gguf + umt5-xxl-encoder-Q4_K_M.gguf + wan_2.1_vae.safetensors, LTX-2.5-Distilled-Q4_K_M.gguf, minimax_h3_ref2va_pruned-Q4_K_M.gguf

Summary

Three video paths break differently across recent builds on a 16 GB Vulkan setup — with two distinct regression windows:

master-841 6b3edaa master-864 ca37fad master-866 42d6c0a
Wan 2.2 T2V 14B ❌ (#1946, fixed later) ✅ ❌ (works with --auto-fit off)
LTX-2.5 (distilled, T2V & I2V) ✅ ❌ ❌ (identical failure)
MiniMax-H3 Ref2VA (turbo LoRA) ✅ ❌ ❌ (identical failure)
Flux / SDXL / ESRGAN upscale ✅/✅/✅ ✅/✅/✅ ✅/✅/✅
  • LTX-2.5 and MiniMax-H3 broke in the 841→864 window (the memory-manager rework), not in 864→866. We had initially mislabeled master-864 as the last good build for these two — in fact all our LTX/H3 validations had been done on 6b3edaa and these paths were simply never re-tested after upgrading; the failures below are deterministic (identical computed refusals across builds and days, on freshly booted machines).
  • Wan 2.2 broke in the 864→866 window (suspect: feat: preserve explicit backend assignments during auto-fit #1967), where the DiT graph switches from fully resident (1 segment) to a 42-segment cut that fails mid-sampling.

1. Wan 2.2 T2V A14B — fails at step 5/8 on master-866, works on master-864

Command (identical on both versions, seed fixed):

sd-cli.exe -M vid_gen \
  --diffusion-model Wan2.2-T2V-A14B-LowNoise-Q4_K_M.gguf \
  --t5xxl umt5-xxl-encoder-Q4_K_M.gguf \
  --vae wan_2.1_vae.safetensors \
  -p "a red fox walking through deep snow, winter forest, cinematic" \
  --cfg-scale 6.0 --steps 8 --sampling-method euler \
  -W 832 -H 480 --video-frames 17 --fps 24 \
  --diffusion-fa --temporal-tiling --vae-on-cpu --seed 42

master-864 (working): DiT weights resident (model manager prepared params backend buffers (9211.74 MB, 1095 tensors, 10 blocks, VRAM)); compute graph runs in 1 segment (Wan2.x-T2V-14B compute buffer size: 941.58 MB(VRAM) (peak across 1 segment)); sampling completes 8/8 (~264 s/it); VAE decode on CPU (--vae-on-cpu); webm written (2332 s total).

master-866 (failing): identical weight placement (same 9211.74 MB / 1095 tensors resident in VRAM), but the graph is now cut into 42 segments (compute buffer size: 787.47 MB(VRAM) (peak across 42 segments)), and generation dies at step 5/8 after 1390 s of sampling:

[WARN   ] model_manager.cpp:1767 - model manager cannot make enough memory available on Vulkan0: need 664.34 MB device / 1092.16 MB budget, available 631.30 MB device / 2366.17 MB budget
[ERROR  ] ggml_runner.cpp:877  - Wan2.x-T2V-14B segment 24/42 (wan.blocks.22) failed during weight preparation
[ERROR  ] diffusion_engine.cpp:2421 - diffusion model compute failed

The shape is reminiscent of #1946 (a few dozen MB short during weight preparation), but this is the new segmented path on a graph that master-864 kept fully resident.

Workaround (verified end-to-end on master-866): adding --auto-fit off restores the master-864 behavior — back to 1-segment resident compute (941.58 MB), sampling completes 8/8 (~272 s/it vs 264), webm written (2399 s total). Output is the same scene/trajectory with visible numerical divergence vs the master-864 run (middle-frame mean abs diff 14.8/255), as expected across different graph layouts.

2. LTX-2.5 — broken since master-864, with AND without --auto-fit off (no workaround found)

Recipe: explicit --backend diffusion=vulkan0,te=cpu,vae=cpu (all weights streamed from RAM: text encoder ~9.0 GB, DiT ~14.4 GB, VAE ~1.7 GB — 0 MB of weights in VRAM), 768×512, 33 frames, 8 steps with Lightricks distilled sigmas, euler_a, cfg 1.0. Last verified working on master-841; on both master-864 and master-866 it dies ~20 s in at the first workspace capacity check, with byte-identical numbers:

[WARN   ] model_manager.cpp:1767 - model manager cannot make enough memory available on Vulkan0: need 910.16 MB device / 398.16 MB budget, available 319.37 MB device / 668.90 MB budget
[ERROR  ] ggml_runner.cpp:877  - ltxav segment 1/1 (graph) failed during workspace capacity check
[ERROR  ] video.cpp:1684 - sampling failed after 21.24s

With --auto-fit off the budget becomes "unlimited" but available device stays ~317 MB → same failure at 19 s. Note the paradox: with ~15 GB of VRAM physically free and 0 MB of weights resident in VRAM, the manager still reports ~320 MB "available device".

I2V (65 frames, 832×480, init image) also fails on master-864 with vk::Queue::submit: ErrorOutOfDeviceMemory ~115 s into sampling; interestingly that run did not segment the graph at all, while the same command in 480×832 portrait did segment (50 segments) and still died at submit.

3. MiniMax-H3 Ref2VA — broken since master-864 (same signature as master-866)

Recipe: --backend diffusion=vulkan0,te=cpu,vae=cpu --offload-to-cpu --rng cpu --max-vram 10, turbo LoRA 8 steps, 22 frames, 864×480. On master-841 the graph-cut produced 2 segments and was stable. On both master-864 and master-866, after the CPU-side reference/audio conditioning completes (~15–20 min), the DiT reports 51 segments and dies at submit:

[VERBOSE] ggml_runner.cpp:857  - minimax_h3 using 51 segments
[ERROR   ] ggml_runner.cpp:651  - minimax_h3 graph execution failed on Vulkan0: vk::Queue::submit: ErrorOutOfDeviceMemory
[ERROR   ] video.cpp:1684 - sampling failed after 262.86s   (master-864)  /  358.87s  (master-866)

Control run (master-841, parallel install, same machine 24 h apart, exact same command/seed/assets): PASSES — minimax_h3 graph cut executing segment 1/2: minimax_h3.blocks.0..minimax_h3.blocks.37 + segment 2/2, 8 steps × 2 segments, webm 864×480 + PCM 32 kHz stereo written in 2279 s (historical timing). This isolates the regression to the build, not the environment.

What still works on master-866

  • Flux schnell smoke (512×512, 4 steps, seed 42): clean image
  • SDXL (monolithic checkpoint, 512×512, 20 steps): clean image
  • ESRGAN upscale mode (4x-UltraSharp): fine

Current mitigation on our workstation (until fixes land)

Two side-by-side installs, same GPU/driver, same models directory (C:\Modeles_LLM\), only the binaries differ (sd-cli.exe --version verified before each run):

  1. C:\SD\ = master-864 ca37fad (main install) — used for:
    • image paths: Flux schnell/dev, SDXL, ESRGAN upscale (-M upscale);
    • Wan 2.2 T2V and I2V (A14B LowNoise Q4_K_M): ≤ ~20 frames with --vae-on-cpu mandatory on 16 GB (resident GPU VAE decode of 17 frames requests 19.4 GB; 33 frames gets a clean 2.2 GB workspace refusal). Validated end-to-end at 17 frames.
  2. C:\SD-6b3edaa\ = master-841 6b3edaa (parallel install, binaries restored from our archive into a separate directory, never mixed with the main install) — used for:
    • MiniMax-H3 Ref2VA production: re-validated 2026-09-15 with the exact turbo recipe (2-segment graph cut × 8 steps, webm 864×480 + PCM 32 kHz stereo in 2279 s);
    • LTX-2.5 production: last validated on this build (33-frame benchmark 2026-09-10, 65-frame I2V production chains 2026-09-12/13); not re-run since the A/B, same binary as the H3 control.

Coverage: every validated recipe on this workstation has a working binary again. What we lose meanwhile: the 841→864 memory-manager improvements (fix #1946 for Wan/Flux, streaming perf) on the LTX/H3 paths — which is why a fix in the 841→864 window would let us collapse back to a single install.

Suspects

Happy to run any experiment that helps (bisect builds for the 841→864 window, extra logs, other flag combinations).

Activity

  1. added a commit that references this issue on Sep 14, 2026
  2. changed the title [-]Regression master-864 → master-866 (16 GB Vulkan): Wan 2.2 / LTX-2.5 / MiniMax-H3 video generation all fail — memory manager segmentation/budget behavior changed (--auto-fit off only rescues Wan)[/-] [+]Video regressions on 16 GB Vulkan: LTX-2.5 + MiniMax-H3 broken since master-864, Wan 2.2 broken by master-866 (graph segmentation / weight-budget behavior) — --auto-fit off only rescues Wan[/+] on Sep 15, 2026
  3. laurentvv commented on Sep 23, 2026

    @laurentvv
    Author

    Follow-up: we tested master-899 28b454b (2026-09-22 release, includes #2019 and #2020) on the same machine, same models, same commands. Full matrix below — net result: the Wan 2.2 regression is gone, LTX-2.5 now passes but is VRAM-margin sensitive, MiniMax-H3 is still broken with the identical signature.

    Results on master-899 (freshly booted, desktop VRAM usage 0.4/16 GB)

    Path master-841 6b3edaa master-864/866 master-899 28b454b
    Flux 1024² (dev, 4 steps) ✅ ✅ ✅
    Wan 2.2 T2V 17f (exact recipe above) ❌ (#1946) ✅ (864) / ❌ (866) ✅ — sampling 2023 s vs 2116 s on 864 (−4 %), no --auto-fit off needed
    LTX-2.5 T2V 33f benchmark (2 steps, 768×512) ✅ ❌ ✅ on fresh boot — 53.75 s sampling, webm written — but ❌ on the same build before a reboot, see below
    MiniMax H3 Ref2VA turbo (exact recipe above) ✅ (2 segments) ❌ ❌ unchanged: minimax_h3 using 51 segments + vk::Queue::submit: ErrorOutOfDeviceMemory, sampling failed after 276.68 s (864: 262.86 s, 866: 358.87 s) — deterministic, reproduced on a freshly rebooted machine

    1. LTX-2.5 on master-899: passes, but the capacity check is desktop-VRAM-margin sensitive

    Same command, same build, same machine, ~40 minutes apart:

    Before reboot (desktop/apps holding ~0.8 GB VRAM) — fails at 21.5 s:

    [WARN   ] model_manager.cpp:1809 - model manager cannot make enough memory available on Vulkan0: need 908.75 MB device / 396.75 MB budget, available 787.26 MB device / 668.90 MB budget
    [ERROR  ] ggml_runner.cpp:880  - ltxav segment 1/1 (graph) failed during workspace capacity check
    [ERROR  ] video.cpp:1715 - sampling failed after 21.51s
    

    After reboot (desktop VRAM usage 0.8 → 0.4 GB) — passes end-to-end (53.75 s sampling, webm written, one non-fatal Failed to allocate pinned memory WARN).

    So the failure margin here was ~121 MB of "available device". Two observations worth flagging:

    • The available-device under-reporting seems greatly reduced by fix: handle GPU memory reports and LLM encoding failures #2020. In our master-864/866 runs (original report above) the manager claimed available 319.37 MB device while ~15 GB was physically free and 0 MB of weights were resident in VRAM. On master-899 the same staging reports available 787.26 MB device / total 16368.00 MB, tracked weights 14388.75 MB — consistent with the actual free VRAM minus non-tracked overhead. Whatever ghost reservation produced the 319 MB figure appears gone (matches fix: handle GPU memory reports and LLM encoding failures #2020 "handle GPU memory reports").
    • The remaining ~120 MB shortfall was real desktop VRAM held by other processes. Practical consequence: on a 16 GB card the LTX-2.5 33-frame capacity check now passes but with little headroom — anything else holding a few hundred MB of VRAM flips it back to a refusal. We keep our master-841 parallel install for LTX/H3 production for that reason (our LTX I2V 65-frame production chains are also not re-validated on 899 yet).

    2. MiniMax-H3: still the blocker

    The 841→864 regression is not fixed by #2019/#2020: the DiT graph is still cut into 51 segments (control build: 2) and dies at submit with ErrorOutOfDeviceMemory several minutes into sampling, deterministically, on a fresh boot. H3 production on our side still runs on the master-841 parallel install. Happy to test any candidate fix or bisect the 841→864 window for this path — we have byte-stable A/B repro (same command/seed/assets, control webm from the 841 binary).

  4. laurentvv commented on Sep 24, 2026

    @laurentvv
    Author

    Follow-up: we tested master-908 88411ef (2026-09-23 release, includes #1900 "add graph cuts for MiniMax-H3 text conditioning") on the same machine, same models, same commands as the matrix above, freshly rebooted (desktop VRAM usage 0.4/16 GB). Net result: Flux/LTX/Wan are stable or slightly faster, but MiniMax-H3 is still broken — the graph is now cut into 54 segments (was 51 on 864/866/899, control = 2 on 841) and still dies at submit with ErrorOutOfDeviceMemory.

    Results on master-908 vs previous builds

    Path master-841 6b3edaa master-864→899 master-908 88411ef
    Flux 1024² (dev, 4 steps) ✅ ✅ 26.98 s (899) ✅ 27.00 s
    LTX-2.5 T2V 33f benchmark (2 steps, 768×512) ✅ ✅ on fresh boot — 53.75 s (899) ✅ 54.21 s, 1 segment, no capacity-check refusal, same non-fatal Failed to allocate pinned memory WARN
    Wan 2.2 T2V 17f (exact recipe above) ✅ ✅ 2023 s (899, −4 % vs 864) ✅ 1999.5 s (−1.2 % vs 899), 1 resident segment (9211.74 MB weights, identical to 864/899), VAE CPU decode 209.5 s
    MiniMax H3 Ref2VA turbo (exact recipe above) ✅ (2 segments) ❌ using 51 segments + OOM at submit (899: 276.68 s) ❌ using 54 segments + OOM at submit, sampling failed after 257.37 s

    #1900 changes the H3 graph cut (51 → 54 segments) but does not restore the healthy plan

    Full failure sequence on master-908 (same command/seed/assets as the control webm produced by the 841 binary):

    [VERBOSE] vae.hpp:262  - computing vae encode graph completed, taking 855.82s
    [VERBOSE] conditioner.hpp:3139 - computing condition graph completed, taking 92884 ms
    [VERBOSE] ggml_runner.cpp:864  - minimax_h3 using 54 segments
    [ERROR  ] ggml_runner.cpp:654  - minimax_h3 graph execution failed on Vulkan0: vk::Queue::submit: ErrorOutOfDeviceMemory
    [ERROR  ] diffusion_engine.cpp:2662 - Diffusion model sampling failed
    [ERROR  ] video.cpp:1709 - sampling failed after 257.37s
    

    Two observations:

    • The segment count changed across the range that includes fix: reduce MiniMax-H3 VRAM spikes and redundant token refinemen #1900 (51 on 864/866/899 → 54 on 908), so the commit does affect how the H3 DiT graph is partitioned — but the result is still a fully fragmented plan (54 segments) versus the healthy 2-segment cut on master-841 (segment 1/2: blocks.0..37 + segment 2/2), and the deterministic ErrorOutOfDeviceMemory at submit remains. The 841→864 regression reported here is still open on master-908.
    • Everything else in the matrix is unchanged or marginally better (Wan −1.2 % sampling vs 899, LTX passes on a fresh boot with the same ~120 MB desktop-VRAM sensitivity noted above), so we kept 908 pinned — H3/LTX production on our side still runs on the master-841 parallel install.

    Happy to test any candidate fix — we have a byte-stable A/B repro (same command/seed/assets, control webm + logs from the 841 binary).

  5. laurentvv commented on Oct 1, 2026

    @laurentvv
    Author

    Follow-up: we tested master-929 3f8527a on the same machine (RX 6950 XT 16 GB,
    RDNA2, no matrix cores), same models, same commands as the matrix above, machine idle
    (gate-checked). Net result: Wan 2.2 stays fixed and faster; the Wan I2V full path now
    also passes; LTX-2.5 33f now fails EARLY and EXPLICITLY (the new memory manager reports
    the numbers instead of crashing); MiniMax-H3 is still broken with the identical
    54-segment OOM signature.

    Results on master-929

    Path master-841 6b3edaa master-864→908 master-929 3f8527a
    Flux 1024² (dev, 4 steps) ✅ ✅ ✅
    Wan 2.2 T2V 17f ✅ ✅ (fixed by 899) ✅ (+8.4% vs 841)
    Wan I2V full path (25 steps, --vae-on-cpu) ❌ 0xC0000409 (Sep 25) ? ✅ PASS (52 min, healthy)
    LTX-2.5 33f distilled (Q4_K_M, 768×512) ❌ (machine-state) ❌ ❌ explicit refusal (below)
    MiniMax-H3 turbo Ref2VA (validated recipe, 22f + LoRA) ✅ 2 segments ❌ 51→54 segments, OOM ❌ still 54 segments, OOM at submit

    LTX-2.5 33f on master-929 — no more silent crash, the memory manager now refuses
    cleanly with the full picture (this looks like a 16 GB fit problem, not a crash bug):
    stages 14,388.75 MB (4,349 tensors, 15 blocks) onto Vulkan0, leaving 776.09 MB free,
    then: model manager cannot make enough memory available on Vulkan0: need 908.75 MB device / 396.75 MB budget, available 776.09 MB device / 668.90 MB budget →
    ltxav segment 1/1 (graph) failed during workspace capacity check at 21 s.
    Question: should --offload-to-cpu stream the LTX DiT weights for vid_gen (it does
    rescue Wan 14B for us), or is 33f @ Q4_K_M simply above the 16 GB envelope?

    MiniMax-H3 turbo on master-929 — retested today with the full validated recipe
    (--turbo LoRA 8 steps, 864×480, 22 frames @ 24 fps, --offload-to-cpu --max-vram 10 --rng cpu, te+vae on CPU, ref tail + LoRA): minimax_h3 using 54 segments →
    graph execution failed on Vulkan0: vk::Queue::submit: ErrorOutOfDeviceMemory (exit
    0xC0000409). Control master-841 = 2 segments, runs fine (re-validated on the
    parallel build). The #1900 graph cuts (51→54 segments) did not close the gap.
    The question from the matrix still stands: is 16 GB below the supported envelope for
    H3 even with offload + max-vram, or should the graph segmentation eventually make it
    fit? Happy to test any candidate build — this machine is our standard 16 GB Vulkan
    datapoint.

  6. nebkmb commented on Oct 2, 2026

    @nebkmb

    I have similar problems, spent hours testing various configurations, various sd.cpp builds. Minimax H3 support is hopelessly broken and memory inefficient with stable-diffusion.cpp. Can only generate short clips on low res even with the 3090. Comfy can do 15 sec 1344x768 (H3's specified max) clips with H3 even on my 12 gb laptop 4080 without additional memory optimization nodes.

  7. laurentvv commented on Oct 2, 2026

    @laurentvv
    Author

    Confirming @nebkmb's findings with our own matrix (RX 6950 XT 16 GB, RDNA2, no matrix cores, machine gate-checked idle):

    • MiniMax-H3: full validated recipe on master-929 3f8527a -> minimax_h3 using 54 segments -> vk::Queue::submit: ErrorOutOfDeviceMemory, exit 0xC0000409. Same signature on master-841. Since it also fails on your 24 GB 3090, this is not a 16 GB envelope issue — the segmentation path itself looks broken/memory-inefficient in sd.cpp (ComfyUI doing 15 s @ 1344x768 on a 12 GB 4080 corroborates that).
    • LTX-2.5: on master-929 the memory manager stages 14,388.75 MB (4,349 tensors) onto Vulkan0, leaves 838.65 MB free, then refuses cleanly: cannot make enough memory available on Vulkan0: need 908.56 MB device / 396.56 MB budget -> ltxav segment 1/1 (graph) failed during workspace capacity check (33 frames @ 768x512, 2 steps). LTX-2.5 33f does not fit on 16 GB with 929, while it runs fine on master-841.
    • Tested today: --params-backend disk does not change the LTX outcome on 929 — same staging onto Vulkan0, same workspace refusal.

    Net on RDNA2 16 GB: Wan 2.1/2.2 (incl. I2V + VACE) fixed and faster on 929; LTX-2.5 fits only on 841; H3 broken on both. Happy to test candidate patches/builds — the H3 envelope and the LTX offload behavior remain the two open questions from our previous comment.

  8. laurentvv commented on Oct 8, 2026

    @laurentvv
    Author

    Follow-up: we tested master-945 a1ded76 (2026-10-07 release, first build carrying the
    H3-path change #2103 "keep MiniMax-H3 VAE weights resident across temporal chunks") on the
    same machine (RX 6950 XT 16 GB, RDNA2, no matrix cores), same models, same commands as the
    matrix above, machine gate-checked idle. Net result: MiniMax-H3 is still broken with the
    identical signature, and LTX-2.5 still fails the same workspace-capacity refusal.

    Results on master-945

    Path master-841 6b3edaa master-864→929 master-945 a1ded76
    MiniMax-H3 Ref2VA 22 f turbo (factory full path, --turbo) ✅ (2 segments) ❌ 51→54 segments + ErrorOutOfDeviceMemory ❌ identical: minimax_h3 using 54 segments → vk::Queue::submit: ErrorOutOfDeviceMemory at step ~10/28, exit 0xC0000409
    LTX-2.5 T2V 33 f (2-step distilled benchmark) ✅ ❌ early explicit refusal (908 MB / 776 free on 929) ❌ same class: ltxav segment 1/1 failed during workspace capacity check — tracked weights 14,388.75 MB on Vulkan0, need 893.75 MB device, reported free 694.57 MB
    Flux 1024² ✅ ✅ ✅ (real-path smoke, 160 s)

    Notes:

    • fix: keep MiniMax-H3 VAE weights resident across temporal chunks #2103 did not change the H3 graph segmentation on our envelope: still 54 segments
      (control = 2 on 841). The VAE-residency change may help other configurations, but the
      DiT graph cut is unchanged and the submit OOM is byte-for-byte the 929 signature.
    • LTX numbers on 945 are the same failure mode as reported for 929 in the 10-01 comment;
      --params-backend disk was already tested negative on 929 (10-02).
    • Production on our side remains the pinned master-841 parallel install for H3 + LTX.

    Not closing — the issue stays accurate for both regressions on 16 GB Vulkan. Happy to run
    any experiment on 945+ (bisect offers from the original comment still stand).

  9. laurentvv commented on Oct 11, 2026

    @laurentvv
    Author

    Follow-up 2: we tested master-956 1b0ba10 (2026-10-10 release, carrying #2119 "demote resident params to disk residency when memory reclamation fails") on the same machine (RX 6950 XT 16 GB, RDNA2, no matrix cores), same models, same commands as the matrix above (exact validated recipe: h3_ref2va --turbo, 864x480, 22 frames, seed 42, ref = 12 frames + 0.5 s WAV), machine gate-checked idle.

    MiniMax-H3 is still broken with the identical signature: minimax_h3 using 54 segments -> sampling failed after 325.43 s (submit OOM) -> exit 3221226505 (0xC0000409). Neither #2103 (945) nor #2119 (956) changes the DiT graph cut — the healthy control stays master-841 6b3edaa (2 segments), still our production binary for H3/LTX.

    Updated row:

    Path master-841 6b3edaa master-864→945 master-956 1b0ba10
    MiniMax-H3 Ref2VA 22 f turbo (factory full path, --turbo) ✅ (2 segments) ❌ 51→54 segments + ErrorOutOfDeviceMemory ❌ identical: 54 segments → submit OOM at ~325 s, exit 0xC0000409

    The issue stays accurate for the H3/LTX regressions. Separately, we hit what looks like a new master-956-specific regression on the Qwen-Image path (not #1976-related) — filing it as its own issue and linking it here if it turns out to share the memory-management root cause.

  10. laurentvv commented on Oct 11, 2026

    @laurentvv
    Author

    Addendum to the follow-up above: the Qwen-Image slowdown I mentioned (and filed as #2127, now closed) turned out to be environmental on our side — the models sit on an external USB 3.0 spinning drive and the post-graph-cut tensor re-read is disk/cache-bound, not build-related. No master-956 qwen regression; apologies for the cross-noise. The H3 datapoint in the table above (54 segments + submit OOM on 956) is unaffected by this and stands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions