Skip to content

MiniMax-H3 I2V/Ref2VA conditioning hangs GPU on gfx1100 ROCm 7.2.4 & 7.14 (works on CPU backend; Unsloth fork 13b9d92) #1985

Description

@kaishi00

Summary

MiniMax-H3 Ref2VA / keyframe I2V conditioning hangs the GPU when run on the ROCm backend of the Unsloth fork build, while the identical workload completes on the CPU backend. Tested across two ROCm versions (7.2.4 and 7.14), multiple resolutions (384×640, 640×384, 544×960, 960×544), and both Studio-spawned and shell-launched invocations. Text-to-video works fine on the same installation; only the image/ref-conditioning path is affected.

Environment

  • HW: AMD Radeon RX 7900 XTX (gfx1100), 24 GB; Ryzen 9 9950X, 64 GB RAM
  • OS: Ubuntu 24.04 (LXC, GPU passthrough via /dev/kfd + /dev/dri)
  • Binary: self-compiled Unsloth fork 13b9d92 (H3 fixes) — sd-cli --mode vid_gen
  • ROCm runtime tested: 7.2.4 (segfault variant) and 7.14.0 (hang variant)
  • Models: minimax_h3_ref2va_pruned-Q5_0.gguf + minimax_h3_fl2va_pruned-Q5_0.gguf + qwen3vl 32B encoder + fp16 video VAE + fp32 audio VAE

Repro

sd-cli --mode vid_gen --offload-to-cpu \
  --backend "vae=cpu,clip=cpu,diffusion=ROCm0" --vae-conv-direct \
  --diffusion-model .../minimax_h3_ref2va_pruned-Q5_0.gguf \
  --vae .../minimax_h3_video_vae_fp16.safetensors \
  --audio-vae .../minimax_h3_audio_vae_fp32.safetensors \
  --llm .../qwen3vl_32b_minimax_h3-Q4_K_M.gguf \
  --ref-image <ref.png> \
  --prompt "lantern" --cfg-scale 1 --width 640 --height 384 --rng cpu --fps 24 \
  --video-frames 33 --steps 8 --flow-shift 12 --seed 42 \
  --output out.webm

Behavior matrix (all tested on this host)

ROCm Backend I2V (keyframe/ref) Result
7.2.4 GPU T2V 640-960 ✅ works
7.2.4 GPU I2V 640×384 ❌ SIGSEGV (exit 139) inside libamdhip64, first sampling step
7.2.4 GPU I2V 960×544 ❌ SIGSEGV
7.14 GPU I2V any resolution ❌ HANG: main thread busy-loops sched_yield() inside libamdhip64; HSA threads wait on GPU events; 0% GPU, VRAM flat, killed after 25+ min
7.14 CPU backend I2V ✅ completes (proves conditioning logic + weights are sound)

gdb backtrace (7.14 hang, main thread)

#0  __GI_sched_yield ()
#1-#11 ?? () from /opt/rocm-7.14/rocm/core-7.14/lib/libamdhip64.so.7

Thread 1 spins in a HIP busy-wait; HSA async threads wait on GPU events that never fire. A GPU kernel launched by the keyframe/ref conditioning encode never completes on gfx1100.

What we ruled out

  • Resolution (640→960, portrait and landscape)
  • Quantization (Q3_K_XL, Q4_K, Q5_0, Q6_K)
  • Image format / staged-file presence (valid PNG verified by md5 before each run)
  • Studio wrapper flags (--diffusion-fa, --max-vram -1, --stream-layers — all removed, same result)
  • --end-img presence (FLF2V full contract — same crash)
  • VRAM pressure (hang occurs at 1.5 GB used, before weights finish loading)
  • The Unsloth fork's H3 fixes (CPU backend runs the same conditioning fine — logic is sound, kernel launch is what dies)

Impact

Any I2V / Ref2VA use on gfx1100 with ROCm is blocked. The workaround is running the diffusion model on CPU too (--offload-to-cpu without any GPU backend assignment), which costs ~97 min for a 5-second clip at 544×960/30 steps vs an expected few minutes on GPU.

Happy to provide additional traces, HSA logs (HSA_ENABLE_SDMA=0 was also tested — no change), or test builds.

Activity

  1. wbruna commented on Sep 17, 2026

    @wbruna
    Contributor

    Binary: self-compiled Unsloth fork 13b9d92 (H3 fixes) — sd-cli --mode vid_gen

    Apart from 13b9d92 being on a fork, it's more than a month old. Can you reproduce this issue on latest master?

  2. kaishi00 commented on Sep 17, 2026

    @kaishi00
    Author

    Result of testing latest master (cc515a0, self-built against gfx1100):

    Short answer: the hang reproduces on master — but master also improved the conditioning stage, which changes the failure signature.

    Test env: same host as the OP (RX 7900 XTX, Ubuntu 24.04 LXC), two runtimes:

    1) master + ROCm 7.14

    • Conditioning (Ref2VA image encode with the ref image) now completes on GPU — no hang there anymore, reaches sampling 1/8 at ~2 it/s
    • Then hard hang at first sampling step: VRAM flatlines at 1.5 GB, GPU busy 0%, no further steps, killed after 12+ min. Same signature as the OP's fork builds.
    • Notable: master's auto-fit now assigns te=cpu,vae=cpu params automatically (--backend "diffusion=ROCm0,te=ROCm0,vae=ROCm0" --params-backend "te=cpu,vae=cpu"), and even with VAE compute nominally on ROCm the stall point is identical.

    2) master + ROCm 7.2.4 (the distro /opt/rocm stack)

    • Reaches sampling 1/8, then SIGSEGV — backtrace is a heap corruption signature (_int_malloc → operator new ← deep libamdhip64.so.7 recursion), i.e. the same crash window as the OP, but the failure now lands after conditioning instead of during it.

    So on current master: 7.2.4 → segfault, 7.14 → hang. Both die in/around the first Ref2VA sampling step; the CPU-backend run of the identical command still completes (weights/logic exonerated). Whatever upstream fixed in the conditioning path since the Sep 13 fork base didn't reach the underlying ROCm interaction.

    Happy to test any patch/branch, or provide full logs + gdb traces from both runs (logs preserved at /root/master_714.log, /root/master_full.log on the test box).

  3. kaishi00 commented on Sep 17, 2026

    @kaishi00
    Author

    Follow-up: reproduced the 7.14 hang on a completely fresh boot — container rebooted, zero GPU holders, VRAM 0.0 GB at start, nothing else loaded. Master cc515a0 + ROCm 7.14, same 640×384 Ref2VA command:

    • conditioning completes, sampling starts at 1.48 it/s
    • pins at 1/8 permanently: VRAM flatlines 1.5 GB, GPU busy 0%, RSS frozen at 1,056,392 kB (byte-identical across minutes), 4+ min CPU spin before kill

    So this is deterministic on the GPU path, not accumulated-state — a fresh process/host state does not clear it. Logs: /root/master_coldboot.log.

  4. kaishi00 commented on Sep 17, 2026

    @kaishi00
    Author

    Breakthrough: the same Ref2VA workload completes ALL-GPU in 75s on the Vulkan backend (master cc515a0, self-built -DGGML_VULKAN=ON, Mesa RADV, RX 7900 XTX) — no hang, no segfault, no CPU offload of the diffusion model:

    auto-fit: DiT 13245 MiB -> compute Vulkan1, params Vulkan1
    generate_video completed in 74.96s
    save result video to '/root/vulkan_gpu_i2v_test.avi'
    

    Output verified: real pixels, correct dimensions, audio present. Same command that hangs >12 min at 1/8 sampling on ROCm 7.14 (and SIGSEGVs on 7.2.4).

    So the defect is specific to the ROCm/HIP backend path on gfx1100 — the Vulkan stack on the same GPU handles the identical graph fine. This also gives AMD users a same-GPU escape hatch: build sd.cpp with GGML_VULKAN and point --backend diffusion=Vulkan1.

    Happy to test any ROCm-side patch; meanwhile Vulkan is a viable workaround that keeps diffusion fully on the 7900 XTX.

  5. wbruna commented on Sep 17, 2026

    @wbruna
    Contributor

    This could be #1947 , and the 'hang' actually an extreme slowdown; that issue mentions steps taking more than 13 minutes. One attempt to debug this further would be applying the logging patch from that issue, and if it behaves the same a HIP/rocBLAS trace.

  6. farawayso commented on Sep 17, 2026

    @farawayso

    Same gfx1100 hardware (RX 7900 XTX, 24GB), and I'm running the exact same binary commit as @kaishi00's latest-master test — cc515a0 — so I can corroborate the ROCm hang on H3 I2VA/Ref2VA conditioning from a matching setup. Two additions the matrix doesn't cover yet:

    1. ROCm version axis. I'm on ROCm 10.0 (Windows/MSVC build, HIP 7.15), well past the 7.2.4 / 7.14 in the repro. I2VA/Ref2VA conditioning on the ROCm backend is still blocked here — so this looks gfx1100 + ROCm-independent rather than a 7.x regression. (Different OS though: I'm Windows, the OP is Linux LXC, so I'm not claiming the 7.x hang reproduces on 10.x, only that the ROCm path stays unusable for H3 conditioning on the same GPU + commit.)

    2. The Vulkan escape hatch is confirmed on this same GPU + cc515a0. I don't just manually flip --backend diffusion=Vulkan1; I enforce it at the wrapper layer — i2va/fl2va/ref2va semantics all route to Vulkan unconditionally, leaving only T2VA (decode-only) on ROCm. That's because ROCm VAE encode on gfx1100 is its own separate crash (Access Violation 0xC00000FD, upstream #7691, still open), distinct from this sampling-stage hang.

    For triage: the failure point here is after conditioning completes — first sampling step, VRAM flatlines ~1.5GB, GPU 0%, main thread busy-looping sched_yield() in libamdhip64. That's a diffusion-kernel/ROCm hang, not the VAE-encode 0xC00000FD stack overflow from #7691. So I'd keep it as a separate gfx1100+ROCm issue rather than a duplicate of #7691 (but it's the same family: "H3 conditioning path dies on ROCm/gfx1100").

    Happily test any patch/branch on ROCm 10.0 + Windows or drop HIP/rocBLAS traces — let me know the logging patch from #1947 you want applied and I'll run it.

  7. kaishi00 commented on Sep 17, 2026

    @kaishi00
    Author

    Regression-hunt result (Outcome B): does NOT reproduce on our environment — 6b3edaa falls back to CPU at the same resolutions as current master.

    Tested the reported pre-regression revision on our RX 7900 XTX / Ubuntu 24.04 / Mesa RADV 25.2.8 / Vulkan loader 1.3.275:

    Resolution old 6b3edaa current master cc515a0
    640×384 33f ✅ GPU compute confirmed (GPU 100% sustained), 100 s/it, completed ✅ GPU confirmed, 0.5 s/it, 75 s total
    864×480 33f ⚠️ mixed — completed, 360 s/it, ggml_vec_dot_f32 in hot stacks, ~11 min CPU burn/step, GPU bursts 26-100% alternating with 0%, pinned memory allocation failed ×1 ❌ CPU fallback (from earlier comment)
    960×544 73f ❌ CPU fallback confirmed — 1448 s/it (24 min/step), 17 min CPU/step, ggml_vec_dot_f32/mul_mat in hot stacks, VRAM residency 25.6GB while GPU 0-3% ❌ same (earlier comment)

    Key observations:

    1. The old revision is not better here — it degrades at the same graph sizes, just with slower, older shaders when it does use the GPU (100 s/it vs 0.5 s/it at 640×384).
    2. The failure correlates with RADV's pinned-memory allocation failure ("Requested buffer size exceeds device buffer size limit") appearing right as the compute falls into CPU ggml_vec_dot_f32. VRAM can still hold the weights (25.6GB reported) while the matmuls silently run on CPU — residency ≠ compute, as the earlier comments established.
    3. So on our hardware the large-graph CPU fallback looks like a ggml-Vulkan/RADV buffer-limit interaction, not a stable-diffusion.cpp code regression. The upstream reporter's success at ~864×480 on the old revision likely reflects different buffer limits (NVIDIA, different GPU, or different driver) rather than a code change between 6b3edaa and master.

    Classification per the issue's framework: Outcome B — not reproduced. We're not pinning old code; our working all-GPU path is current master Vulkan at ≤640×384-class graphs, and larger graphs need either graph chunking in ggml-vulkan or a maxAllocationSize accommodation.

  8. kaishi00 commented on Sep 17, 2026

    @kaishi00
    Author

    Follow-up with the #1947 diagnostic patch + HIP tracing: this does not match #1947's finite first-step slowdown.

    I instrumented the same cc515a0 / ROCm 7.14 / gfx1100 Ref2VA repro and left it running for 3h30m after it reached sampling 1/8. It never reached 2/8.

    Conditioning is healthy:

    VAE prep:    0.30s
    VAE compute: 0.49s on ROCm0
    outputs:     0.00s
    

    The failure localizes to the first Ref2VA denoiser graph launch. The last HIP trace sequence is:

    hipGraphExecUpdate: Returned hipSuccess
    hipGraphLaunch
    GraphExec::Run max_streams: 1
    UpdateStreams failed for device id:0
    

    After that:

    • no subsequent useful kernel activity is observed
    • GPU remains 0% busy
    • VRAM remains flat at ~1.5 GB
    • the main thread spins in sched_yield() inside libamdhip64.so.7
    • all-thread gdb backtrace contains the corresponding HSA/HIP wait stack
    • the compute call never returns

    The process was finally terminated after 3h30m30s and exited cleanly on SIGTERM.

    So this appears materially different from #1947: there, the pathological first compute eventually returned after ~731s and later steps ran normally. Here, the first denoiser compute is non-returning and the trace ends immediately after the HIP graph stream-update failure.

    I'm treating UpdateStreams failed for device id:0 as the failure location, not claiming it is the deeper root cause.

    I have the full artifacts preserved:

    • instrumented ggml_runner.cpp patch
    • complete debug log / AMD trace
    • command/environment/build metadata
    • main-thread and all-thread gdb backtraces
    • timeline samples

    The full trace is ~200 MB, so I can upload a compressed archive or any narrower subset that would be more useful.

  9. farawayso commented on Sep 17, 2026

    @farawayso

    Thanks for the follow-up. We have a complementary data point from the same gfx1100 hardware, and the two cases line up nicely:

    • Env: gfx1100 (RX 7900 XTX 24GB), ROCm 10.0, sdcpp at recent commit, Ref2VA/I2VA paths
    • Our first-step trace shows the same API-call overhead pattern, but it eventually returns after ~732s (vs. your non-returning case). The fast/slow traces have identical ~1.3M-event API call structure (PAL-VMM heap and launch-config events each ~31%, GEMM only ~0.7%) — so the 700s delta is pure HIP API call-interval overhead, not compute.
    • This suggests your hard hang (UpdateStreams failed for device id:0 → main thread spinning in sched_yield) is the same pathological HIP graph/stream-update path, just without the eventual-return behavior.

    We have the full fast (37s) vs. slow (732s) traces captured locally. Happy to share a compressed subset (or specific rocm-instrument / timeline windows) if it helps narrow down whether the 7.x regression is in the graph-update path or the stream-sync fallback. Also relevant to our side: we have a runtime guard that demotes I2VA to the Vulkan backend on ROCm, so the hang doesn't block generation — if your artifacts include the instrumented ggml_runner.cpp patch, that would let us A/B it against our guard.

    Note: this analysis was drafted with help from Hermes Agent running on a free-tier LLM — the trace numbers were verified locally, but the interpretation may not be fully accurate, so please treat the root-cause claims as hypotheses.

  10. basementops commented on Oct 6, 2026

    @basementops

    Same family on gfx1151 (Strix Halo), but in the MiniMax-H3 video VAE encode, sporadic, and avoided by disabling HIP graphs

    Env: Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151, UMA), Debian 13, ROCm 7.2.4, built with -DSD_HIPBLAS=ON -DAMDGPU_TARGETS=gfx1151 -DGGML_HIP_GRAPHS=ON.
    Command: I2VA, --init-img, 704×704, --video-frames 49, --steps 2, --cfg-scale 1 --diffusion-fa --offload-to-cpu --rng cpu, minimax_h3_fl2va_pruned-Q4_K.gguf + qwen3vl_32b_minimax_h3-Q4_K_M.gguf + fp16 video VAE + fp32 audio VAE.

    • Hangs sporadically during the VAE encode of the init image, right after ggml_backend_cuda_graph_compute: CUDA graph warmup complete, at varying tiles (2/16, 12/16). GPU busy 100 %, one CPU core spinning. Main thread: sched_yield() in libamdhip64 ← ggml_cuda_graph_evaluate_and_capture ← MiniMaxH3VideoVAERunner::encode (full backtrace available).
    • Rate with graphs enabled: 11/31 runs on b167b94, 1/10 on current master a1ded76 (same location; master tiles 3×3).
    • With GGML_CUDA_DISABLE_GRAPHS=1: 0/10 hangs, interleaved with graph-enabled runs that did hang. Vulkan build: 0/5.
    • T2VA is unaffected.

    Happy to run a patch or collect a HIP trace.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions