Repository navigation
MiniMax-H3 I2V/Ref2VA conditioning hangs GPU on gfx1100 ROCm 7.2.4 & 7.14 (works on CPU backend; Unsloth fork 13b9d92) #1985
Description
Activity
Binary: self-compiled Unsloth fork
13b9d92(H3 fixes) —sd-cli --mode vid_genApart from 13b9d92 being on a fork, it's more than a month old. Can you reproduce this issue on latest master?
Result of testing latest master (cc515a0, self-built against gfx1100):
Short answer: the hang reproduces on master — but master also improved the conditioning stage, which changes the failure signature.
Test env: same host as the OP (RX 7900 XTX, Ubuntu 24.04 LXC), two runtimes:
1) master + ROCm 7.14
- Conditioning (Ref2VA image encode with the ref image) now completes on GPU — no hang there anymore, reaches sampling
1/8at ~2 it/s - Then hard hang at first sampling step: VRAM flatlines at 1.5 GB, GPU busy 0%, no further steps, killed after 12+ min. Same signature as the OP's fork builds.
- Notable: master's auto-fit now assigns
te=cpu,vae=cpuparams automatically (--backend "diffusion=ROCm0,te=ROCm0,vae=ROCm0" --params-backend "te=cpu,vae=cpu"), and even with VAE compute nominally on ROCm the stall point is identical.
2) master + ROCm 7.2.4 (the distro
/opt/rocmstack)- Reaches sampling
1/8, then SIGSEGV — backtrace is a heap corruption signature (_int_malloc→operator new← deeplibamdhip64.so.7recursion), i.e. the same crash window as the OP, but the failure now lands after conditioning instead of during it.
So on current master: 7.2.4 → segfault, 7.14 → hang. Both die in/around the first Ref2VA sampling step; the CPU-backend run of the identical command still completes (weights/logic exonerated). Whatever upstream fixed in the conditioning path since the Sep 13 fork base didn't reach the underlying ROCm interaction.
Happy to test any patch/branch, or provide full logs + gdb traces from both runs (logs preserved at
/root/master_714.log,/root/master_full.logon the test box).- Conditioning (Ref2VA image encode with the ref image) now completes on GPU — no hang there anymore, reaches sampling
Follow-up: reproduced the 7.14 hang on a completely fresh boot — container rebooted, zero GPU holders, VRAM 0.0 GB at start, nothing else loaded. Master cc515a0 + ROCm 7.14, same 640×384 Ref2VA command:
- conditioning completes, sampling starts at 1.48 it/s
- pins at
1/8permanently: VRAM flatlines 1.5 GB, GPU busy 0%, RSS frozen at 1,056,392 kB (byte-identical across minutes), 4+ min CPU spin before kill
So this is deterministic on the GPU path, not accumulated-state — a fresh process/host state does not clear it. Logs:
/root/master_coldboot.log.Breakthrough: the same Ref2VA workload completes ALL-GPU in 75s on the Vulkan backend (master cc515a0, self-built
-DGGML_VULKAN=ON, Mesa RADV, RX 7900 XTX) — no hang, no segfault, no CPU offload of the diffusion model:auto-fit: DiT 13245 MiB -> compute Vulkan1, params Vulkan1 generate_video completed in 74.96s save result video to '/root/vulkan_gpu_i2v_test.avi'Output verified: real pixels, correct dimensions, audio present. Same command that hangs >12 min at
1/8sampling on ROCm 7.14 (and SIGSEGVs on 7.2.4).So the defect is specific to the ROCm/HIP backend path on gfx1100 — the Vulkan stack on the same GPU handles the identical graph fine. This also gives AMD users a same-GPU escape hatch: build sd.cpp with GGML_VULKAN and point
--backend diffusion=Vulkan1.Happy to test any ROCm-side patch; meanwhile Vulkan is a viable workaround that keeps diffusion fully on the 7900 XTX.
This could be #1947 , and the 'hang' actually an extreme slowdown; that issue mentions steps taking more than 13 minutes. One attempt to debug this further would be applying the logging patch from that issue, and if it behaves the same a HIP/rocBLAS trace.
Same gfx1100 hardware (RX 7900 XTX, 24GB), and I'm running the exact same binary commit as @kaishi00's latest-master test — cc515a0 — so I can corroborate the ROCm hang on H3 I2VA/Ref2VA conditioning from a matching setup. Two additions the matrix doesn't cover yet:
-
ROCm version axis. I'm on ROCm 10.0 (Windows/MSVC build, HIP 7.15), well past the 7.2.4 / 7.14 in the repro. I2VA/Ref2VA conditioning on the ROCm backend is still blocked here — so this looks gfx1100 + ROCm-independent rather than a 7.x regression. (Different OS though: I'm Windows, the OP is Linux LXC, so I'm not claiming the 7.x hang reproduces on 10.x, only that the ROCm path stays unusable for H3 conditioning on the same GPU + commit.)
-
The Vulkan escape hatch is confirmed on this same GPU + cc515a0. I don't just manually flip --backend diffusion=Vulkan1; I enforce it at the wrapper layer — i2va/fl2va/ref2va semantics all route to Vulkan unconditionally, leaving only T2VA (decode-only) on ROCm. That's because ROCm VAE encode on gfx1100 is its own separate crash (Access Violation 0xC00000FD, upstream #7691, still open), distinct from this sampling-stage hang.
For triage: the failure point here is after conditioning completes — first sampling step, VRAM flatlines ~1.5GB, GPU 0%, main thread busy-looping sched_yield() in libamdhip64. That's a diffusion-kernel/ROCm hang, not the VAE-encode 0xC00000FD stack overflow from #7691. So I'd keep it as a separate gfx1100+ROCm issue rather than a duplicate of #7691 (but it's the same family: "H3 conditioning path dies on ROCm/gfx1100").
Happily test any patch/branch on ROCm 10.0 + Windows or drop HIP/rocBLAS traces — let me know the logging patch from #1947 you want applied and I'll run it.
-
Regression-hunt result (Outcome B): does NOT reproduce on our environment — 6b3edaa falls back to CPU at the same resolutions as current master.
Tested the reported pre-regression revision on our RX 7900 XTX / Ubuntu 24.04 / Mesa RADV 25.2.8 / Vulkan loader 1.3.275:
- Worktree:
6b3edaaf32cc("generalize temporal tiling across video VAEs", feat: generalize temporal tiling across video VAEs #1926, Aug 31) self-built-DGGML_VULKAN=ON, same toolchain as our master build
Resolution old 6b3edaa current master cc515a0 640×384 33f ✅ GPU compute confirmed (GPU 100% sustained), 100 s/it, completed ✅ GPU confirmed, 0.5 s/it, 75 s total 864×480 33f ⚠️ mixed — completed, 360 s/it,ggml_vec_dot_f32in hot stacks, ~11 min CPU burn/step, GPU bursts 26-100% alternating with 0%,pinned memory allocation failed×1❌ CPU fallback (from earlier comment) 960×544 73f ❌ CPU fallback confirmed — 1448 s/it (24 min/step), 17 min CPU/step, ggml_vec_dot_f32/mul_matin hot stacks, VRAM residency 25.6GB while GPU 0-3%❌ same (earlier comment) Key observations:
- The old revision is not better here — it degrades at the same graph sizes, just with slower, older shaders when it does use the GPU (100 s/it vs 0.5 s/it at 640×384).
- The failure correlates with RADV's pinned-memory allocation failure ("Requested buffer size exceeds device buffer size limit") appearing right as the compute falls into CPU
ggml_vec_dot_f32. VRAM can still hold the weights (25.6GB reported) while the matmuls silently run on CPU — residency ≠ compute, as the earlier comments established. - So on our hardware the large-graph CPU fallback looks like a ggml-Vulkan/RADV buffer-limit interaction, not a stable-diffusion.cpp code regression. The upstream reporter's success at ~864×480 on the old revision likely reflects different buffer limits (NVIDIA, different GPU, or different driver) rather than a code change between 6b3edaa and master.
Classification per the issue's framework: Outcome B — not reproduced. We're not pinning old code; our working all-GPU path is current master Vulkan at ≤640×384-class graphs, and larger graphs need either graph chunking in ggml-vulkan or a maxAllocationSize accommodation.
- Worktree:
kaishi00 commented
on Sep 17, 2026 AuthorMore actionsFollow-up with the #1947 diagnostic patch + HIP tracing: this does not match #1947's finite first-step slowdown.
I instrumented the same
cc515a0/ ROCm 7.14 / gfx1100 Ref2VA repro and left it running for 3h30m after it reached sampling1/8. It never reached2/8.Conditioning is healthy:
VAE prep: 0.30s VAE compute: 0.49s on ROCm0 outputs: 0.00sThe failure localizes to the first Ref2VA denoiser graph launch. The last HIP trace sequence is:
hipGraphExecUpdate: Returned hipSuccess hipGraphLaunch GraphExec::Run max_streams: 1 UpdateStreams failed for device id:0After that:
- no subsequent useful kernel activity is observed
- GPU remains 0% busy
- VRAM remains flat at ~1.5 GB
- the main thread spins in
sched_yield()insidelibamdhip64.so.7 - all-thread gdb backtrace contains the corresponding HSA/HIP wait stack
- the compute call never returns
The process was finally terminated after 3h30m30s and exited cleanly on SIGTERM.
So this appears materially different from #1947: there, the pathological first compute eventually returned after ~731s and later steps ran normally. Here, the first denoiser compute is non-returning and the trace ends immediately after the HIP graph stream-update failure.
I'm treating
UpdateStreams failed for device id:0as the failure location, not claiming it is the deeper root cause.I have the full artifacts preserved:
- instrumented
ggml_runner.cpppatch - complete debug log / AMD trace
- command/environment/build metadata
- main-thread and all-thread gdb backtraces
- timeline samples
The full trace is ~200 MB, so I can upload a compressed archive or any narrower subset that would be more useful.
Thanks for the follow-up. We have a complementary data point from the same gfx1100 hardware, and the two cases line up nicely:
- Env: gfx1100 (RX 7900 XTX 24GB), ROCm 10.0, sdcpp at recent commit, Ref2VA/I2VA paths
- Our first-step trace shows the same API-call overhead pattern, but it eventually returns after ~732s (vs. your non-returning case). The fast/slow traces have identical ~1.3M-event API call structure (PAL-VMM heap and launch-config events each ~31%, GEMM only ~0.7%) — so the 700s delta is pure HIP API call-interval overhead, not compute.
- This suggests your hard hang (UpdateStreams failed for device id:0 → main thread spinning in sched_yield) is the same pathological HIP graph/stream-update path, just without the eventual-return behavior.
We have the full fast (37s) vs. slow (732s) traces captured locally. Happy to share a compressed subset (or specific rocm-instrument / timeline windows) if it helps narrow down whether the 7.x regression is in the graph-update path or the stream-sync fallback. Also relevant to our side: we have a runtime guard that demotes I2VA to the Vulkan backend on ROCm, so the hang doesn't block generation — if your artifacts include the instrumented ggml_runner.cpp patch, that would let us A/B it against our guard.
Note: this analysis was drafted with help from Hermes Agent running on a free-tier LLM — the trace numbers were verified locally, but the interpretation may not be fully accurate, so please treat the root-cause claims as hypotheses.
Same family on gfx1151 (Strix Halo), but in the MiniMax-H3 video VAE encode, sporadic, and avoided by disabling HIP graphs
Env: Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151, UMA), Debian 13, ROCm 7.2.4, built with
-DSD_HIPBLAS=ON -DAMDGPU_TARGETS=gfx1151 -DGGML_HIP_GRAPHS=ON.
Command: I2VA,--init-img, 704×704,--video-frames 49,--steps 2,--cfg-scale 1 --diffusion-fa --offload-to-cpu --rng cpu,minimax_h3_fl2va_pruned-Q4_K.gguf+qwen3vl_32b_minimax_h3-Q4_K_M.gguf+ fp16 video VAE + fp32 audio VAE.- Hangs sporadically during the VAE encode of the init image, right after
ggml_backend_cuda_graph_compute: CUDA graph warmup complete, at varying tiles (2/16, 12/16). GPU busy 100 %, one CPU core spinning. Main thread:sched_yield()inlibamdhip64←ggml_cuda_graph_evaluate_and_capture←MiniMaxH3VideoVAERunner::encode(full backtrace available). - Rate with graphs enabled: 11/31 runs on
b167b94, 1/10 on current mastera1ded76(same location; master tiles 3×3). - With
GGML_CUDA_DISABLE_GRAPHS=1: 0/10 hangs, interleaved with graph-enabled runs that did hang. Vulkan build: 0/5. - T2VA is unaffected.
Happy to run a patch or collect a HIP trace.
- Hangs sporadically during the VAE encode of the init image, right after
Summary
MiniMax-H3 Ref2VA / keyframe I2V conditioning hangs the GPU when run on the ROCm backend of the Unsloth fork build, while the identical workload completes on the CPU backend. Tested across two ROCm versions (7.2.4 and 7.14), multiple resolutions (384×640, 640×384, 544×960, 960×544), and both Studio-spawned and shell-launched invocations. Text-to-video works fine on the same installation; only the image/ref-conditioning path is affected.
Environment
13b9d92(H3 fixes) —sd-cli --mode vid_genminimax_h3_ref2va_pruned-Q5_0.gguf+minimax_h3_fl2va_pruned-Q5_0.gguf+ qwen3vl 32B encoder + fp16 video VAE + fp32 audio VAERepro
Behavior matrix (all tested on this host)
sched_yield()inside libamdhip64; HSA threads wait on GPU events; 0% GPU, VRAM flat, killed after 25+ mingdb backtrace (7.14 hang, main thread)
Thread 1 spins in a HIP busy-wait; HSA async threads wait on GPU events that never fire. A GPU kernel launched by the keyframe/ref conditioning encode never completes on gfx1100.
What we ruled out
--diffusion-fa,--max-vram -1,--stream-layers— all removed, same result)--end-imgpresence (FLF2V full contract — same crash)Impact
Any I2V / Ref2VA use on gfx1100 with ROCm is blocked. The workaround is running the diffusion model on CPU too (
--offload-to-cpuwithout any GPU backend assignment), which costs ~97 min for a 5-second clip at 544×960/30 steps vs an expected few minutes on GPU.Happy to provide additional traces, HSA logs (
HSA_ENABLE_SDMA=0was also tested — no change), or test builds.