Repository navigation
[Bug] ROCm/gfx1100: qwen_image models fail with "unspecified launch failure" at first compute after #1994 (works on 6b3edaa) #2008
Description
Activity
Thanks for the detailed report. Krea2 uses a separate model implementation, even though the GGUF metadata says qwen_image, so we haven’t confirmed that #1994 introduced this regression.
Could you please share:
- The exact download links for the GGUF files, especially krea2_turbo-Q8_0.gguf, plus the Qwen3-VL and mmproj files.
- The full log from startup through the crash with --log-level debug.
- Whether the same command still crashes with --disable-prefetch added.
The debug log should help identify which model and execution segment fails. Disabling prefetch will help check whether the regression involves asynchronous weight transfers.
Here are the files, all SHA-256 verified against what the repos serve right now:
krea2_turbo-Q8_0.gguf vantagewithai/Krea-2-Turbo-GGUF 1fa2da08a7a708827c2100d0af41c8371d79efd9f7c5490c23ebc65408536490 13705958688 mmproj-Qwen3VL-4B-Instruct-F16.gguf Qwen/Qwen3-VL-4B-Instruct-GGUF 256f3a43bd4205ffef48d6b92715e1e70b5b0e9aef06522584967513a9985331 836180256 Qwen3VL-4B-Instruct-Q8_0.gguf Qwen/Qwen3-VL-4B-Instruct-GGUF 054721f478bc5fa6beffb7f38eae575d45298f88cbb8d2f83ef675a727863eb1 4280406144 wan_2.1_vae.safetensors Comfy-Org/Wan_2.1_ComfyUI_repackaged (split_files/vae/) 2fc39d31359a4b0a64f55876d8ff7fa8d780956ae2cb13463b0223e15148976b 253815318Resolve URLs: krea2_turbo-Q8_0, mmproj, Qwen3VL-4B-Q8_0, wan_2.1_vae.
One correction to my report, since it affects the #1994 question: the command I pasted is not the one that crashed. The failing run also passed
--max-vram 4.5 --lora-model-dir D:/LLM/Models/Krea, both injected by the launcher's config. On 137f740 without--max-vram, the same four files finish a 1024x1024 request in 35.6s. With it, every request aborts at the same point:[DEBUG ] model_manager_prefetch.cpp:188 - model manager queued segment prefetch (440.06 MB, 13 tensors) to ROCm0 [DEBUG ] ggml_runner.cpp:946 - krea2 executing segment 17/29: krea2.blocks.16 [ERROR ] util.cpp:663 - ROCm error: unspecified launch failure [ERROR ] util.cpp:663 - current device: -1, in function ggml_cuda_kernel_launch at ...\ggml-cuda\common.cuh:1684--max-vram 4.5makes auto-fit emit--params-backend "diffusion=cpu,te=disk,vae=cpu", so 12.99 GB of diffusion weights sit in RAM and get staged per segment (440.06 MB / 13 tensors per block). That is the only thing I can make reproduce the fault so far.Full debug log from startup through the abort is attached (
krea-repro-qmflags.log).--disable-prefetch: yes, still crashes, same segment 17/29 and same request. It only changes which call reports it, fromggml_cuda_kernel_launchtoggml_backend_cuda_buffer_set_tensor+hipStreamSynchronize(((hipStream_t)2)). Log attached askrea-repro-noprefetch.log. So it does not look like the async weight transfers.For scale, 6b3edaa with
--max-vram 4.5works, but it has no auto-fit, keeps all 18.1 GB in VRAM, and never enters that RAM-resident path, so I do not think it settles the #1994 question either way.All runs:
POST /sdapi/v1/txt2img, 8 steps, cfg 1, 1024x1024.Thanks, this narrows it down considerably.
One detail in the no-prefetch log: segment 17 actually finishes and its output is cached. The failure occurs while uploading weights for the next segment. With prefetch enabled, that upload happens earlier, which may explain why the error appears during segment 17.
Both logs fail after cumulative
ROCm_Hostweight allocations increase from approximately 8153 MiB to 8593 MiB. This makes pinned host memory a useful next suspect, though it does not establish an 8 GiB limit.Could you retry the same failing command, keeping
--max-vram 4.5and all other settings unchanged, with this environment variable set before starting the server?$env:GGML_CUDA_NO_PINNED = "1"
Please launch the server from that same PowerShell session. Despite its name, this variable also applies to ROCm and switches host allocations to ordinary CPU memory.
Please share whether the request completes and the new debug log. This should help distinguish a pinned-memory issue from a more general segmented-execution problem.
Yes, it completes. But my first attempt was confounded, so here is the controlled version.
First attempt: your exact instructions (
--max-vram 4.5plus the env var, nothing else). The request completed, but auto-fit chose a different plan under that env var,--params-backend "diffusion=disk,te=cpu,vae=cpu"with the diffusion weights resident in VRAM, so that run never entered the RAM-resident staging path at all. It says nothing about pinning.So I pinned the plan explicitly in both runs, using the values auto-fit had chosen in the crashing run, which disables auto-fit:
--backend "diffusion=ROCm0,te=ROCm0,vae=ROCm0" --params-backend "diffusion=cpu,te=disk,vae=cpu"Both runs report identical placement:
total params memory size = 18110.45MB (VRAM 4873.86MB, RAM 13236.59MB): text_encoders 4873.86MB(VRAM), diffusion_model 12994.49MB(RAM), vae 242.10MB(RAM).Without the env var, same as before: crash, last segment reached 17/29 (
krea2.blocks.16), 18 host allocations logged asmodel manager prepared params backend buffers (...) on ROCm_Host.With
GGML_CUDA_NO_PINNED=1, same flags: the request completes, 1024x1024 in 42.8s, 0 ROCm errors, all 29 segments executed. The same allocation lines now readon CPU(30 of them, none onROCm_Host).So pinned host memory is the difference here. I cannot say from this whether it is the pinning itself or the size of the pool the pinned allocations come out of, and I have no number for the limit; your 8 GiB reading is consistent with what I see but this does not add evidence to it.
Attached: the debug log with the env var set (
krea-repro-nopinnedB.log) and the control run without it (krea-repro-controlA.log), same request both times.
Krea-2-Turbo (
general.architecture = qwen_image) aborts immediately after tensor load on the official Windows ROCm buildsd-master-137f740-bin-win-rocm-7.14.0-x64.zip. The same model, same command line, same GPU works onsd-master-6b3edaa-bin-win-rocm-7.14.0-x64.zip.The process then exits with
0xC0000409(GGML_ABORT). It happens on both txt2img and img2img, so it is not VAE-encode specific.Command
Not affected, on the identical build and runtime
general.architecturelumina2fluxqwen_imageSo this looks confined to the
qwen_imagepath rather than a general ROCm breakage.Environment
Range
6b3edaa...137f740(42 commits). Prime suspect is137f740/ #1994, as the only change touching theqwen_imagegraph in that range. The flash-attention commits (#1987, #1992) and the convolution/VAE ones (#1993, #1996) are all exercised by the two working models above (Z-Image also runs with--diffusion-faand an LLM text encoder).