Skip to content

[Bug] ROCm/gfx1100: qwen_image models fail with "unspecified launch failure" at first compute after #1994 (works on 6b3edaa) #2008

Description

@Radu0120

Krea-2-Turbo (general.architecture = qwen_image) aborts immediately after tensor load on the official Windows ROCm build sd-master-137f740-bin-win-rocm-7.14.0-x64.zip. The same model, same command line, same GPU works on sd-master-6b3edaa-bin-win-rocm-7.14.0-x64.zip.

[INFO  ] model_loader.cpp:1236 - loading tensors completed, taking 0.23s (read: 0.13s, memcpy: 0.00s, convert: 0.00s, copy_to_backend: 0.00s)
[ERROR ] util.cpp:663  - ROCm error: unspecified launch failure
[ERROR ] util.cpp:663  -   current device: -1, in function ggml_cuda_kernel_launch at ggml/src/ggml-cuda/common.cuh:1684
[ERROR ] util.cpp:663  -   hipGetLastError()
ggml/src/ggml-cuda/ggml-cuda.cu:107: ROCm error

The process then exits with 0xC0000409 (GGML_ABORT). It happens on both txt2img and img2img, so it is not VAE-encode specific.

Command

sd-server --diffusion-model krea2_turbo-Q8_0.gguf \
  --vae wan_2.1_vae.safetensors \
  --llm Qwen3VL-4B-Instruct-Q8_0.gguf \
  --llm_vision mmproj-Qwen3VL-4B-Instruct-F16.gguf \
  --diffusion-fa --vae-tiling -t 7 --steps 8 --cfg-scale 1

Not affected, on the identical build and runtime

Model general.architecture Result
Z-Image-Turbo Q8_0 lumina2 works
Flux.2-Klein 9B flux works
Krea-2-Turbo Q8_0 qwen_image crashes (txt2img + img2img)

So this looks confined to the qwen_image path rather than a general ROCm breakage.

Environment

  • RX 7900 XTX (gfx1100), Windows 11
  • ROCm 7.14.0 release assets (both builds)
  • The Windows ROCm zips ship no HIP runtime, so the runtime DLLs are supplied locally. The same runtime files are used for both builds, which controls for that variable.

Range

6b3edaa...137f740 (42 commits). Prime suspect is 137f740 / #1994, as the only change touching the qwen_image graph in that range. The flash-attention commits (#1987, #1992) and the convolution/VAE ones (#1993, #1996) are all exercised by the two working models above (Z-Image also runs with --diffusion-fa and an LLM text encoder).

Activity

  1. leejet commented on Sep 21, 2026

    @leejet
    Owner

    Thanks for the detailed report. Krea2 uses a separate model implementation, even though the GGUF metadata says qwen_image, so we haven’t confirmed that #1994 introduced this regression.

    Could you please share:

    • The exact download links for the GGUF files, especially krea2_turbo-Q8_0.gguf, plus the Qwen3-VL and mmproj files.
    • The full log from startup through the crash with --log-level debug.
    • Whether the same command still crashes with --disable-prefetch added.

    The debug log should help identify which model and execution segment fails. Disabling prefetch will help check whether the regression involves asynchronous weight transfers.

  2. Radu0120 commented on Sep 21, 2026

    @Radu0120
    Author

    Here are the files, all SHA-256 verified against what the repos serve right now:

    krea2_turbo-Q8_0.gguf                 vantagewithai/Krea-2-Turbo-GGUF
      1fa2da08a7a708827c2100d0af41c8371d79efd9f7c5490c23ebc65408536490  13705958688
    mmproj-Qwen3VL-4B-Instruct-F16.gguf   Qwen/Qwen3-VL-4B-Instruct-GGUF
      256f3a43bd4205ffef48d6b92715e1e70b5b0e9aef06522584967513a9985331  836180256
    Qwen3VL-4B-Instruct-Q8_0.gguf         Qwen/Qwen3-VL-4B-Instruct-GGUF
      054721f478bc5fa6beffb7f38eae575d45298f88cbb8d2f83ef675a727863eb1  4280406144
    wan_2.1_vae.safetensors               Comfy-Org/Wan_2.1_ComfyUI_repackaged (split_files/vae/)
      2fc39d31359a4b0a64f55876d8ff7fa8d780956ae2cb13463b0223e15148976b  253815318
    

    Resolve URLs: krea2_turbo-Q8_0, mmproj, Qwen3VL-4B-Q8_0, wan_2.1_vae.

    One correction to my report, since it affects the #1994 question: the command I pasted is not the one that crashed. The failing run also passed --max-vram 4.5 --lora-model-dir D:/LLM/Models/Krea, both injected by the launcher's config. On 137f740 without --max-vram, the same four files finish a 1024x1024 request in 35.6s. With it, every request aborts at the same point:

    [DEBUG  ] model_manager_prefetch.cpp:188  - model manager queued segment prefetch (440.06 MB, 13 tensors) to ROCm0
    [DEBUG  ] ggml_runner.cpp:946  - krea2 executing segment 17/29: krea2.blocks.16
    [ERROR  ] util.cpp:663  - ROCm error: unspecified launch failure
    [ERROR  ] util.cpp:663  -   current device: -1, in function ggml_cuda_kernel_launch at ...\ggml-cuda\common.cuh:1684
    

    --max-vram 4.5 makes auto-fit emit --params-backend "diffusion=cpu,te=disk,vae=cpu", so 12.99 GB of diffusion weights sit in RAM and get staged per segment (440.06 MB / 13 tensors per block). That is the only thing I can make reproduce the fault so far.

    Full debug log from startup through the abort is attached (krea-repro-qmflags.log).

    --disable-prefetch: yes, still crashes, same segment 17/29 and same request. It only changes which call reports it, from ggml_cuda_kernel_launch to ggml_backend_cuda_buffer_set_tensor + hipStreamSynchronize(((hipStream_t)2)). Log attached as krea-repro-noprefetch.log. So it does not look like the async weight transfers.

    For scale, 6b3edaa with --max-vram 4.5 works, but it has no auto-fit, keeps all 18.1 GB in VRAM, and never enters that RAM-resident path, so I do not think it settles the #1994 question either way.

    All runs: POST /sdapi/v1/txt2img, 8 steps, cfg 1, 1024x1024.

    krea-repro-qmflags.log

    krea-repro-noprefetch.log

  3. leejet commented on Sep 21, 2026

    @leejet
    Owner

    Thanks, this narrows it down considerably.

    One detail in the no-prefetch log: segment 17 actually finishes and its output is cached. The failure occurs while uploading weights for the next segment. With prefetch enabled, that upload happens earlier, which may explain why the error appears during segment 17.

    Both logs fail after cumulative ROCm_Host weight allocations increase from approximately 8153 MiB to 8593 MiB. This makes pinned host memory a useful next suspect, though it does not establish an 8 GiB limit.

    Could you retry the same failing command, keeping --max-vram 4.5 and all other settings unchanged, with this environment variable set before starting the server?

    $env:GGML_CUDA_NO_PINNED = "1"

    Please launch the server from that same PowerShell session. Despite its name, this variable also applies to ROCm and switches host allocations to ordinary CPU memory.

    Please share whether the request completes and the new debug log. This should help distinguish a pinned-memory issue from a more general segmented-execution problem.

  4. Radu0120 commented on Sep 21, 2026

    @Radu0120
    Author

    Yes, it completes. But my first attempt was confounded, so here is the controlled version.

    First attempt: your exact instructions (--max-vram 4.5 plus the env var, nothing else). The request completed, but auto-fit chose a different plan under that env var, --params-backend "diffusion=disk,te=cpu,vae=cpu" with the diffusion weights resident in VRAM, so that run never entered the RAM-resident staging path at all. It says nothing about pinning.

    So I pinned the plan explicitly in both runs, using the values auto-fit had chosen in the crashing run, which disables auto-fit:

    --backend "diffusion=ROCm0,te=ROCm0,vae=ROCm0" --params-backend "diffusion=cpu,te=disk,vae=cpu"
    

    Both runs report identical placement: total params memory size = 18110.45MB (VRAM 4873.86MB, RAM 13236.59MB): text_encoders 4873.86MB(VRAM), diffusion_model 12994.49MB(RAM), vae 242.10MB(RAM).

    Without the env var, same as before: crash, last segment reached 17/29 (krea2.blocks.16), 18 host allocations logged as model manager prepared params backend buffers (...) on ROCm_Host.

    With GGML_CUDA_NO_PINNED=1, same flags: the request completes, 1024x1024 in 42.8s, 0 ROCm errors, all 29 segments executed. The same allocation lines now read on CPU (30 of them, none on ROCm_Host).

    So pinned host memory is the difference here. I cannot say from this whether it is the pinning itself or the size of the pool the pinned allocations come out of, and I have no number for the limit; your 8 GiB reading is consistent with what I see but this does not add evidence to it.

    Attached: the debug log with the env var set (krea-repro-nopinnedB.log) and the control run without it (krea-repro-controlA.log), same request both times.

    krea-repro-nopinnedB.log

    krea-repro-controlA.log

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions