Skip to content

There are two problems, 1: VAE crash, 2:Performance Tuning #747

Description

@mgxhhg

1.VAE crash:

log:

➜ ~ /data/data/com.termux/files/home/llama/bin/sd.sh
ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23
ggml_opencl: vector subgroup broadcast support: true
ggml_opencl: device FP16 support: true
ggml_opencl: mem base addr align: 128
ggml_opencl: max mem alloc size: 1024 MB
ggml_opencl: SVM coarse grain buffer support: true
ggml_opencl: SVM fine grain buffer support: true
ggml_opencl: SVM fine grain system support: false
ggml_opencl: SVM atomics support: true
ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q)
ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS)
ggml_opencl: loading OpenCL kernels.........................................................
ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
[INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors'
[INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format
[INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors'
[INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format
[INFO ] stable-diffusion.cpp:243 - Version: SD 1.x
[INFO ] stable-diffusion.cpp:277 - Weight type: f16
[INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16
[INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16
[INFO ] stable-diffusion.cpp:280 - VAE weight type: f16
|==================================================| 1131/1131 - 500.00it/s
|===================> | 443/1131 - 0.00it/s
[INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 2042.16MB, RAM 0.00MB): clip 307.44MB(VRAM), unet 1640.25MB(VRAM), vae 94.47MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM)
[INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 1.47s
[INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode
[INFO ] stable-diffusion.cpp:2017 - TXT2IMG
[INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs
[INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s
[INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 150 ms
[INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method
[INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 42
|==================================================| 2/2 - 26.07s/it
[INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 52.28s
[INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 52.41s
[INFO ] stable-diffusion.cpp:1817 - decoding 1 latents
ggml_backend_opencl_buffer_type_alloc_buffer: requested size 1664.00 MiB exceeds max allocation size 1024.00 MiB, fallback to host RAM
ggml_gallocr_reserve_n: failed to allocate OpenCL buffer of size 1744830464
[ERROR] ggml_extend.hpp:1161 - vae: failed to allocate the compute buffer

/data/data/com.termux/files/home/llama/bin/sd.sh: line 15: 27015 Segmentation fault /data/data/com.termux/files/home/llama/bin/sd -m /storage/emulated/0/Download/1dm/picxReal_10.safetensors --verbose --type f16 --width 512 --height 512 --steps 2 --vae /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors -t 5 --cfg-scale 7.5 --sampling-method dpm++2m --prompt "girl" --negative-prompt "bad proportions, extra limbs, extra legs, extra legs and arms, disfigured, missing arms, simple background, transparent background, copyright, worst aesthetic, old, early, blurry, lowres, signature, artist name, watermark, twitter username, sketch, logo, furry, text, speech bubble, censored,angry boy, girl smiling, boy shaded face " -o "/data/data/com.termux/files/home/sd_$(date +%s)_$RANDOM.png"

reason:

AI:

The problem you're encountering primarily involves an OpenCL memory allocation failure— ggml_backend_opencl_buffer_type_alloc_buffer: failed to allocate 1664.00 MiB. Even if you have sufficient system memory (RAM) and zRAM, you'll still get this error. This is because:

Cause Summary:
GPU memory ≠ system memory
Qualcomm Adreno GPUs (such as your 830) have a fixed memory limit, typically 1GB to 2GB. Even with a large amount of system RAM, OpenCL's max mem alloc size limit cannot be exceeded.

OpenCL limits the size of a single buffer. Your log entry states:

ggml_opencl: max mem alloc size: 1024 MB

This means that any buffer cannot exceed 1024 MB. However, the attempt to allocate:

failed to allocate 1664.00 MiB
exceeds this limit.

Video memory isn't allocated as needed; it must be allocated in chunks.
Qualcomm GPUs can't flexibly allocate small chunks like desktop NVIDIA does. GGML requires allocating a whole chunk (e.g., 1664MB) for VAEs at once. Failure to do so will result in a crash.

Solution:

Allocate large buffers in chunks

2.Performance Tuning:

Because the application does not limit the GPU utilization, this is not a good thing. It will exhaust the system's graphics resources and cause the system interface to crash. So can we use -ngl and other methods to limit the resource consumption on the GPU like llama.cpp?

Image

@leejet

Activity

  1. Green-Sky commented on Aug 1, 2025

    @Green-Sky
    Contributor
    1. for now, try --vae-tiling to limit the vae allocation size. There is work happening to optimize the convolution operation, which causes the huge memory usage and degraded speed.
    2. no, sd.cpp does not implement partial offloading for a model. You can offload embedders and vae entirely to the cpu though. Also, like 1. @rmatif is working on opencl optimizations that will improve general performance.

    refs: #739 #744
    edit: also probably ggml-org/llama.cpp#14987

  2. rmatif commented on Aug 1, 2025

    @rmatif
    Contributor

    Like @Green-Sky said, either use --vae-tiling or --taesd (you can grab it here).

    Using conv2d_direct will bring you down to ~13.4 s/it for SD 1.x (I assume you're using cfg-scale > 1) on Adreno 830, and it also cuts the VAE buffer size by a factor of 2.6.

    Ironically, SD 2.x and SDXL are much faster:

    • SDXL: 7.2 s/it at 512x512, cfg-scale > 1
    • SD 2.x: 4.6 s/it at 512x512, cfg-scale > 1

    I think this is due to the nature of the compute graph, OpenCL backend performs better when compute-bound.

    Using the upcoming --diffusion-fa, you can trade ~15–20% performance to reduce compute buffer size:

    SD 2.x:

    • 512x512: 367 MB → 70 MB
    • 768x768: 1718 MB → 157 MB

    SDXL:

    • 1024x1024: 829 MB → 230 MB

    Offloading some tensors to the CPU is a bad idea, it introduces significant overhead and memory transfer costs, which are especially sensitive on this type of hardware. More performance optimizations will come over time :)

  3. mgxhhg commented on Aug 1, 2025

    @mgxhhg
    Author

    @rmatif

    There is no need to worry about the overhead of unloading tensors to the CPU and the memory transfer cost, because I found in llama.cpp: after using kleidiai's instruction set optimization, the CPU performance has been significantly improved, and the efficiency of GPU+CPU hybrid computing has also been improved. Even the 8b q8_0 quantized model can still be very fast. Judging from the output speed, this model has already achieved a certain degree of practicality. Another point is that the vram and ram of Android devices are both on the same LPDDR5. In theory, you only need to unify the memory type to reduce the memory transfer cost, such as: --cache-type-k f16 --cache-type-v f16

  4. mgxhhg commented on Aug 2, 2025

    @mgxhhg
    Author

    Adreno OpenCL maximum 1024MB limit has been resolved

    ➜ ~ /data/data/com.termux/files/home/llama/bin/sd.sh
    ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

    ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23
    ggml_opencl: vector subgroup broadcast support: true
    ggml_opencl: device FP16 support: true
    ggml_opencl: mem base addr align: 128
    ggml_opencl: max mem alloc size: 1024 MB
    ggml_opencl: SVM coarse grain buffer support: true
    ggml_opencl: SVM fine grain buffer support: true
    ggml_opencl: SVM fine grain system support: false
    ggml_opencl: SVM atomics support: true
    ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q)
    ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS)
    ggml_opencl: loading OpenCL kernels.........................................................
    ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    [INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:243 - Version: SD 1.x
    [INFO ] stable-diffusion.cpp:277 - Weight type: f16
    [INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16
    [INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16
    [INFO ] stable-diffusion.cpp:280 - VAE weight type: f16
    [INFO ] stable-diffusion.cpp:326 - CLIP: Using CPU backend
    [INFO ] stable-diffusion.cpp:387 - VAE Autoencoder: Using CPU backend
    |==================================================| 1131/1131 - 1000.00it/s
    |===================> | 443/1131 - 0.00it/s
    [INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 1640.25MB, RAM 401.91MB): clip 307.44MB(RAM), unet 1640.25MB(VRAM), vae 94.47MB(RAM), controlnet 0.00MB(VRAM), pmid 0.00MB(RAM)
    [INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 0.75s
    [INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode
    [INFO ] stable-diffusion.cpp:2017 - TXT2IMG
    [INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs
    [INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s
    [INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 699 ms
    [INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method
    [INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 3786717814
    |==================================================| 20/20 - 50.67s/it
    [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 992.41s
    [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 992.45s
    [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents
    [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 87.25s
    [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 87.25s
    [INFO ] stable-diffusion.cpp:2088 - generate_image completed in 1080.40s
    save result PNG image to '/data/data/com.termux/files/home/sd_1754088679_32319.png'

    The principle is to ask AI to help me write a .so, intercept the call to OpenCL, VRAM Allocate large buffers in chunks, and modify the source code, but I have no idea where to start, so this is the only way

  5. rmatif commented on Aug 2, 2025

    @rmatif
    Contributor

    @rmatif

    There is no need to worry about the overhead of unloading tensors to the CPU and the memory transfer cost, because I found in llama.cpp: after using kleidiai's instruction set optimization, the CPU performance has been significantly improved, and the efficiency of GPU+CPU hybrid computing has also been improved. Even the 8b q8_0 quantized model can still be very fast. Judging from the output speed, this model has already achieved a certain degree of practicality. Another point is that the vram and ram of Android devices are both on the same LPDDR5. In theory, you only need to unify the memory type to reduce the memory transfer cost, such as: --cache-type-k f16 --cache-type-v f16

    Even with shared LPDDR5X, the CPU and GPU have separate caches and logical memory spaces. Offloading mandates an explicit data transfer managed by the OpenCL backend. This synchronization introduces significant latency, creating a far greater bottleneck than keeping the data resident on the GPU

    Adreno OpenCL maximum 1024MB limit has been resolved

    It's not resolved, you just kept the VAE on the CPU. The buffer limit is imposed by the OpenCL driver, and there's nothing we can do about it

    The principle is to ask AI to help me write a .so, intercept the call to OpenCL, VRAM Allocate large buffers in chunks, and modify the source code, but I have no idea where to start, so this is the only way

    Honestly, I did try to understand, but I don't get what you mean. If you're referring to processing VAE decoding in chunks, that's exactly what --vae-tiling does. Combined with taesd and the upcoming conv2d_direct and fa for compute buffer, a 1024 MB budget is more than enough to run inference on almost any model at any resolution

  6. mgxhhg commented on Aug 2, 2025

    @mgxhhg
    Author

    You will understand if you look at this. Because I loaded a lot of models, the maximum value of the buffer is exceeded.

    ➜ ~ /data/data/com.termux/files/home/llama/bin/sd.sh
    ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

    ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23
    ggml_opencl: vector subgroup broadcast support: true
    ggml_opencl: device FP16 support: true
    ggml_opencl: mem base addr align: 128
    ggml_opencl: max mem alloc size: 1024 MB
    ggml_opencl: SVM coarse grain buffer support: true
    ggml_opencl: SVM fine grain buffer support: true
    ggml_opencl: SVM fine grain system support: false
    ggml_opencl: SVM atomics support: true
    ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q)
    ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS)
    ggml_opencl: loading OpenCL kernels.........................................................
    ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    [INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:243 - Version: SD 1.x
    [INFO ] stable-diffusion.cpp:277 - Weight type: f16
    [INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16
    [INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16
    [INFO ] stable-diffusion.cpp:280 - VAE weight type: f16
    |==================================================| 1131/1131 - 1000.00it/s
    |===================> | 443/1131 - 0.00it/s
    [INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 2042.16MB, RAM 0.00MB): clip 307.44MB(VRAM), unet 1640.25MB(VRAM), vae 94.47MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM)
    [INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 1.86s
    [INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode
    [INFO ] stable-diffusion.cpp:2017 - TXT2IMG
    [INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs
    [INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s
    [INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 270 ms
    [INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method
    [INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 3786717814
    ggml_backend_opencl_buffer_type_alloc_buffer: failed to allocate 1227.63 MiB
    ggml_gallocr_reserve_n: failed to allocate OpenCL buffer of size 1287263360
    [ERROR] ggml_extend.hpp:1161 - unet: failed to allocate the compute buffer

    /data/data/com.termux/files/home/llama/bin/sd.sh: line 18: 11522 Segmentation fault

    The dynamic library marked in red acts as a relay, dynamically loaded via LD_PRELOAD, and intercepts ggml's calls to OpenCL. The buffer allocation fails because it directly requests a buffer exceeding the maximum size from the driver, and the driver directly informs it that the allocation cannot be made (stable-diffusion.cpp/ggml does not have a solution for buffer allocations exceeding the driver's maximum size. This is a communication issue with the driver and a problem with the Adreno driver's buffer handling method), resulting in the program crash. If it were simply a communication issue, it would be simple: the problem is that ggml requested an excessively large buffer (over 1024MB) at once, and the Adreno driver cannot handle such a large request, resulting in allocation failure and a crash. This library's purpose is to split the large request into smaller 512MB chunks, request them one by one, and then piece them together to notify ggml of the allocation success. It effectively creates an intermediate layer between ggml and the driver.

    Image

    ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

    ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23
    ggml_opencl: vector subgroup broadcast support: true
    ggml_opencl: device FP16 support: true
    ggml_opencl: mem base addr align: 128
    ggml_opencl: max mem alloc size: 1024 MB
    ggml_opencl: SVM coarse grain buffer support: true
    ggml_opencl: SVM fine grain buffer support: true
    ggml_opencl: SVM fine grain system support: false
    ggml_opencl: SVM atomics support: true
    ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q)
    ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS)
    ggml_opencl: loading OpenCL kernels.........................................................
    ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)'
    [INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors'
    [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format
    [INFO ] stable-diffusion.cpp:243 - Version: SD 1.x
    [INFO ] stable-diffusion.cpp:277 - Weight type: f16
    [INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16
    [INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16
    [INFO ] stable-diffusion.cpp:280 - VAE weight type: f16
    |==================================================| 1131/1131 - 500.00it/s
    |===================> | 443/1131 - 0.00it/s
    [INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 2042.16MB, RAM 0.00MB): clip 307.44MB(VRAM), unet 1640.25MB(VRAM), vae 94.47MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM)
    [INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 1.54s
    [INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode
    [INFO ] stable-diffusion.cpp:2017 - TXT2IMG
    [INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs
    [INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s
    [INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 277 ms
    [INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method
    [INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 3786717814

    |==================================================| 20/20 - 51.35s/it
    [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 995.08s
    [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 995.16s
    [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents
    [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 51.02s
    [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 51.03s
    [INFO ] stable-diffusion.cpp:2088 - generate_image completed in 1046.49s
    save result PNG image to '/data/data/com.termux/files/home/sd_1754107956_17407.png'

    @rmatif

  7. chengjingxiang commented on Feb 9, 2026

    @chengjingxiang

    You will understand if you look at this. Because I loaded a lot of models, the maximum value of the buffer is exceeded.

    ➜ ~ /data/data/com.termux/files/home/llama/bin/sd.sh ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

    ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)' ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23 ggml_opencl: vector subgroup broadcast support: true ggml_opencl: device FP16 support: true ggml_opencl: mem base addr align: 128 ggml_opencl: max mem alloc size: 1024 MB ggml_opencl: SVM coarse grain buffer support: true ggml_opencl: SVM fine grain buffer support: true ggml_opencl: SVM fine grain system support: false ggml_opencl: SVM atomics support: true ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q) ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS) ggml_opencl: loading OpenCL kernels......................................................... ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)' [INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format [INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors' [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format [INFO ] stable-diffusion.cpp:243 - Version: SD 1.x [INFO ] stable-diffusion.cpp:277 - Weight type: f16 [INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16 [INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16 [INFO ] stable-diffusion.cpp:280 - VAE weight type: f16 |==================================================| 1131/1131 - 1000.00it/s |===================> | 443/1131 - 0.00it/s [INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 2042.16MB, RAM 0.00MB): clip 307.44MB(VRAM), unet 1640.25MB(VRAM), vae 94.47MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM) [INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 1.86s [INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode [INFO ] stable-diffusion.cpp:2017 - TXT2IMG [INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs [INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s [INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 270 ms [INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method [INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 3786717814 ggml_backend_opencl_buffer_type_alloc_buffer: failed to allocate 1227.63 MiB ggml_gallocr_reserve_n: failed to allocate OpenCL buffer of size 1287263360 [ERROR] ggml_extend.hpp:1161 - unet: failed to allocate the compute buffer

    /data/data/com.termux/files/home/llama/bin/sd.sh: line 18: 11522 Segmentation fault

    The dynamic library marked in red acts as a relay, dynamically loaded via LD_PRELOAD, and intercepts ggml's calls to OpenCL. The buffer allocation fails because it directly requests a buffer exceeding the maximum size from the driver, and the driver directly informs it that the allocation cannot be made (stable-diffusion.cpp/ggml does not have a solution for buffer allocations exceeding the driver's maximum size. This is a communication issue with the driver and a problem with the Adreno driver's buffer handling method), resulting in the program crash. If it were simply a communication issue, it would be simple: the problem is that ggml requested an excessively large buffer (over 1024MB) at once, and the Adreno driver cannot handle such a large request, resulting in allocation failure and a crash. This library's purpose is to split the large request into smaller 512MB chunks, request them one by one, and then piece them together to notify ggml of the allocation success. It effectively creates an intermediate layer between ggml and the driver.

    Image

    ggml_opencl: selected platform: 'QUALCOMM Snapdragon(TM)'

    ggml_opencl: device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)' ggml_opencl: OpenCL driver: OpenCL 3.0 QUALCOMM build: 0800.46 Compiler E031.47.18.23 ggml_opencl: vector subgroup broadcast support: true ggml_opencl: device FP16 support: true ggml_opencl: mem base addr align: 128 ggml_opencl: max mem alloc size: 1024 MB ggml_opencl: SVM coarse grain buffer support: true ggml_opencl: SVM fine grain buffer support: true ggml_opencl: SVM fine grain system support: false ggml_opencl: SVM atomics support: true ggml_opencl: flattening quantized weights representation as struct of arrays (GGML_OPENCL_SOA_Q) ggml_opencl: using kernels optimized for Adreno (GGML_OPENCL_USE_ADRENO_KERNELS) ggml_opencl: loading OpenCL kernels......................................................... ggml_opencl: default device: 'QUALCOMM Adreno(TM) 830 (OpenCL 3.0 Adreno(TM) 830)' [INFO ] stable-diffusion.cpp:192 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/picxReal_10.safetensors using safetensors format [INFO ] stable-diffusion.cpp:231 - loading vae from '/storage/emulated/0/Download/1dm/color101VAE_v1.safetensors' [INFO ] model.cpp:1013 - load /storage/emulated/0/Download/1dm/color101VAE_v1.safetensors using safetensors format [INFO ] stable-diffusion.cpp:243 - Version: SD 1.x [INFO ] stable-diffusion.cpp:277 - Weight type: f16 [INFO ] stable-diffusion.cpp:278 - Conditioner weight type: f16 [INFO ] stable-diffusion.cpp:279 - Diffusion model weight type: f16 [INFO ] stable-diffusion.cpp:280 - VAE weight type: f16 |==================================================| 1131/1131 - 500.00it/s |===================> | 443/1131 - 0.00it/s [INFO ] stable-diffusion.cpp:558 - total params memory size = 2042.16MB (VRAM 2042.16MB, RAM 0.00MB): clip 307.44MB(VRAM), unet 1640.25MB(VRAM), vae 94.47MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM) [INFO ] stable-diffusion.cpp:562 - loading model from '/storage/emulated/0/Download/1dm/picxReal_10.safetensors' completed, taking 1.54s [INFO ] stable-diffusion.cpp:604 - running in eps-prediction mode [INFO ] stable-diffusion.cpp:2017 - TXT2IMG [INFO ] stable-diffusion.cpp:738 - Attempting to apply 0 LoRAs [INFO ] stable-diffusion.cpp:1562 - apply_loras completed, taking 0.00s [INFO ] stable-diffusion.cpp:1696 - get_learned_condition completed, taking 277 ms [INFO ] stable-diffusion.cpp:1719 - sampling using DPM++ (2M) method [INFO ] stable-diffusion.cpp:1768 - generating image: 1/1 - seed 3786717814

    |==================================================| 20/20 - 51.35s/it [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 995.08s [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 995.16s [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 51.02s [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 51.03s [INFO ] stable-diffusion.cpp:2088 - generate_image completed in 1046.49s save result PNG image to '/data/data/com.termux/files/home/sd_1754107956_17407.png'

    @rmatif

    Can I see how you modified it to fix it? I am currently facing the same problem as you. You said you were using AI, but I don't know what your modifications are, so it's easier to share

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions