Skip to content

metal: speed up Qwen3-VL image encoding on large images by ~11% - #21443

Open
Avidanborisov wants to merge 3 commits into
ggml-org:masterfrom
Avidanborisov:metal-img-encode-optim
Open

Avidanborisov wants to merge 3 commits into
ggml-org:masterfrom
Avidanborisov:metal-img-encode-optim

Conversation

@Avidanborisov

@Avidanborisov Avidanborisov commented Apr 4, 2026 •

Copy link
Copy Markdown

Overview

This PR reduces the image encoding time of Qwen3-VL for large images on the Metal backend by ~11%. This issue was first reported here: #20704

Additional information

I've profiled the image encoding path of unsloth/Qwen3.5-9B on large images and noticed that the majority of the runtime is spent in the FlashAttention kernels.

image-20260404204338209

This patch optimizes the FlashAttention runtime in two complementary ways, each of which increases performance independently:

  1. It uses 2x fewer threadgroups and 2x more queries (16) per threadgroup. This reduces the number of K and V reads from device memory, thus reducing the memory bottleneck experienced by the FA kernel, at the cost of lower GPU occupancy and increased work per threadgroup. This tradeoff turns out to be advantageous for Q=16, but less so for Q=24, and not advantageous for Q=32.
  2. It increases per-threadgroup utilization by using 8 simdgroups instead of 4.

To keep the impact of this patch constrained and avoid regressions in untested scenarios, the optimized path is gated to f16 KV types and head sizes of 72 only (matching the Qwen3-VL family). This also means that the decoding phase is unaffected in the Qwen3.5-9B pipeline as well.

The first commit adds proper support for query sizes above 8 in the Metal FlashAttention implementation. The second commit applies the optimizations mentioned above to the Qwen3-VL image encoder path. Each commit was tested independently for correctness with the CI suite, including test-backend-ops -b MTL0 -o FLASH_ATTN_EXT in particular, and by asserting bit-identical embedding outputs.

Benchmark results

The benchmark runs the following command on small and large image inputs and looks for the image slice encoded in {X} ms output:

env MTMD_DEBUG_EMBEDDINGS=1 /bin/llama-mtmd-cli \
  -m models/unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf \
  --mmproj models/unsloth/Qwen3.5-9B-GGUF/mmproj-F16.gguf \
  --image test-inputs/{test-1,test-1-large-4096}.jpeg \
  -p "Describe this image." \
  --seed 42 \
  --temp 0 \
  -n 1 \
  --ctx-size 16384

We run it with env MTMD_DEBUG_EMBEDDINGS=1 to produce logs like this, to verify that the image embeddings are bit-identical.

=== MTMD_DEBUG_EMBEDDINGS ===
Shape: [4096, 4015]
Token 0 (first 16 values): -0.845602 0.362703 0.272525 -0.325752 -0.750353 0.281986 0.004355 -0.247313 -0.219053 -0.975575 1.062003 -0.348723 -0.463675 0.053039 -0.110674 -0.249361
Token 0 (last 16 values):  -0.516369 -0.280803 -0.110585 -0.229355 -0.019050 -0.037337 -0.094557 -0.741450 0.634281 0.572622 -0.103098 -0.384964 0.333557 -0.192545 -0.370187 -0.322191
Stats: mean=0.001318, std=0.122224, min=-2.676455, max=26.066109, sum=21667.273438
=== END MTMD_DEBUG_EMBEDDINGS ===

The reported runtimes are the medians of 3 runs per variant. I've attached the full logs of the runs for more information.

Variant Small Large
Baseline 1189 ms 40129 ms
+ Q=16 support in FA (1st commit, verify no regressions) 1189 ms (0.0%) 40191 ms (+0.2%)
+ use 16 queries and 8 simdgroups (2nd commit) 1188 ms (-0.1%) 35702 ms (-11.0%)

Logs

Technical details

  • Model: unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf
  • MMProj: unsloth/Qwen3.5-9B-GGUF/mmproj-F16.gguf
  • Machine used: Apple M2 MacBook Pro with 16 GB RAM
  • Small image: test-inputs/test-1.jpeg (640 x 488)
  • Large image: test-inputs/test-1-large-4096.jpeg (the same image from the repo, upscaled to 4096 x 3123)

CI results

I ran all the relevant CI tests locally, along with llama-bench and llama-perplexity, and ran test-backend-ops for FLASH_ATTN_EXT.

Requirements

  • I have read and agree with the contributing guidelines: YES

  • AI usage disclosure: YES:

    • I've used AI to familiarize myself with llama.cpp and to profile the image encoding path.
    • I've used AI to help write the optimization implementations after identifying bottlenecks.

This is my first PR to llama.cpp, so I hope I got everything right :)

@Avidanborisov
Avidanborisov requested a review from a team as a code owner April 4, 2026 19:12
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Apr 4, 2026
@ggerganov

Copy link
Copy Markdown
Member

Have you tried increasing the -ub 2048 to improve the baseline?

@Avidanborisov

Copy link
Copy Markdown
Author

The reported metric (image slice encoded in:) measures only mtmd_encode_chunk(), i.e. the vision encoder forward pass, and does not include the llama_decode() step, so ubatch does not directly affect that metric, iiuc.

Re-ran with -ub 2048 anyway to confirm and got overall similar results:

Variant Small Large
Baseline 1275 ms 39381 ms
+ Q=16 support in FA (1st commit) 1257 ms (-1.4%) 39505 ms (+0.3%)
+ use 16 queries and 8 simdgroups (2nd commit) 1248 ms (-2.1%) 35723 ms (-9.3%)

@Avidanborisov

Copy link
Copy Markdown
Author

Hi! Just a gentle bump on this PR, would appreciate any feedback when you have time. Thanks!

@Avidanborisov
Avidanborisov force-pushed the metal-img-encode-optim branch from fd0c693 to 6489117 Compare April 17, 2026 11:29
@Avidanborisov

Copy link
Copy Markdown
Author

Hey @ggerganov, any updates on this PR? Let me know if I need to change anything. Am happy to adjust whatever needed.

@Avidanborisov

Copy link
Copy Markdown
Author

Hey llama team (CC @ggerganov),
Can you let me know how to progress with this PR? Anything missing?
Thanks!

@forforever73

Copy link
Copy Markdown
Contributor

Thanks for the PR and the thorough benchmarks. On my M4 Max I measure ~2.5% regression for the same workload, opposite to your M2 result — the optimal Q here is hardware-family dependent. Look at #23114 (comment)

There's an ongoing effort (with @ggerganov) to add generic per-device FA parameter tuning to the Metal backend, so I'd prefer to bring the optimization in through there rather than merge the hardcoded path now. Your M2 numbers and profiling would be valuable input, and it'd be great to have you involved once the infrastructure lands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants