Repository navigation
llama: fix qwen4exp - #29751
llama: fix qwen4exp#29751
Conversation
5642e6c to
89f409c
Compare
89f409c to
f72ec5a
Compare
|
Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end. |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
I experience the same issue, but with a single GPU (RTX 3090). With this PR my |
|
Hey @am17an, heads up: since this PR, test-recurrent-state-rollback aborts on glm5-next-moe in builds with GGML_SCHED_NO_REALLOC=ON, which is why gpu-rocm is red on master: "unexpected graph reallocation (graph size = 424, nodes = 424, leafs = 79)" on the first 256 token decode. Same node count, so my guess is a glm5-next input in llama-memory-hybrid-idx outgrowing its reserved size. I bisected it on CUDA to 66e0c17 (7677678 passes), so it is not ROCm specific. Repro: cmake -B build -DGGML_CUDA=ON -DGGML_SCHED_NO_REALLOC=ON && cmake --build build --target test-llama-archs test-recurrent-state-rollback Happy to help dig into it if you want! https://github.com/ggml-org/llama.cpp/actions/runs/36868095055/job/110388592119 |
|
@ServeurpersoCom I think this should fix it: #29805 PTAL |
Retried today after the merge: lowered batch from 4096 to 2048 and -np 2 -kvu (from standard, that is 4) and no OOM. |
I don't think it's your fault: that big allocation at the end of the load is likely the graph reserve, which sizes the k-pool re-pool list for every pool the cache can hold, so it grows with the context size times -np. That would explain why fewer parallel slots (or a smaller context) brings it back under the limit. I'll run some checks on my side to confirm, and your start log would still help @am17an if you can post it. |
Upstream merged replacements for ggml-org#28243 (ggml-org#29761 MTP), ggml-org#29030 (TENSOR_READ_LAZY + ggml-org#29599) and ggml-org#28213 (ggml-org#29751 QSA rewrite); the branch now carries only the two server checkpoint PRs plus devops. Assisted-by: opencode
|
Metal: #29751 regresses Qwen3.8-Flash-Next prompt processing at depth (~36% at 80K) Hi @am17an, thanks for the qwen4exp + MTP work. MTP works well here: Q8 drafter, Setup: Apple M5 Ultra (64-core GPU, 256 GB), macOS 27.0, Qwen3.8-Flash-Next UD-Q6_K_XL (Unsloth), no MTP.
Happy to test patches or rerun with other flags ( |
@LucaAmigoni could you give this commit a try with your original settings (-b 4096, default -np)? 9119cef |
Yup, now it is restored as before. With default -np and -b 4096 no OOM, doing some stress test now. |
@ggerganov Thanks — #29824 fixes it on Metal, and then some. Prompt processing at depth now beats v0.5.0. Same machine, model and flags as above: M5 Ultra (64-core GPU, 256 GB), Qwen3.8-Flash-Next UD-Q6_K_XL, no MTP. pp2048 (t/s)
tg64 (t/s)
|
Clean upstream sync past 254b177. Relevant to the fork's deployment: - 4e2713c qwen4exp : optimize mask constructions (ggml-org#29824) - 631109b ggml : add alloc_buffer_n to the buffer type interface (ggml-org#23671) plus sycl / vulkan / opencl kernel work, llama warning and abort cleanups, and a new server /v1/systemone API (ggml-org#29818). Why it matters here: adopting upstream's Qwen4Exp MTP (ggml-org#29761) costs roughly +2.6 GB of compute buffers per device on the Qwen3.8-Flash-Next route, because the fixed attention path (ggml-org#29751) builds a k-pool input per QSA layer and makes the indexer cache track the attention cache cell for cell. ggml-org#29824 reduces the mask construction cost of that path. Fork-own work (RAM prompt-cache retention, selective CUDA P2P transport, server prefill/decode phase isolation, DFlash M-RoPE inference, docs, tests) is unchanged.
Adds the NextN/MTP draft head for Qwen3.8-Flash-Next (qwen4exp): one full-attention QSA block after the trunk, fed by the trunk's hyper-connection-wide residual, with its own hc head mixer, run through --spec-type draft-mtp with the head loaded from a separate MTP GGUF (-md). Ported from upstream ggml-org#29761 "Qwen4Exp: add MTP" (am17an, merged 2026-10-01 as c061df1; design reviewed by ggerganov). Upstream credit: the converter, gguf-py tensor names, graph_mtp, the MTP memory filter, the empty-recurrent-memory handling and the draft-path fix in common/speculative.cpp are theirs; this commit only carries them over. Hand-adapted against this tree, which predates upstream ggml-org#29751 (the QSA rewrite that ggml-org#29761 builds on): - graph_mtp calls build_layer_attn without the kpool input. This tree derives QSA from mctx_hyb->get_idx() and the layer's compress ratio, so the MTP block picks QSA up the same way a trunk layer does. - the indexer_kpool loop-bound hunk and a TODO comment hunk have no target here and are dropped. No SYCL wiring: this commit is exercised on a CPU-only build to measure draft acceptance. The draft context, the multi-token verify path and the host-expert cost on the SYCL backend are deliberately not touched. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AWkHqCXdysC9zf9AVvCiNn
Qwen3.8-Flash-Next (qwen4exp): upstream's own implementation and MTP (ggml-org#29761, ggml-org#29751, ggml-org#29824, ggml-org#29819) replace the copy of the earlier MTP PR (conversion/qwen4exp.py, src/models/qwen4exp.cpp and the qwen4exp hunks of gguf constants / tensor mapping, llama-arch.h, llama-model.h, models.h, llama-memory-recurrent.cpp). Kept from this fork where upstream has no counterpart: the GLM-5.3-Flash MTP additions (indexer_index_share_mtp, the hybrid-memory mtp_dsa accessors), dev_offload, the expert cache. common.cpp: both blocks kept (cache auto setup and the decision types). Compiles CPU-only on pc2; CUDA build and runs pending.
Upstream ggml-org#29751 rebuilds the qwen4exp QSA indexer on the k-pool cache: the indexer scores pooled block keys, takes the top blocks plus the tail, and builds the attention mask as REPEAT(FILL(seed, -inf)) -> SET_ROWS(0 at the list) -> VIEW -> ADD(kq_mask) with one extra cell per row for the n_kv sentinel of padded blocks. Match that chain from the flash attention node and let the node read the selection list and the causal mask directly (fattn-qsa.cpp), as before. Every node of the chain is absorbed, including the seed scalars, so the two [n_kv + 1, n_ubatch] f16 tensors are never reserved. The earlier graph had a score gather and a different mask chain; the fusions for them no longer match anything and are removed: the QSA gather and top-k fusions, the f16 cast + ADD and CONT + ADD fusions, selection bitmap modes 1 and 2 of GGML_SYCL_FUSE_QSA_FA_MASK and GGML_SYCL_FUSE_QSA_MASK. The radix top-k is back to the upstream code; only its select moves to the header for the MoE top-k. GGML_SYCL_FUSE_QSA_FA_MASK is now 0 or 1. Assisted-by: Claude Opus 5.5
Qwen3.8-Flash-Next (qwen4exp): upstream's own implementation and MTP (ggml-org#29761, ggml-org#29751, ggml-org#29824, ggml-org#29819) replace the copy of the earlier MTP PR (conversion/qwen4exp.py, src/models/qwen4exp.cpp and the qwen4exp hunks of gguf constants / tensor mapping, llama-arch.h, llama-model.h, models.h, llama-memory-recurrent.cpp). Kept from this fork where upstream has no counterpart: the GLM-5.3-Flash MTP additions (indexer_index_share_mtp, the hybrid-memory mtp_dsa accessors), dev_offload, the expert cache. common.cpp: both blocks kept (cache auto setup and the decision types). Compiles CPU-only on pc2; CUDA build and runs pending.
The radix top-k picks a value threshold and then gathers. Values that tie with the threshold exactly are the ones the budget cuts through, and which of them the gather keeps is decided by the order its atomics happen to run in, so the selected set is not reproducible between two runs of the same build. Ties are common where the scores are rectified, since those collapse to exactly zero. ggml_top_k leaves the choice among tied values unspecified, so this is valid, but it makes selections and outputs hard to compare between runs. GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied values with a second radix walk, over the column index and from the low bin upwards, and keeps only the tied columns at or below it. The set is then fixed; the order within the reserved slots still depends on the atomics, which the callers do not rely on. The extra walk only covers the bits a column index can use, so it is two passes for a context of up to 64k rather than four. The switch is read once per process and applies to every ggml_top_k that the backend runs through the radix selection, not only to one model's graph. The switch is off by default. Nothing about quality motivates one order among equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does not pay for the extra passes. It is a supported option for reproducible runs, where selections or outputs are compared between runs. R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756 per-layer selections without the switch and agree on all of them with it, and with the switch set a 128-token greedy continuation reproduces byte for byte. test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph slot of the scheduler, so a speculative verification width seen before is made current again instead of being built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a buffer that has to grow); when all slots are in use the least recently used shape is evicted. LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism. Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17 requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB, unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the defaults (24, 32): - Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751): 46.52 -> 51.67 t/s (+11.1 %) - DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %) Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
* llama: fix qwen4exp * qwen4exp: keep kq_mask input the same shape
The radix top-k picks a value threshold and then gathers. Values that tie with the threshold exactly are the ones the budget cuts through, and which of them the gather keeps is decided by the order its atomics happen to run in, so the selected set is not reproducible between two runs of the same build. Ties are common where the scores are rectified, since those collapse to exactly zero. ggml_top_k leaves the choice among tied values unspecified, so this is valid, but it makes selections and outputs hard to compare between runs. GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied values with a second radix walk, over the column index and from the low bin upwards, and keeps only the tied columns at or below it. The set is then fixed; the order within the reserved slots still depends on the atomics, which the callers do not rely on. The extra walk only covers the bits a column index can use, so it is two passes for a context of up to 64k rather than four. The switch is read once per process and applies to every ggml_top_k that the backend runs through the radix selection, not only to one model's graph. The switch is off by default. Nothing about quality motivates one order among equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does not pay for the extra passes. It is a supported option for reproducible runs, where selections or outputs are compared between runs. R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756 per-layer selections without the switch and agree on all of them with it, and with the switch set a 128-token greedy continuation reproduces byte for byte. test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph slot of the scheduler, so a speculative verification width seen before is made current again instead of being built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a buffer that has to grow); when all slots are in use the least recently used shape is evicted. LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism. Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17 requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB, unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the defaults (24, 32): - Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751): 46.52 -> 51.67 t/s (+11.1 %) - DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %) Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
…evert ggml-org#29751) Revert 66e0c17 on top of ec7630a for src/models/qwen4exp.cpp, models.h, llama-hparams.h and hybrid-idx. Kept from 81e39ad: the n_pool_max clamp, with the stable n_new_g padding it clamps (GLM5 path). Kept from c061df1: recurrent is_empty(), speculative.cpp, conversion and gguf-py, arch and tensor names. Dropped from c061df1: the qwen4exp graph_mtp, loader and create_memory filter (fork has its own MTP). Dropped from 159c651 and 4e2713c: their qwen4exp.cpp hunks, which edit build_qsa_sel of the reverted rewrite. The hybrid-idx pad_cell change of 159c651 and the glm5-next change of 4e2713c are kept. Dropped from 889edf4: its qwen4exp.cpp hunks (pool_mask F16 and the lightning-indexer score), which edit the same rewrite-only k-pool code. Its other files are kept.
The radix top-k picks a value threshold and then gathers. Values that tie with the threshold exactly are the ones the budget cuts through, and which of them the gather keeps is decided by the order its atomics happen to run in, so the selected set is not reproducible between two runs of the same build. Ties are common where the scores are rectified, since those collapse to exactly zero. ggml_top_k leaves the choice among tied values unspecified, so this is valid, but it makes selections and outputs hard to compare between runs. GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied values with a second radix walk, over the column index and from the low bin upwards, and keeps only the tied columns at or below it. The set is then fixed; the order within the reserved slots still depends on the atomics, which the callers do not rely on. The extra walk only covers the bits a column index can use, so it is two passes for a context of up to 64k rather than four. The switch is read once per process and applies to every ggml_top_k that the backend runs through the radix selection, not only to one model's graph. The switch is off by default. Nothing about quality motivates one order among equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does not pay for the extra passes. It is a supported option for reproducible runs, where selections or outputs are compared between runs. R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756 per-layer selections without the switch and agree on all of them with it, and with the switch set a 128-token greedy continuation reproduces byte for byte. test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph slot of the scheduler, so a speculative verification width seen before is made current again instead of being built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a buffer that has to grow); when all slots are in use the least recently used shape is evicted. LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism. Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17 requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB, unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the defaults (24, 32): - Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751): 46.52 -> 51.67 t/s (+11.1 %) - DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %) Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
Overview
Qwen4exp is broken on current master. This PR fixes the attention path based on how
llama-memory-hybrid-idxpooling works.Qwen3.8-Flash-Next UD-IQ4_XS, DGX Spark (GB10), -lzm on, b046420
Command:
llama-bench -m <IQ4_XS> -ngl 999 -fa 1 -lzm on -p 2048 -n 64 -ub 512,2048 -b 2048 -d 0,32768,131072 -r 1Additional information
Requirements