Skip to content

llama: fix qwen4exp - #29751

Merged
ggerganov merged 2 commits into
masterfrom
aman/qwen4-opt
Oct 1, 2026
Merged

ggerganov merged 2 commits into
masterfrom
aman/qwen4-opt

Conversation

@am17an

@am17an am17an commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Overview

Qwen4exp is broken on current master. This PR fixes the attention path based on how llama-memory-hybrid-idx pooling works.

Qwen3.8-Flash-Next UD-IQ4_XS, DGX Spark (GB10), -lzm on, b046420

Command: llama-bench -m <IQ4_XS> -ngl 999 -fa 1 -lzm on -p 2048 -n 64 -ub 512,2048 -b 2048 -d 0,32768,131072 -r 1

test ub master (t/s) this PR (t/s) speedup
pp2048 512 885.10 889.60 1.01x
tg64 512 25.61 29.05 1.13x
pp2048 @ d32768 512 650.78 747.06 1.15x
tg64 @ d32768 512 21.91 26.98 1.23x
pp2048 @ d131072 512 356.99 570.56 1.60x
tg64 @ d131072 512 14.08 21.83 1.55x
pp2048 2048 1022.86 1046.71 1.02x
tg64 2048 25.49 28.94 1.14x
pp2048 @ d32768 2048 763.76 887.09 1.16x
tg64 @ d32768 2048 21.80 26.90 1.23x
pp2048 @ d131072 2048 370.71 650.60 1.76x
tg64 @ d131072 2048 14.08 21.20 1.51x

Additional information

Requirements

@ggerganov ggerganov self-assigned this Sep 30, 2026
@github-actions github-actions Bot added the model Model specific label Sep 30, 2026
@am17an
am17an added this pull request to stack #29762 September 30, 2026 17:42
@LucaAmigoni

Copy link
Copy Markdown

Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end.
I could provide the start log tomorrow, if needed.

@Volunteer-1

This comment has been minimized.

@Volunteer-1

This comment has been minimized.

@ggerganov ggerganov added the highlight Changes that will be highlighted in the next release notes label Oct 1, 2026
@drrros drrros mentioned this pull request Oct 1, 2026
@quidow

quidow commented Oct 1, 2026 •

Copy link
Copy Markdown

Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end. I could provide the start log tomorrow, if needed.

I experience the same issue, but with a single GPU (RTX 3090). With this PR my -ot setting fails to load the model, while working perfectly with the build compiled from master. I was able to run the model only after significantly reducing the number of layers offloaded to GPU.

@ggerganov
ggerganov merged commit 66e0c17 into master Oct 1, 2026
20 of 21 checks passed
@ServeurpersoCom

ServeurpersoCom commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Hey @am17an, heads up: since this PR, test-recurrent-state-rollback aborts on glm5-next-moe in builds with GGML_SCHED_NO_REALLOC=ON, which is why gpu-rocm is red on master: "unexpected graph reallocation (graph size = 424, nodes = 424, leafs = 79)" on the first 256 token decode. Same node count, so my guess is a glm5-next input in llama-memory-hybrid-idx outgrowing its reserved size.

I bisected it on CUDA to 66e0c17 (7677678 passes), so it is not ROCm specific. Repro:

cmake -B build -DGGML_CUDA=ON -DGGML_SCHED_NO_REALLOC=ON && cmake --build build --target test-llama-archs test-recurrent-state-rollback
build/bin/test-llama-archs -o models && mkdir g && cp models/glm5-next-moe.gguf g/
build/bin/test-recurrent-state-rollback --models g

Happy to help dig into it if you want!

https://github.com/ggml-org/llama.cpp/actions/runs/36868095055/job/110388592119

@ggerganov

Copy link
Copy Markdown
Member

@ServeurpersoCom I think this should fix it: #29805 PTAL

@LucaAmigoni

Copy link
Copy Markdown

Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end. I could provide the start log tomorrow, if needed.

Retried today after the merge: lowered batch from 4096 to 2048 and -np 2 -kvu (from standard, that is 4) and no OOM.
My fault, I hadn't understood what "Qwen4exp is broken on current master" did mean.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Small follow-up todo: after the qwen4exp changes (#29751 and #29761) the Metal fusion baseline was not regenerated, so the Fusion / metal job fails on qwen4exp ADD+ADD+ADD (2 fused, 1 expected in tests/fusion/MTL.csv).

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end. I could provide the start log tomorrow, if needed.

Retried today after the merge: lowered batch from 4096 to 2048 and -np 2 -kvu (from standard, that is 4) and no OOM. My fault, I hadn't understood what "Qwen4exp is broken on current master" did mean.

I don't think it's your fault: that big allocation at the end of the load is likely the graph reserve, which sizes the k-pool re-pool list for every pool the cache can hold, so it grows with the context size times -np. That would explain why fewer parallel slots (or a smaller context) brings it back under the limit. I'll run some checks on my side to confirm, and your start log would still help @am17an if you can post it.

sourenaraya added a commit to sourenaraya/llama.cpp that referenced this pull request Oct 1, 2026
Upstream merged replacements for ggml-org#28243 (ggml-org#29761 MTP), ggml-org#29030
(TENSOR_READ_LAZY + ggml-org#29599) and ggml-org#28213 (ggml-org#29751 QSA rewrite); the
branch now carries only the two server checkpoint PRs plus devops.

Assisted-by: opencode
@kwalb

kwalb commented Oct 1, 2026

Copy link
Copy Markdown

Metal: #29751 regresses Qwen3.8-Flash-Next prompt processing at depth (~36% at 80K)

Hi @am17an, thanks for the qwen4exp + MTP work. MTP works well here: Q8 drafter, --spec-draft-n-max 2, lossless at greedy.
On Metal, though, git bisect lands on #29751 for a long-context prompt-processing regression.

Setup: Apple M5 Ultra (64-core GPU, 256 GB), macOS 27.0, Qwen3.8-Flash-Next UD-Q6_K_XL (Unsloth), no MTP.
Release build with -DGGML_METAL=ON.

llama-bench -m Qwen3.8-Flash-Next-UD-Q6_K_XL-00001-of-00006.gguf -ngl 999 -fa 1 -ctk q8_0 -ctv q8_0 -p 2048 -n 0 -d 0,81920 -r 1
build pp2048 pp2048 @ d81920
7fe450e (v0.5.0) 1269 857
7677678 (parent of #29751) — 911
66e0c17 (#29751) — 581
d775ebf (master, incl. #29761) 1276 579
  • Bisect over ggml/ + src/, 7fe450e (good) to d775ebf (bad), 7 steps: 66e0c17 is the first bad commit (good ≥800, bad ≤650; every step measured clearly in one camp, 910–914 vs 579–581).
  • Depth 0 is unaffected (1269 vs 1276), and so is a dense model (Qwen3.8-27B Q8_0, 328 vs 325 @ d81920), so it looks specific to the new qwen4exp / llama-memory-hybrid-idx attention path on Metal, not Metal-wide.
  • End to end through llama-server, a ~97K-token prompt goes from 910 to 716 t/s (−21%), deterministic over 2 runs.
  • Unlike the DGX Spark table in the llama: fix qwen4exp #29751 description, the parent commit was not slow on Metal for us, so the "broken master" state the PR fixes may be backend- or -lzm-specific.

Happy to test patches or rerun with other flags (-ub, -lzm, f16 KV) if useful.

@ggerganov

Copy link
Copy Markdown
Member

@kwalb Could you compare the performance with #29824 and provide the numbers?

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Not on the rig now, but this PR fails to load the model in dual GPU because of an huge GPU0 mem alloc at the load end. I could provide the start log tomorrow, if needed.

Retried today after the merge: lowered batch from 4096 to 2048 and -np 2 -kvu (from standard, that is 4) and no OOM. My fault, I hadn't understood what "Qwen4exp is broken on current master" did mean.

@LucaAmigoni could you give this commit a try with your original settings (-b 4096, default -np)? 9119cef
The qwen4exp indexer was holding the scores of all heads at once plus a full copy for the relu, about 4 GB of intermediates at 131k context and -b 4096 for a 0.5 GB result. It now scores one head at a time, in place: same computation and same top-k selection, half the compute buffer. That waste is most likely what hit your GPU0 at the end of the load, so the "compute buffer size" lines from the load log before and after would be great to compare.

@LucaAmigoni

Copy link
Copy Markdown

@LucaAmigoni could you give this commit a try with your original settings (-b 4096, default -np)? 9119cef The qwen4exp indexer was holding the scores of all heads at once plus a full copy for the relu, about 4 GB of intermediates at 131k context and -b 4096 for a 0.5 GB result. It now scores one head at a time, in place: same computation and same top-k selection, half the compute buffer. That waste is most likely what hit your GPU0 at the end of the load, so the "compute buffer size" lines from the load log before and after would be great to compare.

Yup, now it is restored as before. With default -np and -b 4096 no OOM, doing some stress test now.

@kwalb

kwalb commented Oct 1, 2026

Copy link
Copy Markdown

@kwalb Could you compare the performance with #29824 and provide the numbers?

@ggerganov Thanks — #29824 fixes it on Metal, and then some. Prompt processing at depth now beats v0.5.0.

Same machine, model and flags as above: M5 Ultra (64-core GPU, 256 GB), Qwen3.8-Flash-Next UD-Q6_K_XL, no MTP.

llama-bench -m Qwen3.8-Flash-Next-UD-Q6_K_XL-00001-of-00006.gguf -ngl 999 -fa 1 -ctk q8_0 -ctv q8_0 -p 2048 -n 64 -d 0,32768,81920,131072 -r 1

pp2048 (t/s)

build d0 d32768 d81920 d131072
v0.5.0 (7fe450e) 1275 1006 857 766
master (e358d59) 1313 832 581 450
#29819 (169ce4f) 1260 801 565 441
#29819 + #29824 (ce5b0e0) 1324 1128 1063 1027

tg64 (t/s)

build d0 d32768 d81920 d131072
v0.5.0 (7fe450e) 51.7 42.2 33.8 27.0
master (e358d59) 53.7 45.8 45.8 40.3
#29819 (169ce4f) 51.6 48.0 42.6 40.2
#29819 + #29824 (ce5b0e0) 51.4 49.1 44.9 40.1

novkien added a commit to novkien/llama.cpp-fork that referenced this pull request Oct 2, 2026
Clean upstream sync past 254b177. Relevant to the fork's deployment:

- 4e2713c qwen4exp : optimize mask constructions (ggml-org#29824)
- 631109b ggml : add alloc_buffer_n to the buffer type interface (ggml-org#23671)

plus sycl / vulkan / opencl kernel work, llama warning and abort cleanups,
and a new server /v1/systemone API (ggml-org#29818).

Why it matters here: adopting upstream's Qwen4Exp MTP (ggml-org#29761) costs roughly
+2.6 GB of compute buffers per device on the Qwen3.8-Flash-Next route, because
the fixed attention path (ggml-org#29751) builds a k-pool input per QSA layer and makes
the indexer cache track the attention cache cell for cell. ggml-org#29824 reduces the
mask construction cost of that path.

Fork-own work (RAM prompt-cache retention, selective CUDA P2P transport, server
prefill/decode phase isolation, DFlash M-RoPE inference, docs, tests) is
unchanged.
kainlan added a commit to kainlan/llama.cpp-intel-optimizations that referenced this pull request Oct 3, 2026
Adds the NextN/MTP draft head for Qwen3.8-Flash-Next (qwen4exp):
one full-attention QSA block after the trunk, fed by the trunk's
hyper-connection-wide residual, with its own hc head mixer, run through
--spec-type draft-mtp with the head loaded from a separate MTP GGUF (-md).

Ported from upstream ggml-org#29761 "Qwen4Exp: add MTP" (am17an,
merged 2026-10-01 as c061df1; design reviewed by ggerganov). Upstream
credit: the converter, gguf-py tensor names, graph_mtp, the MTP memory
filter, the empty-recurrent-memory handling and the draft-path fix in
common/speculative.cpp are theirs; this commit only carries them over.

Hand-adapted against this tree, which predates upstream ggml-org#29751 (the QSA
rewrite that ggml-org#29761 builds on):
- graph_mtp calls build_layer_attn without the kpool input. This tree
  derives QSA from mctx_hyb->get_idx() and the layer's compress ratio, so
  the MTP block picks QSA up the same way a trunk layer does.
- the indexer_kpool loop-bound hunk and a TODO comment hunk have no target
  here and are dropped.

No SYCL wiring: this commit is exercised on a CPU-only build to measure
draft acceptance. The draft context, the multi-token verify path and the
host-expert cost on the SYCL backend are deliberately not touched.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AWkHqCXdysC9zf9AVvCiNn
neurall added a commit to neurall/llama.cpp that referenced this pull request Oct 3, 2026
Qwen3.8-Flash-Next (qwen4exp): upstream's own implementation and MTP (ggml-org#29761, ggml-org#29751, ggml-org#29824, ggml-org#29819) replace the copy of the earlier MTP PR
(conversion/qwen4exp.py, src/models/qwen4exp.cpp and the qwen4exp hunks of gguf constants / tensor mapping, llama-arch.h, llama-model.h, models.h,
llama-memory-recurrent.cpp). Kept from this fork where upstream has no counterpart: the GLM-5.3-Flash MTP additions (indexer_index_share_mtp,
the hybrid-memory mtp_dsa accessors), dev_offload, the expert cache. common.cpp: both blocks kept (cache auto setup and the decision types).
Compiles CPU-only on pc2; CUDA build and runs pending.
cwriter pushed a commit to cwriter/llama.cpp that referenced this pull request Oct 4, 2026
Upstream ggml-org#29751 rebuilds the qwen4exp QSA indexer on the k-pool cache: the
indexer scores pooled block keys, takes the top blocks plus the tail, and
builds the attention mask as
  REPEAT(FILL(seed, -inf)) -> SET_ROWS(0 at the list) -> VIEW -> ADD(kq_mask)
with one extra cell per row for the n_kv sentinel of padded blocks.

Match that chain from the flash attention node and let the node read the
selection list and the causal mask directly (fattn-qsa.cpp), as before. Every
node of the chain is absorbed, including the seed scalars, so the two
[n_kv + 1, n_ubatch] f16 tensors are never reserved.

The earlier graph had a score gather and a different mask chain; the fusions
for them no longer match anything and are removed: the QSA gather and top-k
fusions, the f16 cast + ADD and CONT + ADD fusions, selection bitmap modes 1
and 2 of GGML_SYCL_FUSE_QSA_FA_MASK and GGML_SYCL_FUSE_QSA_MASK. The radix
top-k is back to the upstream code; only its select moves to the header for
the MoE top-k. GGML_SYCL_FUSE_QSA_FA_MASK is now 0 or 1.

Assisted-by: Claude Opus 5.5
neurall added a commit to neurall/llama.cpp that referenced this pull request Oct 4, 2026
Qwen3.8-Flash-Next (qwen4exp): upstream's own implementation and MTP (ggml-org#29761, ggml-org#29751, ggml-org#29824, ggml-org#29819) replace the copy of the earlier MTP PR
(conversion/qwen4exp.py, src/models/qwen4exp.cpp and the qwen4exp hunks of gguf constants / tensor mapping, llama-arch.h, llama-model.h, models.h,
llama-memory-recurrent.cpp). Kept from this fork where upstream has no counterpart: the GLM-5.3-Flash MTP additions (indexer_index_share_mtp,
the hybrid-memory mtp_dsa accessors), dev_offload, the expert cache. common.cpp: both blocks kept (cache auto setup and the decision types).
Compiles CPU-only on pc2; CUDA build and runs pending.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 5, 2026
The radix top-k picks a value threshold and then gathers. Values that tie with
the threshold exactly are the ones the budget cuts through, and which of them
the gather keeps is decided by the order its atomics happen to run in, so the
selected set is not reproducible between two runs of the same build. Ties are
common where the scores are rectified, since those collapse to exactly zero.
ggml_top_k leaves the choice among tied values unspecified, so this is valid,
but it makes selections and outputs hard to compare between runs.

GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied
values with a second radix walk, over the column index and from the low bin
upwards, and keeps only the tied columns at or below it. The set is then fixed;
the order within the reserved slots still depends on the atomics, which the
callers do not rely on. The extra walk only covers the bits a column index can
use, so it is two passes for a context of up to 64k rather than four.

The switch is read once per process and applies to every ggml_top_k that the
backend runs through the radix selection, not only to one model's graph.

The switch is off by default. Nothing about quality motivates one order among
equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does
not pay for the extra passes. It is a supported option for reproducible runs,
where selections or outputs are compared between runs.

R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the
time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756
per-layer selections without the switch and agree on all of them with it, and
with the switch set a 128-token greedy continuation reproduces byte for byte.
test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 5, 2026
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph
slot of the scheduler, so a speculative verification width seen before is made current again instead of being
built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a
buffer that has to grow); when all slots are in use the least recently used shape is evicted.

LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification
widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as
before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is
disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism.

Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17
requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB,
unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the
defaults (24, 32):
- Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751):
  46.52 -> 51.67 t/s (+11.1 %)
- DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %)
Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell
from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* llama: fix qwen4exp

* qwen4exp: keep kq_mask input the same shape
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 6, 2026
The radix top-k picks a value threshold and then gathers. Values that tie with
the threshold exactly are the ones the budget cuts through, and which of them
the gather keeps is decided by the order its atomics happen to run in, so the
selected set is not reproducible between two runs of the same build. Ties are
common where the scores are rectified, since those collapse to exactly zero.
ggml_top_k leaves the choice among tied values unspecified, so this is valid,
but it makes selections and outputs hard to compare between runs.

GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied
values with a second radix walk, over the column index and from the low bin
upwards, and keeps only the tied columns at or below it. The set is then fixed;
the order within the reserved slots still depends on the atomics, which the
callers do not rely on. The extra walk only covers the bits a column index can
use, so it is two passes for a context of up to 64k rather than four.

The switch is read once per process and applies to every ggml_top_k that the
backend runs through the radix selection, not only to one model's graph.

The switch is off by default. Nothing about quality motivates one order among
equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does
not pay for the extra passes. It is a supported option for reproducible runs,
where selections or outputs are compared between runs.

R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the
time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756
per-layer selections without the switch and agree on all of them with it, and
with the switch set a 128-token greedy continuation reproduces byte for byte.
test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 6, 2026
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph
slot of the scheduler, so a speculative verification width seen before is made current again instead of being
built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a
buffer that has to grow); when all slots are in use the least recently used shape is evicted.

LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification
widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as
before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is
disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism.

Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17
requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB,
unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the
defaults (24, 32):
- Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751):
  46.52 -> 51.67 t/s (+11.1 %)
- DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %)
Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell
from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
sodre90 added a commit to sodre90/llama.cpp that referenced this pull request Oct 8, 2026
…evert ggml-org#29751)

Revert 66e0c17 on top of ec7630a for src/models/qwen4exp.cpp, models.h, llama-hparams.h and hybrid-idx.
Kept from 81e39ad: the n_pool_max clamp, with the stable n_new_g padding it clamps (GLM5 path).
Kept from c061df1: recurrent is_empty(), speculative.cpp, conversion and gguf-py, arch and tensor names.
Dropped from c061df1: the qwen4exp graph_mtp, loader and create_memory filter (fork has its own MTP).
Dropped from 159c651 and 4e2713c: their qwen4exp.cpp hunks, which edit build_qsa_sel of the reverted rewrite. The hybrid-idx pad_cell change of 159c651 and the glm5-next change of 4e2713c are kept.
Dropped from 889edf4: its qwen4exp.cpp hunks (pool_mask F16 and the lightning-indexer score), which edit the same rewrite-only k-pool code. Its other files are kept.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 9, 2026
The radix top-k picks a value threshold and then gathers. Values that tie with
the threshold exactly are the ones the budget cuts through, and which of them
the gather keeps is decided by the order its atomics happen to run in, so the
selected set is not reproducible between two runs of the same build. Ties are
common where the scores are rectified, since those collapse to exactly zero.
ggml_top_k leaves the choice among tied values unspecified, so this is valid,
but it makes selections and outputs hard to compare between runs.

GGML_CUDA_TOP_K_STABLE_TIES=1 selects the smallest columns among the tied
values with a second radix walk, over the column index and from the low bin
upwards, and keeps only the tied columns at or below it. The set is then fixed;
the order within the reserved slots still depends on the atomics, which the
callers do not rely on. The extra walk only covers the bits a column index can
use, so it is two passes for a context of up to 64k rather than four.

The switch is read once per process and applies to every ggml_top_k that the
backend runs through the radix selection, not only to one model's graph.

The switch is off by default. Nothing about quality motivates one order among
equal scores, so normal use keeps the behaviour of ggml_top_k as it is and does
not pay for the extra passes. It is a supported option for reproducible runs,
where selections or outputs are compared between runs.

R9700, Qwen3.8-Flash-Next UD-Q4_K_XL, 32k prefill, with the Qwen indexer of the
time (before upstream ggml-org#29751): two runs of the same build disagree on 497 of 756
per-layer selections without the switch and agree on all of them with it, and
with the switch set a 128-token greedy continuation reproduces byte for byte.
test-backend-ops -o TOP_K passes 525/525 with the switch unset and set.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 9, 2026
A context keeps the built graph of each batch shape together with its scheduler split and allocation in a graph
slot of the scheduler, so a speculative verification width seen before is made current again instead of being
built, split and allocated anew. The slots are dropped when the compute buffers are planned again (a reserve, or a
buffer that has to grow); when all slots are in use the least recently used shape is evicted.

LLAMA_GRAPH_REUSE_SHAPES: not set = up to 24 shapes per context (a speculative server uses up to 16 verification
widths, and the rest leaves room for a second graph variant of a width), 0 = off (one graph per output class, as
before), 1 = 8, N > 1 = up to N (at most 64). It switches itself off, and logs the reason, when graph reuse is
disabled, for a model loaded with no_alloc, with GGML_CUDA_GRAPH_OPT=1 and with pipeline parallelism.

Measured with both switches set by environment on one build, R9700, 128 GB host memory, English roleplay (17
requests), MTP --spec-draft-n-max 7 --spec-draft-p-min 0.7 with the joint cache, exclusive expert cache 20480 MiB,
unlimited host tier, warm, one run per row, off (LLAMA_GRAPH_REUSE_SHAPES=0, backend bound 16) against the
defaults (24, 32):
- Qwen3.8-Flash-Next UD-Q4_K_XL (MTP head of upstream pull request ggml-org#28243, Qwen indexer before upstream ggml-org#29751):
  46.52 -> 51.67 t/s (+11.1 %)
- DeepSeek V4 Flash UD-IQ3_XXS: 28.97 -> 31.45 t/s (+8.6 %)
Replies 17/17 identical and acceptance unchanged in both. The extra time of the call after a width change fell
from +3.6..+9.8 ms to about 0 (within 0.8 ms). Peak private bytes +0.85..0.97 GiB, dedicated VRAM unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

highlight Changes that will be highlighted in the next release notes model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants