graph: fix CI realloc abort by gathering the recurrent states once - #29856
Merged
ServeurpersoCom merged 1 commit intoOct 3, 2026
Merged
ServeurpersoCom merged 1 commit into
ServeurpersoCom merged 1 commit into
Conversation
…plit build_rs gathered the extra states (n_rs - n_seqs rows) with their own get_rows. The worst-case reserve has n_rs == n_seqs, so that node was sized at zero rows, and any ubatch whose cells are not contiguous forced a graph reallocation at an unchanged node count, which aborts under GGML_SCHED_NO_REALLOC. A single get_rows now gathers the n_rs states: the ubatch states and the extra states are views of it, and its size only depends on n_rs, which the reserve already sets to the maximum. A custom getter (mamba ssm_scan) gathers from the second state, so a single sequence ubatch copies no state. The views are built once per graph in the input to keep the host overhead of the graph unchanged.
ggerganov
approved these changes
Oct 2, 2026
CISC
approved these changes
Oct 2, 2026
ravi9
added a commit
to ravi9/llama.cpp
that referenced
this pull request
Oct 5, 2026
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views.
5 tasks
ravi9
added a commit
to ravi9/llama.cpp
that referenced
this pull request
Oct 6, 2026
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views.
ggerganov
pushed a commit
that referenced
this pull request
Oct 6, 2026
* ggml-openvino: skip unselected graph branches and support DUP Upstream #29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend. * ggml-openvino: make inp_scale_rows token dim dynamic #29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path. * ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in #29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU. * ggml-openvino: create FILL in the output type translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant. * ggml-openvino: reject CONCAT with a quantized type Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type. * ggml-openvino: handle the single recurrent state gather of build_rs #29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views. * openvino: align eltwise operand ranks to work around a GPU-plugin defect * openvino: match the MoE fusion on the rank-3 stateful graph * ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 the lower-rank operand of the post-attention residual add is the norm output, and unsqueezing it makes the GPU plugin compute the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution. Skip the rewrite when the lower-rank operand is an RMS norm output. * docs : update OpenVINO validated models --------- Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fix CI / graph
Hybrid models keep a recurrent state per sequence, and build_rs sometimes has to move the states of other sequences back into place. The worst-case reserve did not account for those moved states, so the first time one appeared the graph was reallocated mid-run: an abort under GGML_SCHED_NO_REALLOC (test_systemone since #29818) and a silent reallocation otherwise.
build_rs now gathers all the states at once in a buffer that the reserve always sizes for the maximum, so this reallocation no longer happens. The output is bit-exact and decode speed is unchanged.
Additional information
Follow-up #29818
Requirements