Skip to content

Ggml integer overflow - #29384

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
apach301:ggml-integer-overflow
Sep 30, 2026
Merged

ggerganov merged 3 commits into
ggml-org:masterfrom
apach301:ggml-integer-overflow

Conversation

@apach301

Copy link
Copy Markdown
Contributor

Fixes #29383

Overview

While fuzzing llama.cpp, I found several uncaught integer overflows during tensor parsing.

First one is improper int-overflow guard at gguf_init_from_reader(). The existing validation attempts
to ensure that the total number of elements is representable:

if (ok && ggml_nelements(&info.t) > 0 &&
    ((INT64_MAX/info.t.ne[1] <= info.t.ne[0]) ||
     (INT64_MAX/info.t.ne[2] <= info.t.ne[0]*info.t.ne[1]) ||
     (INT64_MAX/info.t.ne[3] <= info.t.ne[0]*info.t.ne[1]*info.t.ne[2]))) {
    ...
}

However, the arithmetic performed by the validation itself can overflow in ggml_nelements(), before result is checked. Additionally,
if one of the elements is zero (or multiplication overflows to zero), the condition becomes false and validation skipped entirely.

Another one is missing guard for possible integer overflows in ggml_new_tensor_impl(), that cause many overflow errors in ggml.c.
One of the UBSAN reports:

   "/llama.cpp/ggml/src/ggml.c:1319:42: runtime error: unsigned integer overflow: 6917530127152709631 * 103903848824832 cannot be represented in type 'unsigned long'",
    "    #0 0x12ab637 in ggml_nbytes /llama.cpp/ggml/src/ggml.c:1319:42",
    "    #1 0x14afbe3 in gguf_init_from_reader(gguf_reader const&, gguf_init_params) /llama.cpp/ggml/src/gguf.cpp:786:34",
    "    #2 0x14b36c3 in gguf_init_from_file_ptr /llama.cpp/ggml/src/gguf.cpp:956:12",
    "    #3 0x14b36c3 in gguf_init_from_file /llama.cpp/ggml/src/gguf.cpp:1000:36",
    "    #4 0x900dbb in llama_model_loader::llama_model_loader(gguf_context*, void (*)(ggml_tensor*, void*), void*, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >&, _IO_FILE*, llama_load_mode, bool, bool, bool, llama_model_kv_override const*, llama_model_tensor_buft_override const*) /llama.cpp/src/llama-model-loader.cpp:565:28",
    "    #5 0x52bfcf in llama_model_load(gguf_context*, void (*)(ggml_tensor*, void*), void*, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >&, _IO_FILE*, llama_model_params&) /llama.cpp/src/llama.cpp:318:28",
    "    #6 0x529065 in llama_model_load_from_file_impl(gguf_context*, void (*)(ggml_tensor*, void*), void*, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >&, _IO_FILE*, llama_model_params) /llama.cpp/src/llama.cpp:425:34",
    "    #7 0x529832 in llama_model_load_from_file /llama.cpp/src/llama.cpp:466:12",
    "    #8 0x4f11b1 in LLVMFuzzerTestOneInput /llama.cpp/fuzzers/fuzz_inference.cpp:63:19",
    "    #9 0x1ba05f0 in fuzzer::Fuzzer::ExecuteCallback(unsigned char const*, unsigned long) (/out_fuzz/fuzz_inference_fuzz+0x1ba05f0)",
    "    #10 0x1b8a822 in fuzzer::RunOneTest(fuzzer::Fuzzer*, char const*, unsigned long) (/out_fuzz/fuzz_inference_fuzz+0x1b8a822)",
    "    #11 0x1b9038a in fuzzer::FuzzerDriver(int*, char***, int (*)(unsigned char const*, unsigned long)) (/out_fuzz/fuzz_inference_fuzz+0x1b9038a)",
    "    #12 0x1b8a542 in main (/out_fuzz/fuzz_inference_fuzz+0x1b8a542)",
    "    #13 0x14a1d5f3c082 in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x24082) (BuildId: 5792732f783158c66fb4f3756458ca24e46e827d)",
    "    #14 0x43021d in _start (/out_fuzz/fuzz_inference_fuzz+0x43021d)",
    "",
    "SUMMARY: UndefinedBehaviorSanitizer: unsigned-integer-overflow /llama.cpp/ggml/src/ggml.c:1319:42 in"

In both cases malformed tensor does not appear to be directly exploitable through this issue alone: subsequent checks (check_tensor_dims, asserts) are succesfully caught errors and bail out.

Additional information

Requirements

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 24, 2026

Copy link
Copy Markdown

Hi @apach301, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 24, 2026
@github-actions
github-actions Bot marked this pull request as draft September 24, 2026 17:11
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning and removed draft PR will be changed to draft by github-actions bot labels Sep 24, 2026
@ggerganov

Copy link
Copy Markdown
Member

@apach301 Please address the CI failures.

@uis246

This comment was marked as low quality.

@apach301
apach301 marked this pull request as ready for review September 25, 2026 10:53
@apach301

Copy link
Copy Markdown
Contributor Author

@ggerganov I made a fix, 2 webgpu failures seems unrelated (/Users/ggml/actions-runner/_work/llama.cpp/llama.cpp/ggml/src/ggml-webgpu/ggml-webgpu.cpp:4200: ggml_webgpu: Device error! Reason: 2, Message: BufferOffset (28930) is not a multiple of 4.)

@ggerganov
ggerganov merged commit 2149c00 into ggml-org:master Sep 30, 2026
25 of 27 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print

(cherry picked from commit 2149c00)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Misc. bug: uncaught integer overflows during model loading

4 participants