Skip to content

ggml : add alloc_buffer_n to buffer type interface - #23671

Merged
ggerganov merged 9 commits into
masterfrom
gg/ggml-alloc-refactor-meta
Oct 2, 2026
Merged

ggerganov merged 9 commits into
masterfrom
gg/ggml-alloc-refactor-meta

Conversation

@ggerganov

@ggerganov ggerganov commented May 25, 2026 •

Copy link
Copy Markdown
Member

Overview

cont #19378

The ggml_backend_meta_alloc_ctx_tensors_from_buft was a temporary workaround. This patch avoids the function by extending the buffer type interface with an API that allocates a backend buffer from a list of tensors. This decouples the ggml-alloc from meta backend specifics and allows more flexible buffer allocations in the future.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp + pi:llama.cpp/Qwen3.8-27B

@github-actions github-actions Bot added Nvidia GPU Issues specific to Nvidia GPUs Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs OpenCL Issues specific to the OpenCL backend IBM zDNN issues specific to IBM zDNN Accelerator Hexagon OpenVINO WebGPU labels May 25, 2026
michaelw9999 added a commit to michaelw9999/advanced-gguf-quantizer that referenced this pull request May 26, 2026
Port ggml-org/llama.cpp#23671 into advanced-gguf-quantizer as one local integration commit.

Includes upstream commits 0743dce657dc9901ee56734dbce571cf0a6abb8d (add alloc_buffer_n to buffer type interface) and 3f9bcab883c84c6336e100ce5d7408bb4bf47ac9 (fix cur_buf_size after buffer flush).
@ggerganov
ggerganov force-pushed the gg/ggml-alloc-refactor-meta branch from 3f9bcab to 7c7be0f Compare June 15, 2026 13:06
@github-actions github-actions Bot added the CUDA Related to the CUDA backend label Jun 15, 2026
@ggerganov
ggerganov force-pushed the gg/ggml-alloc-refactor-meta branch 2 times, most recently from d2ebf49 to a75c427 Compare September 19, 2026 10:11
@github-actions github-actions Bot added the testing Everything test related label Sep 19, 2026
@ggerganov
ggerganov marked this pull request as ready for review September 20, 2026 07:06
@ggerganov
ggerganov requested review from a team, marty1885 and wine99 as code owners September 20, 2026 07:06
Comment thread ggml/src/ggml-alloc.c Outdated

static ggml_backend_buffer_t ggml_backend_alloc_ctx_tensors_from_buft_impl(
struct ggml_context * ctx, ggml_backend_buffer_type_t buft, size_t * nbytes_total, bool no_alloc) {
// TODO [TAG_ALLOC_SHARED_BUFFER_SPLIT]: reuse shared buffer-splitting logic from ggml_backend_buft_alloc_buffer_n_default

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we allow ggml backends to implement their own logic for allocating ggml tensors then I think we have to also make the corresponding function for projecting the memory use part of the ggml backend API. Otherwise a backend implementing its own memory allocation will automatically break -fit. This is in essence why as of right now ggml_backend_buft_get_alloc_size is broken for the meta backend and why -fit and -sm tensor are incompatible. To the code in fit.cpp it is currently not clear how e.g. a logical meta buffer maps to physical buffers.

The API to make this work will I think need to be that a logical ggml backend buffer type gets some property like n_physical to represent how many physical ggml backend buffer types it contains (1 by default). When fetching the (projected) allocated size, add functions that return n_physical values, so effectively a breakdown by physical buffer types. We can keep the current functions for convenience but change their semantic meaning to be that they return the aggregate size of all physical buffer types.

@ggerganov ggerganov Sep 21, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And I guess for the fit + -sm tensor we cannot work with just knowing the total size across all the physical buffers? Because if we could, we can just add the respective get_alloc_size_n.

Can the fit logic not determine the physical split of the total size based on the split states information that it should have knowledge of?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on how large our tolerance for imprecision is. We can approximate the memory use per backend from the total sum + the tensor split. For most cases I expect this to be within ~5% and in principle workable (since we have a 1 GiB margin by default). But we will not be able to give users information with MiB precision as we do otherwise.

I see this as less of a llama.cpp issue and more as a ggml issue. If we have a function that is supposed to tell user code how much memory will be allocated I think it's a comparatively poorer experience for developers if the result they get is only approximately correct.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can in principle also treat the meta backend as a special case that needs to be treated separately in the user code. The question is whether we expect to have more such backend buffer types in the future that would warrant a more general API.

@ggerganov ggerganov Sep 21, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding a get_alloc_size_n call is fine and is basically the right thing to do in order to make the ggml_backend_alloc_ctx_tensors_from_buft_size work correctly in general, as you correctly noted. So I think I'll follow up on that.

Whether the ggml API should provide a fine-grained information about the sub-buffers - I am not yet convinced. Querying for n_physical values loses generality and does not seem like a good design. A better option would be to have a call that basically asks "how much is the allocation size of these tensors on device X using this buffer type?". But still, I am not sure we really want such level of specificity.

If we can make the fit logic work with a small tolerance, I would say it's good enough for now. And as the meta backend gets more adopted, we can reconsider.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's for now only add the functionality that we are sure we will definitely need.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In principle, (almost) all of the functionality that I would need in fit.cpp is already exposed in ggml-backend-impl.h. I think we at some point had a header like ggml-ext.h for unstable ggml functionality that we removed again for some reason. What was the problem again?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The ggml-ext.h is not compatible in cases where a downstream builds nightly llama.cpp and uses a system-wide ggml binaries (instead of the local ggml copy from llama.cpp). This was the case for Homebrew back then: #21869. Now that we have a semantic versioning in place, this is no longer a problem, since the downstreams are supposed to build from the stable tags. For example, Homebrew is already doing that.

@max-krasnyansky max-krasnyansky left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me as a general improvement.

For ggml-hexagon, it would be useful to allocate different buffers for different tensors from the list. With the addition of 64-bit DMA support, we now have two buffer flavors: regular and extended (64-bit mapping). Currently, all buffers with usage=weight are mapped as extended, which forces kernels to use DMA for every tensor allocated within them. I’ve already updated most kernels to handle this, so it’s not a blocker, but it would be helpful to have per-tensor flexibility—e.g., if a tensor is too small or its kernel isn't ready for DMA, we could route it to a different buffer rather than bundling it with the rest.

@ggerganov
ggerganov force-pushed the gg/ggml-alloc-refactor-meta branch 2 times, most recently from 514cd02 to 6662b12 Compare September 23, 2026 19:14
@max-krasnyansky

Copy link
Copy Markdown
Member

Ah. I see that it's possible to return the multi-buffer where different tensors map to different sub-buffer. Thanks for adding that alloc-plan based implementation. Let me know if I misread that :)
Looks good to me. I'll play with that option in ggml-hexagon.

@ggerganov
ggerganov force-pushed the gg/ggml-alloc-refactor-meta branch from 6662b12 to aa7ee0b Compare September 24, 2026 14:03
@ggerganov

Copy link
Copy Markdown
Member Author

Ah. I see that it's possible to return the multi-buffer where different tensors map to different sub-buffer. Thanks for adding that alloc-plan based implementation. Let me know if I misread that :)

Yes. We don't currently have a specific application of this new API apart from the meta backend. So would be nice in case you find a use case in the hexagon backend in order to give us more confidence that this change is worth it.

@max-krasnyansky

Copy link
Copy Markdown
Member

Ah. I see that it's possible to return the multi-buffer where different tensors map to different sub-buffer. Thanks for adding that alloc-plan based implementation. Let me know if I misread that :)

Yes. We don't currently have a specific application of this new API apart from the meta backend. So would be nice in case you find a use case in the hexagon backend in order to give us more confidence that this change is worth it.

Sounds good. I'll give that a shot a bit later today.

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you.

Comment thread ggml/src/ggml-alloc.c
}
struct ggml_tensor ** tensors = (struct ggml_tensor **) malloc(n * sizeof(struct ggml_tensor *));
if (tensors == NULL) {
return NULL;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we return NULL here ggml_backend_alloc_ctx_tensors_from_buft_size will return 0 which is likely to cause weird bugs in user code. I think this needs a print to one of the ggml logs to make debugging easier. Preferably also document in the header that NULL and 0 are returned in case of an error.

@max-krasnyansky

Copy link
Copy Markdown
Member

Ah. I see that it's possible to return the multi-buffer where different tensors map to different sub-buffer. Thanks for adding that alloc-plan based implementation. Let me know if I misread that :)

Yes. We don't currently have a specific application of this new API apart from the meta backend. So would be nice in case you find a use case in the hexagon backend in order to give us more confidence that this change is worth it.

Sounds good. I'll give that a shot a bit later today.

Sorry for the delay. Seems to work as expected.

Here is the first cut where we split large tensors into standalone buffersmaxk/hexagon-alloc-buffer-n.
Feel free to pull that branch into this PR or we can merge this and I'll start a separate one.

@max-krasnyansky

Copy link
Copy Markdown
Member

@ggerganov @JohannesGaessler any reason we're not merging this yet?
I'm going to add some more logic to the ggml-hexagon alloc_buffer_n.
So it'd be good to merge this.

Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.

- Default implementation in ggml-backend.cpp handles multi-buffer
  splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
  per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
  into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
  interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)

Assisted-by: llama.cpp:local pi
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Assisted-by: pi:llama.cpp/Qwen3.8-27B
- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
  default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
  ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo

Assisted-by: pi:llama.cpp/Qwen3.8-27B
- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
@ggerganov
ggerganov force-pushed the gg/ggml-alloc-refactor-meta branch from aa7ee0b to d211167 Compare October 2, 2026 06:23
@ggerganov

ggerganov commented Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Sorry for the delay - I wanted to also implement get_alloc_size_n for the meta backend, but will do it in a follow-up PR to not block this. Merging on green.

@ggerganov
ggerganov merged commit 631109b into master Oct 2, 2026
66 of 69 checks passed
novkien added a commit to novkien/llama.cpp-fork that referenced this pull request Oct 2, 2026
Clean upstream sync past 254b177. Relevant to the fork's deployment:

- 4e2713c qwen4exp : optimize mask constructions (ggml-org#29824)
- 631109b ggml : add alloc_buffer_n to the buffer type interface (ggml-org#23671)

plus sycl / vulkan / opencl kernel work, llama warning and abort cleanups,
and a new server /v1/systemone API (ggml-org#29818).

Why it matters here: adopting upstream's Qwen4Exp MTP (ggml-org#29761) costs roughly
+2.6 GB of compute buffers per device on the Qwen3.8-Flash-Next route, because
the fixed attention path (ggml-org#29751) builds a k-pool input per QSA layer and makes
the indexer cache track the attention cache cell for cell. ggml-org#29824 reduces the
mask construction cost of that path.

Fork-own work (RAM prompt-cache retention, selective CUDA P2P transport, server
prefill/decode phase isolation, DFlash M-RoPE inference, docs, tests) is
unchanged.
hoivb612 pushed a commit to hoivb612/llama.cpp that referenced this pull request Oct 3, 2026
Merge the latest ggml-dx12-main work (linalg / flash-attn shader set,
autotune header, linalg-bench tooling) and adapt all three DX12 backends
to the buffer-type interface change from 631109b (ggml-org#23671), which
inserted the optional alloc_buffer_n and get_alloc_size_n members and
bumped GGML_BACKEND_API_VERSION from 2 to 3.

- ggml-dx12, ggml-dx12x, ggml-dx12-main: add the two new nullptr slots to
  dx12_buffer_type_interface so the positional initializers line up with
  the new member order. Both new members are optional with documented
  defaults, so behaviour is unchanged.
- ggml-dx12-main: forward-declare ggml_backend_dx12_set_env_refresh and
  ggml_backend_dx12_set_flag_sink ahead of dx12_reg_get_proc_address.
  They are defined later in the same file and are deliberately absent
  from ggml/include/ggml-dx12.h, being reached via get_proc_address.

graph_optimize is nullptr in ggml-dx12 and ggml-dx12x, so upstream's
signature change to that callback affected only ggml-dx12-main.

Build-verified with Ninja/Release under MSVC: GGML_DX12_MAIN=ON,
GGML_DX12=ON and GGML_DX12X=ON each configure and compile cleanly with
zero warnings in the DX12 sources. Not yet exercised on device.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d7f2fcd-a611-433c-a522-a341fb127413
hoivb612 pushed a commit to hoivb612/llama.cpp that referenced this pull request Oct 3, 2026
Flow-down of b43b8c1, which merges the latest ggml-dx12-main work and
adapts all three DX12 backends to the buffer-type interface change from
631109b (ggml-org#23671) that bumped GGML_BACKEND_API_VERSION from 2 to 3.

ggml/src/ggml-dx12, ggml-dx12-main and ggml-dx12x are maintained on
hv/b612_100326 and mirrored here verbatim, so the three directories are
replaced wholesale rather than merged. Their tree hashes match b43b8c1
exactly. No files outside ggml/src/ggml-dx12* are touched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d7f2fcd-a611-433c-a522-a341fb127413
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* ggml : add `alloc_buffer_n` to buffer type interface

Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.

- Default implementation in ggml-backend.cpp handles multi-buffer
  splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
  per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
  into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
  interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)

Assisted-by: llama.cpp:local pi

* cont : fix `cur_buf_size` init after flushing a buffer

* ggml : add TODO tag for shared buffer split logic

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* tests : add alloc_buffer_n coverage

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : fix compile warnings

* tests : add descriptions for alloc_buffer_n tests

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : address review comments on alloc_buffer_n

- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
  default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
  ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : add get_alloc_size_n to buffer type interface

- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : report malloc failure
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning Hexagon IBM zDNN issues specific to IBM zDNN Accelerator Nvidia GPU Issues specific to Nvidia GPUs OpenCL Issues specific to the OpenCL backend OpenVINO SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related Vulkan Issues specific to the Vulkan backend WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants