Skip to content

ci: add Hexagon NPU backend build for Windows Arm64 - #29052

Open
tfenster wants to merge 8 commits into
ggml-org:masterfrom
tfenster:master
Open

tfenster wants to merge 8 commits into
ggml-org:masterfrom
tfenster:master

Conversation

@tfenster

Copy link
Copy Markdown

Overview

This PR implements the missing Hexagon NPU build

Additional information

fixes #26877

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. I asked GitHub Copilot to implement the missing CI build and then asked Claude to review it. Afterwards I reviewed it myself

Copilot AI lite review requested due to automatic review settings September 17, 2026 21:43
@tfenster
tfenster requested a review from a team as a code owner September 17, 2026 21:43
@github-actions github-actions Bot added the devops improvements to build systems and github actions label Sep 17, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Critical workflow, SDK environment, packaging, and signing issues block reliable builds.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Adds CI and release packaging for Windows Arm64 Hexagon NPU binaries.

Changes:

  • Installs Snapdragon SDKs and builds Hexagon artifacts.
  • Packages and uploads the Hexagon build.
  • Integrates the artifact into release dependencies and documentation links.
File summaries
File Description
.github/workflows/release.yml Adds the Windows Arm64 Hexagon build and release integration.
Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 6
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/release.yml Outdated
Comment thread .github/workflows/release.yml Outdated
Comment thread .github/workflows/release.yml Outdated
Comment thread .github/workflows/release.yml Outdated
Comment thread .github/workflows/release.yml Outdated
Comment thread .github/workflows/release.yml Outdated
@taronaeo

Copy link
Copy Markdown
Member

cc: @ggml-org/ggml-hexagon

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The unsigned artifact lacks the required catalog/signing path, and the unused OpenCL setup adds release risk.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

.github/workflows/release.yml:1135

  • This installs and configures the OpenCL SDK, but the build target list only builds ggml-hexagon and the HTP skels, so ggml-opencl.dll never reaches this archive. That adds an unnecessary external download to every release and another failure point. Remove the OpenCL setup/options, or explicitly build and package ggml-opencl if this is intended to be a dual-backend artifact.
  • Files reviewed: 1/1 changed files
  • Comments generated: 1
  • Review effort level: Balanced

Comment thread .github/workflows/release.yml Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The linked documentation does not explain how users can sign and install the unsigned release artifact.

Get a fresh assessment by requesting another Copilot review.

Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 1
  • Review effort level: Balanced

Comment thread .github/workflows/release.yml Outdated
@tfenster
tfenster requested a review from a team as a code owner September 18, 2026 05:49
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 18, 2026
@tfenster
tfenster requested a balanced review from Copilot September 18, 2026 05:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The release-critical Windows toolchain and driver-package workflow was not exercised by the available PR checks.

Review details
  • Files reviewed: 2/2 changed files
  • Comments generated: 1
  • Review effort level: Balanced

Comment thread docs/backend/snapdragon/windows.md Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Catalog OS compatibility and package upgrade behavior must be corrected before release.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (1)

Previously missed (1) — in code that hasn't changed since the last review.

docs/backend/snapdragon/windows.md:158

  • This install command is not a reliable upgrade path for recurring releases. The bundled INF still has the fixed DriverVer = 01/01/2026,1.0.0.0 and marks every skel with COPYFLG_NO_OVERWRITE (ggml/src/ggml-hexagon/libggml-htp.inf:6,29-32), so a machine that installed an earlier release can retain that same-ranked package and its old HTP libraries. Generate/bump the INF version for each release, or provide an explicit uninstall/replacement procedure before publishing this as the release installation flow.
  • Files reviewed: 2/2 changed files
  • Comments generated: 1
  • Review effort level: Balanced

> $certificate="c:\Users\MyUser\Certs\ggml-htp-v1.pfx"
> $windowsSdkBin="c:\Program Files (x86)\Windows Kits\10\bin\10.0.26100.0"

> & "$windowsSdkBin\arm64\Inf2Cat.exe" /driver:"$package" /os:10_25H2_ARM64
@tfenster

Copy link
Copy Markdown
Author

can you PTAL @max-krasnyansky @CISC? @ericcurtin mentioned that you would be the best to ping

@max-krasnyansky

Copy link
Copy Markdown
Member

can you PTAL @max-krasnyansky @CISC? @ericcurtin mentioned that you would be the best to ping

Looks good overall.
Interesting idea to deal with the signing requirement on the released package.
I'll take another look a bit later today and add more specific comments.

btw we also have signed releases available via https://github.com/qualcomm/GenieX/releases/
For example, if you grabgeniex-bench-windows-arm64-v0.7.0.zip
It contains signed libggml-htp libraries.

@zhiyuan8 Just FYI.

@ericcurtin

ericcurtin commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

can you PTAL @max-krasnyansky @CISC? @ericcurtin mentioned that you would be the best to ping

Looks good overall. Interesting idea to deal with the signing requirement on the released package. I'll take another look a bit later today and add more specific comments.

btw we also have signed releases available via https://github.com/qualcomm/GenieX/releases/ For example, if you grabgeniex-bench-windows-arm64-v0.7.0.zip It contains signed libggml-htp libraries.

@zhiyuan8 Just FYI.

I think the main problem is the lack of a signed llama-server binary, that one is particular is kinda useful (as well has having binaries/libs per llama.cpp upstream release). @max-krasnyansky would it be possible to have signed releases here also?

@max-krasnyansky

Copy link
Copy Markdown
Member

can you PTAL @max-krasnyansky @CISC? @ericcurtin mentioned that you would be the best to ping

Looks good overall. Interesting idea to deal with the signing requirement on the released package. I'll take another look a bit later today and add more specific comments.
btw we also have signed releases available via https://github.com/qualcomm/GenieX/releases/ For example, if you grabgeniex-bench-windows-arm64-v0.7.0.zip It contains signed libggml-htp libraries.
@zhiyuan8 Just FYI.

I think the main problem is the lack of a signed llama-server binary, that one is particular is kinda useful (as well has having binaries/libs per llama.cpp upstream release). @max-krasnyansky would it be possible to have signed releases here also?

Hmm. Not sure what you mean. There is no need to sign the host side apps and tools.
For WoS (Windows on Snapdragon) we just need signed .cat file for libggml-htp-v73 (X-Elite) and libggml-htp-v81 (X2-Elite).

@ericcurtin

ericcurtin commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

can you PTAL @max-krasnyansky @CISC? @ericcurtin mentioned that you would be the best to ping

Looks good overall. Interesting idea to deal with the signing requirement on the released package. I'll take another look a bit later today and add more specific comments.
btw we also have signed releases available via https://github.com/qualcomm/GenieX/releases/ For example, if you grabgeniex-bench-windows-arm64-v0.7.0.zip It contains signed libggml-htp libraries.
@zhiyuan8 Just FYI.

I think the main problem is the lack of a signed llama-server binary, that one is particular is kinda useful (as well has having binaries/libs per llama.cpp upstream release). @max-krasnyansky would it be possible to have signed releases here also?

Hmm. Not sure what you mean. There is no need to sign the host side apps and tools. For WoS (Windows on Snapdragon) we just need signed .cat file for libggml-htp-v73 (X-Elite) and libggml-htp-v81 (X2-Elite).

I guess it would work... FWIW I don't have one of these devices :) So I haven't tinkered around with what does and doesn't need to be signed... The driving feature is this I guess llmmanorg/llmman#516 ... If we can pull from llama.cpp releases and it works... Everything is good...

@CISC CISC left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nvm. :)

Comment thread .github/workflows/release.yml Outdated
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
@tfenster

Copy link
Copy Markdown
Author

@CISC can you PTAL again? I have applied your suggestion

CISC
CISC previously approved these changes Sep 22, 2026
@CISC

CISC commented Sep 22, 2026

Copy link
Copy Markdown
Member

@tfenster Can you run release on your fork so we can see that the build succeeds?

@tfenster

tfenster commented Sep 22, 2026 •

Copy link
Copy Markdown
Author

https://github.com/tfenster/llama.cpp/actions/runs/35724025761

Update: Fails :( Need to take a closer look to figure out why

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Sep 22, 2026
set(CMAKE_FIND_ROOT_PATH_MODE_INCLUDE ONLY)
set(CMAKE_FIND_ROOT_PATH_MODE_PACKAGE ONLY)
set(CUSTOM_RUNELF_PATH "")
set(CMAKE_TRY_COMPILE_TARGET_TYPE STATIC_LIBRARY)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm. Why do we need this?
Did you test with our docker toolchains and scripts/snapdragon/build.py?

@tfenster tfenster Sep 23, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I ran the release pipeline as suggested here and the first run failed. After the change it worked. As it only affects the test build (if I understand it correctly), I thought it was worth a try

@max-krasnyansky max-krasnyansky Sep 24, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When you get the chance please try ./scripts/snapdragon/build.py --target adb. --target ubuntu also works.

You don't need anything other than the working docker setup. It'll run local build in snapdragon-toolchain docker container.
The change looks like it's not needed to me. And actually looks wrong. We don't build anything static on HTP. I'll take another look why your first build has failed.

@ericcurtin

Copy link
Copy Markdown
Collaborator

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

os: ubuntu-22.04
- build: 'arm64'
os: ubuntu-24.04-arm
- build: 's390x'

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this a mistake? @tfenster

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm guessing it's because they are not using a branch, but PRing directly from master, so in order to test they need to temporary disable to allow action to finish.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

indeed

@ericcurtin

Copy link
Copy Markdown
Collaborator

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

Actually on this, maybe a Dockerfile makes more sense for Linux...

@max-krasnyansky

max-krasnyansky commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

Why "future" proofing ;) the future is here now! :)

This kit runs Ubuntu, all NPU stuff is available via APT. IQ9 with dual-NPU.
https://www.qualcomm.com/developer/hardware/qualcomm-iq-9075-evaluation-kit-evk/hardware

Arduino Ventuno-Q is around the corner as well. Also Ubuntu. Single NPU.
https://www.arduino.cc/product-ventuno-q

IQ10 kits should start showing up soon as well. IQ10 with quad-NPU.

All fully supported with Tensor and Row split implementations.

So yeah, it'd be worthwhile packaging Ubuntu arm64 builds now.
The toolchain docker and things are all available.

@ericcurtin

Copy link
Copy Markdown
Collaborator

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

Why "future" proofing ;) the future is here now! :)

This kit runs Ubuntu, all NPU stuff is available via APT. IQ9 with dual-NPU. https://www.qualcomm.com/developer/hardware/qualcomm-iq-9075-evaluation-kit-evk/hardware

Arduino Ventuno-Q is around the corner as well. Also Ubuntu. Single NPU. https://www.arduino.cc/product-ventuno-q

IQ10 kits should start showing up soon as well. IQ10 with quad-NPU.

All fully supported with Tensor and Row split implementations.

So yeah, it'd be worthwhile packaging Ubuntu arm64 builds now. The toolchain docker and things are all available.

You see this issue with doing out of tree is for integrations like llmman, we have a release on a per commit basis for ROCM, CUDA, opencl, vulkan, etc. (it's a long list) at the per-commit level... So you can have the exact same version of llama.cpp/ggml across all the GPU types... And "llama-server" is the key binary other than the libs... Doing one out of tree is messy and prone to drift...

@max-krasnyansky

Copy link
Copy Markdown
Member

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

Why "future" proofing ;) the future is here now! :)
This kit runs Ubuntu, all NPU stuff is available via APT. IQ9 with dual-NPU. https://www.qualcomm.com/developer/hardware/qualcomm-iq-9075-evaluation-kit-evk/hardware
Arduino Ventuno-Q is around the corner as well. Also Ubuntu. Single NPU. https://www.arduino.cc/product-ventuno-q
IQ10 kits should start showing up soon as well. IQ10 with quad-NPU.
All fully supported with Tensor and Row split implementations.
So yeah, it'd be worthwhile packaging Ubuntu arm64 builds now. The toolchain docker and things are all available.

You see this issue with doing out of tree is for integrations like llmman, we have a release on a per commit basis for ROCM, CUDA, opencl, vulkan, etc. (it's a long list) at the per-commit level... So you can have the exact same version of llama.cpp/ggml across all the GPU types... And "llama-server" is the key binary other than the libs... Doing one out of tree is messy and prone to drift...

I'm suggesting (or ACKing your suggestion) to do in-tree.
The APT stuff I mentioned above is for the NPU drivers & FW
I think we're on the same page -- let's do WoS and Ubuntu ARM64 releases here.

@ericcurtin

Copy link
Copy Markdown
Collaborator

I wonder should we do Linux as a follow on PR, just for future proofing... And I'm sure a lot of the inference nerds like myself want to run on Linux anyway

Why "future" proofing ;) the future is here now! :)
This kit runs Ubuntu, all NPU stuff is available via APT. IQ9 with dual-NPU. https://www.qualcomm.com/developer/hardware/qualcomm-iq-9075-evaluation-kit-evk/hardware
Arduino Ventuno-Q is around the corner as well. Also Ubuntu. Single NPU. https://www.arduino.cc/product-ventuno-q
IQ10 kits should start showing up soon as well. IQ10 with quad-NPU.
All fully supported with Tensor and Row split implementations.
So yeah, it'd be worthwhile packaging Ubuntu arm64 builds now. The toolchain docker and things are all available.

You see this issue with doing out of tree is for integrations like llmman, we have a release on a per commit basis for ROCM, CUDA, opencl, vulkan, etc. (it's a long list) at the per-commit level... So you can have the exact same version of llama.cpp/ggml across all the GPU types... And "llama-server" is the key binary other than the libs... Doing one out of tree is messy and prone to drift...

I'm suggesting (or ACKing your suggestion) to do in-tree. The APT stuff I mentioned above is for the NPU drivers & FW I think we're on the same page -- let's do WoS and Ubuntu ARM64 releases here.

FWIW I do keep an eye on the geniex ecosystem (heck I'm a contributor there) and I am also open to Qualcomm AI Engine Direct (qairt) backend in llmman as an alternate to llama.cpp . I actually don't have a qualcomm machine. Although strangely enough a detailed background with Qualcomm Automotive boards which are quite similar :)

@max-krasnyansky

Copy link
Copy Markdown
Member

FWIW I do keep an eye on the geniex ecosystem (heck I'm a contributor there) and I am also open to Qualcomm AI Engine Direct (qairt) backend in llmman as an alternate to llama.cpp . I actually don't have a qualcomm machine. Although strangely enough a detailed background with Qualcomm Automotive boards which are quite similar :)

Alternative to llama.cpp? Nooo -- llama.cpp is the BEST!
Jokes aside (though llama.cpp is the best), sure, GenieX would be a great source for things like that.

@ericcurtin

ericcurtin commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

FWIW I do keep an eye on the geniex ecosystem (heck I'm a contributor there) and I am also open to Qualcomm AI Engine Direct (qairt) backend in llmman as an alternate to llama.cpp . I actually don't have a qualcomm machine. Although strangely enough a detailed background with Qualcomm Automotive boards which are quite similar :)

Alternative to llama.cpp? Nooo -- llama.cpp is the BEST! Jokes aside (though llama.cpp is the best), sure, GenieX would be a great source for things like that.

Oh I prefer llama.cpp you don't have to sell me on that :)

But if people have special model types that need to run on an alternate runtime such as qairt in:

https://github.com/llmmanorg/llmman

that's fine too :)

Heck this is a special runtime about to be merged soon:

llmmanorg/llmman#547

llama.cpp don't have interest in becoming an omni/multi-modal type engine in terms of outputs at least:

#28540

@tfenster

Copy link
Copy Markdown
Author

Here's where I am now: This build https://github.com/tfenster/llama.cpp/actions/runs/35978701970 created a working llama-server.exe and after copying in the suggested geniex files, I could run it. The issue now is that somehow the NPU is not reporting memory. So if I just run it directly, it ignores the NPU

.\llama-server.exe `
>>   -m "C:\Users\tobia\AppData\Local\llmman\store\blobs\sha256\8334b850b7bd46238c16b0c550df2138f0889bf433809008cc17a8b05761863e"
0.00.042.230 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.042.558 W srv  llama_server: -----------------
0.00.042.562 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
0.00.042.563 W srv  llama_server: this can be a security risk (cross-origin attacks)
0.00.042.563 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
0.00.042.563 W srv  llama_server: -----------------
0.00.049.955 I srv    load_model: loading model 'C:\Users\tobia\AppData\Local\llmman\store\blobs\sha256\8334b850b7bd46238c16b0c550df2138f0889bf433809008cc17a8b05761863e'
0.00.529.303 W common_get_device_memory_data_impl: device HTP0 did not report memory; --fit will not use it
0.00.878.599 W common_get_device_memory_data_impl: device HTP0 did not report memory; --fit will not use it
0.01.233.552 W common_get_device_memory_data_impl: device HTP0 did not report memory; --fit will not use it
0.01.591.665 W common_get_device_memory_data_impl: device HTP0 did not report memory; --fit will not use it

With --fit off it starts to use the NPU

.\llama-server.exe `
>>   -m "C:\Users\tobia\AppData\Local\llmman\store\blobs\sha256\8334b850b7bd46238c16b0c550df2138f0889bf433809008cc17a8b05761863e" `
>>   -fit off
0.00.041.510 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.041.980 W srv  llama_server: -----------------
0.00.041.984 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
0.00.041.985 W srv  llama_server: this can be a security risk (cross-origin attacks)
0.00.041.985 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
0.00.041.986 W srv  llama_server: -----------------
0.00.047.759 I srv    load_model: loading model 'C:\Users\tobia\AppData\Local\llmman\store\blobs\sha256\8334b850b7bd46238c16b0c550df2138f0889bf433809008cc17a8b05761863e'
0.06.101.320 I cmn          init: llama threadpool init, n_threads = 12
0.08.703.575 I srv    load_model: initializing, n_slots = 4, n_ctx_slot = 65536, kv_unified = 'true'
0.08.708.700 I srv  llama_server: model loaded
0.08.708.706 I srv  llama_server: listening on http://127.0.0.1:8080
0.08.708.708 W srv  llama_server: NOTICE: server default port will be changed to :9931 in a future release
0.08.708.709 W srv  llama_server:         ref: https://github.com/ggml-org/llama.cpp/pull/26508
1.03.554.822 I slot get_availabl: id  3 | task -1 | selected slot by LRU, t_last = -1
1.03.555.127 I slot launch_slot_: id  3 | task 0 | processing task, is_child = 0
1.11.003.285 I slot print_timing: id  3 | task 0 | prompt eval time =    2462.59 ms /   792 tokens (    3.11 ms per token,   321.61 tokens per second)
1.11.003.289 I slot print_timing: id  3 | task 0 |        eval time =    4985.30 ms /    61 tokens (   83.09 ms per token,    12.04 tokens per second)
1.11.003.291 I slot print_timing: id  3 | task 0 |       total time =    7447.90 ms /   853 tokens
1.11.003.292 I slot print_timing: id  3 | task 0 |    graphs reused =         60
1.11.003.475 I slot      release: id  3 | task 0 | stop processing: n_tokens = 852, truncated = 0
image

Any idea on what might cause that memory issue and how to fix it?

@ericcurtin

ericcurtin commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Without reading the code, this seems like a bug:

-fit,  --fit [on|off]                   whether to adjust unset arguments to fit in device memory ('on' or
                                        'off', default: 'on')

I suspect it's looking at the wrong memory, if it's the hexagon backend, it should look at NPU memory (NPU being the device in this case), I suspect it's looking at the wrong memory (just a hypothesis)

@tfenster

Copy link
Copy Markdown
Author

Hm, listing the devices also shows no memory for the NPU. Which in a way is true as it only has shared memory, but the same is true for the GPU, which does show memory

.\llama-server.exe --list-devices
Available devices:
  GPUOpenCL: Qualcomm(R) Adreno(TM) X1-85 GPU (16162 MiB, 15138 MiB free)
  HTP0: Hexagon (0 MiB, 0 MiB free)

@taronaeo

Copy link
Copy Markdown
Member

The lack of memory reporting for NPU is "correct" because it is declared as an accelerator device type compared to a GPU (with dedicated VRAM).

I've tried proposing a Memory Aware API so that devices can explicitly declare if they have memory information to report, while being accurate to the device type they are, but it's currently stuck: #22949

@ericcurtin

ericcurtin commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Hm, listing the devices also shows no memory for the NPU. Which in a way is true as it only has shared memory, but the same is true for the GPU, which does show memory

.\llama-server.exe --list-devices
Available devices:
  GPUOpenCL: Qualcomm(R) Adreno(TM) X1-85 GPU (16162 MiB, 15138 MiB free)
  HTP0: Hexagon (0 MiB, 0 MiB free)

If we open a PR to get this from 0 to the correct value we might be lucky and everything might just work. I don't have a device though 😅

@max-krasnyansky

Copy link
Copy Markdown
Member

The lack of memory reporting for NPU is "correct" because it is declared as an accelerator device type compared to a GPU (with dedicated VRAM).

I've tried proposing a Memory Aware API so that devices can explicitly declare if they have memory information to report, while being accurate to the device type they are, but it's currently stuck: #22949

Ah. Yes. Sorry been meaning to follow up on this. Sorry for the delay.
Re-added close to the top of TODO.

@max-krasnyansky

Copy link
Copy Markdown
Member

Hm, listing the devices also shows no memory for the NPU. Which in a way is true as it only has shared memory, but the same is true for the GPU, which does show memory

.\llama-server.exe --list-devices
Available devices:
  GPUOpenCL: Qualcomm(R) Adreno(TM) X1-85 GPU (16162 MiB, 15138 MiB free)
  HTP0: Hexagon (0 MiB, 0 MiB free)

If we open a PR to get this from 0 to the correct value we might be lucky and everything might just work. I don't have a device though 😅

The correct value is all available RAM. There is no fixed limit.

@tfenster

Copy link
Copy Markdown
Author

Windows task manager claims it is 16G, compared to 32G overall on the device (see screenshot)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops improvements to build systems and github actions documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature Request: Formal release of Windows Arm Hexagon NPU builds

6 participants