Skip to content

Where can I download wheel for Cuda 12.8? Trying to install llama.cpp to use with ComfyUI custom nodes. #2068

Description

@overallbit

Trying to install this for comfyui, it install with pip install llama-cpp-python and works but it is loading the gguf models to System RAM, I'm trying to make it load on GPU VRAM.

Asked chatgpt for help and it want me to download prebuilt wheel for my Python version.... and run this on comfyui venv: pip install path\to\llama_cpp_python-cu128.whl

I read from some user that to install llama-cpp-python with CUDA support need to use: CMAKE_ARGS="-DGGML_CUDA=ON -DLLAMA_LLAVA=OFF" pip install llama-cpp-python

but it gives me error: 'CMAKE_ARGS' is not recognized as an internal or external command,
operable program or batch file.

RTX 2060
Cuda Toolkit 12.8
Python 3.12

Activity

  1. sergey21000 commented on Sep 17, 2025

    @sergey21000
    Contributor

    but it gives me error: 'CMAKE_ARGS' is not recognized as an internal or external command, operable program or batch file.

    Correct command for PowerShell

    $env:CMAKE_ARGS="-DGGML_CUDA=on -DLLAMA_LLAVA=OFF"
    pip install llama-cpp-python

    For CMD (run each command separately)

    set CMAKE_ARGS=-DGGML_CUDA=on -DLLAMA_LLAVA=OFF
    pip install llama-cpp-python

    To verify the installation

    python -c "import llama_cpp; print(llama_cpp.llama_print_system_info())"
    
    # Example output for CPU support:
    # CPU : SSE3 = 1 | SSSE3 = 1 | ...
    
    # Example output for CUDA support:
    # ggml_cuda_init: found 1 CUDA devices: ...
    # CUDA : ARCHS = 700,750,800,860,890,900 | FORCE_MMQ = 1 ...
  2. overallbit commented on Sep 17, 2025

    @overallbit
    Author

    but it gives me error: 'CMAKE_ARGS' is not recognized as an internal or external command, operable program or batch file.

    Correct command for PowerShell

    $env:CMAKE_ARGS="-DGGML_CUDA=on -DLLAMA_LLAVA=OFF"
    pip install llama-cpp-python
    For CMD (run each command separately)

    set CMAKE_ARGS=-DGGML_CUDA=on -DLLAMA_LLAVA=OFF
    pip install llama-cpp-python
    To verify the installation

    python -c "import llama_cpp; print(llama_cpp.llama_print_system_info())"

    Example output for CPU support:

    CPU : SSE3 = 1 | SSSE3 = 1 | ...

    Example output for CUDA support:

    ggml_cuda_init: found 1 CUDA devices: ...

    CUDA : ARCHS = 700,750,800,860,890,900 | FORCE_MMQ = 1 ...

    Thanks but I solved it by using another llama-cpp fork.

  3. rudolphos commented on Oct 12, 2025

    @rudolphos
    $env:CMAKE_ARGS="-DGGML_CUDA=on -DLLAMA_LLAVA=OFF"
    pip install llama-cpp-python
    

    It gives an error because pip install would require compiling and - tons of dependencies. It doesn't work like that simply out of box on Win unfortunately..

    PS C:\WINDOWS\system32> $env:CMAKE_ARGS="-DGGML_CUDA=on -DLLAMA_LLAVA=OFF"
    PS C:\WINDOWS\system32> pip install llama-cpp-python
    Collecting llama-cpp-python
      Using cached llama_cpp_python-0.3.16.tar.gz (50.7 MB)
      Installing build dependencies ... done
      Getting requirements to build wheel ... done
      Installing backend dependencies ... done
      Preparing metadata (pyproject.toml) ... done
    Requirement already satisfied: typing-extensions>=4.5.0 in c:\users\user\appdata\local\programs\python\python312\lib\site-packages (from llama-cpp-python) (4.15.0)
    Requirement already satisfied: numpy>=1.20.0 in c:\users\user\appdata\local\programs\python\python312\lib\site-packages (from llama-cpp-python) (2.3.3)
    Collecting diskcache>=5.6.1 (from llama-cpp-python)
      Using cached diskcache-5.6.3-py3-none-any.whl.metadata (20 kB)
    Requirement already satisfied: jinja2>=2.11.3 in c:\users\user\appdata\local\programs\python\python312\lib\site-packages (from llama-cpp-python) (3.1.6)
    Requirement already satisfied: MarkupSafe>=2.0 in c:\users\user\appdata\local\programs\python\python312\lib\site-packages (from jinja2>=2.11.3->llama-cpp-python) (3.0.3)
    Using cached diskcache-5.6.3-py3-none-any.whl (45 kB)
    Building wheels for collected packages: llama-cpp-python
      Building wheel for llama-cpp-python (pyproject.toml) ... error
      error: subprocess-exited-with-error
    
      × Building wheel for llama-cpp-python (pyproject.toml) did not run successfully.
      │ exit code: 1
      ╰─> [20 lines of output]
          *** scikit-build-core 0.11.6 using CMake 4.1.0 (wheel)
          *** Configuring CMake...
          2025-10-13 02:33:22,897 - scikit_build_core - WARNING - Can't find a Python library, got libdir=None, ldlibrary=None, multiarch=None, masd=None
          loading initial cache file C:\Users\User\AppData\Local\Temp\tmp3b9yn4wc\build\CMakeInit.txt
          -- Building for: NMake Makefiles
          CMake Error at CMakeLists.txt:3 (project):
            Running
    
             'nmake' '-?'
    
            failed with:
    
             no such file or directory
    
    
          CMake Error: CMAKE_C_COMPILER not set, after EnableLanguage
          CMake Error: CMAKE_CXX_COMPILER not set, after EnableLanguage
          -- Configuring incomplete, errors occurred!
    
          *** CMake configuration failed
          [end of output]
    
      note: This error originates from a subprocess, and is likely not a problem with pip.
      ERROR: Failed building wheel for llama-cpp-python
    Failed to build llama-cpp-python
    error: failed-wheel-build-for-install
    
    × Failed to build installable wheels for some pyproject.toml based projects
    ╰─> llama-cpp-python
    

    RTX 4050
    CUDA Toolkit 12.8
    Python 3.12.10

  4. rudolphos commented on Oct 12, 2025

    @rudolphos

    @overallbit

    Thanks but I solved it by using another llama-cpp fork.

    Can you give a link?

    nvm, found it:
    https://github.com/boneylizard/llama-cpp-python-cu128-gemma3/releases

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions