Note
This fork adds native support for Kev System One decision models. A Kev GGUF loads like any other model, and llama-server exposes the TypeSafe-compatible POST /v1/systemone endpoint plus a /studio page for editing state and questions. No Python at inference time. Everything else is stock upstream llama.cpp.
Jump to Kev in 5 minutes for install, model download and a first run, or try it right now in your browser at espetro.github.io/llama.cpp (Kev-0.8B compiled to WebAssembly, nothing leaves the tab).
Kev answers typed questions about a piece of state and returns calibrated probabilities instead of text (prefill only, nothing generated). Useful for classification, routing, moderation, scoring and tool-call gating. Full reference: docs/kev.md.
Stock llama.cpp packages (brew, winget, conda-forge) do not include Kev support. Use one of these:
# Linux x64 - latest pre-built release (see the releases page for macOS/Windows/arm64/Vulkan/CUDA/SYCL assets)
TAG=$(curl -s https://api.github.com/repos/espetro/llama.cpp/releases | grep -m1 '"tag_name"' | cut -d'"' -f4)
curl -L -o llama-kev.tar.gz https://github.com/espetro/llama.cpp/releases/download/$TAG/llama-$TAG-bin-ubuntu-x64.tar.gz
tar xf llama-kev.tar.gz && export PATH="$PWD/llama-$TAG:$PATH"# any platform - mise; kev releases are pre-releases, so prerelease=true is required
mise use -g "github:espetro/llama.cpp[asset_pattern=llama-*-bin-ubuntu-x64.tar.gz,prerelease=true]@latest"
# macOS arm64: asset_pattern=llama-*-bin-macos-arm64.tar.gz Windows: llama-*-bin-win-cpu-x64.zip
# mise hides releases younger than 24 h; pass an explicit @kev-<tag> to take a fresh one# from source
git clone -b kev https://github.com/espetro/llama.cpp && cd llama.cpp
cmake -B build && cmake --build build -j --target llama-server llama-decide llama-quantizeReleased assets: macOS arm64/x64, Linux x64/arm64 (CPU, Vulkan, CUDA 12.8 and 13.4), SYCL, OpenVINO, Snapdragon, Windows (CPU, Vulkan, CUDA, SYCL, OpenCL), Android, iOS xcframework. The CPU builds are the ones tested with Kev so far. On macOS, a tarball downloaded with a browser needs xattr -d com.apple.quarantine. All kev-* releases are on the releases page.
Pre-packed GGUFs of Kev v1.0 (pointer head baked in) are on Hugging Face — -hf downloads and loads in one step. Each repo carries two files: q8_0 for inference and f16 as the re-quantization source / zero-drift reference:
llama-server -hf espetro/kev-0.8b-gguf:Q8_0 # also kev-4b-gguf and kev-9b-ggufOr fetch the file yourself:
pip install -U huggingface_hub
hf download espetro/kev-0.8b-gguf kev-0.8b-q8_0.gguf --local-dir kev-0.8b
llama-server -m kev-0.8b/kev-0.8b-q8_0.ggufThe raw v0 gojev bundles (F16 backbone + separate head.json, no packing) remain at taigrr/kev-0.8b-gguf (also kev-4b-gguf, kev-9b-gguf) — run them with -m model-f16.gguf --kev-head head.json. The v1.0 sources are the jaredpalmer/kev-* adapter + head.pt repos; see "pack a checkpoint yourself" below for the pipeline.
For the smallest downloads, each size also has a demo quant (q4_k_m + importance-matrix calibration, 0-2 answer flips vs F16 on a 17-question probe): espetro/kev-0.8b-demo-gguf (466 MB), espetro/kev-4b-demo-gguf (2.5 GB), espetro/kev-9b-demo-gguf (5.6 GB). Great for a first look; ship the q8_0 in production.
If you started the server with -hf above it is already running; otherwise:
llama-server -m kev-0.8b/kev-0.8b-q8_0.ggufcurl localhost:8080/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"refund": {"type": "noul", "instructions": "Should we refund?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Refunds, exchanges", "shipping": "Delays", "billing": "Charges"}}
}
}'{"answers":{"refund":{"type":"noul","noul":0.4431},
"department":{"type":"choice","choice":"shipping","confidence":0.3088,
"probabilities":{"returns":0.2019,"shipping":0.5392,"billing":0.2589}}},
"latency_ms":344.2}Same request from the CLI, without a server:
llama-decide -hf espetro/kev-0.8b-gguf:Q8_0 --json request.jsonOpen http://localhost:8080/studio while the server runs: edit the state, add noul / choice / score questions, re-run on every change, and read per-option probability bars with confidence labels (automate / review / escalate). The page also shows the matching curl and Python snippets for the request you built.
The espetro/kev-*-gguf repos are produced exactly like this — merge the v1.0 LoRA adapter into the pinned Qwen3.5 base, convert to GGUF, fold head.pt in, then quantize:
hf download jaredpalmer/kev-0.8b --local-dir kev-0.8b-src
hf download Qwen/Qwen3.5-0.8B-Base --revision dc7cdfe2 --local-dir kev-0.8b-base
python tools/kev/kev_v10_merge.py kev-0.8b-src kev-0.8b-base
python convert_hf_to_gguf.py kev-0.8b-base --outtype f16 --outfile model-f16.gguf --no-mtp
python tools/kev/kev_head.py kev-0.8b-src/head.pt head.json
python tools/kev/kev_pack.py --gguf model-f16.gguf --head head.json --manifest manifest.json --out kev-0.8b-f16.gguf
llama-quantize kev-0.8b-f16.gguf kev-0.8b-q8_0.gguf q8_0
llama-server -m kev-0.8b-q8_0.ggufThe merge applies the adapter in fp32 (scale = alpha/r = 2.0) because peft cannot resolve the composite multimodal checkpoint layout itself; --no-mtp skips the ~90 MB of unused NextN draft tensors.
The head tensors stay F32 through quantization. Measured against Kev's Python reference on the 0.8B fixtures: max probability delta 0.0005 (F16) / 0.011 (Q8_0), 0 argmax flips.
The 0.8B model also runs client side, compiled with emscripten (tools/kev/wasm/build.sh, about 1 GB live in the tab, 2.1 s for 3 questions with 4 threads). examples/kev-web is the static page for it: the runtime and the GGUF are fetched only when you press Load, then cached by the browser. A hosted copy runs at espetro.github.io/llama.cpp. See docs/kev.md.
- espetro.github.io/llama.cpp - this fork compiled to WebAssembly, Kev-0.8B runs fully in the tab. Loads the 466 MB demo quant by default (2 near-tie flips vs F16 on the probe set, max drift ~0.2);
?model=selects the q8_0. - huggingface.co/spaces/jaredpalmer/kev - Kev's authors' hosted Gradio demo on free ZeroGPU with the original Python stack (0.8B and 4B, ready-made examples). Good for a first look at 4B; this fork is the path for running Kev yourself.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

