Skip to content
 
 

Latest commit

 

History

11,141 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama.cpp

Note

This fork adds native support for Kev System One decision models. A Kev GGUF loads like any other model, and llama-server exposes the TypeSafe-compatible POST /v1/systemone endpoint plus a /studio page for editing state and questions. No Python at inference time. Everything else is stock upstream llama.cpp.

Jump to Kev in 5 minutes for install, model download and a first run, or try it right now in your browser at espetro.github.io/llama.cpp (Kev-0.8B compiled to WebAssembly, nothing leaves the tab).

Kev in 5 minutes

Kev answers typed questions about a piece of state and returns calibrated probabilities instead of text (prefill only, nothing generated). Useful for classification, routing, moderation, scoring and tool-call gating. Full reference: docs/kev.md.

1. Install this fork

Stock llama.cpp packages (brew, winget, conda-forge) do not include Kev support. Use one of these:

# Linux x64 - latest pre-built release (see the releases page for macOS/Windows/arm64/Vulkan/CUDA/SYCL assets)
TAG=$(curl -s https://api.github.com/repos/espetro/llama.cpp/releases | grep -m1 '"tag_name"' | cut -d'"' -f4)
curl -L -o llama-kev.tar.gz https://github.com/espetro/llama.cpp/releases/download/$TAG/llama-$TAG-bin-ubuntu-x64.tar.gz
tar xf llama-kev.tar.gz && export PATH="$PWD/llama-$TAG:$PATH"
# any platform - mise; kev releases are pre-releases, so prerelease=true is required
mise use -g "github:espetro/llama.cpp[asset_pattern=llama-*-bin-ubuntu-x64.tar.gz,prerelease=true]@latest"
# macOS arm64: asset_pattern=llama-*-bin-macos-arm64.tar.gz    Windows: llama-*-bin-win-cpu-x64.zip
# mise hides releases younger than 24 h; pass an explicit @kev-<tag> to take a fresh one
# from source
git clone -b kev https://github.com/espetro/llama.cpp && cd llama.cpp
cmake -B build && cmake --build build -j --target llama-server llama-decide llama-quantize

Released assets: macOS arm64/x64, Linux x64/arm64 (CPU, Vulkan, CUDA 12.8 and 13.4), SYCL, OpenVINO, Snapdragon, Windows (CPU, Vulkan, CUDA, SYCL, OpenCL), Android, iOS xcframework. The CPU builds are the ones tested with Kev so far. On macOS, a tarball downloaded with a browser needs xattr -d com.apple.quarantine. All kev-* releases are on the releases page.

2. Get a model

Pre-packed GGUFs of Kev v1.0 (pointer head baked in) are on Hugging Face — -hf downloads and loads in one step. Each repo carries two files: q8_0 for inference and f16 as the re-quantization source / zero-drift reference:

llama-server -hf espetro/kev-0.8b-gguf:Q8_0   # also kev-4b-gguf and kev-9b-gguf

Or fetch the file yourself:

pip install -U huggingface_hub
hf download espetro/kev-0.8b-gguf kev-0.8b-q8_0.gguf --local-dir kev-0.8b
llama-server -m kev-0.8b/kev-0.8b-q8_0.gguf

The raw v0 gojev bundles (F16 backbone + separate head.json, no packing) remain at taigrr/kev-0.8b-gguf (also kev-4b-gguf, kev-9b-gguf) — run them with -m model-f16.gguf --kev-head head.json. The v1.0 sources are the jaredpalmer/kev-* adapter + head.pt repos; see "pack a checkpoint yourself" below for the pipeline.

For the smallest downloads, each size also has a demo quant (q4_k_m + importance-matrix calibration, 0-2 answer flips vs F16 on a 17-question probe): espetro/kev-0.8b-demo-gguf (466 MB), espetro/kev-4b-demo-gguf (2.5 GB), espetro/kev-9b-demo-gguf (5.6 GB). Great for a first look; ship the q8_0 in production.

3. Run it

If you started the server with -hf above it is already running; otherwise:

llama-server -m kev-0.8b/kev-0.8b-q8_0.gguf
curl localhost:8080/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
  "questions": {
    "refund":     {"type": "noul",   "instructions": "Should we refund?"},
    "department": {"type": "choice", "instructions": "Which team should handle this?",
                   "criteria": {"returns": "Refunds, exchanges", "shipping": "Delays", "billing": "Charges"}}
  }
}'
{"answers":{"refund":{"type":"noul","noul":0.4431},
            "department":{"type":"choice","choice":"shipping","confidence":0.3088,
                          "probabilities":{"returns":0.2019,"shipping":0.5392,"billing":0.2589}}},
 "latency_ms":344.2}

Same request from the CLI, without a server:

llama-decide -hf espetro/kev-0.8b-gguf:Q8_0 --json request.json

4. Play with it in the browser

Open http://localhost:8080/studio while the server runs: edit the state, add noul / choice / score questions, re-run on every change, and read per-option probability bars with confidence labels (automate / review / escalate). The page also shows the matching curl and Python snippets for the request you built.

5. Optional: pack a checkpoint yourself

The espetro/kev-*-gguf repos are produced exactly like this — merge the v1.0 LoRA adapter into the pinned Qwen3.5 base, convert to GGUF, fold head.pt in, then quantize:

hf download jaredpalmer/kev-0.8b --local-dir kev-0.8b-src
hf download Qwen/Qwen3.5-0.8B-Base --revision dc7cdfe2 --local-dir kev-0.8b-base
python tools/kev/kev_v10_merge.py kev-0.8b-src kev-0.8b-base
python convert_hf_to_gguf.py kev-0.8b-base --outtype f16 --outfile model-f16.gguf --no-mtp
python tools/kev/kev_head.py kev-0.8b-src/head.pt head.json
python tools/kev/kev_pack.py --gguf model-f16.gguf --head head.json --manifest manifest.json --out kev-0.8b-f16.gguf
llama-quantize kev-0.8b-f16.gguf kev-0.8b-q8_0.gguf q8_0
llama-server -m kev-0.8b-q8_0.gguf

The merge applies the adapter in fp32 (scale = alpha/r = 2.0) because peft cannot resolve the composite multimodal checkpoint layout itself; --no-mtp skips the ~90 MB of unused NextN draft tensors.

The head tensors stay F32 through quantization. Measured against Kev's Python reference on the 0.8B fixtures: max probability delta 0.0005 (F16) / 0.011 (Q8_0), 0 argmax flips.

6. Optional: no server at all

The 0.8B model also runs client side, compiled with emscripten (tools/kev/wasm/build.sh, about 1 GB live in the tab, 2.1 s for 3 questions with 4 threads). examples/kev-web is the static page for it: the runtime and the GGUF are fetched only when you press Load, then cached by the browser. A hosted copy runs at espetro.github.io/llama.cpp. See docs/kev.md.

7. Try Kev without installing anything

  • espetro.github.io/llama.cpp - this fork compiled to WebAssembly, Kev-0.8B runs fully in the tab. Loads the 466 MB demo quant by default (2 near-tie flips vs F16 on the probe set, max drift ~0.2); ?model= selects the q8_0.
  • huggingface.co/spaces/jaredpalmer/kev - Kev's authors' hosted Gradio demo on free ZeroGPU with the original Python stack (0.8B and 4B, ready-made examples). Good for a first look at 4B; this fork is the path for running Kev yourself.

llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

LLM & Kev inference in C/C++

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages