RAdeon Narrative llaMA.cpp — a Radeon-optimized llama.cpp fork for roleplay and narrative inference.
This is a hobby project maintained by a single developer. Issues are welcome. Pull requests are accepted on the terms of CONTRIBUTING.md. See Project policy.
RANMA.cpp is llama.cpp with a small, curated set of changes aimed at one workload: interactive roleplay and narrative generation on consumer AMD Radeon GPUs, running on Windows.
That workload looks different from the general-purpose serving that upstream optimizes for:
- One user (or a handful of slots), not dozens of concurrent requests.
- Long, growing multi-turn contexts where most of the prompt was already seen last turn.
- Decode throughput at long context matters more than batch prefill throughput.
- Sessions stay open for hours with idle gaps between turns.
- Large MoE models that do not fit in VRAM, or even in RAM.
Some of the optimizations that pay off here trade away multi-request throughput or generality, so they are not appropriate for upstream. Others are general improvements and may be submitted upstream. This fork is where both kinds live together in a tested state.
It can also serve one person's own development work in a single session, such as a coding assistant on the same machine.
It is not a general replacement for llama.cpp. The expert cache does not change its placement while a request is processed: it installs a new plan only when every slot is idle, so when several sessions send requests in turn it rarely gets the chance. If you serve many users, use upstream.
- Expert cache - for a MoE model whose routed experts do not fit in VRAM, the server profiles which
experts the router selects and keeps the most valuable set in a VRAM budget (
--expert-l1-mib); the other experts are read from host memory over PCIe, and with--expert-l2-mibfrom the GGUF file when they do not fit in RAM either. Needs host-direct and--load-mode none. docs/ranma/expert-cache.md, docs/ranma/expert-cache-l2.md- Give the slot count with
-np(-np 1for one user). Without it the server runs four slots: the state kept per slot grows with them (for Qwen3.8-Flash-Next at a 256K context the recurrent state and the compute buffers grow by several GiB), which can overflow VRAM, and a finite host tier is refused withfinite L2 needs positive batch/parallel bounds.
- Give the slot count with
- Smart MTP draft length (
--spec-smart) - with--spec-type draft-mtp, llama-server chooses the draft length at every step from the measured verification time per width and a calibrated acceptance of the draft probabilities, without per-model thresholds. On by default withdraft-mtp;--spec-smart-store PATHkeeps the estimates across restarts. docs/ranma/spec-smart.md - EXL3 weights - ExLlamaV3 EXL3 quantizations run from GGUF files, on the CPU and on HIP (GEMV for
decode, WMMA GEMM on RDNA4 for prompts), with host-direct expert banks and the expert cache. They need
GGUF files in this fork's format: an ordinary GGUF or an EXL3 safetensors checkpoint is not read
directly. Take the converted files from Hugging Face
(Qwen3.8-Flash-Next,
DeepSeek-V4-Flash-0731).
For another EXL3 model, convert its EXL3 checkpoint with
convert_hf_to_gguf.py. Several architectures have been converted and run (checked models), but not every model is guaranteed to work yet. These files use tensor types of this fork and do not open in upstream llama.cpp, LM Studio, Ollama or other llama.cpp-based tools. docs/ranma/exl3.md
Every other change, one line each with its page: docs/ranma/README.md.
The release ranma_20261009 is based on upstream 462524043. Its release table puts three models side by side on six systems:
Qwen3.8-Flash-Next, DeepSeek V4 Flash and GLM-5.3 Flash, each in the EXL3 conversion and in a general GGUF, on the R9700
and the RX 9070 XT emulation with 128, 64 and 32 GB. The 64 and 32 GB rows take the physical memory away, not only the
host-tier budget. The table shows llama-bench PP512 / TG128 t/s at depth 0, warm. Upstream runs the general GGUF with
-ncmoe. Its experts outside VRAM do not fit in the host memory of the 64 and 32 GB systems, so it cannot load there
(record, with the 0 / 8K / 32K rows, the server scenarios with MTP and
memory).
| model | build | R9700 128 GB | R9700 64 GB | R9700 32 GB | 9070 XT 128 GB | 9070 XT 64 GB | 9070 XT 32 GB |
|---|---|---|---|---|---|---|---|
| Qwen3.8-Flash-Next EXL3 4.05 bpw | ranma_20261009 | 1125 / 50.1 | 1189 / 50.4 | 1223 / 44.0 | 775 / 39.9 | 790 / 37.8 | 341 / 29.5 |
| Qwen3.8-Flash-Next UD-Q4_K_XL | ranma_20261009 | 1066 / 48.9 | 1067 / 46.0 | 797 / 39.5 | 623 / 37.5 | 651 / 33.4 | 167 / 18.5 |
| upstream | 541 / 15.6 | not testable | not testable | 471 / 12.2 | not testable | not testable | |
| DeepSeek V4 Flash EXL3 3.04 bpw | ranma_20261009 | 363 / 33.1 | 291 / 29.0 | 134 / 17.9 | 294 / 24.3 | 167 / 19.2 | 98 / 10.2 |
| DeepSeek V4 Flash UD-IQ3_XXS | ranma_20261009 | 344 / 35.0 | 282 / 31.1 | 125 / 18.0 | 285 / 25.7 | 163 / 20.5 | 95 / 9.7 |
| upstream | 265 / 10.1 | not testable | not testable | 232 / 9.0 | not testable | not testable | |
| GLM-5.3 Flash EXL3 3.05 bpw | ranma_20261009 | 309 / 25.0 | 168 / 18.5 | 101 / 10.3 | 258 / 17.2 | 120 / 11.7 | 79 / 6.5 |
| GLM-5.3 Flash UD-IQ3_XXS | ranma_20261009 | 272 / 26.1 | 159 / 19.3 | 88 / 10.5 | 234 / 18.4 | 109 / 12.3 | 70 / 6.3 |
| upstream | 210 / 8.9 | not testable | not testable | failed (ROCm error) | not testable | not testable |
The earlier records keep the release each number was measured at: Qwen3.8-Flash-Next in
EXL3 against UD-Q4_K_XL is from ranma_20261005 (upstream 42d958167), the MTP draft length from ranma_20261001, the
ratios against upstream from ranma_20260928 and the full llama-bench set from ranma_20260922. Everything was
measured on the Radeon AI PRO R9700;
docs/ranma/benchmarks has every number with its release, and
the method.
EXL3 against UD-Q4_K_XL, measured for ranma_20261005: Qwen3.8-Flash-Next EXL3 3.05 bpw and 4.05 bpw on two systems,
the R9700 with 128 GB (exclusive expert cache of 20480 MiB, unlimited host tier) and the RX 9070 XT emulation with
64 GB (3072 MiB, host tier of 40960 MiB). llama-bench TG128 at depth 0, warm with the prefill swap, and server
decode in the five scenarios with MTP and --spec-smart, warm, against UD-Q4_K_XL on the same system
(record, with the 64 GB rows, the prompt rates, cold and memory).
| model | system | TG128 @0 t/s | TG128 @0 against UD-Q4_K_XL | server decode with MTP against UD-Q4_K_XL |
|---|---|---|---|---|
| EXL3 3.05 bpw | R9700, 128 GB | 53.5 | +8.6 % | +17.2 to +23.1 % |
| EXL3 4.05 bpw | R9700, 128 GB | 49.6 | +0.7 % | +3.9 to +10.0 % |
| UD-Q4_K_XL | R9700, 128 GB | 49.3 | ||
| EXL3 3.05 bpw | RX 9070 XT emulation, 64 GB | 43.2 | +20.5 % | +50.6 to +85.9 % |
| EXL3 4.05 bpw | RX 9070 XT emulation, 64 GB | 37.0 | +3.3 % | +15.4 to +28.1 % |
| UD-Q4_K_XL | RX 9070 XT emulation, 64 GB | 35.8 |
Decode against upstream, measured at ranma_20260928 against its upstream base: the server scenarios that the project
runs (roleplay in four languages and coding, one slot, greedy decoding), release warm against upstream. Upstream keeps
the routed experts of the first 35 layers in host memory (-ncmoe 35; -ncmoe 36 with the DeepSeek MTP head on the
GPU); the release runs the exclusive expert cache with a VRAM budget of 20480 MiB (18432 MiB with the DeepSeek MTP
head outside the cache) and an unlimited host tier
(record, with the t/s, the cold and 64 GB rows and the prompt times).
| model | English roleplay | coding | Korean roleplay | Japanese roleplay | Chinese roleplay |
|---|---|---|---|---|---|
| DeepSeek V4 Flash UD-IQ3_XXS, MTP n1 | x2.79 | x2.52 | x2.89 | x2.95 | x2.92 |
| DeepSeek V4 Flash UD-IQ3_XXS, no MTP | x2.97 | x2.70 | x3.07 | x3.11 | x3.09 |
| Qwen3.8-Flash-Next UD-Q4_K_XL | x2.78 | x2.57 | x2.78 | x2.79 | x2.79 |
MTP draft length, measured for ranma_20261001: decode with no MTP, with the single-argument values tuned on this
machine for these two models (Qwen3.8-Flash-Next --spec-draft-n-max 2 --spec-draft-p-min 0, DeepSeek V4 Flash
--spec-draft-n-max 1 --spec-draft-p-min 0), and with --spec-smart, the default of draft-mtp now, from a cold
start and from a warm store. Six scenarios per model (coding, roleplay in four languages, a mixed scenario), the MTP
head on the GPU in the joint expert cache (record).
| model | English roleplay t/s: no MTP / single / smart cold / smart warm | single against no MTP | smart cold against single | smart warm against single |
|---|---|---|---|---|
| Qwen3.8-Flash-Next UD-Q4_K_XL | 46.10 / 57.13 / 56.99 / 57.54 | +19.7 to +29.5 % | -4.3 to +0.6 % | -1.9 to +2.6 % |
| DeepSeek V4 Flash UD-IQ3_XXS | 31.17 / 35.00 / 34.07 / 34.55 | +9.8 to +14.6 % | -3.8 to +1.0 % | -3.2 to +1.5 % |
Smart needs no per-model values and stays within a few percent of the tuned values; on DeepSeek V4 Flash it is 1.3 to
3.8 % slower in every roleplay scenario. --no-spec-smart or the thresholds give the single-argument rule back.
llama-bench PP512 / TG128 in t/s at depth 0, 8192 and 65536, measured at ranma_20260922 against its upstream base on
two systems: the R9700 (32 GiB) with 128 GiB of host memory, and an RX 9070 XT (16 GiB) with 64 GiB, emulated on the
R9700 by the placement and the host-tier budget. Upstream runs -ncmoe 35; ranma_20260922 runs the exclusive expert
cache (20480 MiB on the R9700, 3072 MiB for Qwen and 4096 MiB for DeepSeek on the emulated RX 9070 XT) with prefill
swap where that is faster. The last column is the peak VRAM / host commit of the process in GiB
(record, with the 32 GB rows and Gemma 4 31B).
| model | system | @0 | @8192 | @65536 | VRAM / host |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next UD-Q4_K_XL | upstream, R9700, 128 GB | 335 / 15.9 | 316 / 15.2 | 307 / 13.2 | 29.9 / 80.7 |
| ranma_20260922, R9700, 128 GB | 1014 / 49.7 | 949 / 47.9 | 692 / 44.8 | 29.1 / 78.2 | |
| ranma_20260922, RX 9070 XT emulation, 64 GB | 625 / 36.7 | 596 / 38.2 | 480 / 35.2 | 12.0 / 48.8 | |
| DeepSeek V4 Flash UD-IQ3_XXS | upstream, R9700, 128 GB | 248 / 9.8 | 209 / 9.7 | (measured to 8192) | 27.8 / 98.7 |
| ranma_20260922, R9700, 128 GB | 310 / 31.8 | 255 / 30.9 | 108 / 28.1 | 27.9 / 90.3 | |
| ranma_20260922, RX 9070 XT emulation, 64 GB | 129 / 18.8 | 128 / 20.3 | 78 / 18.9 | 11.8 / 49.0 |
Planned, not done yet. Nothing here is promised or scheduled.
- DeepSeek V4 Flash: try further performance work. It may not pay off.
- DeepSeek V4.1 Flash: take the upstream pull request for this model while it is not merged yet, then add EXL3 support on top. If upstream merges a different implementation, this fork follows upstream, and GGUF files made for the earlier one, EXL3 GGUF files in particular, may stop loading and need to be converted again.
- Experimental RDNA3 and NVIDIA (CUDA) builds: open the paths that do not depend on RDNA4 instructions and check only that they compile. Whether they work is not guaranteed; reports and pull requests are welcome.
| Primary target | |
|---|---|
| OS | Windows 11 |
| GPU | AMD Radeon RDNA4, gfx1201; developed and measured on a Radeon AI PRO R9700 |
| Backend | HIP / ROCm (GGML_HIP=ON) |
| Use case | llama-server driving a roleplay client, single or few slots |
The maintainer develops and tests every change only on this environment. Some changes rely on Windows APIs or on gfx1201-specific behavior; on other platforms they may have no effect or may misbehave. Other GPUs, backends and operating systems may work, but nothing is guaranteed there: releases are made without testing them. Reports and pull requests are welcome (CONTRIBUTING.md).
| feature | RDNA4 (gfx1201) | RDNA3 (gfx1100/1101/1102) |
|---|---|---|
| Host-direct MoE weights | ✔ | ◇ ranma_20261005 |
| Expert cache, VRAM tier | ✔ | ◇ ranma_20261005 |
| Expert cache, finite host tier with file backing | ✔ | ◇ ranma_20261005 |
| Smart MTP draft length | ✔ | ? |
| Qwen sparse attention, selected cells only | ✔ | ? |
| DeepSeek V4 selected-cell attention | ✔ | ? |
| DeepSeek V4 graph ops (hyper-connection coefficients, KV compressor) | ✔ | ? |
| GLM-5.3 Flash decode kernels (hyper-connection mixes, KDA conv and state, MLA decode attention) | ✔ | ? |
| EXL3 decode (GEMV) | ✔ | ○ |
| EXL3 prompt processing (WMMA GEMM) | ✔ | ✕ (prompts run on the GEMV ○) |
| Qwen3.8 hyper-connection fusion | ✔ | ○ |
✔ tested by the maintainer at the current release. ◇ reported by a user who tested it on their own machine at the release named; not tested by the maintainer. ○ the HIP device code compiles for gfx1100 (the EXL3 kernels also for gfx1101 and gfx1102); not run. ✕ not implemented for that GPU yet; it takes a slower general path. ? not confirmed. Optimizations that only make sense with RDNA4 instructions or RDNA4 tuning, such as the RDNA4 matmul selection and the shared expert unit, are not listed.
- Follow upstream closely. The
masterbranch is the unmodified upstreamggml-org/llama.cppcommit that the latest release is based on. Theranmabranch is the latest release:masterplus the curated patch set, rebuilt for every release on a recent upstream. That means the history ofranmais rewritten regularly: commit hashes change from release to release, and a commit whose reason has disappeared (upstream landed an equivalent, or the feature stopped paying for itself) is dropped rather than carried along. Do not build long-lived work onranmahashes. Every release is a tag (ranma_YYYYMMDD) that never moves and keeps the exact commits that thedocs/ranma/pages cite. See docs/ranma/releases.md. - Curate, do not accumulate. Every change must be explainable in a few sentences, measurable on the target workload, and small enough to carry across upstream updates. Features that stop paying for themselves are removed.
- Be honest about trade-offs. Each feature documents what it costs: concurrency, VRAM, generality, or complexity.
- Upstream what belongs upstream. Changes that are general improvements may be submitted to llama.cpp rather than kept only here.
RANMA.cpp builds exactly like upstream. For the primary target, follow the HIP section of docs/build.md on Windows: install the ROCm SDK, open an x64 Native Tools Command Prompt for VS, then:
set PATH=%HIP_PATH%\bin;%PATH%
cmake -S . -B build -G Ninja -DGGML_HIP=ON -DGPU_TARGETS=gfx1201 ^
-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++ -DCMAKE_BUILD_TYPE=Release
cmake --build buildPrebuilt Windows binaries may be published on the releases page once the first curated features land.
- Issues: welcome. Bug reports on the target environment and reproducible performance regressions are most useful. Feature requests are read but not promised.
- Pull requests: accepted on conditions, against
ranma_upstream. There is no promise of review, merge or response time; the pull request carries its own verification and test environment; an accepted change may be squashed, reordered, rewritten or dropped in a later release. See CONTRIBUTING.md. - Sponsorship: GitHub Sponsors. It does not buy review, features, or support.
- Upstream contributions: selected features from this fork may be submitted to llama.cpp under the upstream contribution rules.
| Ref | Purpose |
|---|---|
ranma |
The latest release. This is the default branch. |
ranma_YYYYMMDD |
A release, as a tag that never moves, so the revisions cited in docs/ranma/ still resolve after ranma has moved on. The three latest releases are also kept as branches of the same name. |
master |
The unmodified upstream ggml-org/llama.cpp master commit that the latest release is based on. Never commit here. |
ranma_upstream |
The latest release rebased onto a newer upstream commit with its overlaps with upstream resolved: the baseline of further work, the starting point of the next release and the target of pull requests. Not benchmarked again. |
features/* |
Changes prepared for upstream pull requests. |
How a release is made and how far each ref is verified: docs/ranma/releases.md. What each release changed: docs/ranma/changelog.md.
Upstream is tracked as the git remote upstream. When the series is rebased, upstream
changes to this README and other fork-owned files are reviewed by hand and applied or
dropped as appropriate.
Everything not covered above is unchanged from llama.cpp:
- Build guide · Install · Docker
- Supported models · Multimodal
- llama-server · Function calling
- Speculative decoding · Multi-GPU
RANMA.cpp is distributed under the MIT License, the same license as llama.cpp.
All of the heavy lifting is the work of ggml-org/llama.cpp and its contributors. This fork exists only to carry a narrow set of workload-specific changes on top of it.
The EXL3 format support follows ExLlamaV3 (MIT License, turboderp); see NOTICE.