Skip to content

ggml : add E4M3 (fp8) CPU quantization type - #25336

Open
RapidMark wants to merge 1 commit into
ggml-org:masterfrom
RapidMark:cloudhands/fp8-cpu
Open

RapidMark wants to merge 1 commit into
ggml-org:masterfrom
RapidMark:cloudhands/fp8-cpu

Conversation

@RapidMark

Copy link
Copy Markdown
Contributor

Overview

Add GGML_TYPE_E4M3 (OCP e4m3fn), a CPU fp8 quantization type, so native-fp8 weights load and run as fp8 instead of being dequantized to f16 (double the memory) or approximated as q8_0. Goal is fp8 for CPU and GPU; per the CPU-first rule this is the CPU side, and the GPU backends follow in separate PRs.

Block is q8_0-shaped (fp16 scale + 32 fp8 bytes per block), dotting against q8_0 activations.

Additional information

What's in here:

  • GGML_TYPE_E4M3 type, block struct, e4m3fn conversion + reference quant/dequant
  • CPU: scalar, AVX2, NEON vec_dot; repacked 8x8 gemv/gemm for faster prefill
  • gguf-py: type constant, numpy quant/dequant, endian convert
  • tests: test-backend-ops + test-quantize-fns

Test model (small, runnable): https://huggingface.co/RapidMark/Qwen2.5-1.5B-E4M3-GGUF
This is the exact e4m3 gguf the numbers below were measured on. The f16 baseline and q8_0 comparison are regenerated from the base model (Qwen/Qwen2.5-1.5B-Instruct) with the quantize command in Reproduction.

Evaluation. Qwen2.5-1.5B-Instruct, wikitext-2-raw test (full, 584 chunks, n_ctx 512), pure CPU. The similar-size comparison type is Q8_0 (both are one fp16 scale + 32 one-byte elements per block).

Perplexity and KL-divergence vs f16:

type size PPL vs f16 Mean KLD Median KLD Max KLD 99% KLD same-top-token
f16 2.88 GiB 10.134 - - - - - -
q8_0 1.53 GiB 10.168 +0.33% 0.00200 0.00158 0.213 0.00969 97.33%
e4m3 1.48 GiB 10.248 +1.12% 0.01146 0.00937 0.846 0.05113 93.64%

Pure-CPU throughput (llama-bench, 64 threads):

type pp512 (t/s) tg128 (t/s)
f16 528.3 19.2
q8_0 499.3 33.1
e4m3 253.2 34.6

e4m3 lands a little behind q8_0 on accuracy, which is expected: q8_0's linear int8 is near-lossless for these weights, and e4m3 spends half its bits on the exponent. On CPU its value is faithful fp8 at half the size of f16, not beating q8_0.

The repacked gemv/gemm is in this PR (not just vec_dot) because it's a 2.5x prefill speedup: on the same model and box, pp512 is 269.8 t/s with the repacked gemm vs 106.4 t/s on the plain vec_dot (tg128 33.1 vs 27.9).

Reproduction.

# f16 baseline from the base model, then the two quants (e4m3 also downloadable from the HF link above)
llama-quantize qwen2.5-1.5b-f16.gguf qwen2.5-1.5b-q8_0.gguf Q8_0
llama-quantize qwen2.5-1.5b-f16.gguf qwen2.5-1.5b-e4m3.gguf E4M3

# perplexity + KL-divergence vs f16
llama-perplexity -m qwen2.5-1.5b-f16.gguf  -f wiki.test.raw --kl-divergence-base kld-base.dat -t 64
llama-perplexity -m qwen2.5-1.5b-q8_0.gguf -f wiki.test.raw --kl-divergence --kl-divergence-base kld-base.dat -t 64
llama-perplexity -m qwen2.5-1.5b-e4m3.gguf -f wiki.test.raw --kl-divergence --kl-divergence-base kld-base.dat -t 64

# pure-CPU throughput
llama-bench -m qwen2.5-1.5b-f16.gguf -m qwen2.5-1.5b-q8_0.gguf -m qwen2.5-1.5b-e4m3.gguf -t 64 -p 512 -n 128

Validation. test-backend-ops -b CPU passes 16176/16176 (e4m3 MUL_MAT included). Verified on x86 (MSVC and Linux GCC), ARM NEON, and Apple (Metal correctly declines e4m3 and falls back to the CPU path). The gguf-py numpy class is bit-exact against the C library (test_quants.py).

Environment.

  • CPU: AMD Threadripper 3970X (Zen2), 64 threads
  • OS / compiler: Windows 11, MSVC (Release)
  • Cross-checked: Linux (GCC), aarch64 (NEON), Apple Silicon (Metal)
  • Commit: 136a83339

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. AI was used to understand the codebase, run the test suite across the available hardware, provide peer review, and offer suggestions. The design and code are mine and I can explain any line.

@RapidMark
RapidMark requested review from a team, CISC and ggerganov as code owners July 5, 2026 23:59
@github-actions github-actions Bot added testing Everything test related examples ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion labels Jul 6, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Jul 6, 2026

Copy link
Copy Markdown

Hi @RapidMark, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

  • Multiple backend changes in one PR: When adding support for a new model or feature, focus on CPU support only in the initial PR. Add support for other backends like CUDA in follow-up PRs. If you have a good reason to modify multiple backends in one PR, please explain it.

  • Large PR: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@RapidMark

Copy link
Copy Markdown
Contributor Author

Multiple backend changes in one PR:

Actually.. it's just CPU... but I needed to add m2 lines to ggml-metal-device.m to extend the existing GGML_TYPE_NVFP4 exclusions in ggml_metal_device_supports_op (for MUL_MAT/MUL_MAT_ID and GET_ROWS) to also cover E4M3.

@jeffbolznv

Copy link
Copy Markdown
Contributor

How did you arrive at using an fp16 scale factor? I wonder if it should be e8m0 instead to make this MXFP8. But I also worry about poor alignment with this block size, generally I would prefer to have at least 4B aligned blocks.

@RapidMark

Copy link
Copy Markdown
Contributor Author

I acknowledge the e8m0/mxfp4 are already in the tree, so fp16 was a deliberate choice, not for lack of e8m0.

We mirrored q8_0's block so the q8_0 activation path and the repack gemv/gemm would apply directly. ALso fp16 is a finer per-block scale than e8m0 (e8m0 is power-of-2 only, no mantissa). While fp16 has less exponent range, that doesn't matter (AFAIK) for normalized weight blocks.

We wanted to carry native fp8 as fp8. Native fp8 today is usually per-tensor/per-channel (MXFP8's per-block e8m0 is the other real variant), and what our block preserves faithfully is the e4e3 elements vs q8_0, the e4e3 elements carry exponent range that linear int8 doesn't, so they match the fp8 distributions the model was trained in; the goal is representing fp8 models, not beating q8_0.

As fro block alignment, 34 bytes isn;t 4-byte-aligned, but we felt e8m0 alone would make it worse; 1 byte scale + 32 bytes = 33 bytes. So switching to MXFP8 does not fix the alignment (by itself), which would need a block redesign and would make larger blocks like nvfp4's 64 with grouped scales (or extra padding).

Basically we chose q8_0 reuse + finer scale.

I am working on a consumer only path, which means I also need to supprot AMD as much as NVIDIA is supported (which is a major task, as you know), so everywhere I run across a descripency, I try to fill it. It also means I'm learning alot, but also means I'm coming from an untainted view.

Are you steering me toward MXFP8 (e8m0) and if so... what model/format should we target (MXFP8 vs per-tensor)?

also.. 34 bytes is not aligned, but e8m0 would make it 33 bytes so doing a 4-byte/16-byte layout would need a significant redesign. Do you have a preferred block size/padding?

@RapidMark
RapidMark force-pushed the cloudhands/fp8-cpu branch from 136a833 to ead2709 Compare July 6, 2026 15:19
@jeffbolznv

Copy link
Copy Markdown
Contributor

I understand the appeal of the f16 scale factor. I don't have a strong opinion on what we do here. If MXFP8 becomes widely used, then we'd obviously want to support that natively, and it also allows for hardware acceleration that an f16 scale won't.

Regarding alignment, yeah, I'd generally prefer to use a larger block size with multiple scales, packed similar to NVFP4. Some models may have layers that are not a multiple of the native block size, but maybe we could just pad with zeros?

@RapidMark

Copy link
Copy Markdown
Contributor Author

My focus is on consumer hardware, which is limited by memory. You run it to fit the model in 12-24 GB.

The goal is to make it fit, and still give the highest quality. I don't want to degrade the memory-forced consumer twice, once by dropping to fp8 and again with a coarse scale.

From my point of view e8m0's only real upside (today) is Blackwell's fused microscaling. RDNA4 and Ada, don't have that. They apply the block scale in software either way, so e8m0 buys no acceleration, just less quality.

Also, e4m3 is cross-vendor because the OCP formats get consumed on both AMD and NVIDIA fp8 units. Same stored weights accelerate on RDNA4 and on Ada/Blackwell. One format, both vendors.

I do agree MXFP8 is the right format for MX hardware (Blackwell, the CDNA4 parts) and most likely future consumer HW, but I also think we can add MXFP8 when consumer HW catches up.

Overall, datacenter's, while they can do FP8, will likely use FP16, unless they're looking for speed over quality.

This is just my point of view, but overall... I kinow we'll need MXFP8 in the future as well... so we can go either way... I just feel like we're giving up on existing hardware in hopes of future hardware.

Once we deicde on a direction... I can go back in and align blocks, groupped scales, padding, etc.

Lastly... I beleive, the only thing we're deciding is the per-block scale, fp16 (ours) vs e8m0 (MXFP8). Same fp8 e4m3 data, different scale.

@RapidMark

RapidMark commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Thinking about fp8 vs MXFP8, storage format, and acceleration...

While this PR is CPU only... it's clear we want to move onto GPU next and I was wondering... since e8m0 has loss anyway... what would the impact be if we stored in fp8 (higher quality) and had the backend convert to MXFP8 at load to use the fused instruction. What would the loss be, how long would it take to convert, could we support mxfp8 accel (where it exists) and also supprt fp8 natively. Basically, the same stored fp8 could accelerates on both paths.

I measured the convert cost on Qwen2.5-1.5B (wikitext, vs f16) and the convert-at-load path only costs (an additional) +0.44% PPL over native MXFP8.

  • our e4m3 (fp16 scale): +1.12% PPL, mean KLD 0.011
  • MXFP8 direct (f16 -> e8m0): +7.56% PPL, mean KLD 0.066
  • MXFP8 via convert (f16 -> our e4m3 -> e8m0): +8.03% PPL, mean KLD 0.069

Down-converting the stored fp8 covers both cases: full fp16-scale quality on non-MX hardware, and the MX fast path (at MXFP8 quality) on Blackwell/CDNA4, without a second file.

So f16 scale is a higher-quality source we down-convert when the hardware wants MXFP8. On MX hardware you take the quality drop on purpose, but on everything else you keep the better scale.

I do not have specific numbers on converting in memory yet (as I have not written it...) but on a consumer GPU running ~500 GB/s - 1TB/s, ~1.5GB to ~3GB should be about 3-6 ms, and 7-14GB should be about 15-30 ms.

A native-MXFP8 storage type still makes sense down the line, but mainly to carry natively-MXFP8-trained models faithfully, not for PTQ, where the convert-at-load already covers it.

Goal was to add fp8 for CPU/GPU, but directed to only do CPU First.
This adds GGML_TYPE_E4M3 (OCP e4m3fn), q8_0-shaped, with fp16 scale + 32 fp8 bytes per block.
Hand optimized with lookup tables and unrolled loops for a speed increase of ~60% (AVX2) to ~140% (NEON) after initial working version (the optimized paths are the vec_dot and the repacked 8x8 gemv/gemm).
We have tested on Windows, Linux, Mac, Arm, Intel, AMD, with speedup for AVX2, NEON, etc (test-backend-ops 16176/16176 on CPU).
AI was used to understand the codebase, provide peer reviews, and offer suggestions, as well as run tests on all available hardware.
@RapidMark
RapidMark force-pushed the cloudhands/fp8-cpu branch from ead2709 to e22b279 Compare July 7, 2026 20:32
@michaelw9999

Copy link
Copy Markdown
Contributor

I can get you some real numbers on that for the CUDA side. I've been sitting on fully working CUDA MXFP8 for a while and was just waiting for my MXFP6 PRs to get done before sharing anything.

@michaelw9999

Copy link
Copy Markdown
Contributor

what would the impact be if we stored in fp8 (higher quality) and had the backend convert to MXFP8 at load to use the fused instruction. What would the loss be, how long would it take to convert, could we support mxfp8 accel (where it exists) and also supprt fp8 natively. Basically, the same stored fp8 could accelerates on both paths.

I was not yet able to find a path where I could get converted FP8 to MXFP8 to be as fast or as high in quality compared to the native MXFP8 format just on its own. I think it may still be feasible. However, I think for most cases on Blackwell it would be worth keeping MXFP8 separate and on its own native format, and not converting from the FP8 format, because the extra loss makes it even farther away from what the original FP8 to MXFP8 loss is to begin with, hurting MXFP8 a bit as a viable format. Then in cases where higher quality is warranted, even on Blackwell it would still be better to keep using FP8 in that case (even if it's slower), however we don't have a CUDA FP8 kernel yet to compare to the MXFP8 speeds (which I posted below).

On Qwen2.5-1.5B-MXFP8-GGUF:

Model Quantized codes Scale bytes Total Bytes CUDA buffer
Native FP8 1,310,195,712 81,887,232 FP16 1,392,082,944 1,584,101,888 B
Native MXFP8 1,310,195,712 40,943,616 E8M0 1,351,139,328 1,543,158,272 B
Repack FP8 to MXFP8 1,310,195,712 81,887,232 duplicated E8M0 1,392,082,944 1,584,101,888 B

On MX hardware you take the quality drop on purpose

Here are the numbers I got in testing conversion. I did not see such a huge ppl/kld loss on MXFP8 from FP8 in my testing:

Quantization PPL PPL vs BF16 Mean KLD vs BF16 Median KLD P99 KLD Same top p
FP8 (offline) 10.081803 +0.62% 0.007007 0.005401 0.039124 95.050%
Native MXFP8 (execution) 10.153583 +1.34% 0.015255 0.010634 0.100271 93.060%
Repack FP8 to MXFP8 (execution) 10.199294 +1.80% 0.019295 0.012868 0.130996 92.460%

I tried to reduce the FP8 to MXFP8 conversion loss as best as I could and even tried a a +/-5 scale search, but that increased load time (to 34 ms from 10.7 ms), and only had around a tiny 0.001ish ppl improvement, so that would not be worth doing.

I do not have specific numbers on converting in memory yet (as I have not written it...) but on a consumer GPU running ~500 GB/s - 1TB/s, ~1.5GB to ~3GB should be about 3-6 ms, and 7-14GB should be about 15-30 ms.

Convert time for 5090 (Laptop) Repack on Qwen2.5-1.5B (196 tensors):

Measurement Repack time H2D with Repack
From PR 7.703 ms* 98.886 ms
Tested first load 11.257 ms 100.063 ms
Tested steady median, n=4 10.687 ms 100.681 ms
  • used different rounding (tied to even)
Metric Median time
H2D CUDA time 90.028 ms
Repack CUDA time 10.687 ms
H2D with Repack 100.681 ms
Host kernel-enqueue calls (196) 0.786 ms
Stream synchronization wait 15.516 ms
Full model load time 287.045 ms

A native-MXFP8 storage type still makes sense down the line, but mainly to carry natively-MXFP8-trained models faithfully, not for PTQ, where the convert-at-load already covers it.

From the experiments I ran, converting from BF16 to the native MXFP8 format with convert_hf_to_gguf directly will maintain better quality than converting from FP8 to MXFP8, so that would still be a viable path forward. I also could not get the converted/repack kernel to be as fast as with the native format, even when converting it offline and not as a repack.

On 5090 Laptop:
Qwen2.5-1.5B:

Type pp32 tk/s pp128 tk/s pp512 tk/s tg8 tk/s tg32 tk/s tg128 tk/s
Q8_0 1,983.21 7,555.18 17,256.68 102.80 104.57 105.10
Native MXFP8 2,028.24 7,942.10 19,291.59 100.91 102.69 103.21
Converted E4M3 to MXFP8 1,826.48 7,081.20 16,291.88 100.73 102.54 103.07

Qwen3.5-4B-MXFP8:

Type pp32 tk/s pp128 tk/s pp512 tk/s tg8 tk/s tg32 tk/s tg128 tk/s
Q8_0 1,987.61 5,612.11 8,005.55 134.80 145.81 144.71
Native MXFP8 2,105.52 6,505.06 9,302.49 138.28 147.60 148.07

So I think there's a place for FP8 and for MXFP8 to stand separately (and still MXFP6 too).
If there's any consensus yet for a PR for MXFP8, I can post that up to start as CPU-only.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion examples ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants