Repository navigation
Conversation
|
Hi @RapidMark, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Multiple backend changes in one PR: Actually.. it's just CPU... but I needed to add m2 lines to ggml-metal-device.m to extend the existing GGML_TYPE_NVFP4 exclusions in ggml_metal_device_supports_op (for MUL_MAT/MUL_MAT_ID and GET_ROWS) to also cover E4M3. |
|
How did you arrive at using an fp16 scale factor? I wonder if it should be e8m0 instead to make this MXFP8. But I also worry about poor alignment with this block size, generally I would prefer to have at least 4B aligned blocks. |
|
I acknowledge the e8m0/mxfp4 are already in the tree, so fp16 was a deliberate choice, not for lack of e8m0. We mirrored q8_0's block so the q8_0 activation path and the repack gemv/gemm would apply directly. ALso fp16 is a finer per-block scale than e8m0 (e8m0 is power-of-2 only, no mantissa). While fp16 has less exponent range, that doesn't matter (AFAIK) for normalized weight blocks. We wanted to carry native fp8 as fp8. Native fp8 today is usually per-tensor/per-channel (MXFP8's per-block e8m0 is the other real variant), and what our block preserves faithfully is the e4e3 elements vs q8_0, the e4e3 elements carry exponent range that linear int8 doesn't, so they match the fp8 distributions the model was trained in; the goal is representing fp8 models, not beating q8_0. As fro block alignment, 34 bytes isn;t 4-byte-aligned, but we felt e8m0 alone would make it worse; 1 byte scale + 32 bytes = 33 bytes. So switching to MXFP8 does not fix the alignment (by itself), which would need a block redesign and would make larger blocks like nvfp4's 64 with grouped scales (or extra padding). Basically we chose q8_0 reuse + finer scale. I am working on a consumer only path, which means I also need to supprot AMD as much as NVIDIA is supported (which is a major task, as you know), so everywhere I run across a descripency, I try to fill it. It also means I'm learning alot, but also means I'm coming from an untainted view. Are you steering me toward MXFP8 (e8m0) and if so... what model/format should we target (MXFP8 vs per-tensor)? also.. 34 bytes is not aligned, but e8m0 would make it 33 bytes so doing a 4-byte/16-byte layout would need a significant redesign. Do you have a preferred block size/padding? |
136a833 to
ead2709
Compare
|
I understand the appeal of the f16 scale factor. I don't have a strong opinion on what we do here. If MXFP8 becomes widely used, then we'd obviously want to support that natively, and it also allows for hardware acceleration that an f16 scale won't. Regarding alignment, yeah, I'd generally prefer to use a larger block size with multiple scales, packed similar to NVFP4. Some models may have layers that are not a multiple of the native block size, but maybe we could just pad with zeros? |
|
My focus is on consumer hardware, which is limited by memory. You run it to fit the model in 12-24 GB. The goal is to make it fit, and still give the highest quality. I don't want to degrade the memory-forced consumer twice, once by dropping to fp8 and again with a coarse scale. From my point of view e8m0's only real upside (today) is Blackwell's fused microscaling. RDNA4 and Ada, don't have that. They apply the block scale in software either way, so e8m0 buys no acceleration, just less quality. Also, e4m3 is cross-vendor because the OCP formats get consumed on both AMD and NVIDIA fp8 units. Same stored weights accelerate on RDNA4 and on Ada/Blackwell. One format, both vendors. I do agree MXFP8 is the right format for MX hardware (Blackwell, the CDNA4 parts) and most likely future consumer HW, but I also think we can add MXFP8 when consumer HW catches up. Overall, datacenter's, while they can do FP8, will likely use FP16, unless they're looking for speed over quality. This is just my point of view, but overall... I kinow we'll need MXFP8 in the future as well... so we can go either way... I just feel like we're giving up on existing hardware in hopes of future hardware. Once we deicde on a direction... I can go back in and align blocks, groupped scales, padding, etc. Lastly... I beleive, the only thing we're deciding is the per-block scale, fp16 (ours) vs e8m0 (MXFP8). Same fp8 e4m3 data, different scale. |
|
Thinking about fp8 vs MXFP8, storage format, and acceleration... While this PR is CPU only... it's clear we want to move onto GPU next and I was wondering... since e8m0 has loss anyway... what would the impact be if we stored in fp8 (higher quality) and had the backend convert to MXFP8 at load to use the fused instruction. What would the loss be, how long would it take to convert, could we support mxfp8 accel (where it exists) and also supprt fp8 natively. Basically, the same stored fp8 could accelerates on both paths. I measured the convert cost on Qwen2.5-1.5B (wikitext, vs f16) and the convert-at-load path only costs (an additional) +0.44% PPL over native MXFP8.
Down-converting the stored fp8 covers both cases: full fp16-scale quality on non-MX hardware, and the MX fast path (at MXFP8 quality) on Blackwell/CDNA4, without a second file. So f16 scale is a higher-quality source we down-convert when the hardware wants MXFP8. On MX hardware you take the quality drop on purpose, but on everything else you keep the better scale. I do not have specific numbers on converting in memory yet (as I have not written it...) but on a consumer GPU running ~500 GB/s - 1TB/s, ~1.5GB to ~3GB should be about 3-6 ms, and 7-14GB should be about 15-30 ms. A native-MXFP8 storage type still makes sense down the line, but mainly to carry natively-MXFP8-trained models faithfully, not for PTQ, where the convert-at-load already covers it. |
Goal was to add fp8 for CPU/GPU, but directed to only do CPU First. This adds GGML_TYPE_E4M3 (OCP e4m3fn), q8_0-shaped, with fp16 scale + 32 fp8 bytes per block. Hand optimized with lookup tables and unrolled loops for a speed increase of ~60% (AVX2) to ~140% (NEON) after initial working version (the optimized paths are the vec_dot and the repacked 8x8 gemv/gemm). We have tested on Windows, Linux, Mac, Arm, Intel, AMD, with speedup for AVX2, NEON, etc (test-backend-ops 16176/16176 on CPU). AI was used to understand the codebase, provide peer reviews, and offer suggestions, as well as run tests on all available hardware.
ead2709 to
e22b279
Compare
|
I can get you some real numbers on that for the CUDA side. I've been sitting on fully working CUDA MXFP8 for a while and was just waiting for my MXFP6 PRs to get done before sharing anything. |
I was not yet able to find a path where I could get converted FP8 to MXFP8 to be as fast or as high in quality compared to the native MXFP8 format just on its own. I think it may still be feasible. However, I think for most cases on Blackwell it would be worth keeping MXFP8 separate and on its own native format, and not converting from the FP8 format, because the extra loss makes it even farther away from what the original FP8 to MXFP8 loss is to begin with, hurting MXFP8 a bit as a viable format. Then in cases where higher quality is warranted, even on Blackwell it would still be better to keep using FP8 in that case (even if it's slower), however we don't have a CUDA FP8 kernel yet to compare to the MXFP8 speeds (which I posted below).
Here are the numbers I got in testing conversion. I did not see such a huge ppl/kld loss on MXFP8 from FP8 in my testing:
I tried to reduce the FP8 to MXFP8 conversion loss as best as I could and even tried a a +/-5 scale search, but that increased load time (to 34 ms from 10.7 ms), and only had around a tiny 0.001ish ppl improvement, so that would not be worth doing.
Convert time for 5090 (Laptop) Repack on Qwen2.5-1.5B (196 tensors):
From the experiments I ran, converting from BF16 to the native MXFP8 format with On 5090 Laptop:
So I think there's a place for FP8 and for MXFP8 to stand separately (and still MXFP6 too). |
Overview
Add GGML_TYPE_E4M3 (OCP e4m3fn), a CPU fp8 quantization type, so native-fp8 weights load and run as fp8 instead of being dequantized to f16 (double the memory) or approximated as q8_0. Goal is fp8 for CPU and GPU; per the CPU-first rule this is the CPU side, and the GPU backends follow in separate PRs.
Block is q8_0-shaped (fp16 scale + 32 fp8 bytes per block), dotting against q8_0 activations.
Additional information
What's in here:
Test model (small, runnable): https://huggingface.co/RapidMark/Qwen2.5-1.5B-E4M3-GGUF
This is the exact e4m3 gguf the numbers below were measured on. The f16 baseline and q8_0 comparison are regenerated from the base model (Qwen/Qwen2.5-1.5B-Instruct) with the quantize command in Reproduction.
Evaluation. Qwen2.5-1.5B-Instruct, wikitext-2-raw test (full, 584 chunks, n_ctx 512), pure CPU. The similar-size comparison type is Q8_0 (both are one fp16 scale + 32 one-byte elements per block).
Perplexity and KL-divergence vs f16:
Pure-CPU throughput (llama-bench, 64 threads):
e4m3 lands a little behind q8_0 on accuracy, which is expected: q8_0's linear int8 is near-lossless for these weights, and e4m3 spends half its bits on the exponent. On CPU its value is faithful fp8 at half the size of f16, not beating q8_0.
The repacked gemv/gemm is in this PR (not just vec_dot) because it's a 2.5x prefill speedup: on the same model and box, pp512 is 269.8 t/s with the repacked gemm vs 106.4 t/s on the plain vec_dot (tg128 33.1 vs 27.9).
Reproduction.
Validation. test-backend-ops -b CPU passes 16176/16176 (e4m3 MUL_MAT included). Verified on x86 (MSVC and Linux GCC), ARM NEON, and Apple (Metal correctly declines e4m3 and falls back to the CPU path). The gguf-py numpy class is bit-exact against the C library (test_quants.py).
Environment.
Requirements