Repository navigation
Conversation
…t load) Quantize the bf16 lm_head to fp8 block-scale at load so decode reads fewer bytes from the big [vocab, hidden] matrix. Reuses the existing fp8-block triton GEMV end-to-end; no new kernel. Opt-in via FREETOKEN_FP8_LM_HEAD=1 (default OFF -> zero change to logit quality). - scheme.py: QuantKind.FP8_BLOCK_QAT + fp8_block_qat_scheme (weight role only, no scale in the checkpoint). - kernel/triton/fp8_block_linear.py: per_block_quant_fp8 (bf16 -> fp8 + bf16 128-block scale), the inverse of dequant_block_fp8. - linear/fp8_block.py: Fp8BlockQuantizeAtLoadLinearMethod + its kernel's finalize() block-quantizes the bf16 weight post-load. - configs/fp8.py: flag-gated scheme_for_name returns BLOCK_QAT for lm_head before the default not_convert exclusion. - qwen3_5_moe/weight.py: the two loader validations exempt FP8_BLOCK_QAT (the checkpoint ships a bf16 weight with no scale companion).
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 10, 2026
…for the lm_head
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 10, 2026
…kernel skip the fp8_block-only scheme check FREETOKEN_FP8_LM_HEAD=1 asserted while building the lm_head: the kernel inherited unusable_reason from the float-scale fp8-block kernel, which calls fp8_block_size on a scheme that is FP8_BLOCK_QAT. Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This change adds an opt-in path that quantizes the bf16
lm_headto fp8 block-scale at load time. Thelm_headis a large[vocab, hidden]matrix, so decoding one token reads a lot of bytes from it; storing it in fp8 halves that traffic. The feature reuses the existing fp8-block triton GEMV end-to-end and adds no new kernel. It is disabled by default and is turned on withFREETOKEN_FP8_LM_HEAD=1, so the default behavior and logit quality are unchanged.How it works
The checkpoint still ships a bf16
lm_head. When the flag is set, the loader assigns it a new quantization scheme whose weight is quantized to fp8 with a bf16 128-block scale after the weight is loaded, in the linear method'sfinalizestep. From then on it is served by the same fp8-block GEMV used for every other fp8-block linear.Changes
layers/quantization/scheme.py: addsQuantKind.FP8_BLOCK_QATandfp8_block_qat_scheme. The scheme declares only aweightrole, because the checkpoint carries no scale for it.kernel/triton/fp8_block_linear.py: addsper_block_quant_fp8, which quantizes a bf16 weight to fp8 plus a bf16 128-block scale. It is the inverse of the existingdequant_block_fp8.layers/quantization/linear/fp8_block.py: addsFp8BlockQuantizeAtLoadLinearMethodand its kernel. The kernel'sfinalizeblock-quantizes the bf16 weight to fp8 after load; the method otherwise reuses the standard fp8-block GEMV.layers/quantization/configs/fp8.py: when the flag is set,scheme_for_namereturns theBLOCK_QATscheme forlm_headbefore the default exclusion that would have kept it in bf16.models/qwen3_5_moe/weight.py: two loader validations are relaxed forFP8_BLOCK_QAT, which ships a bf16 weight with no scale companion, so the bf16 tensor is accepted instead of being rejected against an fp8 scheme.Validation
per_block_quant_fp8anddequant_block_fp8round-trip on CPU with a relative error of about 2.25%, which matches fp8 block granularity.lm_head): bf16lm_headgives 86.9 tok/s and fp8lm_headgives 88.9 tok/s, a 2.3% decode improvement, with byte-identical greedy output (the fp8 path reuses the same triton kernel, so greedy decoding is unchanged).Notes
Test plan
per_block_quant_fp8/dequant_block_fp8round-trip, rel-err ~2.25%