Repository navigation
Conversation
added 2 commits
October 3, 2026 19:17
…t_scale ModelOpt exports MXFP8 block scales as '<proj>.weight_scale' while the scheme loads them as 'weight_scale_inv': the reader skipped the suffix and the module KeyError'd mid WeightLoad. The dialect's storage table now carries the truth (MXFP8 -> .weight_scale) and _rename consults it instead of blanket-skipping, so the scale fuses and validates on the existing block-scale path. Covers the split GDN layout with b|a quantized (local-inference-lab/Qwen3.8-Flash-Next-NVFP4).
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 3, 2026
…QAD checkpoints (NVFP4 PLE tables, MXFP8 dense) Conflict: main had split load_ple_table into the validating scan_ple_table plus a host-RAM admission check. The NVFP4 table (packed rows, per-shard block scales, global weight_scale_2) now goes through that scan with the same checks, and load_ple_table fills both banks. Assisted-by: Claude Opus 5.5
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 3, 2026
…fig's rows and index the oob lookup per row The table scan now checks the row count against the n-gram config, and the cuda gather test assigned 8 ids into a 16-wide row. Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Makes
local-inference-lab/Qwen3.8-Flash-Next-NVFP4(a ModelOpt QAD requant ofQwen3.8-Flash-Next) load and serve out of the box. On
mainthe loader dies in two places:U8, two e2m1 codes per byte + an fp8-e4m3block scale per 16 elements + a global scalar under
.weight_scale_2) andload_ple_tableonly accepts
F8_E4M3;<proj>.weight_scalewhile the scheme loadsthem as
weight_scale_inv, so_renameskipped them and the moduleKeyErrors mid load.Two commits, one per defect:
uint8bank plus a pinned fp8 block-scale bank; the gather kernel gains an NVFP4 branch (LUT unpack,
block scale x global scale, int64 row offsets at 320M rows). Reuses the canonical
_e2m1_lutfrom
kernel/triton/nvfp4_dequant.pyand thee4m3_compathelpers; no new dequant math.Packed rows are served by the
pinnedPLE backend only; the disk backend raisesNotImplementedErrorwith a reason instead of misreading rows.storage table carries the truth (
MXFP8 -> weight_scale) and_renamerenames through itinstead of blanket-skipping scale suffixes, so the scales fuse and validate on the existing
block-scale path. Covers the split GDN layout with
in_proj_b|aquantized.Tests
CPU-only unit tests (no GPU needed), green on this tree:
pytest tests/models/qwen4_exp -m "not slow".test_ple_nvfp4.py: synthetic packed checkpoint (real safetensors files in tmp_path) loadedthrough
load_ple_table; the pinned gather is bitwise-checked against a torch reference thatmirrors the checkpoint formula (e2m1 code x block scale x global scale, fp32, stored bf16);
out-of-range ids store zeros; prefetch path included.
scales come out under
weight_scale_invfused across qkv/z/ba parts.Scope: what is model-specific and what is not
modelopt.py) is engine-wide: every checkpoint loading through theModelOpt dialect with an MXFP8 scheme now resolves its block scales correctly, not just this
checkpoint.
_renamewiring is local toqwen4_expbecause that dialect has its own weight readerthat blanket-skipped scale suffixes; other dialect readers already consult the storage table.
n-gram PLE table; the unpack itself reuses the shared E2M1 LUT and
e4m3_compathelpers.Evaluated on
ft serve --model local-inference-lab/Qwen3.8-Flash-Next-NVFP4 --max-seq-len-override 262144 --memory-ratio 0.94 --kv-reserve-tokens 262144 --ple-backend pinned(env
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True;0.94is tuned for 32 GB cards, lower it elsewhere).RadixArk/Qwen3.8-Flash-Next-NVFP4checkpoint measuredby us on the same machine and the same pinned engine stack; the QAD side additionally carries
exactly the two commits of this PR - without them the checkpoint does not load at all.
Same-day A/B, one box, fixed seeds, both sides at
--memory-ratio 0.93; the0.94in thecommand above is the battle configuration we also run this checkpoint on (the bench rows below
are directly comparable):
Decode, from passive engine-side bs=1 step windows in the logs (long agent sessions, 130-190k
context; same log category both columns, different days - treat as indicative, the bench rows
above are the controlled comparison):
Read of the whole table: quality is parity within noise (the two benchmark rows differ by 1 task
in each direction, and the four differing ARC items split 2/2), while the QAD checkpoint is
27 GiB smaller and faster - MXFP8 dense + NVFP4 PLE put fewer bytes on a bandwidth-bound decode
path, and the bench wall clocks above are the same-day proof that speed did not cost accuracy.
Notes
modelopt.pyone-liner changes MXFP8 scale storage for the ModelOpt dialect as a whole.It matches the exporter's naming in this checkpoint family;
FP8_BLOCKkeepsweight_scale_inv. If another ModelOpt MXFP8 export stores the scales underweight_scale_inv, the rename falls through and the missing key surfaces as it did beforethis PR - no silent wrong-scale load.
reasoning_effort/tool-calling contract is unchanged (same architecture, same chat template).