Repository navigation
feat(moe): serve NVFP4 and MXFP4 experts fused through the expert banks - #601
Merged
Merged
Conversation
Resident experts skip the state dict and load through the offload expert-bank loader (HF pieces or FTW banks) into per-layer GPU banks, so fused serves bf16, fp8_block, NVFP4 and MXFP4 alike. Unified-memory GPUs auto-select fused for these formats and ignore offload-only flags with a warning. BREAKING CHANGE: ft checkpoint always writes MoE experts as banks; an FTW converted with --moe-backend fused or triton (dense experts) no longer loads and must be reconverted.
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 4, 2026
…xperts fused through the expert banks # Conflicts: # docs/cli.md # python/freetoken/checkpoint/ftw.py # python/freetoken/layers/quantization/moe/nvfp4.py # python/freetoken/moe/expert_banks.py
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 4, 2026
…bout --disable-moe-prefill-hit-d2d Prefill hit-D2D is on by default on next (fix/moe-ttft), so the unified-memory inert-flag check warned on every start; name the flag that turns it off instead. Assisted-by: Claude Opus 5.5
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 4, 2026
FlashML-org#601 adopted with fixups, FlashML-org#599 adopted, FlashML-org#602 deferred, FlashML-org#596 own. Assisted-by: Claude Opus 5.5
pretor
pushed a commit
to pretor/FreeToken
that referenced
this pull request
Oct 5, 2026
…ks) and PR FlashML-org#599 (kv page size telemetry)
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 5, 2026
FlashML-org#601's inert-flag list still named --moe-prefill-hit-d2d, which is now the default, so every unified-memory start would have warned about it. Assisted-by: Claude Opus 5.5
zhangnju
added a commit
to zhangnju/FreeToken
that referenced
this pull request
Oct 6, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto the local rank before pack(): gate/up families split the intermediate row axis, down and its block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in the local banks instead of overflowing them.
zhangnju
added a commit
to zhangnju/FreeToken
that referenced
this pull request
Oct 8, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto the local rank before pack(): gate/up families split the intermediate row axis, down and its block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in the local banks instead of overflowing them.
zhangnju
added a commit
to zhangnju/FreeToken
that referenced
this pull request
Oct 8, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto the local rank before pack(): gate/up families split the intermediate row axis, down and its block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in the local banks instead of overflowing them.
zhangnju
added a commit
to zhangnju/FreeToken
that referenced
this pull request
Oct 9, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto the local rank before pack(): gate/up families split the intermediate row axis, down and its block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in the local banks instead of overflowing them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resident (
--moe-strategy fused) experts now load through the offload expert-bank loader instead of the state dict: HF pieces or FTW banks are packed by the kernel'spack()straight into per-layer GPU banks. Fused therefore serves bf16, fp8_block, NVFP4 and MXFP4 (gpt-oss) alike; GGUF q4_0 and DeepSeek-V4 experts stay offload-only.create_weights/resident_viewoverrides are gone;MoEMethodhas one default for every format.autopicks fused for these formats. Offload-only flags are ignored there with a warning; an explicit--moe-strategy offloadstill honors them.ft checkpointalways writes MoE experts as banks, since every strategy reads them.Breaking: an FTW converted with
--moe-backend fusedortritonkeeps its experts dense and no longer loads; reconvert it withft checkpoint.Follows up #445 and removes its
_fused_resident_okFIXME. Part of #369: this implements the resident NVFP4 / gpt-oss MXFP4 steps proposed in this comment. The sm_121 backend-probe audit and the install docs from #369 remain open.Validation
H100 80GB, greedy, four bs=1 prompts plus one batched run, token ids compared with
main:cpu,hybridand--moe-cpu-layers.main; gpt-oss FTW now serves fused (KeyError onmain). The new NVFP4 fused (Qwen3.6, Gemma-4, MiniMax) is identical to offload onmain.ft serve --moe-strategy fusedwith Qwen3.6-35B-A3B-NVFP4 answers a chat request correctly.tests/moe tests/engine tests/models tests/checkpoint tests/layers: no new failures versusmain.Not tested: Marlin / b12x resident, TP > 1, a real GB10.