Skip to content

feat(moe): serve NVFP4 and MXFP4 experts fused through the expert banks - #601

Merged
jason-fxz merged 2 commits into
mainfrom
feat/fused-fp4-resident
Oct 4, 2026
Merged

jason-fxz merged 2 commits into
mainfrom
feat/fused-fp4-resident

Conversation

@jason-fxz

@jason-fxz jason-fxz commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Resident (--moe-strategy fused) experts now load through the offload expert-bank loader instead of the state dict: HF pieces or FTW banks are packed by the kernel's pack() straight into per-layer GPU banks. Fused therefore serves bf16, fp8_block, NVFP4 and MXFP4 (gpt-oss) alike; GGUF q4_0 and DeepSeek-V4 experts stay offload-only.

  • The per-format create_weights / resident_view overrides are gone; MoEMethod has one default for every format.
  • On unified-memory GPUs (GB10) auto picks fused for these formats. Offload-only flags are ignored there with a warning; an explicit --moe-strategy offload still honors them.
  • ft checkpoint always writes MoE experts as banks, since every strategy reads them.

Breaking: an FTW converted with --moe-backend fused or triton keeps its experts dense and no longer loads; reconvert it with ft checkpoint.

Follows up #445 and removes its _fused_resident_ok FIXME. Part of #369: this implements the resident NVFP4 / gpt-oss MXFP4 steps proposed in this comment. The sm_121 backend-probe audit and the install docs from #369 remain open.

Validation

H100 80GB, greedy, four bs=1 prompts plus one batched run, token ids compared with main:

  • Offload family unchanged on 15 configs: Qwen3-30B, Qwen3.6 bf16 / FP8 / NVFP4, gpt-oss HF and FTW, Gemma-4 bf16 / NVFP4, truncated GLM-4.7 / MiniMax-M2.5 / DeepSeek-V4, plus cpu, hybrid and --moe-cpu-layers.
  • Fused bf16 / fp16 / FP8 / gpt-oss identical to fused on main; gpt-oss FTW now serves fused (KeyError on main). The new NVFP4 fused (Qwen3.6, Gemma-4, MiniMax) is identical to offload on main.
  • ft serve --moe-strategy fused with Qwen3.6-35B-A3B-NVFP4 answers a chat request correctly.
  • tests/moe tests/engine tests/models tests/checkpoint tests/layers: no new failures versus main.

Not tested: Marlin / b12x resident, TP > 1, a real GB10.

Resident experts skip the state dict and load through the offload expert-bank loader (HF pieces or FTW banks) into per-layer GPU banks, so fused serves bf16, fp8_block, NVFP4 and MXFP4 alike. Unified-memory GPUs auto-select fused for these formats and ignore offload-only flags with a warning.

BREAKING CHANGE: ft checkpoint always writes MoE experts as banks; an FTW converted with --moe-backend fused or triton (dense experts) no longer loads and must be reconverted.
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 4, 2026
…xperts fused through the expert banks

# Conflicts:
#	docs/cli.md
#	python/freetoken/checkpoint/ftw.py
#	python/freetoken/layers/quantization/moe/nvfp4.py
#	python/freetoken/moe/expert_banks.py
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 4, 2026
…bout --disable-moe-prefill-hit-d2d

Prefill hit-D2D is on by default on next (fix/moe-ttft), so the
unified-memory inert-flag check warned on every start; name the flag
that turns it off instead.

Assisted-by: Claude Opus 5.5
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 4, 2026
FlashML-org#601 adopted with fixups, FlashML-org#599 adopted, FlashML-org#602 deferred, FlashML-org#596 own.

Assisted-by: Claude Opus 5.5
@jason-fxz
jason-fxz merged commit 189c8ad into main Oct 4, 2026
pretor pushed a commit to pretor/FreeToken that referenced this pull request Oct 5, 2026
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 5, 2026
FlashML-org#601's inert-flag list still named --moe-prefill-hit-d2d, which is now the
default, so every unified-memory start would have warned about it.

Assisted-by: Claude Opus 5.5
zhangnju added a commit to zhangnju/FreeToken that referenced this pull request Oct 6, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local
for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto
the local rank before pack(): gate/up families split the intermediate row axis, down and its
block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every
expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in
the local banks instead of overflowing them.
zhangnju added a commit to zhangnju/FreeToken that referenced this pull request Oct 8, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local
for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto
the local rank before pack(): gate/up families split the intermediate row axis, down and its
block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every
expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in
the local banks instead of overflowing them.
zhangnju added a commit to zhangnju/FreeToken that referenced this pull request Oct 8, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local
for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto
the local rank before pack(): gate/up families split the intermediate row axis, down and its
block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every
expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in
the local banks instead of overflowing them.
zhangnju added a commit to zhangnju/FreeToken that referenced this pull request Oct 9, 2026
The FlashML-org#601 bank system sizes each kernel's banks from layout(), which is already TP-local
for the unquantized method and now for fp8-block too. Shard the full checkpoint pieces onto
the local rank before pack(): gate/up families split the intermediate row axis, down and its
block scale split the column axis, per-output-row globals replicate. No-op at tp=1, so every
expert kind that reaches build_expert_banks under TP (its kernel must declare tp_ok) lands in
the local banks instead of overflowing them.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant