Skip to content

perf(moe): avoid intra-op fanout while filling NVFP4 host banks - #635

Open
earlvanze wants to merge 1 commit into
FlashML-org:mainfrom
earlvanze:perf/nvfp4-host-copy-threading-20261008
Open

earlvanze wants to merge 1 commit into
FlashML-org:mainfrom
earlvanze:perf/nvfp4-host-copy-threading-20261008

Conversation

@earlvanze

Copy link
Copy Markdown
Contributor

Change

Port the NVFP4 host-bank copy optimization to the current expert-piece loader. During CPU host-bank placement, temporarily use one PyTorch intra-op thread and restore the prior setting on success or failure. The scope is NVFP4 host banks; resident GPU packing and other expert formats retain their current threading.

The previous optimization lived in a model-specific loader, which upstream has since replaced with build_expert_banks. This patch applies at that shared placement point. It also isolates the loader optimization from the broader GLM-5.3 Flash PR (#292).

Evidence

On this host, 100 warmed 128 KiB CPU copy_ calls took 3.03–4.38 seconds with 12 intra-op threads and 0.0003 seconds with one thread. This is a microbenchmark of the copy operation, not a model startup or inference benchmark.

CUDA_VISIBLE_DEVICES='' PYTHONPATH=python /home/umbrel/.venv/bin/python -m pytest tests/moe/test_expert_bank_threading.py tests/moe/test_offload.py tests/moe/test_nvfp4_backends.py -q -m 'not needs_weights': 23 passed, 14 skipped (GPU-dependent on this host).

Focused Ruff F checks and git diff --check pass. The new test verifies bank contents, one-thread placement, and restoration after a packing error.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant