Skip to content
View original post on X: vLLMOfficial· 28/100AI score28/100

vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x

AISummary

vLLM's locality-aware MoE shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. Early results show MiniMax M3 MoE layers up to 1.2x faster in decode. This is post 4 of a 5-post thread.

Post on XView on X
vLLMVerified on X
@vllm_project

Part of a thread · earlier post

Since Ampere, NVIDIA GPUs have featured non-uniform global memory accesses. From CUDA 13.4, programmers can leverage this feature through locality domain, where HBM is partitioned and SMs within each partition can read its global memory fastest. vLLM’s locality-aware MoE shards FC1 and FC2 weights column-wise and uses locality domains with Green Contexts to launch one kernel on each domain’s SMs, so each SM reads only local HBM. Early results show up to 1.2x faster MiniMax M3 MoE layers in decode.

🧵 4/5

Source: vLLM · x.comPublished · added here