vLLM and Inferact report up to 1.2x faster MoE decode using CUDA locality domains
Overview
vLLM says its locality-aware MoE implementation shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. In a post in a five-post thread, vLLM reports early results showing MiniMax M3 MoE layers up to 1.2x faster in decode.
Inferact, working with NVIDIA, separately says it led the performance tuning and locality-aware MoE work on Rubin GPUs, using the CUDA locality domain feature so each SM reads MoE weights from its own domain. Inferact also says this makes MoE layers up to 1.2x faster in decode. Neither source reports a broader benchmark.
Written by AI from the articles below · updated Oct 9, 10:06 PM ET
Check the sources:
Developments
2 developments
- Oct 9, 9:53 PM ET · 1 articleInferact and NVIDIA tune MoE layers for Rubin GPU, up to 1.2x faster decodeInferact: Inferact and NVIDIA tune MoE layers for Rubin GPUs with locality-aware kernels
- Oct 9, 9:50 PM ET · 1 articlevLLM adds locality-aware MoE using CUDA 13.4 locality domains and Green ContextsvLLM: vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x
Article timeline
The articles in this story. Times are ET.
Inferact@inferactOfficialInferact and NVIDIA tune MoE layers for Rubin GPUs with locality-aware kernelsAIInferact, working with NVIDIA, led performance tuning and integrated the MiniMax Sparse Attention kernel. The team uses the CUDA locality domain feature on Rubin GPUs so each SM reads MoE weights from its own locality domain, making MoE layers up to 1.2x faster in decode.
vLLM@vllm_projectOfficialvLLM uses NVIDIA locality domains to speed MoE decode up to 1.2xAIvLLM's locality-aware MoE shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. Early results show MiniMax M3 MoE layers up to 1.2x faster in decode. This is post 4 of a 5-post thread.

- vLLM BlogOfficialPickvLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200
AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.
Heat trend
Not enough continuous observations to show a trend yet.