Skip to content
Trending storyDeveloping

vLLM and Inferact report up to 1.2x faster MoE decode using CUDA locality domains

3 articles3 sourcessince Oct 9Last article 2h ago ·

Overview

AISummary of 3 articles

vLLM says its locality-aware MoE implementation shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. In a post in a five-post thread, vLLM reports early results showing MiniMax M3 MoE layers up to 1.2x faster in decode.

Inferact, working with NVIDIA, separately says it led the performance tuning and locality-aware MoE work on Rubin GPUs, using the CUDA locality domain feature so each SM reads MoE weights from its own domain. Inferact also says this makes MoE layers up to 1.2x faster in decode. Neither source reports a broader benchmark.

Written by AI from the articles below · updated Oct 9, 10:06 PM ET

Check the sources:

Developments

2 developments

  1. Oct 9, 9:53 PM ET · 1 article
    Inferact and NVIDIA tune MoE layers for Rubin GPU, up to 1.2x faster decode
    Inferact: Inferact and NVIDIA tune MoE layers for Rubin GPUs with locality-aware kernels
  2. Oct 9, 9:50 PM ET · 1 article
    vLLM adds locality-aware MoE using CUDA 13.4 locality domains and Green Contexts
    vLLM: vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. InferactOfficial
    Inferact and NVIDIA tune MoE layers for Rubin GPUs with locality-aware kernels

    AIInferact, working with NVIDIA, led performance tuning and integrated the MiniMax Sparse Attention kernel. The team uses the CUDA locality domain feature on Rubin GPUs so each SM reads MoE weights from its own locality domain, making MoE layers up to 1.2x faster in decode.

  2. vLLMOfficial
    vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x

    AIvLLM's locality-aware MoE shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. Early results show MiniMax M3 MoE layers up to 1.2x faster in decode. This is post 4 of a 5-post thread.

    GIF from @vllm_project's post
Oct 8
  1. vLLM BlogOfficialPick
    vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

    AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

Heat trend

Not enough continuous observations to show a trend yet.