Since Ampere, NVIDIA GPUs have featured non-uniform global memory accesses. From CUDA 13.4, programmers can leverage this feature through locality domain, where HBM is partitioned and SMs within each partition can read its global memory fastest. vLLM’s locality-aware MoE shards FC1 and FC2 weights column-wise and uses locality domains with Green Contexts to launch one kernel on each domain’s SMs, so each SM reads only local HBM. Early results show up to 1.2x faster MiniMax M3 MoE layers in decode.
🧵 4/5
