Skip to content
View original post on X: InferactOfficial· 32/100AI score32/100

Inferact and NVIDIA tune MoE layers for Rubin GPUs with locality-aware kernels

AISummary

Inferact, working with NVIDIA, led performance tuning and integrated the MiniMax Sparse Attention kernel. The team uses the CUDA locality domain feature on Rubin GPUs so each SM reads MoE weights from its own locality domain, making MoE layers up to 1.2x faster in decode.

Post on XView on X
InferactVerified on X
@inferact

Part of a thread · earlier post

Together with @NVIDIAAI, Inferact led performance tuning, MiniMax Sparse Attention kernel integration, and locality-aware MoE.

We use CUDA locality domain feature to optimize MoE layer on Rubin GPU so each SMs can read MoE weights from their own locality domain. That makes MoE layers up to 1.2x faster in decode.

🧵 2/3

Source: Inferact · x.comPublished · added here