Together with @NVIDIAAI, Inferact led performance tuning, MiniMax Sparse Attention kernel integration, and locality-aware MoE.
We use CUDA locality domain feature to optimize MoE layer on Rubin GPU so each SMs can read MoE weights from their own locality domain. That makes MoE layers up to 1.2x faster in decode.
🧵 2/3
