Prime Intellect releases Prime Flash MoE kernels for faster Blackwell inference
Original titlePrime Flash MoE - Faster MoE Kernels optimized for Blackwell
Prime Intellect has released Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for mixture-of-experts feed-forward layers.
The kernels are up to 2.4× faster than PyTorch grouped GEMM and deliver about 2.3× speedup across the 4k–128k token range, and are integrated into its prime-rl framework.
Two pipelines are offered: a fused single-kernel path for small problem sizes and a split three-kernel path for larger ones, supporting both bf16 and MXFP8.
The post explains how fusing MoE expert computation on Blackwell hardware avoids intermediate memory traffic, with benchmarks showing where fused and split pipelines each win.
Source: Prime Intellect Blog · primeintellect.aiPublished · added here