SGLang brings Kimi K3 inference speedups to NVIDIA Vera Rubin
AISGLang optimized attention, MoE, and speculative verification kernels for Kimi K3 NVFP4 on early-access NVIDIA Vera Rubin hardware. Reported gains include up to 20% faster FP8 MLA at batch 1 and 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion. Miles, from RadixArk, uses SGLang for rollouts in end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
