SGLang adds Rubin optimizations that speed up Kimi K3 inference
AISGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Rubin hardware. It reports up to 20% faster FP8 MLA at batch 1 with 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. The post also says SGLang powers Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.





