Skip to content
View original post on X: LMSYS OrgOfficial· 50/100AI score50/100

SGLang brings Kimi K3 inference speedups to NVIDIA Vera Rubin

AISummary

SGLang optimized attention, MoE, and speculative verification kernels for Kimi K3 NVFP4 on early-access NVIDIA Vera Rubin hardware.

Reported gains include up to 20% faster FP8 MLA at batch 1 and 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion.

Miles, from RadixArk, uses SGLang for rollouts in end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

Post on XView on X
LMSYS OrgVerified on X
@lmsysorg

SGLang brings optimized inference to @nvidia Vera Rubin, with performance gains across attention, MoE, and speculative verification for Kimi K3 NVFP4.

Miles by @radixark takes this further with end-to-end RL training, using SGLang for rollouts and the Vera CPU for concurrent agent sandboxes.

The teams share their early results, benchmarks, and the engineering behind them 👇

SGLang@sgl_project
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference. Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware. Highlights: • Up to 20% faster FP8 MLA at batch 1 / 128K context • 20% faster KDA verification, with bitwise-identical output • 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU. Full results and engineering details 👉 https://www.lmsys.org/blog/2026-10-09-vera-rubin
View quoted post on X

Source: LMSYS Org · x.comPublished · added here