SGLang and NVIDIA speed up Kimi K3 inference on early-access Vera Rubin hardware
Overview
SGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Vera Rubin hardware.
It reports up to 20% faster FP8 MLA at batch 1 with 128K context, and a 5.9% end-to-end inference speedup from MoE tail fusion, which removes 276 kernel launches per decode step. These figures come from SGLang's own X post and have not been independently verified.
RadixArk says its Miles framework runs reinforcement learning end to end on Vera Rubin, using SGLang rollouts, Megatron training and one container image. It says agentic RL runs 64 concurrent sandboxes on the Vera CPU alongside the GPUs.
Written by AI from the articles below · updated Oct 9, 10:44 PM ET
Check the sources:
Developments
2 developments
- Oct 9, 10:24 PM ET · 1 articleRadixArk runs Miles end-to-end RL on NVIDIA Vera Rubin with SGLang rollouts and Megatron trainingRadixArk: RadixArk's Miles runs end-to-end RL on NVIDIA Vera Rubin with SGLang
- Oct 9, 7:23 PM ET · 1 articleSGLang optimizes Kimi K3 inference kernels on NVIDIA Vera Rubin early-access hardwareSGLang: SGLang adds Rubin optimizations that speed up Kimi K3 inference
Article timeline
The articles in this story. Times are ET.
RadixArk@radixarkOfficialPickRadixArk's Miles runs end-to-end RL on NVIDIA Vera Rubin with SGLangAIRadixArk says Miles runs reinforcement learning end to end on NVIDIA Vera Rubin, using SGLang rollouts, Megatron training and one container image. Agentic RL runs 64 concurrent sandboxes on the Vera CPU next to the GPUs. The linked SGLang post reports that Kimi K3 inference gained up to 20% faster FP8 MLA at batch 1 with 128K context, and a 5.9% end-to-end speedup from MoE tail fusion.
LMSYS Org@lmsysorgOfficialSGLang brings Kimi K3 inference speedups to NVIDIA Vera RubinAISGLang optimized attention, MoE, and speculative verification kernels for Kimi K3 NVFP4 on early-access NVIDIA Vera Rubin hardware. Reported gains include up to 20% faster FP8 MLA at batch 1 and 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion. Miles, from RadixArk, uses SGLang for rollouts in end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
SGLang@sgl_projectOfficialMiles runs full RL loop on Rubin GPUs with SGLang and MegatronAIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.
SGLang@sgl_projectOfficialSGLang reports inference speedups for MLA, MoE, and KDA kernelsAISGLang says restructured MLA decode kernels on Rubin, which fit a deeper pipeline in 327 KiB of shared memory, deliver a 16% speedup at batch 16 with 128K context and bit-identical output. The post also reports 20% faster full FP8 MLA at batch 1 and 20% faster KDA verify kernels after keeping weights in registers and reducing synchronization. It additionally covers fusing MoE finalization, the shared expert, 8-GPU all-reduce, and RMSNorm into one collective kernel.
SGLang@sgl_projectOfficialSGLang adds Rubin optimizations that speed up Kimi K3 inferenceAISGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Rubin hardware. It reports up to 20% faster FP8 MLA at batch 1 with 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. The post also says SGLang powers Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Heat trend
Not enough continuous observations to show a trend yet.