Skip to content
View original post on X: RadixArkOfficial· Pick60/100AI score60/100

RadixArk's Miles runs end-to-end RL on NVIDIA Vera Rubin with SGLang

AISummary

RadixArk says Miles runs reinforcement learning end to end on NVIDIA Vera Rubin, using SGLang rollouts, Megatron training and one container image. Agentic RL runs 64 concurrent sandboxes on the Vera CPU next to the GPUs. The linked SGLang post reports that Kimi K3 inference gained up to 20% faster FP8 MLA at batch 1 with 128K context, and a 5.9% end-to-end speedup from MoE tail fusion.

AIWhy it matters

The post gives concrete speedup figures and a specific RL setup on early-access Rubin hardware, useful for engineers comparing inference and training stacks.

Post on XView on X
RadixArkVerified on X
@radixark

Miles runs RL end to end on NVIDIA Vera Rubin, out of the box: SGLang rollouts, Megatron training, one container image. And agentic RL with 64 concurrent sandboxes on the Vera CPU, right next to the GPUs. Thanks to @NVIDIAAI for the early access. Details in the blog 👇

SGLang@sgl_project
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference. Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware. Highlights: • Up to 20% faster FP8 MLA at batch 1 / 128K context • 20% faster KDA verification, with bitwise-identical output • 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU. Full results and engineering details 👉 https://www.lmsys.org/blog/2026-10-09-vera-rubin
View quoted post on X

Source: RadixArk · x.comPublished · added here