Miles runs RL end to end on NVIDIA Vera Rubin, out of the box: SGLang rollouts, Megatron training, one container image. And agentic RL with 64 concurrent sandboxes on the Vera CPU, right next to the GPUs. Thanks to @NVIDIAAI for the early access. Details in the blog 👇
RadixArk's Miles runs end-to-end RL on NVIDIA Vera Rubin with SGLang
AISummary
RadixArk says Miles runs reinforcement learning end to end on NVIDIA Vera Rubin, using SGLang rollouts, Megatron training and one container image. Agentic RL runs 64 concurrent sandboxes on the Vera CPU next to the GPUs. The linked SGLang post reports that Kimi K3 inference gained up to 20% faster FP8 MLA at batch 1 with 128K context, and a 5.9% end-to-end speedup from MoE tail fusion.
AIWhy it matters
The post gives concrete speedup figures and a specific RL setup on early-access Rubin hardware, useful for engineers comparing inference and training stacks.
Post on XView on X
RadixArkVerified on X
@radixark
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference. Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware. Highlights: • Up to 20% faster FP8 MLA at batch 1 / 128K context • 20% faster KDA verification, with bitwise-identical output • 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU. Full results and engineering details 👉 https://www.lmsys.org/blog/2026-10-09-vera-rubinView quoted post on X
Source: RadixArk · x.comPublished · added here
