SGLang brings optimized inference to @nvidia Vera Rubin, with performance gains across attention, MoE, and speculative verification for Kimi K3 NVFP4.
Miles by @radixark takes this further with end-to-end RL training, using SGLang for rollouts and the Vera CPU for concurrent agent sandboxes.
The teams share their early results, benchmarks, and the engineering behind them 👇
