vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3
AIInferact says vLLM now supports NVIDIA Vera Rubin, with early results showing more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity on AgentX. The team says it integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.









