vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200
AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.
Why it matters: The post gives specific hardware specs, kernel configurations, and benchmark figures for running vLLM on Vera Rubin NVL72, useful for teams planning deployments.

