Read the original: vLLM BlogOfficial· vLLM Team, Inferact, Red Hat, and NVIDIA· Pick62/100AI score62/100
vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200
Original titlevLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72
AISummary
vLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.
AIWhy it matters
The post gives specific hardware specs, kernel configurations, and benchmark figures for running vLLM on Vera Rubin NVL72, useful for teams planning deployments.
Source: vLLM Blog · vllm.aiPublished · added here