Skip to content
Read the original: vLLM BlogOfficial· Pick62/100AI score62/100

vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

Original titlevLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72

AISummary

vLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

AIWhy it matters

The post gives specific hardware specs, kernel configurations, and benchmark figures for running vLLM on Vera Rubin NVL72, useful for teams planning deployments.

Read the original vllm.ai

Source: vLLM Blog · vllm.aiPublished · added here