Skip to content
View original post on X: vLLMOfficial· 34/100AI score34/100

vLLM reports 7.8x and 3.7x throughput gains on two benchmarks

AISummary

vLLM says it delivered 7.8x or more the throughput of GB200 on SemiAnalysis AgentX, serving MiniMax M3 on Vera Rubin NVL72 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL-235B-A22B.

Post on XView on X
vLLMVerified on X
@vllm_project

Part of a thread · earlier post

The numbers come from two benchmarks. On @SemiAnalysis_ AgentX, vLLM serving MiniMax M3 on Vera Rubin NVL72 delivers 7.8x+ the throughput of GB200 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL-235B-A22B.

🧵 2/5

Source: vLLM · x.comPublished · added here