The numbers come from two benchmarks. On @SemiAnalysis_ AgentX, vLLM serving MiniMax M3 on Vera Rubin NVL72 delivers 7.8x+ the throughput of GB200 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL-235B-A22B.
🧵 2/5
