Skip to content
View original post on X: InferactOfficial· 52/100AI score52/100

vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3

AISummary

Inferact says vLLM now supports NVIDIA Vera Rubin, with early results showing more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity on AgentX. The team says it integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.

Post on XView on X
InferactVerified on X
@inferact

vLLM on NVIDIA Vera Rubin NVL72 delivers more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity, measured on AgentX.

Inferact’s team has been co-leading this effort since Rubin was announced. We have integrated Rubin optimized MSA prefill kernel and tuned AgentX performance with locality aware optimizations targeted for Rubin GPU. These are early results.

🧵 1/3

vLLM@vllm_project
vLLM now supports NVIDIA Vera Rubin. The early results show more than 7.8x the throughput of GB200 on MiniMax M3 on AgentX. @inferact, @NVIDIAAI, @RedHat_AI, and the vLLM community have been bringing vLLM up on Rubin since it was announced. Here is where things stand. 🧵 1/5
View quoted post on X

Source: Inferact · x.comPublished · added here