vLLM on NVIDIA Vera Rubin NVL72 delivers more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity, measured on AgentX.
Inferact’s team has been co-leading this effort since Rubin was announced. We have integrated Rubin optimized MSA prefill kernel and tuned AgentX performance with locality aware optimizations targeted for Rubin GPU. These are early results.
🧵 1/3
