vLLM adds NVIDIA Vera Rubin support and a nightly container for DeepSeek, Kimi, GLM and MiniMax
Overview
vLLM says it now supports NVIDIA Vera Rubin and that users can try it today with the vllm/vllm-openai:cu134-nightly image, which runs DeepSeek, Kimi, GLM and MiniMax models.
The project says it expects further improvements for vLLM on Vera Rubin.
In a separate post, Inferact, which is also named in the vLLM thread, reports early results showing more than 7.8x the throughput of GB200 running MiniMax M3 on AgentX at matched interactivity. Inferact says the team integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.
The vLLM thread says the Rubin bring-up involved Inferact, NVIDIA, Red Hat AI and the vLLM community.
Written by AI from the articles below · updated Oct 9, 10:07 PM ET
Check the sources:
Developments
2 developments
- Oct 9, 9:50 PM ET · 1 articlevLLM lets users try DeepSeek, Kimi, GLM and MiniMax models on Vera Rubin via a nightly containervLLM: vLLM now runs DeepSeek, Kimi, GLM, and MiniMax on Vera Rubin
- Oct 9, 9:49 PM ET · 2 articlesvLLM adds support for NVIDIA Vera Rubin, reporting 7.8x GB200 throughput on MiniMax M3Inferact: vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3
Article timeline
The articles in this story. Times are ET.
Inferact@inferactOfficialvLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3AIInferact says vLLM now supports NVIDIA Vera Rubin, with early results showing more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity on AgentX. The team says it integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.

vLLM@vllm_projectOfficialvLLM now runs DeepSeek, Kimi, GLM, and MiniMax on Vera RubinAIvLLM says users can try its Vera Rubin support today with the vllm/vllm-openai:cu134-nightly image, which runs DeepSeek, Kimi, GLM, and MiniMax models. The project thanked contributors and said it expects further improvements for vLLM on Vera Rubin.
vLLM@vllm_projectOfficialvLLM adds NVIDIA Vera Rubin support, reaching 7.8x GB200 throughput on MiniMax M3AIvLLM now supports NVIDIA Vera Rubin, and early results show more than 7.8x the throughput of GB200 running MiniMax M3 on AgentX. The post is the first in a five-part thread on the Rubin bring-up, which involved Inferact, NVIDIA, Red Hat AI, and the vLLM community.

- vLLM BlogOfficialPickvLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200
AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.
Heat trend
Not enough continuous observations to show a trend yet.