Skip to content
Trending storyDeveloping

vLLM adds NVIDIA Vera Rubin support and a nightly container for DeepSeek, Kimi, GLM and MiniMax

4 articles3 sourcessince Oct 9Last article 2h ago ·

Overview

AISummary of 4 articles

vLLM says it now supports NVIDIA Vera Rubin and that users can try it today with the vllm/vllm-openai:cu134-nightly image, which runs DeepSeek, Kimi, GLM and MiniMax models.

The project says it expects further improvements for vLLM on Vera Rubin.

In a separate post, Inferact, which is also named in the vLLM thread, reports early results showing more than 7.8x the throughput of GB200 running MiniMax M3 on AgentX at matched interactivity. Inferact says the team integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.

The vLLM thread says the Rubin bring-up involved Inferact, NVIDIA, Red Hat AI and the vLLM community.

Written by AI from the articles below · updated Oct 9, 10:07 PM ET

Check the sources:

Developments

2 developments

  1. Oct 9, 9:50 PM ET · 1 article
    vLLM lets users try DeepSeek, Kimi, GLM and MiniMax models on Vera Rubin via a nightly container
    vLLM: vLLM now runs DeepSeek, Kimi, GLM, and MiniMax on Vera Rubin
  2. Oct 9, 9:49 PM ET · 2 articles
    vLLM adds support for NVIDIA Vera Rubin, reporting 7.8x GB200 throughput on MiniMax M3
    Inferact: vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. InferactOfficial
    vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3

    AIInferact says vLLM now supports NVIDIA Vera Rubin, with early results showing more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity on AgentX. The team says it integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.

    Image from @inferact's post
  2. vLLMOfficial
    vLLM now runs DeepSeek, Kimi, GLM, and MiniMax on Vera Rubin

    AIvLLM says users can try its Vera Rubin support today with the vllm/vllm-openai:cu134-nightly image, which runs DeepSeek, Kimi, GLM, and MiniMax models. The project thanked contributors and said it expects further improvements for vLLM on Vera Rubin.

  3. vLLMOfficial
    vLLM adds NVIDIA Vera Rubin support, reaching 7.8x GB200 throughput on MiniMax M3

    AIvLLM now supports NVIDIA Vera Rubin, and early results show more than 7.8x the throughput of GB200 running MiniMax M3 on AgentX. The post is the first in a five-part thread on the Rubin bring-up, which involved Inferact, NVIDIA, Red Hat AI, and the vLLM community.

    Video from @vllm_project's post
Oct 8
  1. vLLM BlogOfficialPick
    vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

    AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

Heat trend

Not enough continuous observations to show a trend yet.