Skip to content
Trending storyDeveloping

vLLM says Blackwell kernels run unmodified on NVIDIA Rubin GPUs

1 article1 sourcesince Oct 9Last article 3h ago ·

Overview

AISummary of 1 article

vLLM (@vllm_project) says its Blackwell kernels run on NVIDIA Rubin GPUs without modification, and that FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

The post cites a Rubin GPU with 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. vLLM says the faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. These figures are vLLM's claims; the post does not include independent benchmark results.

Written by AI from the articles below · updated Oct 9, 10:02 PM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. vLLMOfficial
    vLLM says Rubin GPUs run its Blackwell kernels unmodified

    AIvLLM says a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. The faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

    Image from @vllm_project's post

Heat trend

Not enough continuous observations to show a trend yet.