vLLM says Blackwell kernels run unmodified on NVIDIA Rubin GPUs
Overview
vLLM (@vllm_project) says its Blackwell kernels run on NVIDIA Rubin GPUs without modification, and that FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.
The post cites a Rubin GPU with 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. vLLM says the faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. These figures are vLLM's claims; the post does not include independent benchmark results.
Written by AI from the articles below · updated Oct 9, 10:02 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
vLLM@vllm_projectOfficialvLLM says Rubin GPUs run its Blackwell kernels unmodifiedAIvLLM says a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. The faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

Heat trend
Not enough continuous observations to show a trend yet.