Skip to content
View original post on X: vLLMOfficial· 33/100AI score33/100

vLLM says Rubin GPUs run its Blackwell kernels unmodified

AISummary

vLLM says a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. The faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

Post on XView on X
vLLMVerified on X
@vllm_project

Part of a thread · earlier post

Compared with GB200, a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, 1.7x the NVLink bandwidth, and 2 to 4x faster softmax exponentials. The faster NVLink cuts the cost of AllReduce and all-to-all in MoE serving. vLLM’s Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

🧵 3/5

Source: vLLM · x.comPublished · added here