Compared with GB200, a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, 1.7x the NVLink bandwidth, and 2 to 4x faster softmax exponentials. The faster NVLink cuts the cost of AllReduce and all-to-all in MoE serving. vLLM’s Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.
🧵 3/5
