HOLY ALERT🚨: NVIDIA vLLM INFERENCE RUBIN IS 3.2x BETTER IN PROFIT💰️ PER GIGAWATT & HAS UP TO 🚀 10x BETTER PERF PER DOLLAR THAN EVEN GB300 NVL72. This is on the widely used production LLM engine called @vllm_project. We explain below.👇️(1/3)🧵
NVIDIA says Rubin beats GB300 NVL72 in vLLM inference economics
AISummary
SemiAnalysis says NVIDIA's Rubin delivers 3.2x better profit per gigawatt and up to 10x better performance per dollar than GB300 NVL72 on the vLLM production LLM engine. The post is the first of a three-part thread that will explain the results.
Post on XView on X
Source: SemiAnalysis · x.comPublished · added here
