vLLM maintainers show TPUv7 megakernels beat GB200 NVL72 on Kimi K3
AIInferact says vLLM maintainers used megakernel optimization to reach 700 tokens per second per user on TPUv7 running Kimi K3. SemiAnalysis, which shared the work, reports this is 56% better performance than Nvidia's GB200 NVL72. Inferact links a full technical breakdown of the TPU megakernel work on its blog.



