Congrats to @cognition — the first customer in production on @nvidia Vera Rubin, powered by @CoreWeave, and serving SWE-2 on it with SGLang since September 🎉 4.8x per GPU at matched interactivity over GB200. Excited to see what the community builds on Rubin.
Cognition first in production on NVIDIA Vera Rubin, 4.8x per-GPU over GB200
AISummary
Cognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, and serving SWE-2 on it with SGLang since September. SGLang says Rubin delivers 4.8x per-GPU throughput over GB200 at matched interactivity.
Post on XView on X
SGLangVerified on X
@sgl_project
Cognition is the first customer running @NVIDIA Vera Rubin, powered by @CoreWeave! Hopper in 2022. Blackwell in 2024. Now Vera Rubin: On SWE-2 inference, the new chips deliver about 4.8x more token throughput than GB200 at the same decode speed. More compute, better agents!
Source: SGLang · x.comPublished
