Skip to content
View original post on X: SGLangOfficial· 42/100AI score42/100

Cognition first in production on NVIDIA Vera Rubin, 4.8x per-GPU over GB200

AISummary

Cognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, and serving SWE-2 on it with SGLang since September. SGLang says Rubin delivers 4.8x per-GPU throughput over GB200 at matched interactivity.

Post on XView on X
SGLangVerified on X
@sgl_project

Congrats to @cognition — the first customer in production on @nvidia Vera Rubin, powered by @CoreWeave, and serving SWE-2 on it with SGLang since September 🎉 4.8x per GPU at matched interactivity over GB200. Excited to see what the community builds on Rubin.

Cognition@cognition
Cognition is the first customer running @NVIDIA Vera Rubin, powered by @CoreWeave! Hopper in 2022. Blackwell in 2024. Now Vera Rubin: On SWE-2 inference, the new chips deliver about 4.8x more token throughput than GB200 at the same decode speed. More compute, better agents!
View quoted post on X

Source: SGLang · x.comPublished