Cognition runs SWE-2 on NVIDIA Vera Rubin, citing 4.8x per-GPU throughput over GB200
Overview
Cognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, according to SGLang's October 9, 2026 post on X. SGLang says Cognition has served its SWE-2 model on Rubin with SGLang since September.
SGLang says Rubin delivers about 4.8x more token throughput per GPU than GB200 on SWE-2 inference, at the same decode speed. The figure comes from SGLang's post and has not been independently verified.
Written by AI from the articles below · updated Oct 9, 3:50 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
SGLang@sgl_projectOfficialCognition first in production on NVIDIA Vera Rubin, 4.8x per-GPU over GB200AICognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, and serving SWE-2 on it with SGLang since September. SGLang says Rubin delivers 4.8x per-GPU throughput over GB200 at matched interactivity.
Heat trend
Not enough continuous observations to show a trend yet.