Skip to content
Trending storyDeveloping

Cognition runs SWE-2 on NVIDIA Vera Rubin, citing 4.8x per-GPU throughput over GB200

1 article1 sourcesince Oct 9Last article 2h ago ·

Overview

AISummary of 1 article

Cognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, according to SGLang's October 9, 2026 post on X. SGLang says Cognition has served its SWE-2 model on Rubin with SGLang since September.

SGLang says Rubin delivers about 4.8x more token throughput per GPU than GB200 on SWE-2 inference, at the same decode speed. The figure comes from SGLang's post and has not been independently verified.

Written by AI from the articles below · updated Oct 9, 3:50 PM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. SGLangOfficial
    Cognition first in production on NVIDIA Vera Rubin, 4.8x per-GPU over GB200

    AICognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, and serving SWE-2 on it with SGLang since September. SGLang says Rubin delivers 4.8x per-GPU throughput over GB200 at matched interactivity.

Heat trend

Not enough continuous observations to show a trend yet.