GLM-5.3 served on GB200 NVL72 at 100+ tokens/s per user
Original titleWe served GLM-5.3 on GB200 NVL72. Our interactivity target was 100+ e2e tok/s per user while serving as many concurrent agent task as pos...
AISummary
Prime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.
Source: Prime Intellect · x.comPublished · added here