Skip to content
Read the original: vLLM· Published 42/100AI score42/100

TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

Original titleGreat work from the @TileRT_AI and @AIatAMD teams, who got GLM-5.3 to 469 tok/s single-user decode with vLLM on 8× MI355X on @SemiAnalysi...

AISummary

The TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

Read the original x.com

Source: vLLM · x.comPublished · added here