Skip to content
Read the original: SemiAnalysis· Published 43/100AI score43/100

How GLM-5.3 Sparse Attention Affects HBM and Serving Costs on GB200, GB300, and MI355X

Original titleHow GLM5.3 Sparse Attention Affects HBM Memory Usage

AISummary

Sparse attention cuts per-operation KV cache reads but does not reduce overall memory capacity, so top-k cache misses still depend on HBM. SemiAnalysis's InferenceX estimates GB200 at about $0.044 per million total tokens at 150 tokens per second, roughly 12% below MI355X running ATOM at $0.049.

Neither system holds a uniform cost advantage across the tested 100, 125, and 150 tokens-per-second targets.

Read the original newsletter.semianalysis.com

Source: SemiAnalysis · newsletter.semianalysis.comPublished · added here