Skip to content
View original post on X: RadixArk· 40/100AI score40/100

RadixArk releases experimental NVFP4 checkpoint for GLM-5.3

AISummary

RadixArk has published an experimental NVFP4 checkpoint for Zhipu's GLM-5.3 on Hugging Face, and says Miles support for GLM-5.3 is on the way.

The company says GLM-5.2 is already serving hundreds of thousands of people in production with its partners on SGLang.

Background from SGLang reports day-0 serving support for GLM-5.3, with 537.6 tok/s/user on NVFP4 and 413 tok/s/user on FP8 at BS=1 with TP8 on 8x B300.

Post on XView on X
@radixark

GLM-5.2 on SGLang was battle-tested and is currently serving hundreds of thousands of people in production with our partners. GLM-5.3 carries all the great things we built for it👏

We've added an experimental NVFP4 checkpoint for GLM-5.3, and Miles support for GLM-5.3 is on the way!

http://huggingface.co/RadixArk/GLM-5.3-NVFP4

SGLang@sgl_project
GLM-5.3 weights from @Zai_org are live, with SGLang powering day-0 serving support! GLM-5.3 inherits every optimization and feature we battle-tested for GLM-5.2 over the past months. On real-world multi-turn agentic workloads, we measured 537.6 tok/s/user on NVFP4 and 413 tok/s/user on FP8, at BS=1 with TP8 on 8x B300. It's fast, efficient, and production-ready today on @NVIDIAAI Blackwell and Hopper, and @AIatAMD MI300X/325X/355X. We believe this is a big step forward for GLM-Series in agentic tasks, with a path toward mythos-class cyber capability. Can't wait to see what people build with GLM-5.3 and SGLang🚀 Cookbook👇
View quoted post on X

Source: RadixArk · x.comPublished · added here