Skip to content
Read the original: LMSYS Org· Published 60/100AI score60/100

SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

Original titleSGLang brings Day-0 support for Qwen 3.8-Flash-Next, an early preview of the Qwen4 architecture!

AISummary

SGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

Read the original x.com

Source: LMSYS Org · x.comPublished · added here