SGLang v0.5.21 landed! Native decisions API is here 🎉
Some of our favorite updates:
- Decisions API turns an LLM/VLM into a low-latency classifier and scorer
- /v1/score can now rerank search or RAG results in one go
- PD instances can switch between prefill and decode with no restart needed
- DeepSeek-V4.1 Flash gets 22% faster first token on long prompts
- Kimi K3 gets 20.6% higher prefill throughput in PD serving
- GLM-5.3-Flash now runs on AMD MI355X with FP8 / MXFP4 MoE and MTP
- You can now run MiniMax H3 inside @ComfyUI with SGLang-Diffusion backend
New models include DeepSeek-V4.1 Flash, GigaChat 3.5, MiMo-V2.6, Ling-3.0-flash-VL, IQuest-Q1, Qwen-Image 2.1, FLUX 3 Action, and more.
Full release notes👇
