Skip to content
Read the original: LMSYS Org· Published 38/100AI score38/100

SGLang adds multi-item scoring for faster decision model serving

Original title🚀 New blog: Scaling JEV-like decision models with SGLang

AISummary

SGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

Read the original x.com

Source: LMSYS Org · x.comPublished · added here