Skip to content
View original post on X: SGLangOfficial· 44/100AI score44/100

NexRT launches as a low-latency decode engine for Nex-N2.5-Pro

AISummary

NexRT, the decode engine from the Nex ecosystem, launches for single-request inference on Nex-N2.5-Pro. SGLang handles prefill while NexRT handles decode, and the two engines share routed MoE weights. The source reports over 1,000 tokens/s at 8K context on 8×H100 with DFlash, and says source code will follow.

Post on XView on X
SGLangVerified on X
@sgl_project

Congrats to the @NexEcosystem on launching NexRT!

SGLang runs prefill and NexRT takes over decode for low-latency single-request inference on Nex-N2.5-Pro.
- The two engines share routed MoE weights
- DFlash builds on target feature capture in SGLang v0.5.15

Try NexRT with SGLang: https://github.com/nex-agi/NexRT

Nex@NexEcosystem
⚡ Meet NexRT, the low-latency decode engine behind Nex-N2.5. 💡 Built for single-request inference on Nex-N2.5-Pro. • 1,000+ tokens/s at 8K context on 8×H100 with DFlash* • Standard decoding, MTP and DFlash • Context parallelism for long contexts, up to 256K SGLang handles prefill; NexRT handles decode. The two engines share routed MoE weights. Custom CUDA kernels, full decode CUDA Graphs and direct GPU communication help reduce latency throughout the decode path. Performance report and demo now available. Source code will follow. *DFlash throughput assumes an average of 5.5 committed tokens per round; actual throughput depends on draft acceptance. 🌐 Explore Nex-N2.5 and NexRT: https://nex.sii.edu.cn/ 🔗 GitHub:https://github.com/nex-agi/NexRT
View quoted post on X

Source: SGLang · x.comPublished