NexRT decode engine launches for Nex-N2.5-Pro with SGLang prefill
Overview
SGLang (@sgl_project) says NexRT, the decode engine from the Nex ecosystem, has launched for single-request inference on Nex-N2.5-Pro.
In this setup SGLang handles prefill while NexRT handles decode, and the two engines share routed MoE weights.
SGLang reports over 1,000 tokens/s at 8K context on 8×H100 using DFlash. Source code is not yet released; SGLang says it will follow.
Written by AI from the articles below · updated Oct 10, 12:07 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
SGLang@sgl_projectOfficialNexRT launches as a low-latency decode engine for Nex-N2.5-ProAINexRT, the decode engine from the Nex ecosystem, launches for single-request inference on Nex-N2.5-Pro. SGLang handles prefill while NexRT handles decode, and the two engines share routed MoE weights. The source reports over 1,000 tokens/s at 8K context on 8×H100 with DFlash, and says source code will follow.
Heat trend
Not enough continuous observations to show a trend yet.