Skip to content
Trending storyDeveloping

NexRT decode engine launches for Nex-N2.5-Pro with SGLang prefill

1 article1 sourcesince Oct 10Last article 1h ago ·

Overview

AISummary of 1 article

SGLang (@sgl_project) says NexRT, the decode engine from the Nex ecosystem, has launched for single-request inference on Nex-N2.5-Pro.

In this setup SGLang handles prefill while NexRT handles decode, and the two engines share routed MoE weights.

SGLang reports over 1,000 tokens/s at 8K context on 8×H100 using DFlash. Source code is not yet released; SGLang says it will follow.

Written by AI from the articles below · updated Oct 10, 12:07 PM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 10
  1. SGLangOfficial
    NexRT launches as a low-latency decode engine for Nex-N2.5-Pro

    AINexRT, the decode engine from the Nex ecosystem, launches for single-request inference on Nex-N2.5-Pro. SGLang handles prefill while NexRT handles decode, and the two engines share routed MoE weights. The source reports over 1,000 tokens/s at 8K context on 8×H100 with DFlash, and says source code will follow.

Heat trend

Not enough continuous observations to show a trend yet.