Tsinghua paper introduces TokenRouter, a serving system for token-level LLM routing
Overview
Tsinghua researchers introduce TokenRouter, a serving system for routing individual tokens between small and large language models.
The authors report throughput 2.01 to 64.15 times higher than the stronger existing setup across five routing methods, according to a post by Rohan Paul (@rohanpaul_ai) on X.
The post says current frameworks such as vLLM and SGLang run one model per request, so models sharing an answer wait on each other at every step. TokenRouter instead gives each model its own server, passes partial answers between them, keeps the KV cache, and briefly holds requests to batch work.
Written by AI from the articles below · updated Oct 9, 11:38 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXTokenRouter serves token-level LLM routing up to 64.15x fasterAITsinghua researchers present TokenRouter, a serving system for token-level routing between small and large models that raises throughput 2.01 to 64.15 times over the stronger existing setup across five routing methods. Current frameworks such as vLLM and SGLang run one model per request, so models sharing an answer wait on each other at every step. TokenRouter gives each model its own server and passes partial answers between them, keeping the KV cache and holding requests briefly to batch work.

Heat trend
Not enough continuous observations to show a trend yet.