Skip to content
vLLM Blog· Helen Zhao, Fynn Schmitt-Ulms, Yuchen Fama, Antonio J. Dominguez, and Kevin Li·· 24d agoPickAI score62

How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72

AI summary

The vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

Why it matters

The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

Source: vLLM Blog · vllm.ai