vLLM v0.31.0 released with DeepSeek-V4.1-Flash support and serving updates
Overview
The vLLM project released v0.31.0, according to its official X account, adding support for DeepSeek-V4.1-Flash.
The announcement highlights a vllm preload command that keeps model weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. It also lists large-scale serving, scheduling, and HiSparse fixes. The release comprises 717 commits from 307 contributors, 96 of them first-time contributors. Full release notes are linked on GitHub; the announcement itself does not report performance measurements.
AIWritten by AI from the articles below · overview updated Oct 8, 9:05 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- vLLMvLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features
vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.
Heat trend
Not enough continuous observations to show a trend yet.