Skip to content
Trending storyDeveloping

vLLM v0.31.0 released with DeepSeek-V4.1-Flash support and serving updates

1 article1 sourceLast article 18h ago

Overview

AI overview

The vLLM project released v0.31.0, according to its official X account, adding support for DeepSeek-V4.1-Flash.

The announcement highlights a vllm preload command that keeps model weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. It also lists large-scale serving, scheduling, and HiSparse fixes. The release comprises 717 commits from 307 contributors, 96 of them first-time contributors. Full release notes are linked on GitHub; the announcement itself does not report performance measurements.

AIWritten by AI from the articles below · overview updated Oct 8, 9:05 PM ET

Check the sources:

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. vLLM
    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

Heat trend

Not enough continuous observations to show a trend yet.