vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations
Original titleDeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
AISummary
Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks.
Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.
AIWhy it matters
The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.
Source: vLLM Blog · vllm.ai