Skip to content
Read the original: vLLM Blog· Inferact and the vLLM Team·Published PickAI score62/100

vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations

Original titleDeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0

AISummary

Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks.

Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.

AIWhy it matters

The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.

Read the original vllm.ai

Source: vLLM Blog · vllm.ai