vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations
Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks. Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.
AIWhy it matters: The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.