DeepSeek-V4.1-Flash on vLLM gains 1.9× speed and 5.3× throughput
Original title1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS pe...
AISummary
Three weeks after launch, vLLM's serving of DeepSeek-V4.1-Flash runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on SemiAnalysis AgentX. The post offers interactive figures explaining how these gains were achieved.
Source: vLLM · x.comPublished · added here