DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context
Original titledeepseek-ai/DeepSeek-V4.1-Flash
AISummary
DeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.
AIWhy it matters
The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.
Source: DeepSeek · new models on Hugging Face · huggingface.co