Skip to content
Read the original: Baseten Blog·Published· 27d agoPickAI score62

DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

DeepSeek-V4.1-Flash: more efficient prefill for coding agents

AISummary

DeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

AIWhy it matters

The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

Read the original baseten.co

Source: Baseten Blog · baseten.co