Skip to content

Companies & models

DeepSeek Latest news

Follow DeepSeek model releases, open weights, technical reports, and API pricing changes.

14 picksPast 30 days: 5 itemsTotal: 80 items

Latest pick Key moments ↓

DeepSeek top picks

Oct 6

Oct 6TueItems 1–14
  1. vLLM BlogAI score62

    vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations

    Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks. Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.

    AIWhy it matters: The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.

Sep 11

Sep 11Fri
  1. Baseten BlogAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    DeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    AIWhy it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

Sep 10

Sep 10Thu
  1. DeepSeekAI score72

    DeepSeek V4.1-Flash goes live on its API with native multimodal support

    DeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.

  2. DeepSeek API NewsAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    DeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    AIWhy it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    DeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    AIWhy it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

Aug 31

Aug 31Mon
  1. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    DeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    AIWhy it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 21

Aug 21Fri
  1. DeepSeek API NewsAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    DeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    AIWhy it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 13

Aug 13Thu
  1. DeepSeekAI score68

    DeepSeek Harness v0.1 enters Developer Preview as an open-source agent harness

    DeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.

    AIWhy it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.

  2. DeepSeek API NewsAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    DeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    AIWhy it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek releases DeepSeek-V4-Pro-0813 with stronger agentic benchmark results

    DeepSeek has released DeepSeek-V4-Pro-0813 as the official version superseding the V4-Pro preview, built on the preview structure with a DSpark speculative decoding module. The model scores higher than the preview on the listed benchmarks, including Terminal Bench 2.1 at 87.9 and DeepSWE at 62.7, and the weights are under the MIT License.

    AIWhy it matters: The release reports agent benchmark gains over the preview and lists vLLM and SGLang setup, useful for judging deployment cost and fit.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    DeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    AIWhy it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    DeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    AIWhy it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    Ahmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    AIWhy it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

  2. DeepSeek API NewsAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    The DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    AIWhy it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Key moments

Since 2023
  1. ModelDeepSeek V4.1 Flash released with open weights
  2. ModelDeepSeek V4 Pro update released
  3. ModelDeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities
  4. ResearchDeepSeek-OCR released
  5. ModelDeepSeek-V3.2-Exp introduces sparse attention
  1. CompanyFounded in Hangzhou, backed by High-Flyer
  2. ModelDeepSeek Coder and DeepSeek LLM released
  3. ModelDeepSeek-V2 released at very low prices
  4. ModelDeepSeek-V3 released
  5. ModelDeepSeek-R1 released as open weights
  6. ProductDeepSeek app tops the US App Store
  7. ProductOpen Source Week publishes core infrastructure code
  8. ModelDeepSeek-V3-0324 update released
  9. ModelDeepSeek-R1-0528 update released
  10. ModelDeepSeek-V3.1 released
  11. ResearchR1 training method published in Nature
  12. ModelDeepSeek-V3.2-Exp introduces sparse attention
  13. ResearchDeepSeek-OCR released
  14. ModelDeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities
  15. ModelDeepSeek V4 Pro update released
  16. ModelDeepSeek V4.1 Flash released with open weights