Skip to contentSkip to stories

Updated

#DeepSeek

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. vLLMOfficialAI score62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    AIvLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

    Image from @vllm_project's post
  2. LeiphoneNewsAI score23

    Former Tencent Hunyuan Vision Lead Hu Han Raises Funds at Hundreds of Millions Valuation

    AIHu Han, former head of Tencent Hunyuan's visual large model algorithm center, is raising tens of millions of dollars for a multimodal startup at a target valuation of several hundred million dollars, with Yuanshi Capital as financial advisor. Investors say Hu has spoken with several firms over the past month and has suggested his model could eventually be sold to large technology companies such as DeepSeek.

Oct 7

Oct 7Wed
  1. InferactOfficialAI score38

    Inferact and partners cut vLLM TTFT nearly 70% at ~100K throughput

    AIInferact, working with DeepSeek, NVIDIA, and SemiAnalysis alongside the vLLM community, says joint work across models, custom kernels, and engine serving cuts time to first token (TTFT) by nearly 70% at ~100K throughput. vLLM is the open-source inference engine, and Inferact optimizes it for enterprise production deployments.

  2. vLLMOfficialAI score22

    vLLM and NVIDIA cut TTFT nearly 70% at ~100K throughput

    AITogether with DeepSeek's model and kernels and NVIDIA's collaboration, vLLM reports that time to first token (TTFT) drops nearly 70% at roughly 100K throughput. The post credits Inferact and the vLLM community for the work and thanks SemiAnalysis for AgentX.

Oct 4

Oct 4Sun

Oct 2

Oct 2Fri
  1. SGLangOfficialAI score58

    SGLang v0.5.21 adds native decisions API and new model support

    AISGLang has released v0.5.21 with a native Decisions API that turns an LLM or VLM into a low-latency classifier and scorer. The release also lets /v1/score rerank search or RAG results in one call, lets PD instances switch between prefill and decode without restarting, and adds support for models including DeepSeek-V4.1 Flash, Kimi K3, and GLM-5.3-Flash on AMD MI355X. The announcement reports a 22% faster first token on long prompts for DeepSeek-V4.1 Flash and 20.6% higher prefill throughput for Kimi K3 in PD serving.

    Image from @sgl_project's post