Skip to contentSkip to stories

Updated

#Agent

Aug 27

Aug 27Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

  3. Qwen · new models on Hugging FaceAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    AITencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  3. Jazzyear · ArticlesAI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    AIIn an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.

  4. Bryan CatanzaroAI score62

    NVIDIA and AWS expand partnership with 2 million more GPUs and Vera CPU for agentic AI

    AINVIDIA and AWS are expanding their partnership across GPUs, CPUs, networking, open models and software. The announcement cites 2 million additional NVIDIA GPUs across AWS infrastructure, the NVIDIA Vera CPU coming to AWS for agentic AI, NVLink Fusion with NVHBM memory, and 100,000 GPUs for U.S. government AI factories on secure AWS infrastructure.

  5. LMSYS OrgAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Z.ai Release NotesAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  4. Andrew NgAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    AIOpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  5. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. Microsoft AI BlogAI score14

    Five Signals Show How Organizations Scale AI Through Security, Governance, and Observability

    AIMicrosoft's AI Blog outlines five signals that trust, not speed alone, lets organizations scale AI from pilots to enterprise-wide use. Its first signal is observability, citing Microsoft's Cyber Pulse AI Security Report finding that 29% of employees use unsanctioned AI agents their security teams cannot see. The post also says security should be built into AI systems by design and governance should be continuous rather than a one-time approval.

Aug 21

Aug 21Fri
  1. Andrew NgAI score31

    Andrew Ng outlines six core skills for building and deploying AI applications

    AIAndrew Ng's AI Engineering Skills Map ranks building and deploying AI applications as the top skill tier, spanning LLM foundations, data grounding, agentic systems, evaluation-driven development, production operations, and machine learning foundations. He explains that AI outputs are less predictable than traditional software, so skilled engineers build iteratively, examining results and deciding next steps based on intermediate outcomes. The skills map was derived from job postings, expert interviews, and survey responses.

  2. Amazon ScienceAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    AIAmazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

  3. DeepSeekAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

  4. DeepSeek API NewsAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu