Skip to content

Formats

Product updates Latest news

New features, redesigns, and commercial changes in AI products and applications.

136 picksPast 30 days: 78 itemsTotal: 2,184 items

Latest pick

Top picks archive · Page 7

May 17

May 17SunItems 121–136
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition launches Auto-Triage, letting Devin investigate alerts and open fixes

    Cognition has released Auto-Triage in Devin Automations, which lets Devin respond to Slack messages, Linear events, GitHub activity, schedules, and webhooks. Devin can investigate with connected observability tools and the codebase, then post a summary, tag an owner, or open a PR. Devin runs in network-sandboxed environments with added protections against prompt injection and data exfiltration, and a limited-time offer gives $200 in credits for a first automation.

    AIWhy it matters: The post shows how an agent handles alerts and bug reports from existing team channels, a practical pattern for teams weighing automated incident response.

Apr 24

Apr 24Fri
  1. DeepSeek API NewsAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    The DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    AIWhy it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Apr 21

Apr 21Tue
  1. Nick TurleyAI score67

    ChatGPT Images 2.0 launches with better instruction following and dense text rendering

    Nick Turley announced ChatGPT Images 2.0 as a major advance in image generation, citing better adherence to detailed instructions, rendering of dense text, and more accurate understanding of the world. He said the model can spend extra time planning and refining outputs for tasks needing more accuracy and clarity, and that users have generated over 1 billion images with ChatGPT.

    AIWhy it matters: The post names concrete gains in instruction following, dense text rendering, and optional extended thinking for image output, which helps readers gauge practical scope.

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    Cognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    AIWhy it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    Cognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    AIWhy it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    Anthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    AIWhy it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

  3. Dario AmodeiAI score72

    Dario Amodei backs Project Glasswing to counter AI-driven cyber threats

    Dario Amodei said many of the world's leading companies have joined Project Glasswing, an effort to address cyber threats posed by increasingly capable AI systems. The initiative was introduced by Anthropic and is powered by its newest frontier model, Claude Mythos Preview, which the quoted post says can find software vulnerabilities better than all but the most skilled humans.

    AIWhy it matters: The post gives a concrete example of how a frontier AI lab is organizing industry partners around AI-driven software vulnerability discovery.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)AI score72

    Devin can now break tasks down and run a team of managed Devins

    Devin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    AIWhy it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.

Mar 17

Mar 17Tue
  1. Xiaomi MiMoAI score68

    Xiaomi releases MiMo-V2-TTS, a speech model with controllable emotion and singing

    Xiaomi has launched MiMo-V2-TTS, a speech synthesis model that lets users describe the desired voice style in plain language. The model also supports dialects, character voices, non-verbal sounds such as coughs and sighs, and singing within one model. It was pretrained on over 100 million hours of speech data and refined with multi-dimensional reinforcement learning.

    AIWhy it matters: The source gives concrete controls for emotion, dialect, singing, and non-verbal sounds, showing how a voice model can be directed through plain-language style prompts.

Jan 7

Jan 7Wed
  1. Nick TurleyAI score72

    OpenAI launches ChatGPT Health for connecting medical records

    OpenAI is launching ChatGPT Health, a dedicated and private space where users can securely connect apps and medical records. The launch starts with a small group of users from the waitlist, with access expanding over the coming weeks.

    AIWhy it matters: The post names the access path and a dedicated space for health records, which matters for judging how sensitive data would be handled.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    Cognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    AIWhy it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

May 21, 2025

May 21, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition launches official DeepWiki MCP server for indexed GitHub repos

    Cognition launched the official DeepWiki Model Context Protocol server, which is free and requires no login or authentication. It gives programmatic access to ask_question, read_wiki_contents, and read_wiki_structure for GitHub repositories indexed on DeepWiki.com. Private repositories require a Devin account with GitHub connected, and open-source maintainers can apply for $500 in Devin credits.

    AIWhy it matters: The source names the three tools and the access path, showing how indexed GitHub repositories can be queried programmatically without login.

May 14, 2025

May 14, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin 2.1 adds confidence ratings and built-in codebase intelligence

    Cognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.

    AIWhy it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.

Apr 2, 2025

Apr 2, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score75

    Cognition launches Devin 2.0 with agent-native IDE and new planning tools

    Cognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.

    AIWhy it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.

Dec 9, 2024

Dec 9, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score67

    Cognition makes Devin generally available to engineering teams from $500 a month

    Cognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.

    AIWhy it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    Cognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    AIWhy it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.