Skip to content

Formats · Latest news

Product updates

New features, redesigns, and commercial changes in AI products and applications.

225 top picks all-time · 140 in the past 30 days · chosen from 2,776 items collected all-time

Latest pick

Top picks archive · Page 11

Top picks 201–220 of 225

May 6

May 6Wed
  1. Nick TurleyXAI score62

    OpenAI rolls out GPT-5.5 Instant to ChatGPT with better factuality

    AIOpenAI has shipped GPT-5.5 Instant to ChatGPT, rolling out to everyone over the next couple of days. The quoted post says the model focuses on factuality, reducing hacks, and improving baseline intelligence, and is significantly less likely to hallucinate.

    Why it matters: The quoted post gives concrete targets for the update, factuality and hallucination reduction, useful for judging whether the default ChatGPT model changed in practice.

Apr 24

Apr 24Fri
  1. DeepSeek API NewsOfficialAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    AIThe DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    Why it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Apr 22

Apr 22Wed
  1. Fidji SimoXAI score62

    OpenAI launches ChatGPT for Clinicians and HealthBench Professional

    AIOpenAI announced two health-focused launches: ChatGPT for Clinicians, a free version of ChatGPT designed for clinical work, and HealthBench Professional, a new benchmark for evaluating real clinician chat tasks. The author, Fidji Simo, wrote that she is excited about what these launches can unlock for care.

    Why it matters: The post names two health launches, a free clinician-focused ChatGPT version and a clinician chat benchmark, which shows how OpenAI is targeting medical workflows.

Apr 21

Apr 21Tue
  1. Nick TurleyXAI score67

    ChatGPT Images 2.0 launches with better instruction following and dense text rendering

    AINick Turley announced ChatGPT Images 2.0 as a major advance in image generation, citing better adherence to detailed instructions, rendering of dense text, and more accurate understanding of the world. He said the model can spend extra time planning and refining outputs for tasks needing more accuracy and clarity, and that users have generated over 1 billion images with ChatGPT.

    Why it matters: The post names concrete gains in instruction following, dense text rendering, and optional extended thinking for image output, which helps readers gauge practical scope.

    Image from @nickaturley's post

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringOfficialAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    AIAnthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    Why it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

  3. Werner VogelsXAI score62

    Amazon S3 Files lets users mount any S3 bucket as a filesystem

    AIWerner Vogels announced S3 Files, which lets users mount any S3 bucket as a filesystem without making copies, running sync scripts, or choosing between file and object storage. He linked to a detailed post by Andy Warfield on the feature and its design history, including the filerectories concept that did not make the final release.

    Why it matters: The post explains how S3 Files changes access to existing buckets, which matters for teams weighing file and object storage workflows.

  4. Dario AmodeiXAI score72

    Dario Amodei backs Project Glasswing to counter AI-driven cyber threats

    AIDario Amodei said many of the world's leading companies have joined Project Glasswing, an effort to address cyber threats posed by increasingly capable AI systems. The initiative was introduced by Anthropic and is powered by its newest frontier model, Claude Mythos Preview, which the quoted post says can find software vulnerabilities better than all but the most skilled humans.

    Why it matters: The post gives a concrete example of how a frontier AI lab is organizing industry partners around AI-driven software vulnerability discovery.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Devin can now break tasks down and run a team of managed Devins

    AIDevin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    Why it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.

Mar 17

Mar 17Tue
  1. Xiaomi MiMoOfficialAI score68

    Xiaomi releases MiMo-V2-TTS, a speech model with controllable emotion and singing

    AIXiaomi has launched MiMo-V2-TTS, a speech synthesis model that lets users describe the desired voice style in plain language. The model also supports dialects, character voices, non-verbal sounds such as coughs and sighs, and singing within one model. It was pretrained on over 100 million hours of speech data and refined with multi-dimensional reinforcement learning.

    Why it matters: The source gives concrete controls for emotion, dialect, singing, and non-verbal sounds, showing how a voice model can be directed through plain-language style prompts.

Mar 11

Mar 11Wed
  1. Nano Banana 2.1OfficialAI score62

    How to get the most out of Nano Banana 2 for image generation

    AINano Banana 2, also called Gemini 3.1 Flash Image, adds visual grounding with Google Search, 512px resolutions, and extreme aspect ratios of 1:8 and 1:4. The guide advises using it as the default for new projects, with Nano Banana Pro reserved for complex prompts it fails, and keeping Thinking mode off by default.

    Why it matters: The guide compares Nano Banana 1, 2, and Pro with concrete routing advice, which helps developers decide which model to default to and how to control cost.

Mar 5

Mar 5Thu
  1. Nick TurleyXAI score62

    GPT-5.4 Thinking rolls out to ChatGPT with mid-response interrupts

    AIGPT-5.4 Thinking is rolling out to ChatGPT, and users can now interrupt it before it produces the final answer. Users can steer the response while it is still working rather than sending multiple follow-up turns. The update also improves deep web research and long-context reasoning, which the post says helps specific questions arrive faster and stay focused.

    Why it matters: The post names the new interrupt control and the research and long-context gains, showing how this change affects steering responses in ChatGPT.

Mar 3

Mar 3Tue
  1. Nick TurleyXAI score62

    OpenAI rolls out GPT-5.3 Instant in ChatGPT with fewer refusals and disclaimers

    AIOpenAI's Nick Turley announced that GPT-5.3 Instant is rolling out in ChatGPT starting today. The update responds to feedback that GPT-5.2 was sometimes too cautious, over-caveated, and less natural in conversation, with fewer unnecessary refusals, fewer defensive disclaimers, and more direct answers.

    Why it matters: The post names the specific complaints about GPT-5.2 and the behavior changes made in response, which shows how user feedback shaped this update.

Feb 20

Feb 20Fri
  1. Jim FanXAI score75

    DreamDojo: Open-source world model trained on 44K hours of human video

    AIJim Fan announced DreamDojo, an open-source interactive world model that takes robot motor controls and generates future frames in pixels. It is pre-trained on 44K hours of human egocentric video using latent actions, then post-trained onto specific robot hardware, and a real-time version runs at 10 FPS for live teleoperation, policy evaluation, and model-based planning. The author reports a +17% real-world success gain on a fruit packing task, and weights, code, datasets, and the whitepaper are released.

    Why it matters: The post explains how human egocentric videos are converted into latent actions, a method that reduces dependence on robot-collected data for training world models.

    Video from @DrJimFan's post

Feb 4

Feb 4Wed
  1. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Mini Transcribe 2 and realtime transcription with open weights

    AIMistral announces Mini Transcribe 2 via API at $0.003 per minute and a realtime transcription option at $0.006 per minute. The realtime model's open weights are published on Hugging Face, alongside a realtime demo and a blog post.

    Why it matters: The post gives per-minute API prices and open weights for a realtime tier, letting developers compare transcription costs against existing speech-to-text options.

Jan 7

Jan 7Wed
  1. Nick TurleyXAI score72

    OpenAI launches ChatGPT Health for connecting medical records

    AIOpenAI is launching ChatGPT Health, a dedicated and private space where users can securely connect apps and medical records. The launch starts with a small group of users from the waitlist, with access expanding over the coming weeks.

    Why it matters: The post names the access path and a dedicated space for health records, which matters for judging how sensitive data would be handled.

Dec 16, 2025

Dec 16, 2025Tue
  1. Nick TurleyXAI score60

    OpenAI rolls out new ChatGPT Images with faster, more precise editing

    AIOpenAI's new ChatGPT Images is rolling out in ChatGPT starting today. The update offers more precise edits, stronger instruction following, and up to 4x faster generation while preserving lighting, composition, and likeness across edits.

    Why it matters: The source names the specific editing gains and the speed figure, helping readers judge whether the new image tool changes their editing workflow.

    Image from @nickaturley's post

Dec 4, 2025

Dec 4, 2025Thu
  1. Quoc LeXAI score62

    Gemini 3 Deep Think mode goes live in the Gemini app for Ultra users

    AIGoogle's Gemini 3 Deep Think mode is now available in the Gemini app for Ultra users. The post says it uses parallel thinking for difficult coding and scientific tasks and builds on technology that reached gold-medal level at the ICPC World Finals and IMO.

    Why it matters: The post names the access tier and the coding and scientific task focus, which helps readers judge whether the mode fits their work.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.