Skip to content

Formats

Product updates Latest news

New features, redesigns, and commercial changes in AI products and applications.

136 picksPast 30 days: 78 itemsTotal: 2,184 items

Latest pick

Top picks archive · Page 4

Sep 22

Sep 22TueItems 61–80
  1. Noam BrownAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    OpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    AIWhy it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

  2. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    AIWhy it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  3. AnthropicAI score71

    Anthropic releases Claude Opus 5.5, the first model in its Claude 5.5 family

    Anthropic has made Claude Opus 5.5 available today, introducing it as the first model in its new Claude 5.5 family. According to the quoted @claudeai post, it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

    AIWhy it matters: The post gives a concrete cost comparison against Opus 5 and names the model family, helping readers gauge the trade-off between price and performance.

  4. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    Google made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. vLLM BlogAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    vllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    AIWhy it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  2. LMSYS OrgAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    LMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    AIWhy it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

Sep 17

Sep 17Thu
  1. Google AI StudioAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    Google AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    AIWhy it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

Sep 15

Sep 15Tue
  1. Zed BlogAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    Zed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    AIWhy it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. Claude Apps Release NotesAI score72

    Claude Cowork moves into every conversation, adding designs, slides, and docs

    Claude now makes Cowork capabilities available from any conversation without choosing a mode first, with chats, tasks, projects, connectors, and skills carrying over. Users can also create designs, decks, and docs in any conversation, including Claude Code and the Artifacts tab, and edit them with Claude.

    AIWhy it matters: The release merges Cowork tasks into ordinary chats and adds design, slide, and doc creation, changing how Claude users start larger work.

  3. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    Google AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    AIWhy it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  4. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    Google is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    AIWhy it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  5. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    Google has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    AIWhy it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 10

Sep 10Thu
  1. Sherwin WuAI score72

    OpenAI launches Agents API for building cloud agents on Codex harness

    OpenAI has launched the Agents API, a cloud-based way to build agents backed by the Codex harness. Developers can connect their favorite tools and connectors and attach agents to any sandbox. The author says the API lets firms build scaled agents and expose them inside their own internal AI applications.

    AIWhy it matters: The source describes an API for building cloud agents on the Codex harness, useful for teams planning to embed agents in internal applications.

  2. Greg BrockmanAI score72

    GPT-Live-1 becomes available in the OpenAI API for voice agents

    GPT-Live-1 is now available in the API, letting developers bring ChatGPT-style back-and-forth conversation into their apps. The quoted announcement says the voice agents can listen while they speak and can work with the models and harness developers choose.

    AIWhy it matters: The quoted announcement describes a real-time voice model entering the API, which matters for builders weighing voice agents against existing stacks.

  3. DeepSeekAI score72

    DeepSeek V4.1-Flash goes live on its API with native multimodal support

    DeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.

Sep 9

Sep 9Wed
  1. Cursor ChangelogAI score73

    Cursor launches Projects for long-running, multi-agent coding work

    Cursor is launching Projects, a beta feature for larger work such as a feature, migration, or full app, rolling out to all users starting today. A coordinator agent plans the work, delegates it to implementing agents that can run in parallel, and runs on a cloud computer so it continues when the laptop is closed. Each Project keeps shared context files synced across cloud and local machines, and subscriptions let the coordinator act on Slack channels, schedules, or PRs without a prompt.

    AIWhy it matters: The source details how a coordinator agent plans, delegates, and syncs shared context across cloud and local machines, useful for judging how long-running agent work might fit a team's workflow.

  2. Microsoft Foundry BlogAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    Microsoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    AIWhy it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score72

    Google releases ADK for Kotlin 1.0 for building production AI agents

    Google announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.

    AIWhy it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.

  2. AI at MetaAI score67

    Meta introduces Muse, a personal agent powered by Muse Spark 1.3

    Meta announced Muse, a personal AI agent designed to get things done for users across many parts of life. The product is powered by Muse Spark 1.3, and the post links to an app download and a page describing how Muse was built.

    AIWhy it matters: The announcement names Muse Spark 1.3 as the underlying model, giving readers a concrete product and model pairing to track.