Skip to content

Areas

AI agents Latest news

Models that plan, use tools, and complete multistep tasks, from Claude Code and Manus to agent frameworks and evaluations.

201 picksPast 30 days: 88 itemsTotal: 1,570 items

Latest pick

Top picks archive · Page 4

Sep 23

Sep 23WedItems 61–80
  1. eric zakariassonAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    Cursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    AIWhy it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  2. Anthropic NewsroomAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    Anthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    AIWhy it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

  3. Prime Intellect BlogAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    Prime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    AIWhy it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Google Developers BlogAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    Google announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

    AIWhy it matters: The source shows how to run an agent with a local Gemma 4 26B model using LiteRT, plus a hybrid cloud-planner pattern that keeps most tokens on-device.

  2. Tibor BlahoAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    OpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    AIWhy it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

  3. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    AIWhy it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

Sep 21

Sep 21Mon
  1. vLLM BlogAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    vllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    AIWhy it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  2. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    Xiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    AIWhy it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  3. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    Xiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    AIWhy it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  4. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    Xiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    AIWhy it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

Sep 20

Sep 20Sun
  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    xAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    AIWhy it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 17

Sep 17Thu
  1. Google AI StudioAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    Google AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    AIWhy it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

Sep 15

Sep 15Tue
  1. Zed BlogAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    Zed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    AIWhy it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. Claude Apps Release NotesAI score72

    Claude Cowork moves into every conversation, adding designs, slides, and docs

    Claude now makes Cowork capabilities available from any conversation without choosing a mode first, with chats, tasks, projects, connectors, and skills carrying over. Users can also create designs, decks, and docs in any conversation, including Claude Code and the Artifacts tab, and edit them with Claude.

    AIWhy it matters: The release merges Cowork tasks into ordinary chats and adds design, slide, and doc creation, changing how Claude users start larger work.

  3. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    Google AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    AIWhy it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  4. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    Google is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    AIWhy it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  5. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    AIWhy it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  6. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    AIWhy it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  7. Cognition Blog (Devin, Windsurf)AI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    Cognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    AIWhy it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

Sep 14

Sep 14Mon
  1. Google Developers BlogAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    Part 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    AIWhy it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.