Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Sep 16

Sep 16Wed
  1. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  2. Baseten BlogOfficialAI score54

    Baseten launches Hosted Tools with web search for open-source models

    AIBaseten has launched Hosted Tools, starting with Baseten Grounded Inference, a server-side web search capability for models hosted on Baseten. Developers enable it by adding a hosted search tool to a Messages, Chat Completions, or Responses request, and the platform runs the search loop with partners Exa, Keenable, Parallel, and You.com. In Baseten's benchmarks, agents using the hosted tools saw a 15% reduction in end-to-end latency compared with client-side tools, and the feature is in playground preview with 25 RPM rate limits and $2 of free credits.

  3. catXAI score60

    Claude merges Cowork and chat into one product with automatic routing

    AIAnthropic is merging Claude Cowork and chat into one Claude, and Claude Design is integrated so users can ask for slides, designs, or docs without switching apps. Claude decides from the prompt whether to give a quick answer or do deeper agentic work, and users can still stop, redirect, or adjust its effort. The change rolls out to Pro and Max over the next few weeks.

  4. Mike KriegerXAI score46

    Claude Cowork and Chat Merge into One Unified Claude

    AIAnthropic is merging Claude Cowork and Chat into a single Claude starting today, which Mike Krieger says removes the friction of choosing which product to start with. Per the @claudeai announcement, Claude will carry tasks forward even after the laptop is closed, asking for clarification when needed while users keep final say. The rollout to Pro and Max plans will take place over the coming weeks.

  5. Felix RiesebergXAI score28

    Anthropic's Cowork is now Claude, per a team announcement

    AIFelix Rieseberg, of Anthropic, announced that Cowork is now Claude, linking to a blog post. He framed the change as simplifying tools while exposing users to more capability, and invited users to bring large, messy projects.

  6. Felix RiesebergXAI score36

    Anthropic adds Claude Design, Docs, and Slides to conversations

    AIAnthropic says Claude Design, Claude Docs, and Claude Slides are now available within conversations. Users can co-write documents, have Claude draft slide decks that can be edited, presented from Claude, or downloaded as PowerPoint, and use Claude Design to generate interfaces and layouts.

  7. Felix RiesebergXAI score62

    Claude merges Cowork and Chat into one conversation with local and cloud work

    AIAnthropic's Felix Rieseberg announced that Cowork and Chat are now combined into a single Claude experience. Users can ask quick questions and move into serious work in the same conversation, with Claude using local files and apps when the computer is open and its own cloud computer when it is closed.

Sep 15

Sep 15Tue
  1. Noah ZwebenXAI score17

    Anthropic offers Claude Tag office hours for on-call triage feedback

    AIAnthropic is hosting office hours for teams interested in using Claude Tag for on-call work, and it is asking Team or Enterprise plan users to share triage feedback. Claude Tag can start investigating when a Slack alert fires by pulling metrics, diffing deploys, and checking flags to propose a likely cause and fix. Sign-up is through a Google Calendar booking link.

  2. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  3. Claude Apps Release NotesOfficialAI score72

    Claude Cowork moves into every conversation, adding designs, slides, and docs

    AIClaude now makes Cowork capabilities available from any conversation without choosing a mode first, with chats, tasks, projects, connectors, and skills carrying over. Users can also create designs, decks, and docs in any conversation, including Claude Code and the Artifacts tab, and edit them with Claude.

    Why it matters: The release merges Cowork tasks into ordinary chats and adds design, slide, and doc creation, changing how Claude users start larger work.

  4. Jazzyear · InsightsNewsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.

  5. Google Developers BlogOfficialAI score46

    Google Launches Agent Anomaly Detection in Private Preview on Gemini Enterprise Agent Platform

    AIGoogle has put Agent Anomaly Detection into Private Preview on the Gemini Enterprise Agent Platform, a reasoning-based audit layer that reviews agent reasoning traces, tool calls, and execution flow to flag behavioral anomalies and policy violations. It runs asynchronously without adding runtime latency and publishes findings to Security Command Center. The preview requires ADK 1.2 or later.

  6. xAI News (Grok)OfficialAI score43

    Grok Build Adds Memory That Saves Project Notes Between Sessions

    AIGrok Build now has memory, which records conventions, decisions, and project facts after each completed turn and reads them in later sessions. Notes are stored per project plus a global set, and the /dream command organizes them into topic files while /memory opens a read-only browser. The feature is available now and applies to new sessions.

  7. Google AntigravityOfficialAI score37

    Antigravity adds permissions system and sandboxed command execution

    AIGoogle Antigravity says its new permissions system reduces the number of commands users must approve, as improvements to its sandbox let commands run automatically in an isolated environment. The sandbox has no network access by default, keeping the user's machine protected while the agent can do more out of the box. The update is rolling out today on macOS and Linux.

    Image from @antigravity's post
  8. Mark ZuckerbergXAI score30

    Zuckerberg says labs should prioritize alignment and safety as core capabilities.

    AIMark Zuckerberg argues that every AI lab has both the incentive and responsibility to train models safely, since users will reject misaligned agents and labs face liability for harm. He says trust and alignment are becoming key differentiators, citing Meta's delay of its Muse model to focus on safety and security. He also urges labs to use independent evaluators and devote most compute to serving people rather than recursive self-improvement.

  9. Google AI StudioOfficialAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  10. VercelOfficialAI score22

    Delphi ships 100+ deploys daily on Vercel's Python backend

    AIDelphi, which turns experts' knowledge into digital minds, runs its Python backend on Vercel with a 10-person team and no dedicated infrastructure role. The stack uses Vercel Workflows for long-running agents and Vercel Queues for background jobs. The team reports more than 100 production deploys a day.

  11. Microsoft Foundry BlogOfficialAI score32

    Microsoft Launches Foundry Dev Pack to Install Foundry Development Tools in One Command

    AIMicrosoft has launched Foundry Dev Pack, an all-in-one installer that sets up tools for Microsoft Foundry development across the terminal, IDE, and coding agents. Depending on the environment, it installs Azure CLI (az), Azure Developer CLI (azd) with the Microsoft Foundry Extension for azd, the Microsoft Foundry Skill, the Microsoft Foundry Toolkit for Visual Studio Code, and Foundry Canvas (preview), with the last two conditional on VS Code or GitHub Copilot App being present.

  12. Greg BrockmanXAI score46

    ChatGPT Work adds Data agent for dashboards and actions on company data

    AIOpenAI's Greg Brockman says ChatGPT Work can operate over and act on a company's data, including building dashboards, by connecting existing tools such as PowerBI, Tableau, Clickhouse, Oracle BI, and AWS Redshift. The linked ChatGPT announcement describes a Data agent with a Data Plugin that turns company data into answers, interactive dashboards, and actions through conversation.

  13. Google AIOfficialAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  14. Google DeepMindOfficialAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  15. Google AI StudioOfficialAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  16. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

  17. OdysseyOfficialAI score22

    Odyssey-3 aims to enable physical agents that interact with the world

    AIOdyssey says its new Odyssey-3 model could enable physical agents, a new kind of agent that interfaces natively with physical and virtual systems. The company expressed excitement about the promise its models are showing, linking to an introduction page for Odyssey-3.

  18. Lovable BlogOfficialAI score44

    Lovable and Salesforce partner so teams can build apps inside Salesforce workflows

    AILovable and Salesforce are working together so apps and agents built with Lovable can read and write Salesforce data through Headless 360, using each user's own Salesforce permissions. Teams can publish read-only apps into Salesforce, mention @Lovable in Slack to build apps, and install agents into a Slack workspace.

  19. Baseten BlogOfficialAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  20. Air Street PressBlogAI score39

    Air Street Capital leads $40 million Series A in Jack & Jill, an AI career agent platform

    AIAir Street Capital led Jack & Jill's $40 million Series A, with Madrona joining and Creandum and Entrepreneurs First investing again, following a $20 million seed round less than a year earlier. Jack & Jill uses AI agents named Jack, which helps candidates plan career moves and search job postings, and Jill, which helps companies recruit from opted-in candidates. The company says it has arranged 25,000 interviews and plans 5,000 more each month.

  21. Sebastian RaschkaXAI score28

    GPT-5.6 Astra and Qwen3.8 Max take different Paint approaches

    AIIn a Paint recreation test, GPT-5.6 Astra built the image from layered geometric shapes, while Qwen3.8 Max worked pixel by pixel. Qwen's output looks closer to the original, but Raschka argues this single example does not show either model generalizes better or has stronger computer-use or visual understanding, and it illustrates how benchmarks comparing only final results can be misleading.

    Video from @rasbt's post
  22. Kilo (acq. by Anaconda)OfficialAI score22

    Kilo App launches on Product Hunt for iOS and Android

    AIKilo announces that its Kilo App is live on Product Hunt, letting users start coding agents, check sessions, and review pull requests from iOS and Android. The company asks supporters to upvote or comment on its Product Hunt listing.

    Image from @kilocode's post
  23. MiniMax Design (H3)OfficialAI score26

    MiniMax Design canvas runs Astra agent and Blender to produce full scene

    AIMiniMax Design lets a single brief drive a full production workflow on one canvas, with the Astra agent working in Blender through an official connector to build the scene and camera direction. The final video is generated with MiniMax H3 from the same canvas, with outputs syncing directly onto the canvas.

Sep 14

Sep 14Mon
  1. Factory NewsOfficialAI score40

    Factory raises $200M at $5B valuation to scale self-improving enterprise software development

    AIFactory has raised $200M at a $5B valuation from investors including Blackstone, Khosla Ventures, and Sequoia Capital, bringing its total funding to over $400 million. The company says it will use the capital to accelerate research, product, and global go-to-market efforts. Factory says hundreds of thousands of developers use its platform, with customers including Nvidia, Blackstone, and T-Mobile.

  2. Google Developers BlogOfficialAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  3. Claude Apps Release NotesOfficialAI score46

    Anthropic Launches Salesforce Plugin for Claude in Beta

    AIAnthropic has launched a Salesforce plugin for Claude that brings sellers' accounts, opportunities, and pipeline into the Claude app, with 37 pre-built sales skills. The beta is available on all paid plans for organizations Salesforce approves through its beta sign-up.