Skip to contentSkip to stories
Updated

#MCP/Tool use

Oct 9

TodayOct 9Fri
  1. Sierra BlogOfficialAI score62

    Sierra publishes draft Personal Agent Protocol, called Poppy, with 35 new design partners

    AISierra has published a draft of the Personal Agent Protocol, known as Poppy, and named 35 additional design partners, including Adyen, Bank of America, Mastercard, OpenAI, PayPal, and Visa. Under the protocol, companies publish a /.well-known/poppy.json discovery file, and personal agents start sessions, identify themselves, and sign in through OAuth with session tokens limited to approved access. The company says the draft will be followed by design workshops and a reference implementation over the next month.

    Why it matters: The draft specifies how personal agents identify themselves, obtain customer-approved access, and work with company websites, APIs, or agents, which helps readers assess its practical effect on agent-driven transactions.

Oct 8

Oct 8Thu
  1. LangChain BlogOfficialAI score67

    LangChain's Restock agent shows how to build a payment-capable AI agent

    AILangChain built Restock, a sample office-supply agent that runs in Slack on Managed Deep Agents and pays through Stripe's Link wallet. The agent searches products, builds a cart, and pays over the Machine Payments Protocol, with the user approving the purchase in Slack and the payment in Link. The post uses a pens order at $22.18 to show the flow from request to confirmed order.

    Why it matters: The post walks through how an agent handles search, budget limits, Slack review, and Link approval, showing where each control sits outside the model.

Oct 7

Oct 7Wed
  1. LangChain BlogOfficialAI score63

    Managed Deep Agents v0.9 adds agent schedules, per-run configuration, and Slack reactions

    AILangChain released Managed Deep Agents v0.9 in Public Beta, adding a Schedules SDK, per-run agent configuration, and Slack reactions. Agents can create reminders, follow-ups, and recurring tasks mid-conversation, running as the requesting user and posting results back to the originating channel. Per-run configuration lets one deployment choose the model, instructions, skills, MCP servers, and sandbox based on the run's context, and Slack reactions are on by default with a 👀 emoji.

    Why it matters: The release shows how one agent deployment can be configured per run by channel or repo, separating tool access from model instructions.

Oct 6

Oct 6Tue
  1. Sierra BlogOfficialAI score62

    Sierra and Meta announce Personal Agent Protocol, an open standard for personal agents

    AISierra and Meta are developing Personal Agent Protocol, an open standard defining how personal agents interact with businesses, with industry partners including Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. The protocol uses OAuth sessions where consumers choose read-only or write access and companies choose whether agents reach them through websites, APIs via MCP and OpenAPI, or their own agents. The authors plan to publish the v0.1 specification later this month along with a reference implementation.

    Why it matters: The post specifies how personal agents would authenticate and reach businesses through websites, APIs, or company agents, which matters for anyone building agent integrations.

Oct 3

Oct 3Sat
  1. Hugging Face BlogOfficialAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Sep 29

Sep 29Tue
  1. Replit BlogOfficialAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

Sep 28

Sep 28Mon
  1. catXAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

    Video from @_catwu's post

Sep 24

Sep 24Thu
  1. Azure BlogOfficialAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  2. Microsoft Foundry BlogOfficialAI score61

    Microsoft Foundry Routines reach general availability for scheduled and event-driven agents

    AIMicrosoft announced general availability of Routines in Foundry Agent Service, a managed way to run agents on a timer, on a recurring schedule, or in response to GitHub issue events and new Microsoft Teams channel messages. Routines keep the trigger, agent action, identity, connections, and run history in the Foundry project, and each routine can run under the creator's identity or the agent's own Microsoft Entra ID identity. A preview reminder tool lets a Hosted Agent schedule itself to resume later on the same conversation.

    Why it matters: The post explains how scheduled, event-based, and self-reminding agent runs are managed in one place, along with the creator versus agent identity choice for unattended tasks.

Sep 23

Sep 23Wed
  1. Google Developers BlogOfficialAI score62

    Google Cloud API Gateway can now expose REST APIs as MCP tools in preview

    AIGoogle Cloud API Gateway now acts as a remote MCP server in Public Preview, making REST operations in an annotated OpenAPI 3.0.x or 3.1.x spec available as agent-ready MCP tools. Existing JWT or API-key authentication, quotas, and logging apply to MCP calls, so teams do not need a separate MCP server. Current limits include no support for OpenAPI 2.0, a maximum of 1,000 tools per gateway, and no MCP and model routing in the same API config.

    Why it matters: The post shows how an existing OpenAPI spec becomes agent-callable MCP tools, with the same auth and quota policies applied, which helps teams avoid building a separate MCP server.

  2. Greg BrockmanXAI score67

    ChatGPT Voice gains plugins and ChatGPT Work access on web and mobile

    AIChatGPT Voice can now use plugins such as email, calendar, and Slack, and it can be powered by GPT-6 Astra, Sol, and Luna. It is also available in ChatGPT Work on web and mobile for creating docs, decks, sites, and spreadsheets by voice, rolling out globally in the latest app version.

    Why it matters: The quoted OpenAI post names the new voice tool access, supported models, and Work integration, showing how voice now acts across workflows.

Sep 17

Sep 17Thu
  1. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

Sep 15

Sep 15Tue
  1. Google AI StudioOfficialAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

Sep 9

Sep 9Wed
  1. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

That’s everything