Skip to contentSkip to stories

Updated

#MCP/Tool use

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 5

Oct 5Mon
  1. SantiagoAI score47

    Tool generates synthetic companies to test AI agents across business systems

    AIA tool can turn a one-line business description into a complete synthetic company spread across CRM, ticketing, Slack, files, emails, and call recordings. Developers can test agents against this connected data, then reset the company to its initial state and rerun the test when something breaks. The background post describes the product as Era, a free simulated enterprise that connects to Salesforce, Slack, Jira, Zendesk, Gong, and Deel through live MCP and API interfaces.

  2. indigoAI score42

    Five-step Grok Bot method for hiring and managing AI agents

    AIBrian's Grok Bot method treats each bot like a new hire: define the role, test it on text first, run three trials, escalate based on evidence, and add a second agent only after a bottleneck appears. Each bot's role is defined by five fields: a real name with a short label, a one-line job tied to an outcome, what it owns, its inputs, and what it may do freely versus what it must ask before doing. The post frames an Agent Team as the final result of this process, starting with one coordinator and three specialists.

    Image from @indigox's post
  3. meng shaoAI score72

    Uber Designs an MCP Gateway to Expose Thousands of Internal APIs to AI Agents

    AIUber uses a control plane and data plane gateway to automatically convert its internal APIs into MCP tools, with 800+ MCP servers and 5,000+ tools hosted. The design includes an AutoCrawler that generates tool descriptions with an LLM, a default-disabled discover-not-expose security model, and techniques such as Omni MCP, Response Projection, and Code Mode to limit context bloat.

    Image from @shao__meng's post

Oct 4

Oct 4Sun
  1. OpenRouter BlogAI score44

    Server-Side Code Execution Tools for AI Agents, Compared

    AIOpenRouter's shell and bash tools, along with those from OpenAI and Anthropic, run an agent's commands in provider-managed sandboxes during the same API request, so developers don't provision or patch containers. OpenRouter's tools are in beta, with sandbox time billed at $0.0001 per second and a 30-second minimum for a new or sleeping container. The article compares the four providers and notes that self-run sandboxes remain better for custom base images, GPU work, or multi-hour sessions.

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

Oct 2Fri
  1. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  2. DatabricksAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  3. EveryAI score40

    How to Get Better at AI by Asking AI

    AIEvery's senior editor describes moving from single-thread chatbot prompting to delegating complex projects to teams of coordinating subagents, using skills, orchestrator threads, context packets, MCPs, and computer use. He says a subagent workflow verified employee equity costs across multiple grants, strike prices, and vesting schedules, and returned a draft Slack message for approval. The shift was prompted by a June tweet in which Codex placed a colleague at Level 5 of the "Eight Levels of AI Adoption" framework.

  4. Prime Intellect BlogAI score67

    Prime Inference launches serverless and reserved serving for open frontier models

    AIPrime Inference is a serving platform for frontier open-source models, offering serverless endpoints and reserved capacity on Prime's GPU infrastructure across multiple datacenters. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22, and the post reports a near-zero tool-call error rate and 100% uptime since launch. The post also describes GLM-5.3 serving on GB200 NVL72 with prefill/decode disaggregation and NVFP4 KV compression.

    Why it matters: The post separates scheduler, KV-cache, and tool-call fixes, showing concretely which bottlenecks shape production serving of open frontier models.

Oct 1

Oct 1Thu
  1. OpenRouter BlogAI score52

    How agent frameworks handle tool-calling schemas across model providers

    AITool definitions and tool-call responses differ between OpenAI, Anthropic, and Google, so a tool that works on one model may fail on another. The article compares six agent frameworks, including LangChain, CrewAI, and the OpenAI Agents SDK, by where each performs schema translation. It also describes OpenRouter's API-layer normalization, which accepts an OpenAI-style tools array and returns a standard tool_calls response for tool-capable models.

  2. Lydia Hallie ✨AI score62

    Claude Code adds mods that customize behavior and UI via TypeScript plugins

    AIClaude Code can now be modified with mods that change its behavior, customize the UI, and add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app. A TypeScript function can intercept internal events such as tool calls, prompts, model requests, and renders, and add custom UI and commands.

Sep 30

Sep 30Wed
  1. Guillermo RauchAI score30

    Vercel Connect invites services to reach developers and AI agents

    AIGuillermo Rauch invites service providers to add themselves to Vercel Connect to reach over 20 million developers and the agents they build. He argues that connecting services is now the main challenge in building, and that Connect makes it easier and more secure for both agents and apps. Services submit by describing themselves, adding OAuth or API key auth, verifying with a real token, and sending it for review.

  2. Google Cloud · AI & Machine LearningAI score45

    Google Cloud Launches Preview of CLI Remote MCP Server for AI Agents

    AIGoogle Cloud has introduced the Google Cloud CLI remote MCP server in preview, giving AI agents access to gcloud and bq command-line operations through two tools, run_gcloud_command and run_bq_command. The server runs in an isolated, network-restricted execution sandbox on Google Cloud, so teams need no local CLI binaries, and calls are authenticated through Agent Identity, OAuth 2.0, and IAM, with Model Armor screening and Audit Logs available.

  3. Karl's AI WattsAI score38

    Can you keep your session after switching models in magpie?

    AIKarl's AI Watts asks whether a menu-bar tool can switch models while preserving the existing conversation, so users avoid re-explaining their project each time. The post frames this as the reason they want to keep the menu bar tool, which the quoted post describes as magpie, a menu-bar switcher for 20+ agents including Claude Code and Codex that also offers a local gateway.

Sep 29

Sep 29Tue
  1. Replit BlogAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.