Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. Amp NewsOfficialAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  2. Google Developers BlogOfficialAI score62

    Google Cloud API Gateway can now expose REST APIs as MCP tools in preview

    AIGoogle Cloud API Gateway now acts as a remote MCP server in Public Preview, making REST operations in an annotated OpenAPI 3.0.x or 3.1.x spec available as agent-ready MCP tools. Existing JWT or API-key authentication, quotas, and logging apply to MCP calls, so teams do not need a separate MCP server. Current limits include no support for OpenAPI 2.0, a maximum of 1,000 tools per gateway, and no MCP and model routing in the same API config.

    Why it matters: The post shows how an existing OpenAPI spec becomes agent-callable MCP tools, with the same auth and quota policies applied, which helps teams avoid building a separate MCP server.

  3. Google GemmaOfficialAI score60

    Google's Antigravity SDK adds local execution with Gemma 4 and LiteRT

    AIGoogle says the Antigravity SDK now supports running agents entirely on a local machine with Gemma 4 and LiteRT. The post adds support for OpenAI-compatible endpoints, naming Ollama, llama.cpp, and vLLM as options for serving Gemma, and gives the install command pip install google-antigravity litert-lm.

    Why it matters: The post names the specific runtimes and serving endpoints supported, letting developers judge whether their current local setup fits the new SDK path.

    Video from @googlegemma's post
  4. eric zakariassonXAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  5. Google AntigravityOfficialAI score22

    Google Antigravity SDK adds support for local AI models

    AIGoogle announces support for local AI models in the Antigravity SDK, per a linked developer blog post. The post itself offers no further details, so specific features, supported models, or limits cannot be confirmed from this source.

  6. Redwood Research BlogBlogAI score71

    Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues

    AIRedwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.

    Why it matters: The post explains why chain-of-thought is a key oversight tool and how specific latent architectures could weaken it, useful for judging safety tradeoffs in future model design.

  7. LM StudioOfficialAI score28

    Bionic adds a built-in interactive canvas for shared diagrams

    AIBionic now includes a built-in interactive canvas where users can create Excalidraw diagrams that both they and Bionic can view and edit. The canvas supports collaboration on mockups, system designs, and process maps, and users can ask Bionic to implement what is drawn.

    Video from @lmstudio's post
  8. eric zakariassonXAI score36

    Optimizing reading for AI agents cuts context-gathering costs

    AIEric Zakariasson argues that agents spend heavily on reading context before and after work, so optimizing that reading makes a major difference. He recommends the linked guide to builders, or handing it to an agent to implement its findings. Cursor's related post reports 7% lower token costs with no drop in agent quality, achieved through tighter prompts, selective tool loading, better caching, and compressed file reads.

    Image from @ericzakariasson's post
  9. Google GeminiOfficialAI score12

    Gemini can draft and send PandaDoc contracts from a simple prompt

    AIGoogle Gemini now integrates with PandaDoc, letting users draft, customize, and send client-ready contracts. Users simply ask Gemini to create a document, such as a Services Agreement for Acme Corp with a $50,000 contract value, using PandaDoc.

    Image from @GeminiApp's post
  10. Google GeminiOfficialAI score10

    Gemini can create and update Monday.com tasks via natural language commands

    AIGoogle Gemini can manage projects on Monday.com when users give plain-language instructions, such as creating a task named "Finalize Q4 Marketing Plan" on a board and setting its status to "To Do." The post provides this single example and does not describe further capabilities or availability details.

    Image from @GeminiApp's post
  11. Azure BlogOfficialAI score40

    Azure resilience now requires continuous validation, not just architecture diagrams

    AIMicrosoft's Azure Blog argues that resilience drifts as workloads change, so architecture diagrams cannot prove a system is resilient. It says roughly 70 percent of cloud outages are related to change, and that teams need health modeling and resiliency goals measured against live signals. The article is the first in a series on validating resilience at scale.

  12. QwenOfficialAI score10

    Qwen's Mobile Creative Agent Performance Is Highlighted

    AIQwen shared performance results for its Mobile Creative Agent, though the post provides no specific figures or details. The brief announcement offers limited information beyond the performance label.

    Image from @Alibaba_Qwen's post
  13. QwenOfficialAI score17

    Qwen shares performance results for its mobile-use agent

    AIQwen posted performance results for its mobile-use agent, though the post itself provides no benchmark scores or details. The announcement is a brief teaser with no further specifics on the agent's capabilities.

    Image from @Alibaba_Qwen's post
  14. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Why it matters: The post names three mobile agents and their benchmark results, while also releasing the benchmark suite, so readers can check the claims against the reported figures.

    Image from @Alibaba_Qwen's post
  15. ModelScopeOfficialAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Why it matters: The post links benchmark results, parameter scale, and a multi-agent RL training run, giving readers concrete figures to compare against other open models.

    Image from @ModelScope2022's post
  16. Anthropic NewsroomOfficialAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

  17. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Redwood Research BlogBlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. PlatformerBlogAI score15

    Muse Is Having a Moment: Consumer AI Agents Explored

    AIPlatformer asks whether consumer AI agents are the future or a mirage, but the source text provided is only that single question. No further details about Muse's features, performance, or availability are available, so no additional claims can be made.

  3. Google Developers BlogOfficialAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    AIGoogle announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

  4. Tibor BlahoXAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

    Image from @btibor91's post
  5. eric zakariassonXAI score13

    Cursor's scrappy support system evolved into a scalable, self-improving operation

    AICursor's support team says a scrappy version built about 1.5 years ago to handle heavy volume taught them to crawl, walk, then run, and to identify flywheels for continuous improvement. The background post notes the company rebuilt customer support around Grok Bot to handle operations without adding headcount, with the bot responding to customers, resolving tickets, and managing the queue autonomously.

  6. Sierra BlogOfficialAI score34

    Sierra Lets Companies See, Edit, and Export Their AI Agents' Logic and Data

    AISierra says its platform makes enterprise AI agents visible and editable, with journeys, policies, and actions viewable in Agent Studio and testable through Simulations and Experiments before rollout. Customers can export agent logic in a portable structured format, access conversation logs and performance data through export APIs, and manage the agent's code in a Git repository. Sierra agents also connect to existing systems through MCP, REST, GraphQL, or custom integrations.

  7. ChatGPTOfficialAI score62

    OpenAI rolls out GPT-6 Sol and GPT-6 Luna in ChatGPT Work and Codex

    AIOpenAI announced GPT-6 Sol and GPT-6 Luna, rolling out today in ChatGPT Work and Codex. The rollout covers Plus, Pro, Business, Enterprise, and Edu users.

    Why it matters: The post names two new GPT-6 variants and their rollout to specific ChatGPT and Codex plan tiers, which shows how access is being staged.

    Video from @ChatGPT's post