Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Amazon ScienceOfficialAI score38

    Amazon Science explains graph-centric agentic AI for network root cause analysis

    AIAmazon Science says it built a root cause analysis approach that combines a network digital twin graph, cascaded graph algorithms, and agentic AI orchestration. The approach was demonstrated with NTT DOCOMO at Mobile World Congress, where it identified root causes in minutes on commercial networks. The article traces graph-based network modeling from topology and alarm correlation graphs to graph neural networks.

  2. JetBrains AI BlogOfficialAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  3. O'Reilly RadarBlogAI score38

    Conversational AI Interfaces May Matter More Than Full Autonomy for Software

    AIRobert Englander argues that natural language interfaces built on top of deterministic software may prove more valuable than fully autonomous AI agents. He contends that large language models excel at interpreting human intent, while systems of record must still provide the reliability, consistency, and accountability that probabilistic models lack.

  4. One Useful Thing (Ethan Mollick)BlogAI score62

    Ethan Mollick Says Agent Coordination Is Easier Than Expected

    AIEthan Mollick says he was wrong to think coordinating AI agents would require careful human-designed management structures. He points to personal agents like dots and Muse, and to a swarm of thousands of OpenAI agents that solved a Navier-Stokes problem in 88 hours with thin coordination. He argues many management problems stem from human limits, which agents lack, so people should mainly guide direction while agents handle organizing.

  5. Ahmad Al-DahleXAI score10

    Meta's Ahmad Al-Dahle posts brief update announcing new product changes

    AIAhmad Al-Dahle, who leads Meta's Llama work, shared a short post saying the team has released new updates. The post itself gives no concrete features, versions, or figures, and it is quoting a post from Airbnb CEO Brian Chesky that lists trip-idea sharing with friends, AI search, and laundry and baby gear rentals.

  6. Anthropic ResearchOfficialAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  7. Manus BlogOfficialAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

  8. Anthropic NewsroomOfficialAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  9. LangChain BlogOfficialAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed
  1. Hamel HusainXAI score4

    Hamel Husain thanks Lance Martin for updating Claude eval skill

    AIHamel Husain posted a short thank-you emoji reply to Lance Martin's update on an evaluation skill for Claude. Martin says the latest skill now instructs Claude to build a viewer for eval examples, but it does not walk users through the data first, which he agrees could help them prioritize which evals to write.

  2. Sakana AIOfficialAI score33

    Sakana AI's David Ha argues the future of AI lies in orchestrators

    AISakana AI co-founder and CEO David Ha published a Nikkei Asia op-ed titled "The future of AI belongs to the orchestrators." He argues that ever-larger models face limits, as open models close the gap within months and frontier inference costs can exceed the hourly wage of the people they assist. He also contends that sovereignty means supply-chain strength, not national isolation.

  3. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  4. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  5. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  6. Apple Machine Learning ResearchOfficialAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  7. NewcomerBlogAI score38

    Machine Earning Summit Debates Personal AI Agents and Agentic Commerce in San Francisco

    AIPersonal agents dominated the Machine Earning AI Summit in San Francisco, where founders and investors debated how AI agents will reshape finance and commerce. Speakers predicted that people will spend 40% of their digital time using assistants within a year, rising to 90% within five years, according to Town CEO Jean-Denis Greze. Panelists also stressed that consumers remain uncomfortable letting agents make purchases directly, with guardrails such as spend limits still being built.

  8. Amazon ScienceOfficialAI score7

    Amazon researchers present agentic AI and LLM papers at COLM 2026

    AIAmazon researchers will attend COLM 2026 in San Francisco next week with accepted publications on agentic AI, LLM post-training, multi-agent systems, and automated reasoning. The post links to their research and booth schedule.

  9. Yi TayXAI score46

    Gemini 4 Argon launches, reportedly outperforming astra and fable on many tasks

    AIGoogle DeepMind introduced Gemini 4 Argon, a new frontier model built for coding, enterprise knowledge work, and cybersecurity defense, rolling out to trusted testers through its Fairwind Program. Yi Tay says Gemini 4 outperforms astra and fable on many tasks, though the post gives no benchmark figures.

  10. Sundar PichaiXAI score62

    Google previews Gemini 4 Argon with frontier-level benchmark results

    AIGoogle's Sundar Pichai introduces Gemini 4 Argon, which he says shows frontier performance in complex workflows, cyber defense and software engineering. He says teams across Google are using it extensively, from coding to quantum computing, and shares benchmark results in the post.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and two Claude models across many task categories, which helps readers see where each model leads.

    Image from @sundarpichai's post
  11. Google DeepMindOfficialAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  12. Microsoft CopilotOfficialAI score23

    Microsoft's new Copilot combines Home, Code, and Autopilot in one app

    AIMicrosoft's new Copilot brings Home, Code, and Autopilot together in one place for creating custom apps, building decks, automating workflows, and resuming work. Users can start using the Copilot app now and try new features as they become available in Frontier.

    Image from @MSFTCopilot's post
  13. Google · Gemini appOfficialAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  14. eric zakariassonXAI score22

    Cursor promotes engineering bots in Grok bot marketplace

    AICursor's Eric Zakariasson announced engineering bots available in the Grok bot marketplace on x.ai. The bots can hand off coding tasks to Cursor, manage pull requests through GitHub and Origin plugins, and share video demos of what they build.

    Image from @ericzakariasson's post
  15. FireworksOfficialAI score34

    GLM 5.3 Flash now available for training on Fireworks' Serverless API

    AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.

  16. Google Cloud TechOfficialAI score28

    Agent Clinic Ep 3 builds automated eval suite for LangGraph agent

    AITerminal test runs miss multi-turn agent regressions, so Agent Clinic Episode 3 builds an automated eval suite for a LangGraph agent in 60 minutes. The post presents a four-step framework for moving from informal checks to benchmarking AI agents, with a link to the full guide.

    Image from @GoogleCloudTech's post
  17. LlamaIndex 🦙OfficialAI score14

    LlamaIndex hosts document-processing events for AI agents in New York and San Francisco

    AILlamaIndex held Tuesday-night events in New York and San Francisco on document processing for AI agents, with the New York room filling a waitlist and San Francisco drawing almost 600 attendees. The talks focused on the problem that agents often receive document text without its layout, so they must guess which figures, such as a monthly rate versus a total on an invoice, mean what.

    Image from @llama_index's post
  18. OpenClaw🦞OfficialAI score34

    OpenClaw v2026.9.7 adds OpenAI Agents API and ChatGPT sign-in

    AIOpenClaw released v2026.9.7 with faster performance under load, smoother long chats, and update backup and rollback improvements. The release adds OpenAI Agents API and ChatGPT sign-in in Beta, along with better Apple chat and restart recovery. It includes 2,818 PRs from 344 contributors.

  19. Higgsfield AI 🧩OfficialAI score37

    Higgsfield adds computer use via ChatGPT extension inside Codex

    AIHiggsfield's full interface is now available inside Codex, powered by GPT-6.1 Sol, with access to local files and automation skills. The integration lets users automate creative workflows end-to-end within ChatGPT.

    Video from @higgsfield's post
  20. Hacker News · Launch HN, YC launches (10+ points)BlogAI score62

    Magnitude launches an open source inference engine that tunes kernels to local hardware

    AIMagnitude is an open source inference engine for agents that compiles and tunes its kernels on the user's device before running a model. The source claims up to 2x faster decoding than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA, and says one click connects agents such as Pi, OpenCode, Codex, and Claude Code. It supports macOS, Windows, and Linux, and the source states that prompts and models stay on the user's machine.

  21. NVIDIA AIOfficialAI score40

    NVIDIA Shows Visual AI Agent Built in Under 30 Minutes

    AINVIDIA says a single prompt can build and deploy a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports. The method uses the new Build Vision AI skill in NVIDIA VSS Blueprint 3.3, and a tutorial is available for readers who want to build one.

    Video from @NVIDIAAI's post