Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 18

Sep 18Fri
  1. Noam BrownXAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  2. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  3. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

  4. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

Sep 17

Sep 17Thu
  1. KrASIA · Big TechNewsAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  2. Felix RiesebergXAI score40

    Claude builds a multiplayer game from a single request in demo

    AIAnthropic's Felix Rieseberg posted a 60-second demo showing Claude, with built-in Cowork and connected Artifacts, building a multiplayer game when asked. He notes users can also ask Claude to copy a shared Artifact so they can play with others on their own account.

    Image from @felixrieseberg's post
  3. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  4. AI at MetaOfficialAI score44

    Meta's Muse agent now available on Mac for local tasks

    AIMeta is rolling out Muse for Mac today, a personal agent that can complete tasks directly on the user's computer with explicit permission. Examples include organizing the downloads folder, finding lost files, and summarizing messages and notes, with more capabilities coming soon.

  5. LM StudioOfficialAI score44

    LM Studio adds session history search and @ session references

    AILM Studio's new Introspection feature lets its Bionic agent search its own session history, improving handling of long-term context across multiple compactions. Users can also reference other sessions directly in the composer with an @ mention.

    Video from @lmstudio's post
  6. AnthropicOfficialAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  7. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  8. Josh WoodwardXAI score40

    Google Labs launches CC, an AI agent for family logistics

    AIGoogle Labs has announced CC, an AI agent built for families that can be connected to up to 5 members. It syncs schedules and to-dos through shared Google Calendar and Tasks, sends a shared "Your Day Ahead" brief each morning, and handles tasks such as meal plans, shopping lists, and paperwork under user direction. It is available by waitlist or upgrade in the US for users 18 and older.

  9. Google LabsOfficialAI score28

    Google Labs launches CC, a family AI agent for shared logistics

    AIGoogle Labs announced CC, an AI agent built for families to handle scheduling, to-dos, and errands. It supports up to 5 family members, sends a shared "Your Day Ahead" brief email each morning, and syncs schedules and tasks with Google Calendar and Tasks. Access is via waitlist or upgrade, limited to the US and users 18 and older.

    Video from @GoogleLabs's post
  10. Google LabsOfficialAI score44

    Google Labs' CC agent expands to families, sharing one daily brief and calendar across up to six members

    AIGoogle Labs has turned its experimental CC agent into a family and household assistant that supports up to six members, each with a shared view of the day ahead. CC has its own Google account, sees only what members choose to share, and connects to Calendar and Tasks. It is available as an early experiment on web and mobile for U.S. users 18 and older with a personal Google account.

  11. Noah ZwebenXAI score62

    Claude Code adds Projects that run parallel threads from one conversation

    AIAnthropic's Claude Code now runs projects from a single conversation, where Claude directs parallel threads that keep working after the user closes their laptop. The feature is in beta for select Pro and Max users in cloud sessions, with wider availability for all Claude users promised soon.

    Why it matters: The quoted launch replaces scattered sessions with one coordinator that runs parallel threads in the background, a workflow change worth weighing for complex projects.

  12. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

    Why it matters: The post explains how Projects shifts work from managing single sessions to coordinating many parallel tasks, a change that affects how Claude Code users plan and track their work.

  13. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  14. Dwarkesh PatelXAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

    Video from @dwarkesh_sp's post
  15. Dwarkesh PodcastBlogAI score63

    Noam Brown discusses agent swarms, alignment, and recursive self-improvement

    AIDwarkesh Patel interviews OpenAI researcher Noam Brown on multi-agent systems, math progress, and alignment. Brown says a 10,000-agent system solved a Millennium Prize Problem over 88 hours using 130 billion tokens, but he attributes most of that result to the underlying model rather than multi-agent design. The episode also covers the Hugging Face incident, in which agents coordinated in unintended ways, and how alignment might be verified before recursive self-improvement begins.

  16. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  17. Baidu Inc.OfficialAI score22

    Apollo Go plans to seek commercial autonomous driving approval in Hong Kong

    AIBaidu's Apollo Go plans to apply for commercial operation of its autonomous driving service in Hong Kong, citing the HKSAR Government's support in its first Five-Year Plan and the 2026 Policy Address. The company says it builds on fully driverless trials already conducted in the city. Baidu hopes Hong Kong can become a global benchmark for commercial autonomous driving in right-hand-drive markets.

  18. KrASIA · Big TechNewsAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

  19. Gemini API ChangelogOfficialAI score38

    Antigravity Agent 09-2026 replaces 05-2026 with new built-in file and search tools

    AIGoogle released the antigravity-preview-09-2026 agent, which replaces and deprecates antigravity-preview-05-2026. Remote sandbox users reading only output_text or model_output steps need only update the agent string, while local-environment users or those parsing function_call steps must adapt to renamed tools, PascalCase parameters, and line-range file edits. The 05-2026 preview shuts down on October 5, 2026.

Sep 16

Sep 16Wed
  1. hardmaruXAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Google Developers BlogOfficialAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  3. Greg BrockmanXAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

    Why it matters: The post reports internal rollout data on cost and performance, showing how a company manages model access and budgets for engineers at scale.

  4. Latent.SpaceXAI score38

    AIUC cofounder on AI agent risk, insurance, and standards

    AIAI Underwriting Company cofounder Rune Kvist argues that risk and trust may become the main bottlenecks to AI adoption. He discusses stress-testing agents for jailbreaks, hallucinations, and data leaks, why standards and insurance must evolve together, and why AI labs cannot fully act as their own watchdogs.

    Video from @latentspacepod's post
  5. Perplexity DevelopersOfficialAI score34

    Perplexity's Search SDK extracts query-relevant passages from URLs for agents

    AIPerplexity says its Search SDK extracts passages relevant to a query from user-provided URLs. Agents can use those passages instead of full pages, keeping unrelated content out of the model context. The company also points to an Agent Skill for installing the Search SDK in coding agents.

  6. Perplexity DevelopersOfficialAI score21

    Perplexity releases a Search SDK cookbook for coding agents

    AIPerplexity has published a new cookbook for its Search SDK, showing how to run focused searches and filter results to official documentation. The recipe extracts relevant passages and produces a source-linked brief that a coding agent can use.

    Video from @perplexitydevs's post
  7. Google for DevelopersOfficialAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

  8. Alex AlbertXAI score45

    Claude Cowork and chat merge into one unified interface

    AIAnthropic is merging Claude Cowork and chat into a single Claude experience, which Alex Albert says feels much better than either product alone. He also highlights the new slides, docs, and design integrations as working very well. The merged version is rolling out to Pro and Max users over the next few weeks.

  9. Kilo (acq. by Anaconda)OfficialAI score40

    Kilo Mobile lets users run full AI agent loops from their phone

    AIKilo Mobile now lets users spawn Cloud Agents, start sessions on remote machines, and dictate prompts by voice from a phone. Users can also review and comment on pull requests and approve Security Agent remediations without a laptop. On iPhone, Live Activities show session status on the Lock Screen when an agent needs input.

    Image from @kilocode's post
  10. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  11. Baseten BlogOfficialAI score54

    Baseten launches Hosted Tools with web search for open-source models

    AIBaseten has launched Hosted Tools, starting with Baseten Grounded Inference, a server-side web search capability for models hosted on Baseten. Developers enable it by adding a hosted search tool to a Messages, Chat Completions, or Responses request, and the platform runs the search loop with partners Exa, Keenable, Parallel, and You.com. In Baseten's benchmarks, agents using the hosted tools saw a 15% reduction in end-to-end latency compared with client-side tools, and the feature is in playground preview with 25 RPM rate limits and $2 of free credits.

  12. catXAI score60

    Claude merges Cowork and chat into one product with automatic routing

    AIAnthropic is merging Claude Cowork and chat into one Claude, and Claude Design is integrated so users can ask for slides, designs, or docs without switching apps. Claude decides from the prompt whether to give a quick answer or do deeper agentic work, and users can still stop, redirect, or adjust its effort. The change rolls out to Pro and Max over the next few weeks.

    Why it matters: The source states how Claude routes between quick answers and deeper agentic work, useful for anyone deciding which Claude product fits a task.

  13. Mike KriegerXAI score46

    Claude Cowork and Chat Merge into One Unified Claude

    AIAnthropic is merging Claude Cowork and Chat into a single Claude starting today, which Mike Krieger says removes the friction of choosing which product to start with. Per the @claudeai announcement, Claude will carry tasks forward even after the laptop is closed, asking for clarification when needed while users keep final say. The rollout to Pro and Max plans will take place over the coming weeks.