Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Sep 18

Sep 18Fri
  1. Noam BrownXAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  2. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  3. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

  4. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

Sep 17

Sep 17Thu
  1. KrASIA · Big TechNewsAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  2. Felix RiesebergXAI score40

    Claude builds a multiplayer game from a single request in demo

    AIAnthropic's Felix Rieseberg posted a 60-second demo showing Claude, with built-in Cowork and connected Artifacts, building a multiplayer game when asked. He notes users can also ask Claude to copy a shared Artifact so they can play with others on their own account.

    Image from @felixrieseberg's post
  3. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  4. AI at MetaOfficialAI score44

    Meta's Muse agent now available on Mac for local tasks

    AIMeta is rolling out Muse for Mac today, a personal agent that can complete tasks directly on the user's computer with explicit permission. Examples include organizing the downloads folder, finding lost files, and summarizing messages and notes, with more capabilities coming soon.

  5. LM StudioOfficialAI score44

    LM Studio adds session history search and @ session references

    AILM Studio's new Introspection feature lets its Bionic agent search its own session history, improving handling of long-term context across multiple compactions. Users can also reference other sessions directly in the composer with an @ mention.

    Video from @lmstudio's post
  6. AnthropicOfficialAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  7. LlamaIndex 🦙OfficialAI score13

    LlamaIndex's Jerry Liu on document parsing challenges for enterprise agents

    AILlamaIndex CEO Jerry Liu spoke at Connected Stack's Founder Flash Talks about messy, complex documents that general-purpose models struggle to read, a problem enterprise agents eventually face. The company says it is building document infrastructure for agents, which it describes as the new knowledge workers.

    Image from @llama_index's post
  8. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  9. Josh WoodwardXAI score40

    Google Labs launches CC, an AI agent for family logistics

    AIGoogle Labs has announced CC, an AI agent built for families that can be connected to up to 5 members. It syncs schedules and to-dos through shared Google Calendar and Tasks, sends a shared "Your Day Ahead" brief each morning, and handles tasks such as meal plans, shopping lists, and paperwork under user direction. It is available by waitlist or upgrade in the US for users 18 and older.

  10. Google LabsOfficialAI score28

    Google Labs launches CC, a family AI agent for shared logistics

    AIGoogle Labs announced CC, an AI agent built for families to handle scheduling, to-dos, and errands. It supports up to 5 family members, sends a shared "Your Day Ahead" brief email each morning, and syncs schedules and tasks with Google Calendar and Tasks. Access is via waitlist or upgrade, limited to the US and users 18 and older.

    Video from @GoogleLabs's post
  11. Google LabsOfficialAI score44

    Google Labs' CC agent expands to families, sharing one daily brief and calendar across up to six members

    AIGoogle Labs has turned its experimental CC agent into a family and household assistant that supports up to six members, each with a shared view of the day ahead. CC has its own Google account, sees only what members choose to share, and connects to Calendar and Tasks. It is available as an early experiment on web and mobile for U.S. users 18 and older with a personal Google account.

  12. Noah ZwebenXAI score62

    Claude Code adds Projects that run parallel threads from one conversation

    AIAnthropic's Claude Code now runs projects from a single conversation, where Claude directs parallel threads that keep working after the user closes their laptop. The feature is in beta for select Pro and Max users in cloud sessions, with wider availability for all Claude users promised soon.

  13. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

  14. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  15. Dwarkesh PatelXAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

    Video from @dwarkesh_sp's post
  16. Dwarkesh PodcastBlogAI score63

    Noam Brown on Agent Swarms, Alignment, and Recursive Self-Improvement

    AINoam Brown discusses how running many agents in parallel scales test-time compute, citing a 10,000-agent effort on a Millennium Prize Problem. The conversation also covers whether models can be verified as aligned before recursive self-improvement begins, including the Hugging Face incident where agents cooperated in unintended ways.

  17. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  18. Baidu Inc.OfficialAI score22

    Apollo Go plans to seek commercial autonomous driving approval in Hong Kong

    AIBaidu's Apollo Go plans to apply for commercial operation of its autonomous driving service in Hong Kong, citing the HKSAR Government's support in its first Five-Year Plan and the 2026 Policy Address. The company says it builds on fully driverless trials already conducted in the city. Baidu hopes Hong Kong can become a global benchmark for commercial autonomous driving in right-hand-drive markets.

  19. KrASIA · Big TechNewsAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

  20. Gemini API ChangelogOfficialAI score38

    Antigravity Agent 09-2026 replaces 05-2026 with new built-in file and search tools

    AIGoogle released the antigravity-preview-09-2026 agent, which replaces and deprecates antigravity-preview-05-2026. Remote sandbox users reading only output_text or model_output steps need only update the agent string, while local-environment users or those parsing function_call steps must adapt to renamed tools, PascalCase parameters, and line-range file edits. The 05-2026 preview shuts down on October 5, 2026.

Sep 16

Sep 16Wed
  1. hardmaruXAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Google Developers BlogOfficialAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  3. Greg BrockmanXAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

  4. Latent.SpaceXAI score38

    AIUC cofounder on AI agent risk, insurance, and standards

    AIAI Underwriting Company cofounder Rune Kvist argues that risk and trust may become the main bottlenecks to AI adoption. He discusses stress-testing agents for jailbreaks, hallucinations, and data leaks, why standards and insurance must evolve together, and why AI labs cannot fully act as their own watchdogs.

    Video from @latentspacepod's post
  5. Perplexity DevelopersOfficialAI score34

    Perplexity's Search SDK extracts query-relevant passages from URLs for agents

    AIPerplexity says its Search SDK extracts passages relevant to a query from user-provided URLs. Agents can use those passages instead of full pages, keeping unrelated content out of the model context. The company also points to an Agent Skill for installing the Search SDK in coding agents.

  6. Perplexity DevelopersOfficialAI score21

    Perplexity releases a Search SDK cookbook for coding agents

    AIPerplexity has published a new cookbook for its Search SDK, showing how to run focused searches and filter results to official documentation. The recipe extracts relevant passages and produces a source-linked brief that a coding agent can use.

    Video from @perplexitydevs's post
  7. Google for DevelopersOfficialAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

  8. Alex AlbertXAI score45

    Claude Cowork and chat merge into one unified interface

    AIAnthropic is merging Claude Cowork and chat into a single Claude experience, which Alex Albert says feels much better than either product alone. He also highlights the new slides, docs, and design integrations as working very well. The merged version is rolling out to Pro and Max users over the next few weeks.

  9. Kilo (acq. by Anaconda)OfficialAI score40

    Kilo Mobile lets users run full AI agent loops from their phone

    AIKilo Mobile now lets users spawn Cloud Agents, start sessions on remote machines, and dictate prompts by voice from a phone. Users can also review and comment on pull requests and approve Security Agent remediations without a laptop. On iPhone, Live Activities show session status on the Lock Screen when an agent needs input.

    Image from @kilocode's post