Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Oct 5

Oct 5Mon
  1. CognitionOfficialAI score58

    Cognition's Devin adds Dreaming, a nightly memory graph across sessions

    AICognition introduces Dreaming, a feature in which Devin builds a memory graph of how a user likes to work across sessions. At night, Devin self-improves this memory by removing stale records and discovering latent information. Cognition also says it is creating an open-source standard called Agent Memory Repo, linked in the post.

    Video from @cognition's post
  2. Understanding AI (Timothy B. Lee)BlogAI score62

    Agent swarms may be the next scaling law, but speed and capability differ

    AIUnderstanding AI's Dan Kagan-Kans examines whether multi-agent swarms could be a new scaling law after OpenAI's 10,000-agent Navier-Stokes effort. The piece cites Noam Brown saying he would not attribute 10% of the breakthrough's credit to the multi-agent approach. It also cites research suggesting swarms mainly buy speed, with a stepping-on-toes parameter between 0.5 and 0.7, while a Microsoft Research and UC Berkeley paper reports some tasks only teams could complete.

  3. dexXAI score16

    Dex Horthy argues human review still gives AI work its edge

    AIDex Horthy argues there will always be an advantage in reviewing AI output, since unreviewed work tends toward generic "AI slop." He adds that the reviewed object may not always be code, and speculates that once models far surpass humans, human interference could make results worse.

  4. Google AntigravityOfficialAI score8

    Google Antigravity publishes its changelog online

    AIGoogle Antigravity directs users to a new changelog page at antigravity.google/changelog. The post provides only the link and no details about specific updates or features.

  5. Google AntigravityOfficialAI score23

    Google Antigravity adds AlphaGenome Atlas Skill for genomic research workflows

    AIGoogle Antigravity has integrated the AlphaGenome Atlas Skill into its scientific workbench, enabling AI agents to help researchers prioritize genetic variants, generate structural plots, and build testable hypotheses. The company showcases researchers Natasha and Kyle using the tool in a demonstration video.

  6. LiveKitOfficialAI score30

    LiveKit demos Microsoft speech models in a voice support agent

    AILiveKit Agents pairs MAI-Transcribe-2-Streaming for speech-to-text, Gemma 4 on LiveKit Inference for reasoning and tool calls, and MAI-Voice-2.1-Flash for speech output in a demo support call. The post links separate speech-to-text and text-to-speech resources for developers.

    Video from @livekit's post
  7. Elad GilXAI score40

    Era launches free simulated enterprises for testing AI agents

    AIEra, launched by Ofir Ehrlich's team, generates a complete simulated company spanning Salesforce, Slack, Jira, Zendesk, Gong, and Deel, plus cloud databases and storage. Agents interact with it through live MCP and API interfaces, and because Era generated the company, it knows the exact ground truth for testing and benchmarking. The post says the product is live today and free.

  8. CursorOfficialAI score38

    Cursor SDK agents can now be steered while running

    AICursor announced that developers can steer Cursor SDK agents while they run using run.steer(), which adds a message to the next turn. If a subagent is mid-task, it moves to the background and continues working.

    Video from @cursor_ai's post
  9. Amazon Web ServicesOfficialAI score13

    AWS and F1 build agentic AI that resolves race-car issues 86% faster

    AIAWS and F1 built an agentic AI solution that resolves critical issues up to 86% faster, letting engineers focus on development instead of logs. The post notes that each car carries 300 sensors across 24 races with zero margin for error.

    Video from @awscloud's post
  10. IEEE Spectrum · AINewsAI score49

    Human Oversight of AI Agents Could Fail as Approval Processes Push People Out

    AIResearchers Avijit Ghosh, Margaret Mitchell, and Samir Passi argue in a September 6 arXiv paper that current human-in-the-loop designs for AI agents push humans out of meaningful oversight. They say agents are tuned for speed, accuracy, and volume, overwhelming reviewers, and recommend adding friction, such as requiring users to state their own choice first, to counter automation bias and fatigue.

  11. O'Reilly RadarBlogAI score38

    Zero to Agent in 30 Minutes: Building Your First Agent with MCP

    AIBruce Hopkins shows how to wrap an existing stock-data REST API, the Twelve Data API, in a Model Context Protocol (MCP) server so an MCP client can discover and call it. The demo uses Python with FastMCP, exposing current and historical stock-price functions as tools and resources with descriptive prompts. Developers can add an MCP interface around existing capabilities without replacing their underlying application logic.

  12. FireworksOfficialAI score34

    DeepSeek V4.1 Flash now available for training on Fireworks

    AIFireworks AI has made DeepSeek V4.1 Flash available for training on its Dedicated Training API and Managed Training surfaces. The post positions the model as a strong base for agentic coding, terminal automation, and tool use, and notes it is cost-efficient to serve.

  13. MIT Technology Review · AINewsAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    AIA survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

  14. Replit ⠕OfficialAI score16

    Replit adds TikTok Ads MCP integration to its workspace

    AIReplit now offers a TikTok Ads MCP integration, letting users bring advertising workflows into the workspace where they build. Users connect their TikTok Ads account through Replit's integrations page to try it.

    Image from @Replit's post
  15. Replit ⠕OfficialAI score22

    Replit weekly changelog adds GPT-6.1 Sol and Claude Sonnet 5.5 models

    AIReplit shipped a weekly update letting users build with GPT-6.1 Sol and Claude Sonnet 5.5, along with an Ask agent integration with Jev. The release also includes an updated Settings UI and enterprise Workplace controls for company-wide rules and controlled exceptions. Full details are in the Replit changelog.

  16. SantiagoXAI score47

    Santiago highlights Era, a tool that generates simulated enterprise environments for testing agents

    AISantiago shares a post about Era, a tool that generates a complete simulated company across systems such as Salesforce, Slack, Jira, and Zendesk. Agents interact with the environment through live MCP and API interfaces, and because the company is generated, its exact ground truth is known. Users can test and benchmark agents, use failures for post-training, and rerun the same environment to measure impact. The author's own text adds that the company can be reset to its initial state after a run.

  17. Karl's AI WattsXAI score23

    Karl's AI Watts shares a full AI Skills workflow tutorial

    AIKarl's AI Watts publishes the AI workflow he previously shared internally at Tim Studio, covering finding Skills, packaging experience into Skills, combining them into workflows, and batching and scheduling them. The post says viewers could build a local batch video-editing Skill and an end-to-end content pipeline spanning copy, posters, video, and web pages.

    Video from @aiwarts's post
  18. clem 🤗XAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

    Why it matters: The capture proxy lets models train inside real coding harnesses without reimplementing them, with measured gains and a comparison against SFT on the same data.

    Image from @ClementDelangue's post
  19. Guillermo RauchXAI score44

    Guillermo Rauch introduces gdp-ts, a TypeScript library for compile-time authorization proofs

    AIGuillermo Rauch announces gdp-ts, a TypeScript library, linter and AI skill for safer API design. Under this contract, sensitive functions require proofs that the caller performed an authorization check, and the typechecker verifies those proofs at compile time. Rauch says agents now write more code than teams can review, so hard constraints matter more.

    Video from @rauchg's post
  20. Baidu Inc.OfficialAI score13

    Baidu's AI Pulse explores the full stack behind useful, cost-efficient agents

    AIBaidu's latest AI Pulse edition argues that agent experiences depend on a full underlying AI stack, not just the agent itself. It examines what productivity, commerce, and industrial agents need from that stack and how those needs shape it, alongside updates on the AI, Evolving to Miaoda upgrades and the Kooko AI workspace.

  21. Cloudflare Blog · AIOfficialAI score40

    Cloudflare Birthday Week 2026 unveils cf CLI, EmDash CMS, and post-quantum tools

    AICloudflare announced 46 products and updates during Birthday Week 2026, including the cf CLI for the entire Cloudflare API and EmDash, an open-source Astro-based serverless CMS whose plugins run in isolated Worker sandboxes. The company also said it plans to become a public certificate authority that issues free Merkle Tree Certificates for post-quantum authentication.

  22. Import AIBlogAI score47

    Import AI 475 Covers Swarm Scaling, Google DeepMind's SynthID Bio, and AI Science Labs

    AIToby Ord argues that AI agent swarms trade extra tokens for faster completion, needing about twice the total tokens of a single agent for the same performance with four agents, but in half the wall-clock time. He notes swarm scaling shows diminishing returns, with 10x agents yielding roughly 3x to 5x the performance of 10x tokens on one agent. A CSAIP poll found 61% of Americans think voluntary AI industry commitments are "not enough."

  23. O'Reilly RadarBlogAI score45

    How to Build Reliable AI Agent Systems for Production

    AIReliable AI agent systems need deterministic policy checks, not just better prompts or stronger models, because a model's proposed action can succeed at the API level while still updating the wrong account. The article recommends separating the model's proposal from a policy service that checks actions before execution and records an audit trail. It also advises treating agent context as untrusted input, using narrow capabilities instead of broad tokens, and building in stopping rules and idempotent recovery.