Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. SantiagoXAI score32

    Gumloop launches Agent Browsers for AI agents on hosted infrastructure

    AIGumloop's Agent Browsers let agents perform browser-based tasks without an MCP or API, with the browser running on Gumloop's infrastructure rather than the user's desktop. The feature includes built-in account management with a vault, persistent browser profiles, 1Password integration, live viewing, and session replays. Users can also reuse their workflows.

  2. ARC PrizeOfficialAI score28

    Grok 4.7 scores 1.8% on ARC-AGI-3 standard harness

    AIGrok 4.7 scored 1.8% on ARC-AGI-3 in the standard harness, which lets models carry notes between turns, slightly below the 2.1% reported for Grok 4.7 in that setting. In a new provider adapter harness that preserves opaque reasoning and enables auto compaction, the score rose to 10.0%.

  3. Ai2OfficialAI score4

    Ai2 Agents post-training team invites COLM 2026 attendees to connect

    AIAi2 applied scientist Shashank Gupta says he will attend COLM 2026 from Tuesday through Friday and invites people to talk with the Ai2 Agents post-training team. Topics include post-training for coding and long-horizon agents, such as agentic RL, OPD, data and infrastructure, and multi-agent training, plus opportunities at Ai2. He lists an Ai2 booth session Tuesday 1:30–3pm and an Ai2 mixer Tuesday 6–9pm.

  4. KhazixXAI score32

    Khazix builds an enterprise platform replacing Feishu's workspace in two days

    AIThe author spent two days building an internal enterprise platform on all Feishu data and a self-built MCP, replacing Feishu's native workbench to handle Vibe Coding app deployment, security, permissions, and app and skill circulation. A custom configuration interface is planned so employees can use their own Agents to modify their homepages and data pages.

    Image from @Khazix0918's post
  5. SantiagoXAI score34

    Ampersand packages Salesforce integration work for enterprise AI agents

    AIAmpersand lets developers connect an AI agent to a customer's Salesforce by configuring object and field mappings. The platform then handles API calls, authentication, token refreshes, and retries, which the post presents as a major advantage given the difficulty of managing multiple Salesforce accounts.

  6. Google WorkspaceOfficialAI score8

    Google shows how Workspace Studio and Gemini cut busywork

    AIGoogle's productivity advisor Laura Mae Martin explains how Workspace Studio and Gemini can lighten users' workloads by reducing routine tasks. The post promotes a Google article on the topic but gives no specific features, figures, or availability details.

    Image from @GoogleWorkspace's post
  7. Microsoft ResearchOfficialAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  8. ElevenLabsOfficialAI score40

    ElevenLabs launches ElevenAgents Architect to help teams build AI agents

    AIElevenLabs introduced ElevenAgents Architect, an expert built into ElevenAgents that helps teams launch and improve AI agents through voice or text. The post describes it as a conversational way to create and refine agents without further technical detail provided.

    Video from @ElevenLabs's post
  9. 0xMarioNawfalXAI score41

    Open-source OpenWork offers a local alternative to Claude Cowork

    AIAn X post says someone has open-sourced OpenWork, a local alternative to Claude Cowork. The tool lets AI agents work on files directly on a computer using cloud or local models, with support for existing skills, plugins, and MCP servers. The post links to the project's GitHub repository.

  10. Theo OtzXAI score40

    Armature launches agent.reviews, a free review site written by AI agents

    AITheo Otz says Armature has launched agent.reviews, a free site where AI agents review software tools. The post says it already has more than 100,000 reviews of over 6,000 tools, written by agents after real tasks. Reviews are anonymized and contain no code, prompts, or personal data, with checks using Jev from typesafeai plus deterministic and LLM checks. Users can install the skill to read the reviews.

    Video from @Totzenberger's post
  11. Aravind SrinivasXAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

    Video from @AravSrinivas's post
  12. OpenRouter · New modelsBlogAI score62

    Mistral Large 4 is listed on OpenRouter with a 1M-token context window

    AIMistral AI's Mistral Large 4 is listed on OpenRouter as a frontier multimodal model accepting text and image input. The listing says it is built for reasoning, coding, and agentic workloads and offers a 1M-token context window. The feed excerpt is truncated, so further details such as pricing or availability are not confirmed here.

  13. Kilo (acq. by Anaconda)OfficialAI score29

    Kilo launches Kilo Desktop, a unified app for 500+ AI models

    AIKilo has launched Kilo Desktop, a single app offering access to more than 500 models from major labs, including open-source and local models. It includes agents that plan, code, and debug alongside users, plus built-in notebooks, local model support, and conda environments.

    Image from @kilocode's post
  14. Guillaume Lample @ NeurIPS 2024XAI score26

    Mistral model beats GLM 5.3 on STEM, CAD, and finance tasks

    AIOn human evaluation, the model outperforms GLM 5.3 on STEM, CAD, and finance tasks and performs on par on agentic coding. The post is part 5 of a thread, so the model's name and other details come from earlier posts not included here.

    Image from @GuillaumeLample's post
  15. Guillaume Lample @ NeurIPS 2024XAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  16. SantiagoXAI score40

    Gamma 5 adds clarifying questions, PowerPoint import, and app connectors

    AIGamma 5, the latest version of the AI presentation and document tool, now asks questions before building a deck and supports PowerPoint import and export. It also adds connectors to Notion, Slack, and HubSpot for importing content, and lets users change design styles by prompting.

    Video from @svpino's post
  17. Allie K. MillerXAI score22

    Users combine personal AIs for group collaboration and delegation

    AIAllie K. Miller argues that collaboration between people's AIs is an underappreciated feature, with users combining their AIs, delegating across them, and having them sort tasks out. She says this multiplayer AI is already happening, and that Instinct has since added the ability to put a personal Instinct into a group text.

    Image from @alliekmiller's post
  18. The Next PlatformNewsAI score38

    Dell Adds Data Context, Prep, and Storage Features to Its AI Data Platform

    AIDell is adding agentic AI capabilities to its AI Data Platform, including a Unified Semantic Layer with a searchable glossary and an Enterprise Knowledge Graph built with Nvidia's Auto-Ontology open source library. The features are designed to give agents shared context, reducing repeated token generation and compute costs. The platform's layers include the Data Orchestration Engine, Data Engines, and Storage Engines such as PowerScale, ObjectScale, and the Lightning File System.

  19. OpenAI NewsOfficialAI score38

    How Jump Trading is scaling quant research with ChatGPT

    AIJump Trading is using OpenAI to expand its quantitative research, with longer-running AI workflows that combine multiple data sources alongside human review. The source does not give further details on specific models, metrics, or results.

  20. ElevenLabs BlogOfficialAI score41

    ElevenAgents Architect Helps Teams Build and Improve Voice Agents Conversationally

    AIElevenLabs launched ElevenAgents Architect in Alpha, a built-in assistant that helps teams build and improve agents through voice or text conversation. It analyzes transcripts and test failures, proposes changes validated in simulated conversations, and saves them as versioned drafts that require approval before going live. It can also be accessed from Claude, Claude Code, ChatGPT, Cursor, and Grok Bot.

  21. The SequenceBlogAI score62

    Darwin Gödel Machine rewrote its own scaffolding to raise SWE-bench scores

    AIThe Darwin Gödel Machine, a coding agent from Sakana and Jeff Clune's lab, modified its own codebase over roughly eighty iterations without supervision. Its additions included better file viewing, patch validation before submitting fixes, generating and ranking several candidate solutions, and keeping a history of failed attempts. These changes raised its score from 20 to 50 percent on SWE-bench and from 14 to 31 percent on Polyglot.

  22. O'Reilly RadarBlogAI score62

    O'Reilly Radar Trends for October 2026: Models, Agents, and Security

    AIThe roundup covers September 2026 AI developments, including model price cuts and new specialized models from Anthropic, OpenAI, Google, and others. It also tracks agents delegating work to other agents, security incidents involving AI agents, and the author's warning that adopters must remain accountable for what their agents do.

  23. Vaibhav (VB) SrivastavXAI score43

    Auto-review in Codex is now free for ChatGPT-signed-in users

    AIOpenAI has made Auto-review free for all users signed in through a ChatGPT account, and it does not draw usage from their plan. Auto-review uses a second agent to check the primary agent's actions, blocking high-risk moves and actions that drift from user intent, so long tasks can run without constant approval prompts. It can be enabled under settings > permissions > auto-review.

  24. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

  25. Harrison ChaseXAI score20

    Harrison Chase praises a take on agent harnesses

    AIHarrison Chase, founder of LangChain, endorsed a post on harnesses with the brief comment "Good take on harnesses." The post, from @zeeg, argues that general coding harnesses like Codex will be superseded by specialized ones and that local models will handle most daily tasks within five years.

  26. EveryBlogAI score36

    Every launches the Every Agent, an agentic coworker in Slack

    AIEvery has launched the Every Agent, an agentic coworker that lives in Slack and helps teams delegate complex work and share AI experiments. It also sends personalized Frontier Alerts when new models or tools ship, and the company says it charges zero percent markup on tokens, so customers pay what Every pays.

  27. Mastra BlogOfficialAI score67

    Mastra launches Agent Controller GA, a runtime for long-running agent sessions

    AIMastra has released Agent Controller in general availability, a runtime that hosts long-running agent sessions around the agent loop. The team says it was first built for Mastra Code and expanded to support Mastra Factory, which runs many concurrent sessions, and that memory usage in long-running Mastra Code processes dropped from 2–20 GB to 300–750 MB after optimizing UI state snapshots.

    Why it matters: The post explains how the controller evolved from one developer's session to many concurrent sessions, with measured memory and storage changes useful to engineers building multi-user agent apps.

  28. Claude BlogOfficialAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    AIClaude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    Why it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.