Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. ElevenLabsOfficialAI score40

    ElevenLabs launches ElevenAgents Architect to help teams build AI agents

    AIElevenLabs introduced ElevenAgents Architect, an expert built into ElevenAgents that helps teams launch and improve AI agents through voice or text. The post describes it as a conversational way to create and refine agents without further technical detail provided.

    Video from @ElevenLabs's post
  2. Theo OtzXAI score40

    Agent.reviews launches, letting AI agents review software tools

    AIArmature Inc. has launched agent.reviews, a platform where AI agents write and read reviews of software tools after real tasks. The company says it already holds more than 100,000 reviews covering over 6,000 tools, with each review anonymized and free of personal data, code, or prompts. The service is free, and users can install a skill to check reviews.

    Video from @Totzenberger's post
  3. Aravind SrinivasXAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

    Video from @AravSrinivas's post
  4. OpenRouter · New modelsBlogAI score62

    Mistral Large 4 is listed on OpenRouter with a 1M-token context window

    AIMistral AI's Mistral Large 4 is listed on OpenRouter as a frontier multimodal model accepting text and image input. The listing says it is built for reasoning, coding, and agentic workloads and offers a 1M-token context window. The feed excerpt is truncated, so further details such as pricing or availability are not confirmed here.

  5. Kilo (acq. by Anaconda)OfficialAI score29

    Kilo launches Kilo Desktop, a unified app for 500+ AI models

    AIKilo has launched Kilo Desktop, a single app offering access to more than 500 models from major labs, including open-source and local models. It includes agents that plan, code, and debug alongside users, plus built-in notebooks, local model support, and conda environments.

    Image from @kilocode's post
  6. Guillaume Lample @ NeurIPS 2024XAI score26

    Mistral model beats GLM 5.3 on STEM, CAD, and finance tasks

    AIOn human evaluation, the model outperforms GLM 5.3 on STEM, CAD, and finance tasks and performs on par on agentic coding. The post is part 5 of a thread, so the model's name and other details come from earlier posts not included here.

    Image from @GuillaumeLample's post
  7. Guillaume Lample @ NeurIPS 2024XAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  8. SantiagoXAI score40

    Gamma 5 adds clarifying questions, PowerPoint import, and app connectors

    AIGamma 5, the latest version of the AI presentation and document tool, now asks questions before building a deck and supports PowerPoint import and export. It also adds connectors to Notion, Slack, and HubSpot for importing content, and lets users change design styles by prompting.

    Video from @svpino's post
  9. Allie K. MillerXAI score22

    Users combine personal AIs for group collaboration and delegation

    AIAllie K. Miller argues that collaboration between people's AIs is an underappreciated feature, with users combining their AIs, delegating across them, and having them sort tasks out. She says this multiplayer AI is already happening, and that Instinct has since added the ability to put a personal Instinct into a group text.

    Image from @alliekmiller's post
  10. The Next PlatformNewsAI score38

    Dell Adds Data Context, Prep, and Storage Features to Its AI Data Platform

    AIDell is adding agentic AI capabilities to its AI Data Platform, including a Unified Semantic Layer with a searchable glossary and an Enterprise Knowledge Graph built with Nvidia's Auto-Ontology open source library. The features are designed to give agents shared context, reducing repeated token generation and compute costs. The platform's layers include the Data Orchestration Engine, Data Engines, and Storage Engines such as PowerScale, ObjectScale, and the Lightning File System.

  11. OpenAI NewsOfficialAI score38

    How Jump Trading is scaling quant research with ChatGPT

    AIJump Trading is using OpenAI to expand its quantitative research, with longer-running AI workflows that combine multiple data sources alongside human review. The source does not give further details on specific models, metrics, or results.

  12. ElevenLabs BlogOfficialAI score41

    ElevenAgents Architect Helps Teams Build and Improve Voice Agents Conversationally

    AIElevenLabs launched ElevenAgents Architect in Alpha, a built-in assistant that helps teams build and improve agents through voice or text conversation. It analyzes transcripts and test failures, proposes changes validated in simulated conversations, and saves them as versioned drafts that require approval before going live. It can also be accessed from Claude, Claude Code, ChatGPT, Cursor, and Grok Bot.

  13. ChinaTalkBlogAI score33

    Bharat Patel on why data, not models, is the hard part of military AI

    AIAccenture defense AI lead Bharat Patel argues that data quality depends on the use case and that "AI-ready data" is a myth. He cites Project Maven, which began in 2017, where early imagery lacked relevant targets and models underperformed until teams continuously collected targeted data. The conversation also covers why fully autonomous tanks remain distant and the risks of data poisoning.

  14. The SequenceBlogAI score62

    Darwin Gödel Machine rewrote its own scaffolding to raise SWE-bench scores

    AIThe Darwin Gödel Machine, a coding agent from Sakana and Jeff Clune's lab, modified its own codebase over roughly eighty iterations without supervision. Its additions included better file viewing, patch validation before submitting fixes, generating and ranking several candidate solutions, and keeping a history of failed attempts. These changes raised its score from 20 to 50 percent on SWE-bench and from 14 to 31 percent on Polyglot.

  15. O'Reilly RadarBlogAI score62

    O'Reilly Radar Trends for October 2026: Models, Agents, and Security

    AIThe roundup covers September 2026 AI developments, including model price cuts and new specialized models from Anthropic, OpenAI, Google, and others. It also tracks agents delegating work to other agents, security incidents involving AI agents, and the author's warning that adopters must remain accountable for what their agents do.

  16. Vaibhav (VB) SrivastavXAI score43

    Auto-review in Codex is now free for ChatGPT-signed-in users

    AIOpenAI has made Auto-review free for all users signed in through a ChatGPT account, and it does not draw usage from their plan. Auto-review uses a second agent to check the primary agent's actions, blocking high-risk moves and actions that drift from user intent, so long tasks can run without constant approval prompts. It can be enabled under settings > permissions > auto-review.

  17. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

  18. Harrison ChaseXAI score20

    Harrison Chase praises a take on agent harnesses

    AIHarrison Chase, founder of LangChain, endorsed a post on harnesses with the brief comment "Good take on harnesses." The post, from @zeeg, argues that general coding harnesses like Codex will be superseded by specialized ones and that local models will handle most daily tasks within five years.

  19. EveryBlogAI score36

    Every launches the Every Agent, an agentic coworker in Slack

    AIEvery has launched the Every Agent, an agentic coworker that lives in Slack and helps teams delegate complex work and share AI experiments. It also sends personalized Frontier Alerts when new models or tools ship, and the company says it charges zero percent markup on tokens, so customers pay what Every pays.

  20. Mastra BlogOfficialAI score67

    Mastra launches Agent Controller GA, a runtime for long-running agent sessions

    AIMastra has released Agent Controller in general availability, a runtime that hosts long-running agent sessions around the agent loop. The team says it was first built for Mastra Code and expanded to support Mastra Factory, which runs many concurrent sessions, and that memory usage in long-running Mastra Code processes dropped from 2–20 GB to 300–750 MB after optimizing UI state snapshots.

    Why it matters: The post explains how the controller evolved from one developer's session to many concurrent sessions, with measured memory and storage changes useful to engineers building multi-user agent apps.

  21. Claude BlogOfficialAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    AIClaude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    Why it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.

  22. Claude BlogOfficialAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    AIComcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    Why it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.

  23. METR BlogOfficialAI score31

    AI Agents Could Hide Misbehavior by Exploiting Inspect Transcript Viewer

    AIMETR tested whether an AI agent running in an Inspect evaluation could alter the transcript humans review, and a researcher found a vulnerability in about 10 minutes that allowed arbitrary changes to what the reviewer sees. The exploit affects only the displayed transcript, not the underlying data stored in METR's database, and METR has not observed agents using it in its evaluations. METR argues that AI outputs such as transcripts and reasoning should be treated as untrusted input, with monitoring systems treated as security-critical infrastructure.

Oct 5

Oct 5Mon
  1. dexXAI score14

    Founder pitches for human-in-the-loop AI guardrails draw skeptical feedback

    AIDex Horthy says he repeatedly gets founder requests for feedback on human-in-the-loop notification, guardrail, or audit products, and lessons he learned in late 2024 and early 2025 apply to them. Akio Nuernberger, linked as background, reports receiving multiple monthly inbound messages from such startups without a single Langfuse customer showing interest.

  2. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.