Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 6

Oct 6Tue
  1. AnthropicOfficialAI score49

    Anthropic expands Cyber Verification Program for verified security professionals

    AIAnthropic is expanding its Cyber Verification Program to give verified security professionals broader access to its most capable models. Through the program, they can use Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work. New tiers will also allow authorized offensive work such as penetration testing and red-teaming.

  2. Claude Code · GitHub ReleasesOfficialAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    AIClaude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.

  3. Google ResearchOfficialAI score36

    Google Research demos Co-Director for coherent long-form AI video generation

    AIGoogle Research will demo Co-Director, a hierarchical multi-agent framework that optimizes video generation and consistency for long-form storytelling, at the #COLM2026 Google booth #107 today at 1:00 PM PT. The demo showcases interactive cinematic narratives, and the team's blog post details the approach.

    Image from @GoogleResearch's post
  4. elvisXAI score41

    Parsewave audit fixes 206 verifier bugs in AutomationBench

    AIParsewave audited all 600 public tasks in Zapier's AutomationBench and human review confirmed 206 real verifier bugs, all of which were fixed in AutomationBench Verified. Replaying 1,235 Kimi K3 runs on the old and fixed verifiers changed 27.9% of grades, with pass rate rising from 18.8% to 43.8% where verifiers were too strict and falling from 60.2% to 49.7% where they were too lenient.

  5. Matt ShumerXAI score34

    AgentID lets AI agents sign in to apps like Google login

    AIAgentMail has launched AgentID, a sign-in system that lets millions of AI agents with AgentMail accounts log into third-party apps within minutes. The post frames it as the agent-era equivalent of Sign in with Google, and offers three months of free AgentMail plans to developers who integrate it and share a screenshot.

  6. laurenXAI score42

    Developer automates releases and QA with Grok Bot agents in Slack

    AIA developer used Grok Bot to build two Slack team bots, sandcastle for release management and poteto for engineering, automating their release and QA process. Sandcastle DMs contributors PR links, kicks off builds, and runs a fuzz swarm of 10+ Grok 4.7 xhigh agents, while poteto triages and fixes issues via Cursor Projects.

    Image from @poteto's post
  7. Sierra BlogOfficialAI score62

    Sierra and Meta announce Personal Agent Protocol, an open standard for personal agents

    AISierra and Meta are developing Personal Agent Protocol, an open standard defining how personal agents interact with businesses, with industry partners including Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. The protocol uses OAuth sessions where consumers choose read-only or write access and companies choose whether agents reach them through websites, APIs via MCP and OpenAPI, or their own agents. The authors plan to publish the v0.1 specification later this month along with a reference implementation.

    Why it matters: The post specifies how personal agents would authenticate and reach businesses through websites, APIs, or company agents, which matters for anyone building agent integrations.

  8. ClaudeOfficialAI score62

    Claude now works inside Google Docs, Sheets, and Slides

    AIand files from those apps can also be opened within Claude. In Google Workspace, Claude appears in a sidebar next to the open file, reads its contents, and edits it in place, with each edit available for user approval before it is applied.

    Why it matters: The source describes Claude working within Google Workspace files and approving edits, a direct change to how users in those apps can collaborate with the model.

    Video from @claudeai's post
  9. eric zakariassonXAI score32

    Grok turns one product photo into an 8-second vertical video ad

    AIA pipeline built with Grok takes a single product photo and produces an eight-second vertical ad, writing the brief, generating three scenes, selecting the best one, animating it, and adding a voiceover. Each run costs about $1.35, and the code is available in the xai-cookbook repository on GitHub.

    Video from @ericzakariasson's post
  10. eric zakariassonXAI score29

    Cursor adds five SpaceXAI TypeScript SDK demo apps to cookbook

    AICursor has added five apps to its cookbook, all built on the new SpaceXAI TypeScript SDK. The apps cover premise-to-short-film, picture-to-video ads, screenshot-to-React-component, X posts-to-sentiment dashboard, and link-to-podcast conversion. Demos and code are available in the post.

    Video from @ericzakariasson's post
  11. LangChainOfficialAI score46

    LangChain video shows how to build a model router into an agent harness

    AILangChain's Sydney Runkle presents a four-step method for building a model router into a coding agent harness: understanding tasks, understanding models, building the router, and tracking task outcomes. The background post says the router cut costs by 64% without reducing quality by sending tasks that do not need a frontier model to cheaper models.

  12. SantiagoXAI score32

    Gumloop launches Agent Browsers for AI agents on hosted infrastructure

    AIGumloop's Agent Browsers let agents perform browser-based tasks without an MCP or API, with the browser running on Gumloop's infrastructure rather than the user's desktop. The feature includes built-in account management with a vault, persistent browser profiles, 1Password integration, live viewing, and session replays. Users can also reuse their workflows.

  13. ARC PrizeOfficialAI score28

    Grok 4.7 scores 1.8% on ARC-AGI-3 standard harness

    AIGrok 4.7 scored 1.8% on ARC-AGI-3 in the standard harness, which lets models carry notes between turns, slightly below the 2.1% reported for Grok 4.7 in that setting. In a new provider adapter harness that preserves opaque reasoning and enables auto compaction, the score rose to 10.0%.

  14. KhazixXAI score32

    Khazix builds an enterprise platform replacing Feishu's workspace in two days

    AIThe author spent two days building an internal enterprise platform on all Feishu data and a self-built MCP, replacing Feishu's native workbench to handle Vibe Coding app deployment, security, permissions, and app and skill circulation. A custom configuration interface is planned so employees can use their own Agents to modify their homepages and data pages.

    Image from @Khazix0918's post
  15. SantiagoXAI score34

    Ampersand packages Salesforce integration work for enterprise AI agents

    AIAmpersand lets developers connect an AI agent to a customer's Salesforce by configuring object and field mappings. The platform then handles API calls, authentication, token refreshes, and retries, which the post presents as a major advantage given the difficulty of managing multiple Salesforce accounts.

  16. Microsoft ResearchOfficialAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  17. ElevenLabsOfficialAI score40

    ElevenLabs launches ElevenAgents Architect to help teams build AI agents

    AIElevenLabs introduced ElevenAgents Architect, an expert built into ElevenAgents that helps teams launch and improve AI agents through voice or text. The post describes it as a conversational way to create and refine agents without further technical detail provided.

    Video from @ElevenLabs's post
  18. Theo OtzXAI score40

    Agent.reviews launches, letting AI agents review software tools

    AIArmature Inc. has launched agent.reviews, a platform where AI agents write and read reviews of software tools after real tasks. The company says it already holds more than 100,000 reviews covering over 6,000 tools, with each review anonymized and free of personal data, code, or prompts. The service is free, and users can install a skill to check reviews.

    Video from @Totzenberger's post
  19. Aravind SrinivasXAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

    Video from @AravSrinivas's post
  20. OpenRouter · New modelsBlogAI score62

    Mistral Large 4 is listed on OpenRouter with a 1M-token context window

    AIMistral AI's Mistral Large 4 is listed on OpenRouter as a frontier multimodal model accepting text and image input. The listing says it is built for reasoning, coding, and agentic workloads and offers a 1M-token context window. The feed excerpt is truncated, so further details such as pricing or availability are not confirmed here.

  21. Kilo (acq. by Anaconda)OfficialAI score29

    Kilo launches Kilo Desktop, a unified app for 500+ AI models

    AIKilo has launched Kilo Desktop, a single app offering access to more than 500 models from major labs, including open-source and local models. It includes agents that plan, code, and debug alongside users, plus built-in notebooks, local model support, and conda environments.

    Image from @kilocode's post
  22. Guillaume Lample @ NeurIPS 2024XAI score26

    Mistral model beats GLM 5.3 on STEM, CAD, and finance tasks

    AIOn human evaluation, the model outperforms GLM 5.3 on STEM, CAD, and finance tasks and performs on par on agentic coding. The post is part 5 of a thread, so the model's name and other details come from earlier posts not included here.

    Image from @GuillaumeLample's post
  23. Guillaume Lample @ NeurIPS 2024XAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  24. SantiagoXAI score40

    Gamma 5 adds clarifying questions, PowerPoint import, and app connectors

    AIGamma 5, the latest version of the AI presentation and document tool, now asks questions before building a deck and supports PowerPoint import and export. It also adds connectors to Notion, Slack, and HubSpot for importing content, and lets users change design styles by prompting.

    Video from @svpino's post