Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Mastra BlogOfficialAI score29

    Mastra Launches Agency Program with Five Certified Partners to Build Agents

    AIMastra launched the Mastra Agency Program, a network of certified agencies and consultancies that build Mastra agents for clients. The launch includes five partners: Deerfield Group, Blue Drop Labs, Frontleap, Handpicked, and Young Security. Every partner has been vetted by Mastra's FDE team and receives direct access to Mastra's leadership and regular roadmap updates.

  2. Anthropic ResearchOfficialAI score72

    Anthropic launches OSS Scanner, a free AI vulnerability scanner for open-source projects

    AIAnthropic is launching OSS Scanner, an opt-in service that runs periodic security scans of enrolled open-source projects using its strongest models at no cost. Its outputs are fully model-generated without human review, so some reports may be incorrect or invalid, though a pilot found 85 of 97 checked critical and high-severity findings met Anthropic's disclosure bar. Core maintainers of eligible projects can enroll through a GitHub pull request.

  3. Anthropic ResearchOfficialAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

  4. Claude BlogOfficialAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

  5. LangChain BlogOfficialAI score67

    LangChain's Restock agent shows how to build a payment-capable AI agent

    AILangChain built Restock, a sample office-supply agent that runs in Slack on Managed Deep Agents and pays through Stripe's Link wallet. The agent searches products, builds a cart, and pays over the Machine Payments Protocol, with the user approving the purchase in Slack and the payment in Link. The post uses a pens order at $22.18 to show the flow from request to confirmed order.

    Why it matters: The post walks through how an agent handles search, budget limits, Slack review, and Link approval, showing where each control sits outside the model.

Oct 7

Oct 7Wed
  1. elvisXAI score67

    Tool-using multimodal models refuse harmful requests less often, NVIDIA study finds

    AIA NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.

    Image from @omarsar0's post
  2. meng shaoXAI score31

    Stanford publishes lecture 4 and 5 slides for CS 329Z agent course

    AIStanford's CS 329Z: Engineering AI Agents course has published slides for lectures 4 and 5, following the earlier release of lectures 1–3. Lecture 4 covers tool use and is taught by Diyi Yang, while lecture 5 covers frameworks and orchestration, taught by Diyi Yang, Michael Ryan, and John Yang.

    Image from @shao__meng's post
  3. meng shaoXAI score75

    Microsoft positions Windows as the home for hybrid AI agents across four layers

    AIMicrosoft has repositioned Windows as the home for hybrid intelligence, where AI agents can run locally or in the cloud. The announcement covers four layers: MXC reaching general availability for agent isolation, local models such as MAI Code 1.1 Flash, Copilot on Copilot+ PCs gaining local context and actions in coming months, and new hardware including RTX Spark PCs and DGX Station for Windows.

    Image from @shao__meng's post
  4. meng shaoXAI score88

    OpenAI rolls out GPT-6 with Intelligent UI to over 1.2 billion weekly ChatGPT users

    AIOpenAI is rolling out GPT-6 to ChatGPT's over 1.2 billion weekly users, adding Intelligent UI, which lets replies include charts, buttons, forms, and interactive tools. The post's image cites tiered access, with Free/Go and Plus/Pro/Business/Enterprise sharing the Sol and Luna model splits, and says the feature is progressively rendered as the model generates it.

    Image from @shao__meng's post
  5. QbitAINewsAI score30

    Step Terminal to launch STEPX Neo agent-native smartphone at October 13 event

    AIStep Terminal will unveil its first large-model-native agent smartphone, the STEPX Neo, at a "Ready Builder One" launch event in Shanghai on October 13. The company says the device is built agent-native across its model, system and hardware, and the event will also announce the latest progress in its ecosystem partnerships.

  6. Orange AIXAI score34

    Next Token episode 5 covers Personal Agents, open-source software, and hardware projects

    AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.

  7. GeekParkNewsAI score36

    MUZIM L1 Dock, Lumeria Lumoscope, and Other Small-Innovation Gadgets Reviewed

    AIMUZIM L1 is a desktop data dock with up to 24TB of storage, dual SSD slots, and a Vibe Search feature that finds files by natural-language description, with local-first processing rather than default cloud upload. Lumeria Lumoscope is a multispectral skin scope that clips onto a phone, using RGB, ultraviolet, polarized, and near-infrared light, priced at $199 in pre-sale. The article also covers immurok IK-1, a 59-dollar wireless fingerprint key with a 60-day standby battery that authorizes sudo, SSH, and Git actions on Mac, Windows, and Linux.

  8. HorizzonXAI score22

    Solo founder shares Devin Max experience and favorite features

    AIA solo founder writes that after a rocky first trial, they gave Devin Max a second chance and now favor Devin Cloud for shipping real projects. The post praises Cognition and Devin Max's model access and pricing, and covers SWE-1.7 Lightning's speed and its tendency to consume limits quickly. It also compares GPT 6 Astra and Fable 5.1, finding Fable more efficient for bug fixes and features.

  9. Epoch AIOfficialAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  10. Google Developers BlogOfficialAI score62

    Google's AQuA agent diagnoses production failures in a multi-agent travel concierge

    AIGoogle Developers Blog introduces AQuA, an ambient quality agent that runs in a customer's Google Cloud project and samples production sessions to find recurring agent failures. In a 32-session travel-concierge sweep, it verified six issues and traced two of them to specific prompt lines, and a replay after the fixes raised full-session passes from 5/32 to 13/32. The post notes that verification and diagnosis are model-based, and that the tool proposes edits without applying them.

    Why it matters: The post walks through a concrete production workflow, from sweep and verification to a code-anchored fix and replay, that shows how to diagnose silent agent failures.

  11. Hugging Face BlogOfficialAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  12. Teknium 🪽XAI score28

    Teknium posts a "hello" greeting on X

    AITeknium, a Nous Research affiliate, posted only the word "hello" on X. The quoted post from Nous Research announces a Series B raise to advance Hermes Agent and build a mobile app, with investors including NVIDIA and Samsung Next.

  13. IThome · AINewsAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

    This story has a top pick“Anthropic releases Claude Haiku 5.5 as its cheapest, fastest small model”

  14. IThome · AINewsAI score75

    OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users

    AIOpenAI announced on X that GPT-6 and Intelligent UI are now rolling out to all ChatGPT users, after GPT-6 Astra, Sol and Luna were previously limited to ChatGPT Work and Codex. Intelligent UI lets GPT-6 combine text, images and interactive elements such as charts, clickable buttons and forms, with a mahjong learning example shown.

  15. DatabricksOfficialAI score36

    Claude Haiku 5.5 launches on Databricks as a Day 0 release

    AIAnthropic's Claude Haiku 5.5 is available on Databricks from day zero, which Databricks calls its cheapest, fastest, and most capable small model. On Databricks' OfficeQA Pro V1 benchmark, it delivers about 15% higher quality than Haiku 4.5 at a fraction of the cost. Users can run it alongside 60+ other models on data already in Databricks, with Unity Gateway handling governance, monitoring, and security.

    Video from @databricks's post
  16. eric zakariassonXAI score20

    Grok Bot can search X feedback and propose plans, no connector needed

    AIThe post says users can ask the Grok bot to find all feedback about what they are building, summarize it, and propose a plan to address it, without needing an X account or connector. The background post notes that Grok Bot can now search, read, and monitor X.

  17. CognitionOfficialAI score26

    Cognition shares a blog post on Claude Haiku 5.5

    AICognition's X post links to a blog post at devin.ai about Claude Haiku 5.5, but the text provides no further details. The post itself offers no benchmark scores, prices, or capabilities to report.

  18. Vercel DevelopersOfficialAI score34

    Vercel AI Gateway adds Browserbase search and fetch tools

    AIVercel says Browserbase Search and Fetch tools are now available on AI Gateway, letting any model with tool calling search the web and read pages. Browserbase presents the tools as a way to reliably search and extract page contents through an existing Vercel plan.

  19. 🚨 AI News | TestingCatalogXAI score34

    Microsoft brings hybrid local-cloud intelligence to Copilot for Windows

    AIMicrosoft is adding hybrid intelligence to Copilot for Windows, letting it use local PC context and local models for tasks. Per Satya Nadella's quoted post, Windows will route each task to local or cloud models, and Copilot will act on the user's behalf only with permission.

    Video from @testingcatalog's post
  20. MarkTechPostNewsAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    AIAnthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  21. Amjad MasadXAI score40

    Replit building desktop app with Microsoft and Nvidia OpenShell

    AIReplit is building a powerful desktop app with a focus on security and reliability, citing supply-chain attacks and catastrophic agent mistakes as risks of desktop AI apps. The company is partnering with Microsoft and will be an early adopter of Nvidia's OpenShell. A quoted Replit post says the desktop preview runs builds locally on Windows, with each build in its own sandbox powered by Microsoft Execution Containers and OpenShell, and offers a waitlist.

  22. ClaudeDevsOfficialAI score43

    Anthropic adds computer and browser use toolsets to Claude SDKs

    AIAnthropic's Python and TypeScript SDKs now include built-in computer use and browser use toolsets for Claude. The SDKs run the agent loop and send actions to drivers, replacing the custom loop developers previously had to write to map clicks and keystrokes to commands.

    Video from @ClaudeDevs's post
  23. Satya NadellaXAI score38

    Microsoft brings Hybrid Intelligence to Copilot on Windows

    AIMicrosoft is upgrading Copilot on Windows with Hybrid Intelligence, which lets it use context from the user's PC, take actions on the user's behalf, and run local models when appropriate. With the user's permission, the feature aims to add capability while stretching token usage further.

    Video from @satyanadella's post
  24. Replit ⠕OfficialAI score34

    Replit to build and run apps locally on Windows in sandboxes

    AIReplit says it builds and runs apps locally on Windows, with each build executing in its own sandbox powered by Microsoft Execution Containers and Nvidia OpenShell. The announcement was made on stage alongside Microsoft's Pavan Davuluri at 16:29.

    Image from @Replit's post