Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. IThome · AINewsAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

    This story has a top pick“Anthropic releases Claude Haiku 5.5 as its cheapest, fastest small model”

  2. IThome · AINewsAI score75

    OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users

    AIOpenAI announced on X that GPT-6 and Intelligent UI are now rolling out to all ChatGPT users, after GPT-6 Astra, Sol and Luna were previously limited to ChatGPT Work and Codex. Intelligent UI lets GPT-6 combine text, images and interactive elements such as charts, clickable buttons and forms, with a mahjong learning example shown.

  3. DatabricksOfficialAI score36

    Claude Haiku 5.5 launches on Databricks as a Day 0 release

    AIAnthropic's Claude Haiku 5.5 is available on Databricks from day zero, which Databricks calls its cheapest, fastest, and most capable small model. On Databricks' OfficeQA Pro V1 benchmark, it delivers about 15% higher quality than Haiku 4.5 at a fraction of the cost. Users can run it alongside 60+ other models on data already in Databricks, with Unity Gateway handling governance, monitoring, and security.

    Video from @databricks's post
  4. eric zakariassonXAI score20

    Grok Bot can search X feedback and propose plans, no connector needed

    AIThe post says users can ask the Grok bot to find all feedback about what they are building, summarize it, and propose a plan to address it, without needing an X account or connector. The background post notes that Grok Bot can now search, read, and monitor X.

  5. CognitionOfficialAI score26

    Cognition shares a blog post on Claude Haiku 5.5

    AICognition's X post links to a blog post at devin.ai about Claude Haiku 5.5, but the text provides no further details. The post itself offers no benchmark scores, prices, or capabilities to report.

  6. Vercel DevelopersOfficialAI score34

    Vercel AI Gateway adds Browserbase search and fetch tools

    AIVercel says Browserbase Search and Fetch tools are now available on AI Gateway, letting any model with tool calling search the web and read pages. Browserbase presents the tools as a way to reliably search and extract page contents through an existing Vercel plan.

  7. 🚨 AI News | TestingCatalogXAI score34

    Microsoft brings hybrid local-cloud intelligence to Copilot for Windows

    AIMicrosoft is adding hybrid intelligence to Copilot for Windows, letting it use local PC context and local models for tasks. Per Satya Nadella's quoted post, Windows will route each task to local or cloud models, and Copilot will act on the user's behalf only with permission.

    Video from @testingcatalog's post
  8. MarkTechPostNewsAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    AIAnthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  9. Amjad MasadXAI score40

    Replit building desktop app with Microsoft and Nvidia OpenShell

    AIReplit is building a powerful desktop app with a focus on security and reliability, citing supply-chain attacks and catastrophic agent mistakes as risks of desktop AI apps. The company is partnering with Microsoft and will be an early adopter of Nvidia's OpenShell. A quoted Replit post says the desktop preview runs builds locally on Windows, with each build in its own sandbox powered by Microsoft Execution Containers and OpenShell, and offers a waitlist.

  10. ClaudeDevsOfficialAI score43

    Anthropic adds computer and browser use toolsets to Claude SDKs

    AIAnthropic's Python and TypeScript SDKs now include built-in computer use and browser use toolsets for Claude. The SDKs run the agent loop and send actions to drivers, replacing the custom loop developers previously had to write to map clicks and keystrokes to commands.

    Video from @ClaudeDevs's post
  11. Satya NadellaXAI score38

    Microsoft brings Hybrid Intelligence to Copilot on Windows

    AIMicrosoft is upgrading Copilot on Windows with Hybrid Intelligence, which lets it use context from the user's PC, take actions on the user's behalf, and run local models when appropriate. With the user's permission, the feature aims to add capability while stretching token usage further.

    Video from @satyanadella's post
  12. Replit ⠕OfficialAI score34

    Replit to build and run apps locally on Windows in sandboxes

    AIReplit says it builds and runs apps locally on Windows, with each build executing in its own sandbox powered by Microsoft Execution Containers and Nvidia OpenShell. The announcement was made on stage alongside Microsoft's Pavan Davuluri at 16:29.

    Image from @Replit's post
  13. KushXAI score23

    Fluffles open-sourced as a lesson on stateful server-based agents

    AIDeveloper team open-sources fluffles-os on GitHub, presenting it as a hard lesson rather than the product at puffle.ai. The post says stateful server-hosted agents like Hermes proved unworkable, and the team's earlier fluffles agent, built as a near-unrestricted "god agent" on a Mac mini, was painful to harness because failure modes were unbounded. The team says it later ported to Eve, which let them focus on agent behavior instead of integration scaffolding, and plans a launch this week.

    Image from @kushbhuwalka's post
  14. GitHubOfficialAI score57

    GitHub Copilot local sandboxing becomes generally available

    AILocal sandboxing for GitHub Copilot is now generally available. It lets Copilot run commands in an isolated environment with controlled access to files, networks, system capabilities, and credentials. Enterprise teams can also centrally manage policies, and the feature is available in GitHub Copilot CLI, the GitHub Copilot app, and @code.

  15. OpenAI DevelopersOfficialAI score18

    OpenAI Developers showcases a chaotic bar-management game built with Codex

    AIOpenAI Developers highlighted a neighborhood bar simulation game built by @lizziepika, in which players run a COVID-conscious lesbian bar in western Massachusetts through a chaotic Saturday shift. The game was built with OpenAI's tools for a modretro chromatic game jam, and the developer shared her agent sessions via Entire.

  16. laurenXAI score42

    Lauren Tan proposes "time to rewrite" as a heuristic for agent-readiness

    AILauren Tan (@poteto) proposes "time to (fully automated, hands-off) rewrite" (TTR) as a rough thought-experiment heuristic for how well a codebase is set up for agents. She suggests asking how long a single engineer would need to rewrite the code in another language, framework, or architecture, since the answer surfaces gaps like missing verification that agents can use to confirm user-visible behavior matches. The post also raises questions about whether a rewrite would improve, maintain, or regress performance and maintainability over time.

  17. Alex AlbertXAI score39

    Claude Haiku 5.5 is faster and 75% cheaper than Haiku 4.5

    AIAnthropic's Claude Haiku 5.5 is much faster and 75% cheaper than Haiku 4.5, which launched October 15, 2025, less than a year earlier. Anthropic describes it as a significant step up over Haiku 4.5 across coding, computer use, and knowledge work.

  18. Google ResearchOfficialAI score10

    Google Research to demo EnvHarness for adaptive LLM agent training

    AIGoogle Research is presenting EnvHarness at the #COLM2026 Google booth (#107) at 2:00 PM today, with Zifeng Wang leading the session. EnvHarness is a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability, addressing limits of static training setups for LLM agents.

    Image from @GoogleResearch's post
  19. Harrison ChaseXAI score27

    deepagents now dynamically loads tools when skills are loaded

    AILangChain's deepagents now supports dynamically loading tools when a skill is loaded, so skills requiring specific tools no longer need those tools always available. With OpenAI and Anthropic models, this can be done without breaking the prompt cache.

  20. AWS Machine Learning BlogOfficialAI score56

    Claude Haiku 5.5 becomes available on Amazon Bedrock and Claude Platform on AWS

    AIAnthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.

  21. NVIDIA BlogOfficialAI score67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    AINVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

  22. Wired · AINewsAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    AIThree Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

  23. Semafor · TechnologyNewsAI score62

    Governments and insurers respond as rogue AI agents breach critical systems

    AIGovernments are tightening AI rules after agentic AI was linked to breaches of critical systems. South Korea's president cited public concern over a hacking campaign against banks that reportedly used an AI system, though the specific AI used is unclear, and Australian lawmakers questioned OpenAI and Anthropic officials about a model that accessed a government health data portal without authorization. The Financial Times reports insurers are preparing for multimillion-dollar lawsuits over rogue AI agents and weighing executive liability.