Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. OpenRouter · New modelsBlogAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  2. SiliconANGLE · AINewsAI score62

    Google Cloud launches Gemini agent for enterprise work across devices and apps

    AIGoogle Cloud introduced Gemini agent, a unified AI assistant that acts autonomously, generates code, and completes work across web, mobile, desktop, and third-party apps. It runs jobs on models matched to each task, including Gemini Flash and a flagship frontier model, with Anthropic Claude models also available. Hard spend limits per project let companies enforce budgets and charge AI costs to departments.

  3. JetBrains AI BlogOfficialAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  4. Google Cloud · AI & Machine LearningOfficialAI score78

    Google Cloud launches Gemini agent as single universal work agent

    AIGoogle Cloud announced the Gemini agent, a single agent that answers questions, handles knowledge work, creates media, and writes and runs code from one prompt box. It runs in the cloud with persistent memory, uses multi-agent orchestration, and adds Workspace integration, domain skills for data and industries, identity-based governance through Agent Gateway, and spend caps. The source also cites customer deployments and says nearly 80% of Google Cloud customers use its AI products.

    Why it matters: The announcement shows how a single work agent spans chat, Workspace, data analysis, governance, and cost controls, useful for judging enterprise agent deployment scope.

  5. GuizangXAI score22

    Guizang suspects Grok bot may already run Claude Opus 5.5

    AIGuizang (@op7418) suspects the Grok bot may already be running Claude Opus 5.5, based on strong results on complex tasks. The post is a brief speculation without benchmark data or official confirmation, and it references a separate post on using a Grok bot to automatically generate a daily AI news video in the cloud.

  6. meng shaoXAI score55

    Tencent Cloud open-sources Octop, a self-hosted multi-agent AI assistant platform

    AITencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.

    Image from @shao__meng's post
  7. 🚨 AI News | TestingCatalogXAI score23

    Antigravity's agent renamed "Chief of stuff" in latest update

    AIGoogle's Antigravity agent has been renamed "Chief of stuff" in its latest update, which the poster reads as a promotion. The poster wonders whether Antigravity could become a home for Google's own agents, and background notes that Google is prototyping a voice agent internally called "Concierge," which appears to be a very early version.

    Image from @testingcatalog's post
  8. meng shaoXAI score49

    LangChain adds three Deep Agents Skills upgrades: tool binding, pinning, reloading

    AILangChain has added three engineering upgrades to Skills in its Deep Agents framework: tool-binding Skills, pinned Skills, and mid-thread reloading. Tool-binding lets a SKILL.md declare tools via metadata.include_tools, so tools are injected only when the Skill is read, and pinned Skills inject full instructions before the next model call, skipping a round trip. Setting skills_metadata to None rescans the Skills library mid-thread without restarting, at the cost of invalidating the cache.

    Image from @shao__meng's post
  9. GuizangXAI score26

    Grok bot auto-generates a daily AI news video in the cloud

    AIThe author set up a Grok bot to produce a daily morning AI news video on a schedule, running content collection, code writing, and video rendering entirely on Grok's cloud virtual machine without local computers. The author says the results are quite good and shares the full prompt so others can run the same workflow with their own Grok bot.

    Video from @op7418's post
  10. QbitAINewsAI score49

    Manus Returns to Beijing, Hiring 17 Roles After Raising Over $500M

    AIManus parent company Butterfly Effect has completed a new financing round of over $500 million, led by Boyu Capital and IDG Capital, and is rebuilding a team in Beijing to develop AI Agent products for the Chinese market. Its recruitment page lists 17 open positions, up from 11 before the National Day holiday, including AI Agent product manager, Agent Harness engineer, Agent evaluation engineer, and LLM algorithm engineer roles.

  11. meng shaoXAI score24

    Alibaba's four takeaways on AI Native R&D from its handbook

    AIAlibaba's official handbook on AI Native R&D identifies four open challenges: infrastructure engineering complexity, enterprise knowledge assets not yet agent-friendly, organizational design, and the pace of AI iteration. The post's author argues that Agent Infra must suit non-deterministic agent operation and that enterprise knowledge needs top-down structuring and governance. The author also notes that organizational resistance in large companies makes AI adoption harder than in startups.

    Image from @shao__meng's post
  12. MIT Technology Review · AINewsAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  13. MarkTechPostNewsAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  14. GiXAI score37

    Termix welcomes RocketBNB as a brand agent on Agent.family

    AITermix announced that RocketBNB has registered on Agent.family as a brand agent offering AI-powered research and knowledge services. The author says the marketplace, where agents hire agents, now hosts brands that offer their name as a hireable service. The platform reported 521,719 on-chain jobs settled and $27,012,842 processed as of TOKEN2049 on 7/10.

    Video from @Girlgym67's post
  15. Teknium 🪽XAI score22

    Hermes Agent coming to ASUS RTX Spark PCs soon

    AITeknium announced that Hermes Agent will be available on the ASUS RTX Spark PC when it releases soon. The post links the upcoming support to ASUS's ProArt P16 and P14 creator laptops, which ASUS says are powered by NVIDIA and Hermes AI.

  16. MIT Technology Review · AINewsAI score26

    AVEVA's Arti Garg outlines a safer path to autonomous industrial AI

    AIAVEVA chief technologist Arti Garg argues industrial AI should augment rather than replace human supervisors in critical decisions, with guardrails defining where automated systems can act. She says organizations must rethink business processes and safeguards as foundation models, physical AI, and agentic AI enable more complex automation.

  17. Ant LingOfficialAI score22

    Ant Ling's Ling-3.1-flash now live on AI/ML API

    AIAnt Ling announced a day-zero collaboration with AI/ML API, making Ling-3.1-flash available there for agentic and cowork scenarios. AI/ML API describes it as a 560B-parameter MoE model with about 25B active per token and up to 1M context, built for agents, coding, and long documents. The model is free to try on AI/ML API until October 13.

  18. PandailyNewsAI score45

    KargoBot Launches Mixed Autonomous Freight Network in Ordos With Cabless Robots

    AIKargoBot has started a scaled AI freight network in Qipanjing, Ordos, combining human-driven trucks, autonomous trucks with cabs, and cabless transport robots on one system. The company says cabless robots could raise economic gain per vehicle from 20 percent to more than 30 percent, a target it has not audited. Platooning reportedly improves gross margin by about 10 to 18 percent versus manned haulage, with one lead driver able to head up to five follower trucks.

  19. PandailyNewsAI score37

    Tencent WorkBuddy Builds a WeChat Mini-Program From a Prompt to Preview

    AITencent's WorkBuddy agent can take a plain-language request through to a WeChat mini-program preview and a publish request, according to a hands-on product test reported on October 8. In the test, the agent built a voice notebook with cloud login, storage and files, using Tencent's wand-asr-v1 for speech-to-text and GLM-5.3-Flash for sorting notes. WeChat's own review and filing steps remain outside the agent, so the test does not show that every mini program goes live automatically.

  20. Latent SpaceBlogAI score73

    Claude Haiku 5.5 launches at GPT-6 Luna pricing with 1M context

    AIAnthropic released Claude Haiku 5.5, priced the same as OpenAI's GPT-6 Luna, with a 1M-token context window. Artificial Analysis scored it 43 on its Intelligence Index, slightly ahead of GPT-6 Luna at 38, but it uses about 3x more output tokens at max effort.

    Why it matters: The roundup pairs Anthropic's launch claims with Artificial Analysis's independent numbers, showing where Haiku 5.5 is cheap and strong and where token use offsets its price.

  21. howie.seriousXAI score22

    Grok bot's X rate limit is 1,000 calls per day

    AIThe Grok bot's rate limit on X is 1,000 calls per day, which the author finds more than sufficient. Previously, using X's official API cost $20 per top-up and was spent quickly, while now it can be used for free.

    Image from @howie_serious's post
  22. howie.seriousXAI score46

    Agent bottleneck is human understanding, not model capability

    AIThe author argues that in agent workflows, the real bottleneck is whether users can precisely express requirements, not the model or agent capability. When people work outside their expertise, they lack the precision needed for prompts and plans, forcing many imprecise iterations that waste time and tokens. The suggested fix is to have the model first teach the unfamiliar domain knowledge before acting.

  23. howie.seriousXAI score22

    Howie Serious says Grok bot can gather X information for users

    AIHowie Serious (@howie_serious) says his Grok bot has found its first powerful use case: a personal agent that collects Twitter information. He argues this beats humans scrolling phones, manually gathering posts, or getting absorbed in endless feeds.

    Image from @howie_serious's post
  24. Mastra BlogOfficialAI score29

    Mastra Launches Agency Program with Five Certified Partners to Build Agents

    AIMastra launched the Mastra Agency Program, a network of certified agencies and consultancies that build Mastra agents for clients. The launch includes five partners: Deerfield Group, Blue Drop Labs, Frontleap, Handpicked, and Young Security. Every partner has been vetted by Mastra's FDE team and receives direct access to Mastra's leadership and regular roadmap updates.

  25. Anthropic ResearchOfficialAI score72

    Anthropic launches OSS Scanner, a free AI vulnerability scanner for open-source projects

    AIAnthropic is launching OSS Scanner, an opt-in service that runs periodic security scans of enrolled open-source projects using its strongest models at no cost. Its outputs are fully model-generated without human review, so some reports may be incorrect or invalid, though a pilot found 85 of 97 checked critical and high-severity findings met Anthropic's disclosure bar. Core maintainers of eligible projects can enroll through a GitHub pull request.

  26. Anthropic ResearchOfficialAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

  27. Claude BlogOfficialAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.