Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. 🚨 AI News | TestingCatalogAI score46

    Google announces a unified Gemini agent for Gemini Enterprise work

    AIGoogle has announced a single, universal Gemini agent for Gemini Enterprise as part of its Gemini at Work updates. The agent answers questions, handles knowledge work, creates images and media, and writes and runs code. It works inline in Gmail, Drive, Docs, Slides, Sheets, Chat, and Calendar, with new data and analytics skills for plain-language insights and industry-specific tools for financial services and legal teams.

    Image from @testingcatalog's post
  2. SantiagoAI score38

    Teamily AI lets people and agents share one group chat context

    AITeamily AI now lets users add people and AI agents to the same group chat, so everyone works from shared context. The example shows a branding change handled by research, writing, and website-building agents, with a designer's feedback incorporated and the finished page shared in one continuous conversation. The platform's 2.0 release, which the post describes as opening to everyone, adds real-time human–agent collaboration and multi-model routing.

    Image from @svpino's post
  3. a16z NewsAI score45

    CFOs Are Becoming Builders as AI Reshapes Finance Operations

    AIAI-native tools are removing the data bottleneck that long constrained CFOs, shifting the role toward designing the operating systems that turn data into decisions. Finance teams are adopting AI-native software for ERP, forecasting, procurement, and audit, and "finance engineers" are building custom automations and agents. OpenAI's CFO Sarah Friar describes finance moving toward a zero-day close and continuously updated forecasts.

  4. GuizangAI score22

    Grok bot builds and publishes daily AI news videos on a foldable phone

    AIGuizang (@op7418) says a Grok bot paired with a foldable phone lets him chat on one screen while the bot publishes content on the other. Background post: he set Grok to produce a daily morning AI news video on a schedule, running content collection, code writing, and video rendering entirely on Grok's cloud virtual machine rather than his local computer. He shares the full prompt so others can run the same workflow with their own Grok bot.

    Image from @op7418's post
  5. South China Morning Post · TechAI score60

    Can China match Meta's Muse in the race to harness AI agents?

    AIMeta's Muse personal AI agent, launched on September 8, surpassed 2.5 million downloads in its first two weeks and topped free-app rankings on Apple and Google's US app stores. Its popularity has drawn attention to the emerging market for agent harnesses, where China's biggest internet companies are already competing for position. The excerpt does not provide full details on their specific products.

  6. The Verge · AIAI score41

    Meta's Muse and OpenAI's Dots: can consumers trust AI agents with their lives?

    AIMeta's Muse and OpenAI's Dots are always-on AI agents with animated mascots, pitched to consumers and businesses for tasks like restaurant reservations and inbox triage. Muse is free, while Dots is not, and OpenAI also offers "specialist" Dots for marketing, legal work, and accounting. The discussion centers on privacy and security concerns about giving agents access to credit card details and email.

  7. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

    Image from @OpenBMB's post
  8. Gergely OroszAI score26

    Developers working more with AI tools, citing more context switching

    AISoftware developer Gergely Orosz questions why he is working more despite AI tools, quoting Sam Newman's view that AI was meant to free developers from drudgery. Newman says most developers are doing more work, with more context switching and a loss of the big picture. The quoted post adds that AI assistants are not human partners and that pairing with them fragments the shared mental model of a program.

  9. StepFunAI score60

    StepFun's Step 5 Preview is live on OpenRouter with a week of free access

    AIStepFun says Step 5 Preview is now available on OpenRouter, with a week of free access rolling out across OpenCode, Cline, Nous Research, Kilo Code, and other tools. The company describes it as flagship-tier intelligence for agentic and professional work at substantially lower task cost, letting users switch models without changing their workflow.

    Image from @StepFun_ai's post
  10. The DecoderAI score72

    One public AI agent on AWS could take over every other agent in its region

    AIZenity Labs says a single publicly accessible agent on Amazon Bedrock AgentCore could take over all AgentCore agents in the same AWS account and region. A chat prompt let the researchers query the instance metadata service and steal temporary credentials, and AgentCore's default permissions allowed read, write, and delete access across agents. According to Zenity, AWS made IMDSv2 the default for new deployments and changed the default execution role around August.

    Why it matters: The report traces how one public agent's weak isolation exposed credentials and every other agent in the region, showing why default permissions matter for enterprise deployments.

  11. Databricks BlogAI score35

    How to build governed enterprise apps on Databricks with Replit and Lakebase

    AIReplit and Databricks integration, now generally available with native Lakebase support, lets enterprise teams build apps from plain-language prompts using Replit Agent and deploy them as Databricks Apps. Deployed apps inherit automatic user authentication and Unity Catalog access controls, and Replit Agent auto-provisions a managed Lakebase Postgres database for operational data. Lakebase keeps app-written data inside the Databricks perimeter instead of a separate external database.

  12. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  13. SiliconANGLE · AIAI score62

    Google Cloud launches Gemini agent for enterprise work across devices and apps

    AIGoogle Cloud introduced Gemini agent, a unified AI assistant that acts autonomously, generates code, and completes work across web, mobile, desktop, and third-party apps. It runs jobs on models matched to each task, including Gemini Flash and a flagship frontier model, with Anthropic Claude models also available. Hard spend limits per project let companies enforce budgets and charge AI costs to departments.

  14. ZDNet · AIAI score24

    Google Maps adds Ask Maps food ordering via Square and Uber Eats, plus fan-favorite dining list

    AIGoogle Maps now lets users place restaurant orders through its Gemini-powered Ask Maps chat tool, with Square and Uber Eats joining Toast as partners. Google also released its first fan-favorite dining list, which tracks trending food and drink interest across 10 cities, including a 216% rise in cheeseburger interest in Tokyo over the past year.

  15. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  16. Google Cloud · AI & Machine LearningAI score78

    Google Cloud launches Gemini agent as single universal work agent

    AIGoogle Cloud announced the Gemini agent, a single agent that answers questions, handles knowledge work, creates media, and writes and runs code from one prompt box. It runs in the cloud with persistent memory, uses multi-agent orchestration, and adds Workspace integration, domain skills for data and industries, identity-based governance through Agent Gateway, and spend caps. The source also cites customer deployments and says nearly 80% of Google Cloud customers use its AI products.

    Why it matters: The announcement shows how a single work agent spans chat, Workspace, data analysis, governance, and cost controls, useful for judging enterprise agent deployment scope.

  17. GuizangAI score22

    Guizang suspects Grok bot may already run Claude Opus 5.5

    AIGuizang (@op7418) suspects the Grok bot may already be running Claude Opus 5.5, based on strong results on complex tasks. The post is a brief speculation without benchmark data or official confirmation, and it references a separate post on using a Grok bot to automatically generate a daily AI news video in the cloud.

  18. meng shaoAI score55

    Tencent Cloud open-sources Octop, a self-hosted multi-agent AI assistant platform

    AITencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.

    Image from @shao__meng's post
  19. 🚨 AI News | TestingCatalogAI score23

    Antigravity's agent renamed "Chief of stuff" in latest update

    AIGoogle's Antigravity agent has been renamed "Chief of stuff" in its latest update, which the poster reads as a promotion. The poster wonders whether Antigravity could become a home for Google's own agents, and background notes that Google is prototyping a voice agent internally called "Concierge," which appears to be a very early version.

    Image from @testingcatalog's post
  20. meng shaoAI score49

    LangChain adds three Deep Agents Skills upgrades: tool binding, pinning, reloading

    AILangChain has added three engineering upgrades to Skills in its Deep Agents framework: tool-binding Skills, pinned Skills, and mid-thread reloading. Tool-binding lets a SKILL.md declare tools via metadata.include_tools, so tools are injected only when the Skill is read, and pinned Skills inject full instructions before the next model call, skipping a round trip. Setting skills_metadata to None rescans the Skills library mid-thread without restarting, at the cost of invalidating the cache.

    Image from @shao__meng's post
  21. GuizangAI score26

    Grok bot auto-generates a daily AI news video in the cloud

    AIThe author set up a Grok bot to produce a daily morning AI news video on a schedule, running content collection, code writing, and video rendering entirely on Grok's cloud virtual machine without local computers. The author says the results are quite good and shares the full prompt so others can run the same workflow with their own Grok bot.

    Video from @op7418's post
  22. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

  23. QbitAIAI score49

    Manus Returns to Beijing, Hiring 17 Roles After Raising Over $500M

    AIManus parent company Butterfly Effect has completed a new financing round of over $500 million, led by Boyu Capital and IDG Capital, and is rebuilding a team in Beijing to develop AI Agent products for the Chinese market. Its recruitment page lists 17 open positions, up from 11 before the National Day holiday, including AI Agent product manager, Agent Harness engineer, Agent evaluation engineer, and LLM algorithm engineer roles.

  24. meng shaoAI score24

    Alibaba's four takeaways on AI Native R&D from its handbook

    AIAlibaba's official handbook on AI Native R&D identifies four open challenges: infrastructure engineering complexity, enterprise knowledge assets not yet agent-friendly, organizational design, and the pace of AI iteration. The post's author argues that Agent Infra must suit non-deterministic agent operation and that enterprise knowledge needs top-down structuring and governance. The author also notes that organizational resistance in large companies makes AI adoption harder than in startups.

    Image from @shao__meng's post
  25. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  26. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.