Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. JetBrains AI BlogOfficialAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  2. Google Cloud · AI & Machine LearningOfficialAI score78

    Google Cloud launches Gemini agent as single universal work agent

    AIGoogle Cloud announced the Gemini agent, a single agent that answers questions, handles knowledge work, creates media, and writes and runs code from one prompt box. It runs in the cloud with persistent memory, uses multi-agent orchestration, and adds Workspace integration, domain skills for data and industries, identity-based governance through Agent Gateway, and spend caps. The source also cites customer deployments and says nearly 80% of Google Cloud customers use its AI products.

    Why it matters: The announcement shows how a single work agent spans chat, Workspace, data analysis, governance, and cost controls, useful for judging enterprise agent deployment scope.

  3. ElevenLabs BlogOfficialAI score26

    How to build a meeting transcription API with Scribe v2 and Scribe v2 Realtime

    AIElevenLabs explains how to build meeting transcription products using its Scribe v2 and Scribe v2 Realtime models through its API. Real-time transcription suits live captions and in-meeting bots, while batch transcription suits post-meeting notes and records, with Scribe v2 Realtime reporting 150 ms latency and supporting up to 50 key terms for prompting.

  4. GuizangXAI score22

    Guizang suspects Grok bot may already run Claude Opus 5.5

    AIGuizang (@op7418) suspects the Grok bot may already be running Claude Opus 5.5, based on strong results on complex tasks. The post is a brief speculation without benchmark data or official confirmation, and it references a separate post on using a Grok bot to automatically generate a daily AI news video in the cloud.

  5. meng shaoXAI score55

    Tencent Cloud open-sources Octop, a self-hosted multi-agent AI assistant platform

    AITencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.

    Image from @shao__meng's post
  6. Ars Technica · AINewsAI score38

    Nvidia's Halos safety platform extends from robotaxis to humanoid and warehouse robots

    AINvidia's Halos software platform, originally built for autonomous vehicles, has been adapted for robotics, according to Ars Technica. The system monitors hardware and software for failures, isolates safety-critical workloads, and includes simulation tools and an inspection lab for robotics developers. Because safety requirements vary widely between a robotic vacuum and a warehouse forklift, Nvidia made the platform programmable so developers can define custom safety functions.

  7. 🚨 AI News | TestingCatalogXAI score23

    Antigravity's agent renamed "Chief of stuff" in latest update

    AIGoogle's Antigravity agent has been renamed "Chief of stuff" in its latest update, which the poster reads as a promotion. The poster wonders whether Antigravity could become a home for Google's own agents, and background notes that Google is prototyping a voice agent internally called "Concierge," which appears to be a very early version.

    Image from @testingcatalog's post
  8. meng shaoXAI score49

    LangChain adds three Deep Agents Skills upgrades: tool binding, pinning, reloading

    AILangChain has added three engineering upgrades to Skills in its Deep Agents framework: tool-binding Skills, pinned Skills, and mid-thread reloading. Tool-binding lets a SKILL.md declare tools via metadata.include_tools, so tools are injected only when the Skill is read, and pinned Skills inject full instructions before the next model call, skipping a round trip. Setting skills_metadata to None rescans the Skills library mid-thread without restarting, at the cost of invalidating the cache.

    Image from @shao__meng's post
  9. GuizangXAI score26

    Grok bot auto-generates a daily AI news video in the cloud

    AIThe author set up a Grok bot to produce a daily morning AI news video on a schedule, running content collection, code writing, and video rendering entirely on Grok's cloud virtual machine without local computers. The author says the results are quite good and shares the full prompt so others can run the same workflow with their own Grok bot.

    Video from @op7418's post
  10. The SequenceBlogAI score33

    Decision Models Like Jev Aim to Make Routine AI Judgments Cheap

    AIDecision models are small systems built to make fast semantic judgments, such as routing support tickets or deciding whether to escalate, without the cost of a full reasoning model. The article says Amazon, Cloudflare, and OpenAI are competing in this space, with cost economics as a central concern.

  11. The DecoderNewsAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

    This story has a top pick“Anthropic releases Claude Haiku 5.5 as its cheapest, fastest small model”

  12. Lucas Beyer (bl16)XAI score31

    Lucas Beyer shares his annual State of AI report tradition

    AILucas Beyer posted that the State of AI report has become his favourite annual tradition. The post quotes Nathan Benaich's announcement of the 9th annual report, covering AI building better AI, world models, physical AI, geopolitics, cyber defense, and 12-month predictions.

    Image from @giffmana's post
  13. The Guardian · AINewsAI score36

    Altman Says AI Will Cause 'Bad Things' as Columnist Cites Deaths and Lawsuits

    AIOpenAI CEO Sam Altman told Politico that the world should accept some bad things from AI for its benefits, a stance columnist Moustafa Bayoumi calls problematic. The column cites lawsuits over ChatGPT-linked suicides, a February strike on a Minab school that killed at least 120 children with a US military AI system (Palantir's Maven) implicated, and a chatbot error that nearly triggered a military interception.

  14. The DecoderNewsAI score34

    Teen Hiker Needs Helicopter Rescue After Following Claude's Route Advice

    AIA 16-year-old hiker had to be airlifted from a dangerous rock face on Crown Mountain near Vancouver after using Anthropic's Claude to plan a route to the summit. He ended up on the Widowmaker Arete, a steep cliff requiring climbing gear, and called police when he got stuck on a ledge. Rescue manager Paul Markey said Claude has no actual knowledge of locations or terrain and is no substitute for experience and common sense.

  15. X.PINXAI score40

    Tian Keyu's startup raises nearly $30M to build visual-vocabulary video AI

    AITian Keyu's unnamed startup has raised nearly $30M from 5Y Capital and IDG at a $200M post-money valuation, according to Bloomberg. The NeurIPS 2024 award-winning researcher's 10-person team is developing a 200,000-symbol visual vocabulary to help AI process video. Tian claims the approach could cut video-generation costs at least tenfold, with a full model release planned for 2027 and no product yet.

    Image from @thexpin's post
  16. TechRadar · AINewsAI score25

    HP survey finds three in five UK business leaders say AI has created new roles

    AIA HP survey of nearly 20,000 desk-based workers globally found three in five UK business leaders say AI adoption has created new roles or teams in their organizations. Only 10% said AI is primarily replacing or reducing roles, while 52% of UK workers use employer-provided AI tools daily or weekly, up from 38% last year. Still, 38% of UK knowledge workers worry AI could replace their roles.

  17. vLLMOfficialAI score62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    AIvLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

    Why it matters: The release lists concrete changes across serving, scheduling, and model support, which helps operators judge whether the upgrade affects their deployment path.

    Image from @vllm_project's post
  18. QbitAINewsAI score49

    Manus Returns to Beijing, Hiring 17 Roles After Raising Over $500M

    AIManus parent company Butterfly Effect has completed a new financing round of over $500 million, led by Boyu Capital and IDG Capital, and is rebuilding a team in Beijing to develop AI Agent products for the Chinese market. Its recruitment page lists 17 open positions, up from 11 before the National Day holiday, including AI Agent product manager, Agent Harness engineer, Agent evaluation engineer, and LLM algorithm engineer roles.

  19. Wired · AINewsAI score36

    Tristan Harris's Center for Humane Technology lays off about half its staff

    AIThe Center for Humane Technology is laying off about half of its 16 non-founder employees and ending its policy research and litigation work. The organization will refocus on "founder-led" initiatives built around cofounder Tristan Harris, according to WIRED, after its board concluded that operating as both an advocacy group and a think tank had stretched it too thin.

  20. QbitAINewsAI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  21. meng shaoXAI score24

    Alibaba's four takeaways on AI Native R&D from its handbook

    AIAlibaba's official handbook on AI Native R&D identifies four open challenges: infrastructure engineering complexity, enterprise knowledge assets not yet agent-friendly, organizational design, and the pace of AI iteration. The post's author argues that Agent Infra must suit non-deterministic agent operation and that enterprise knowledge needs top-down structuring and governance. The author also notes that organizational resistance in large companies makes AI adoption harder than in startups.

    Image from @shao__meng's post
  22. GuizangXAI score22

    Anthropic's cheaper Haiku 5.5 draws backlash over China rivals

    AIGuizang (@op7418) says he posted news of Anthropic's price cut for Haiku 5.5 and was attacked by commenters who view the pricing as aimed at Chinese models. He compares Haiku 5.5 with DeepSeek-V4.1 flash and Zhipu's GLM 5.3 flash, finding it cheaper than both and scoring one point higher than GLM 5.3 flash on Terminal Bench 4.0 in AA's test.

  23. The DecoderNewsAI score72

    AI hacking tools let a likely single attacker breach multiple South Korean banks

    AIA suspected Chinese-speaking attacker breached several South Korean financial institutions between late September and early October 2026, reportedly stealing over 25,000 records from Shinhan Bank alone. The attacker used ARTEX, a Chinese open-source tool that uses AI language models to automate finding security flaws, and models named in the report include DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6.

    Why it matters: The case shows how AI-driven penetration tools let one attacker breach several banks in a short window, a risk experts had warned about.

  24. QbitAINewsAI score32

    Geely unveils AI-powered Geely Smart Charging with 2250 kW peak charging power

    AIGeely Automobile Group launched its Geely Smart Charging technology on September 23, 2026, reaching a 2250 kW peak single-gun charging power and keeping maximum temperature at or below 65°C. The system, co-developed with StepFun and built on its PowerMind energy model, reportedly raises battery cycle life by more than 20% and targets county-level coverage by the end of 2027.

  25. The Guardian · AINewsAI score42

    Co-op places legal services staff under AI monitoring of customer phone calls

    AIThe Co-op's legal services arm is using AI models to record and score staff phone calls with customers seeking probate and will advice, with a model from OpenAI assessing more than 50 aspects of each call. Managers use pass and fail scores to analyse employee performance, and the Co-op says the system is a support tool, not a decision-maker. Trade unions and a whistleblower have criticised the monitoring as oppressive.

  26. Gizmodo · AINewsAI score45

    OpenAI Reportedly Regains Nearly Half of AI Compute Market Share From Anthropic in 2026

    AIAccording to a Wall Street Journal report citing OpenRouter data, OpenAI's share of AI compute routed through the platform rose from under 25% at the start of 2026 to nearly 50% last month. The data comes largely from AI-native startups, with some legacy tech companies also included. The report comes as OpenAI reportedly shifted focus from Sora and an erotica generator toward productivity and business customers.

  27. MIT Technology Review · AINewsAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  28. MarkTechPostNewsAI score55

    Architect launches Liquid Inference, an auction-based router for LLM inference

    AIArchitect Financial Technologies launched Liquid Inference, an LLM router where providers bid to serve each request and the buyer pays the lowest offer meeting its rules. Developers can switch by changing the base URL, and the first 500 users get $20 of free inference. The source notes that fees, the provider list, and latency data are not yet public.

  29. MarkTechPostNewsAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.