Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. PandailyNewsAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. PandailyNewsAI score52

    Huawei's KV cache storage faces a missing SSD endurance standard

    AIHuawei's OceanStor M900 and Nvidia's CMX move reusable inference KV cache into a shared storage tier, but no agreed SSD specification exists for it. A storage executive said endurance requirements for one design rose from 3 to 9 drive writes per day and could change again. Industry sources expect convergence to take 6 to 12 months, with another year for development and validation.

  3. PandailyNewsAI score40

    Huawei Opens DevEco Studio Public Beta on HarmonyOS PCs with DevEco Code and CLI

    AIHuawei has opened its DevEco Studio for HarmonyOS PCs to public beta, alongside first public betas of the AI tool DevEco Code and the agent toolkit DevEco CLI. The beta requires HarmonyOS 7.0.0.107 or later, at least 16 GB of memory and 100 GB of storage, and runs on several MateBook models and the MatePad Edge. DevEco Code ships with Zhipu AI's GLM-5.3 and GLM-5.1 models and supports third-party model connections.

  4. Tencent HyOfficialAI score47

    Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.

  5. SiliconANGLE · AINewsAI score38

    CoreWeave Adds RL Rollouts and Forge Platform to Target AI Inference Bottlenecks

    AICoreWeave is layering managed services over its infrastructure to address AI inference bottlenecks, including a preview capability called CoreWeave RL Rollouts that improved model reload latency by 15x versus a baseline configuration in testing. The capability is built on Nvidia's Dynamo framework, and the features are packaged into CoreWeave Forge, a platform that is free to start with paid tiers offering additional capabilities.

  6. Sakana AIOfficialAI score36

    Sakana AI's technology powers Iris's physician evidence search tool

    AIIris Inc.'s medical evidence search tool Evidence Finder has adopted Sakana AI's technology for answering physicians' questions. The system searches the literature and generates answers that cite their sources, handling literature comparison, synthesis, and answer generation.

    Image from @SakanaAILabs's post
  7. Teknium 🪽XAI score23

    TinyFish browser backend added to Hermes plugins catalog

    AITeknium announced that TinyFish, a new browser backend, is now available on the Hermes plugins catalog. According to a related post, TinyFish is added as a first-party Hermes plugin that lets agents search and fetch the live web for free, and it can be installed with hermes plugins install tinyfish.

  8. Teknium 🪽XAI score20

    Hermes Desktop Generates Intelligent UI Embeds Unprompted

    AITeknium called a Hermes Desktop demo "pretty sick" after Jonathan Bylos reported that Hermes Agent produced an intelligent UI embed during a design discussion without being asked. Bylos said the feature has been running in Hermes Desktop for a few days.

  9. Higgsfield AI 🧩OfficialAI score36

    Higgsfield's Katana makes a video entirely from Three.js code

    AIA Higgsfield post says a video was made with no video AI model, Blender, or After Effects, using only Three.js code rendered over 12 hours. The video was made with Higgsfield Katana inside Claude, which the post introduces as an AI video editing tool powered by Claude Motion and available via Higgsfield MCP.

    Video from @higgsfield's post
  10. IThome · AINewsAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  11. QbitAINewsAI score62

    Google launches Gemini agent for office work, able to call Claude models

    AIGoogle Cloud introduced the Gemini agent, a general office agent that can search, write emails, build slides, analyze data, run code, and coordinate sub-agents. It can take on an enterprise identity with email, calendar, and account, and it selects underlying models automatically, including Anthropic's Claude. The article presents this alongside OpenAI's Dots and Meta's Muse as competing office and personal agents.

  12. LangChain BlogOfficialAI score42

    Snyk Assist: How Snyk Turned an Internal Support Agent into a Customer Feature

    AISnyk moved its internal support agent, Snyk Assist, into the core Snyk product in September 2026, giving every paying customer access. Built on LangChain and LangGraph with observability in LangSmith, the agent answers questions in plain language and can open support cases or log feature requests. It runs as a single agent behind Slack, web and API surfaces, with tools attached per user permissions.

  13. The Guardian · AINewsAI score42

    Anthropic bans sustained abusive or cruel behavior toward Claude

    AIAnthropic has barred users from exhibiting "sustained and needless abusive or cruel behavior" toward its models, according to a policy change first reported by The Verge. The San Francisco-based company says the ban does not apply to common user frustrations, model testing, or "dark creative themes." The change follows an August feature that lets Claude end conversations when a user is persistently harmful, which Anthropic framed as a safeguard for AI welfare.

  14. CNBC · TechnologyNewsAI score40

    Nvidia-backed Firmus withdraws planned A$11 share IPO citing market volatility

    AIAustralian AI data center operator Firmus, backed by Nvidia, has withdrawn its planned initial public offering, citing market volatility and conditions. Its board concluded the proposed terms did not adequately reflect the company's business strength and long-term growth outlook. Firmus said it will now pursue private market capital and consider other public and private options.

  15. OpenAI · YouTubeOfficialAI score36

    Codex moves from single-player to multiplayer at OpenAI DevDay 2026

    AIOpenAI's DevDay 2026 session demonstrates Codex shifting from a single-user tool to a team-oriented agent. The session shows a persistent personal agent investigating a 2am outage, from the first Slack message through a reviewed fix, using voice, Appshots, plugins, and meeting notes to keep the team informed.

  16. OpenAI · YouTubeOfficialAI score22

    How OpenAI Puts ChatGPT to Work | DevDay 2026

    AIOpenAI's DevDay 2026 session shows how its teams use ChatGPT to manage launches, research competitors, and turn expertise into tools. The source provides no further details on specific features, metrics, or availability.

  17. OpenAI · YouTubeOfficialAI score36

    Stop Overpaying for Intelligence: Cutting AI Costs at DevDay 2026

    AIOpenAI's DevDay 2026 session shows developers how to cut AI costs while preserving quality. It covers measuring cost per completed task, choosing models and reasoning effort, and using prompt caching, Programmatic Tool Calling, and the Batch API.

  18. AWS Machine Learning BlogOfficialAI score40

    Cornerstone cuts database diagnosis time 78% with Orion AI on Amazon Bedrock

    AICornerstone OnDemand built Orion AI, a multi-agent system using Amazon Bedrock and the open source Strands Agents framework, that cut database diagnosis from 45 minutes to 10, a 78% reduction. The system also reduced manual lifecycle steps from more than 10 to a single interaction and filtered redundant alerts by a median of 65%. A three-person team delivered it in six months.

  19. PandailyNewsAI score44

    Tencent Rolls Out Vertical AI Agents, From CraftBuddy to MusicBuddy, Across Its Apps

    AITencent introduced roughly eight vertical AI agent products between July and September, each built on an existing business such as games, music, WeChat or video conferencing. CraftBuddy generates 2D and 3D games from natural-language prompts with more than 40,000 prebuilt art assets, while MusicBuddy adds an AI chat panel to multitrack music editing. The products extend the "Buddy" family alongside WorkBuddy and CodeBuddy.

  20. PandailyNewsAI score38

    Doubao App Launches on Huawei HarmonyOS PCs with Doubao-Seed-2.1 Pro and Turbo

    AIByteDance's Doubao app is now available as a native client on Huawei HarmonyOS PCs, shipping with Doubao-Seed-2.1 Pro and Doubao-Seed-2.1 Turbo. The app offers Chat and Work modes, with Work supporting goal-based tasks, file and folder access, scheduled tasks and parallel browsing tabs. Users can download it by searching for Doubao in the HarmonyOS app store.

  21. LeiphoneNewsAI score58

    Alibaba's Qwen Roadmap Targets 5T to 10T Parameters Amid Self-Improving Model Work

    AIAt the Apsara Conference, Alibaba's Qwen team outlined a roadmap of Qwen4 followed by Qwen4.5 and Qwen5, aiming for 5T to 10T parameters. The article notes that Qwen3.8 reached 2.4T parameters and that Qwen3.8-Flash activates 6B parameters per inference while cutting training cost to one-ninth. It also describes Qwen3.8-Max running model-driven experiments in chip design and inference optimization, and multimodal updates including a video model slated for November.

  22. QbitAINewsAI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    AIAnthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  23. QbitAINewsAI score80

    GPT-6 rolls out to free ChatGPT users with interactive answer interfaces

    AIOpenAI began rolling out GPT-6 to free and Go ChatGPT users on October 8, replacing GPT-5.6 Luna with GPT-6 Luna, while paid users receive GPT-6 Sol. The update adds Intelligent UI, which generates charts, buttons, and interactive tools inside chat answers. OpenAI's safety report shows gains on jailbreak and instruction-hierarchy tests but also regressions in some self-harm, sexual, and emotional-dependence evaluations, including for under-18 users.

    Why it matters: The report pairs the free-tier GPT-6 Luna rollout with its safety evaluations, showing where capability gains and regressions appear for younger users.

  24. SiliconANGLE · AINewsAI score58

    Meta, Walmart, Stripe and Sierra back Personal Agent Protocol for AI shopping agents

    AIMeta, Walmart, Stripe and Sierra Technologies are partnering on a Personal Agent Protocol that would define how AI agents interact with businesses online. Sierra co-founder Bret Taylor said it will handle authentication and give companies visibility into personal agents' activity, and it is open for anyone to implement. The effort responds to concerns from six major banks, which urged the industry last month to set standards around transparency, safety, privacy, choice and interoperability.

  25. SiliconANGLE · AINewsAI score34

    Caterpillar and CoreWeave Shorten Physical AI Learning Loop for Construction Machines

    AICaterpillar and CoreWeave are working to cut the time needed to train construction machines from months to hours, according to Caterpillar's Brandon Hootman. Caterpillar's ecosystem holds about 18 petabytes of federated machine, dealer and customer data, and a single machine can generate terabytes of LiDAR, camera and control data in a day. The partners use AI models, working with Nvidia, to annotate incoming field data so it can feed simulation and training within the same workday.

  26. SiliconANGLE · AINewsAI score47

    Atlassian unveils Agentic Multiplayer Protocol for humans and AI agents to collaborate

    AIAtlassian announced the Agentic Multiplayer Protocol (AMP) at Team '26 Europe, a platform update designed to let humans and AI agents collaborate on shared context and tasks in one governed workspace. Agents receive administrator-assigned identities with defined authority and scope, and the company's Rovo Work mode handles multi-step tasks that humans review and approve. Atlassian's MCP server now exposes 200 tools and handles roughly 15 million tool calls daily.

  27. SiliconANGLE · AINewsAI score47

    Kore.ai launches Autoloop to tune enterprise AI agents after deployment

    AIKore.ai launched Autoloop, an optimization engine that automatically adjusts AI agents built on its Kore.ai Agent Platform to meet business-set goals, including after deployment. The engine scores each proposed change against goals such as task completion, business-rule adherence, accuracy and cost. Autoloop is available now to all customers on the Artemis edition of the Kore.ai Agent Platform.

  28. SiliconANGLE · AINewsAI score23

    Infor pairs industry-specific AI agents with forward-deployed engineers for process automation

    AIInfor is building industry-specific AI agents on its Infor OS foundation and open architecture, according to CEO Kevin Samuelson. Infor says two in three businesses find off-the-shelf AI does not adequately address their industry's needs. The company pairs customers with forward-deployed engineers, and Samuelson says prototypes can now take one to three weeks.