Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

Oct 9Fri
  1. CNBC · TechnologyNewsAI score44

    OpenAI defends firing three safety researchers, citing a breach of trust

    AIOpenAI defended its decision to fire three safety researchers, Jasmine Wang, Tomek Korbak and Mikita Balesni, saying they committed a "significant breach of trust." The company said the dismissals were not about the researchers raising safety concerns, though it agreed with the letter they sent to board members and safety committees about preserving the monitorability of frontier models.

  2. OpenAI NewsroomOfficialAI score45

    OpenAI fires three researchers over sensitive information breach, denies retaliation

    AIOpenAI says it parted ways with researchers Jasmine, Mikita, and Tomek after an internal investigation found they violated policies on handling sensitive information. The company says the decisions were not about raising safety concerns, which it says it encourages, and that it has not terminated any employee for raising concerns. OpenAI also says it is finalizing contracts with third-party safety assessors and will announce details in the coming weeks.

  3. MarkTechPostNewsAI score44

    Google Research RRSI Guide: Mastering Self-Improving AI Agents

    AIMarkTechPost publishes a hands-on tutorial implementing RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent revise its own harness around a frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker benchmarks, but the edit-selection rules are plain Python that the tutorial runs in a simulated environment with a calibrated noise band.

  4. MiniMax Design (H3)OfficialAI score22

    MiniMax H3 Person Remover LoRA erases people from video

    AIA LoRA for MiniMax H3 removes a person from video by tracking them with SAM 3.1 and generating the replacement background in overlapping windows. Users supply the original video and a clean version of its first frame.

  5. X.PINXAI score46

    Manus parent Butterfly Effect raises over $500M led by Boyu Capital

    AIManus parent Butterfly Effect announced funding of over $500M led by Boyu Capital and IDG, with Tencent, HSG and ZhenFund returning, its first disclosed raise since resuming independent operations. The reported $4B post-money valuation was not confirmed. Manus also launched version 2.0 and the personal agent Cue on September 29, and is building a China-focused product team and partnerships with domestic model developers.

    Image from @thexpin's post
  6. X.PINXAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

    Image from @thexpin's post
  7. LeiphoneNewsAI score42

    Doubao Work adds Canvas feature and Doubao 2.1 Lite model

    AIDoubao Work has added a Canvas feature for complex creative tasks, placing materials, design plans and outputs on one infinite canvas where users can keep editing text, colors and layout after images are generated. The update also integrates the lightweight Doubao 2.1 Lite model, aimed at everyday Q&A, document writing, spreadsheets and PPT creation, with optimized response speed and usage consumption.

  8. LeiphoneNewsAI score40

    TRAE Merges TraeWork and TraeCode Into a Full-Chain Development Platform

    AITRAE announced on October 9 that it has merged TraeWork and TraeCode into a single platform offering Agent mode and IDE mode with seamless switching between them. The upgraded product covers desktop, web, and mobile, letting users start tasks on a computer, check progress on mobile, and continue development back on desktop.

  9. ModelScopeOfficialAI score63

    Google releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device search

    AIGoogle released EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0 for private, on-device search and retrieval. It maps text, code, images, video, and audio into one shared space and reports a 9.92-point gain over EmbeddingGemma 1 on MTEB Code. The post lists about 191MB active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro.

    Video from @ModelScope2022's post
  10. QbitAINewsAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

  11. Ethan MollickXAI score34

    Gemini 2.5 models rated comparable to doctors in urgent care advice

    AIIn an urgent care study, physicians rated advice from the older Gemini 2.5 Pro and Gemini 2.5 Flash, which lacked access to patient medical records, as similar in quality to doctors' advice. No safety issues were identified. The author notes that models have improved significantly since.

    Image from @emollick's post
  12. Bloomberg · TechnologyNewsAI score36

    SoftBank Seeks $100 Billion From Gulf Investors for AI Push

    AISoftBank Group Corp. is seeking to raise as much as $100 billion from Gulf investors to expand its AI investments, the Financial Times reported, citing people familiar with the matter. The report did not specify the timing or terms of the proposed fundraising.

  13. Arena.aiOfficialAI score38

    Mistral Large 4 ranks in Agent Arena top 15 at -6.6% net score

    AIMistral Large 4, a preview model from Mistral AI, ranks #43 overall in Agent Arena with a -6.6% net improvement score across more than 5,000 real-world agentic sessions. That is 11 rankings above its predecessor, Mistral Medium 3.5 (-12.60%), and places it in the top 15 labs, the only European lab there. Open weights are expected at the end of October, and at its current score the model would rank #13 among open models.

    Image from @arena's post
  14. X.PINXAI score49

    Apple's HomeHub ships only after LLMs finally improved Siri

    AIAccording to a source familiar with the project, Apple's homeOS was finished years ago, but the HomeHub was held back until large language models made Siri good enough to serve as its voice-driven interface. The hardware team reportedly refused to sign off on a device whose main interface was Siri, given its long-running poor performance. HomeHub is slated for an Oct 13 unveiling alongside a smart-home push with LG.

    Image from @thexpin's post

Oct 8

Oct 8Thu
  1. TechNode · AINewsAI score42

    XPENG names robotaxi service YOYO and opens invite-only public testing

    AIXPENG has named its robotaxi business XPENG YOYO and launched a ride-hailing mini program that lets invited members of the public test the service. The company says its first production robotaxi, based on the flagship GX model, rolled off the line in May 2026, and it has completed more than 2,000 internal test rides in Guangzhou. XPENG says YOYO uses four in-house Turing AI chips delivering 3,000 TOPS, a second-generation VLA model, and a vision-based approach without high-definition maps or LiDAR.

  2. QwenOfficialAI score22

    Free week of Qwen3.8-Max, Qwen3.8-Flash, and Wan3.0 on GMI Cloud

    AIQwen3.8-Max, Qwen3.8-Flash, and Wan3.0 are available free for a week on GMI Cloud, which is extending the offer by seven days and raising rate limits across all three models. GMI Cloud is also running a contest where three winners each receive $200 cash plus $200 in GMI credits for the most creative, most challenging, or most effort-driven projects built with Qwen or Wan.

  3. ClaudeOfficialAI score46

    Anthropic pauses Claude Startups Team and API credit offers amid demand

    AIAnthropic has paused the Claude Team plan and $1,000 API credit offers for its Claude Startups program after underestimating demand, with hundreds of thousands of applicants. Claimed offers will remain in accounts, but some approved applicants who had not yet claimed their offers will lose access as applications are re-reviewed. Startup Stack and Applied AI office hours remain available to accepted members.

  4. Higgsfield AI 🧩OfficialAI score36

    Higgsfield Katana adds community presets for Claude video editing

    AIHiggsfield has released community presets for Higgsfield Katana, its AI video editing tool available inside Claude. Users can pick a preset for motion graphics, 3D animations, product launches, fashion, car, travel, or aura-farming edits, then add their own characters, products, or clothes to recreate it in Claude. More presets are coming soon.

    Video from @higgsfield's post
  5. PandailyNewsAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  6. PandailyNewsAI score40

    Huawei Opens DevEco Studio Public Beta on HarmonyOS PCs with DevEco Code and CLI

    AIHuawei has opened its DevEco Studio for HarmonyOS PCs to public beta, alongside first public betas of the AI tool DevEco Code and the agent toolkit DevEco CLI. The beta requires HarmonyOS 7.0.0.107 or later, at least 16 GB of memory and 100 GB of storage, and runs on several MateBook models and the MatePad Edge. DevEco Code ships with Zhipu AI's GLM-5.3 and GLM-5.1 models and supports third-party model connections.

  7. Tencent HyOfficialAI score47

    Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.

  8. SiliconANGLE · AINewsAI score38

    CoreWeave Adds RL Rollouts and Forge Platform to Target AI Inference Bottlenecks

    AICoreWeave is layering managed services over its infrastructure to address AI inference bottlenecks, including a preview capability called CoreWeave RL Rollouts that improved model reload latency by 15x versus a baseline configuration in testing. The capability is built on Nvidia's Dynamo framework, and the features are packaged into CoreWeave Forge, a platform that is free to start with paid tiers offering additional capabilities.

  9. Sakana AIOfficialAI score36

    Sakana AI's technology powers Iris's physician evidence search tool

    AIIris Inc.'s medical evidence search tool Evidence Finder has adopted Sakana AI's technology for answering physicians' questions. The system searches the literature and generates answers that cite their sources, handling literature comparison, synthesis, and answer generation.

    Image from @SakanaAILabs's post
  10. Teknium 🪽XAI score20

    Hermes Desktop Generates Intelligent UI Embeds Unprompted

    AITeknium called a Hermes Desktop demo "pretty sick" after Jonathan Bylos reported that Hermes Agent produced an intelligent UI embed during a design discussion without being asked. Bylos said the feature has been running in Hermes Desktop for a few days.

  11. Higgsfield AI 🧩OfficialAI score36

    Higgsfield's Katana makes a video entirely from Three.js code

    AIA Higgsfield post says a video was made with no video AI model, Blender, or After Effects, using only Three.js code rendered over 12 hours. The video was made with Higgsfield Katana inside Claude, which the post introduces as an AI video editing tool powered by Claude Motion and available via Higgsfield MCP.

    Video from @higgsfield's post
  12. IThome · AINewsAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  13. QbitAINewsAI score62

    Google launches Gemini agent for office work, able to call Claude models

    AIGoogle Cloud introduced the Gemini agent, a general office agent that can search, write emails, build slides, analyze data, run code, and coordinate sub-agents. It can take on an enterprise identity with email, calendar, and account, and it selects underlying models automatically, including Anthropic's Claude. The article presents this alongside OpenAI's Dots and Meta's Muse as competing office and personal agents.

  14. LangChain BlogOfficialAI score42

    Snyk Assist: How Snyk Turned an Internal Support Agent into a Customer Feature

    AISnyk moved its internal support agent, Snyk Assist, into the core Snyk product in September 2026, giving every paying customer access. Built on LangChain and LangGraph with observability in LangSmith, the agent answers questions in plain language and can open support cases or log feature requests. It runs as a single agent behind Slack, web and API surfaces, with tools attached per user permissions.

  15. The Guardian · AINewsAI score42

    Anthropic bans sustained abusive or cruel behavior toward Claude

    AIAnthropic has barred users from exhibiting "sustained and needless abusive or cruel behavior" toward its models, according to a policy change first reported by The Verge. The San Francisco-based company says the ban does not apply to common user frustrations, model testing, or "dark creative themes." The change follows an August feature that lets Claude end conversations when a user is persistently harmful, which Anthropic framed as a safeguard for AI welfare.

  16. CNBC · TechnologyNewsAI score40

    Nvidia-backed Firmus withdraws planned A$11 share IPO citing market volatility

    AIAustralian AI data center operator Firmus, backed by Nvidia, has withdrawn its planned initial public offering, citing market volatility and conditions. Its board concluded the proposed terms did not adequately reflect the company's business strength and long-term growth outlook. Firmus said it will now pursue private market capital and consider other public and private options.

  17. OpenAI · YouTubeOfficialAI score36

    Codex moves from single-player to multiplayer at OpenAI DevDay 2026

    AIOpenAI's DevDay 2026 session demonstrates Codex shifting from a single-user tool to a team-oriented agent. The session shows a persistent personal agent investigating a 2am outage, from the first Slack message through a reviewed fix, using voice, Appshots, plugins, and meeting notes to keep the team informed.