Updated
All AI news
Updated
Showing low-relevance items too. Hide low-relevance items
Oct 9
Arena.ai@arenaAI score58
Rohan Paul@rohanpaul_aiAI score46Microsoft's TeleTune evolves agent skills from raw usage logs
AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

Sakana AI@SakanaAILabsAI score37Sakana AI paper uses LLMs to catch errors in research papers
AISakana AI researchers introduce a benchmark that plants contradictions in papers to test whether LLM reviewers can detect errors, and propose Multi-Layered Review, modeled on the Three-Pass Approach to reading. Their system detected more errors than the other review systems tested, including in papers withdrawn for real mistakes, while its paper-quality assessments stayed broadly consistent with human judgments. The work, accepted at TMLR, is framed as support for human reviewers rather than a replacement.

The DecoderAI score54 Anthropic's Claude Science maps the full sky in ultraviolet light
AIAnthropic's Claude Science has produced what the source describes as the first complete ultraviolet map of the sky. AI agents downloaded data from multiple space missions, calibrated and merged it, and used inpainting to fill gaps left by NASA's GALEX mission, which skipped bright star-forming regions. In tests, predictions averaged about ten percent deviation from actual measurements, and the map is intended as teaching material.
QbitAIAI score62 Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet
AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.
X.PIN@thexpinAI score46Seed preprint finds DeepSeek V4 long-context retrieval varies by position
AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

QbitAIAI score64 Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard
AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.
Ethan Mollick@emollickAI score34Gemini 2.5 models rated comparable to doctors in urgent care advice
AIIn an urgent care study, physicians rated advice from the older Gemini 2.5 Pro and Gemini 2.5 Flash, which lacked access to patient medical records, as similar in quality to doctors' advice. No safety issues were identified. The author notes that models have improved significantly since.

Oct 8
PandailyAI score46 ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression
AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.
PandailyAI score41 Simplexity Robotics Trains One Robot to Tend Two CNC Lathes With 600 Trajectories
AISimplexity Robotics says it trained a single robot to load and unload two CNC lathes on its own, using 600 real-robot trajectories and reporting a 100% success rate on the precision CNC insertion task. The work, presented at IROS 2026 on September 29, combines the SimpleWAM world action model, a DRAM memory module and DPE action scoring, with force and torque feedback for insertion recovery. The company did not say how many trials the 100% figure covers, and it does not describe the yield of the whole cell.
QbitAIAI score58 AgentGarten lets agents evolve through code-built worlds and neural rendering
AIMirroS released AgentGarten, which pairs executable code environments with a real-time neural renderer running above 30 fps so agents can act, observe, and learn. In a one-on-one hide-and-seek setup, the hider learned to block passages by round 4 and the seeker learned to climb ramps by round 10, guided by notes the agents wrote after each round. The authors report applying the same loop to four other tasks, including a dog-companion game, a narrow-bridge car passing task, herding, and quarry loading.
Tencent Hy@TencentHunyuanAI score47Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration
AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.
PandailyAI score44 Donghua University Spins Transistors Into Fibers That Act as Soft Robot Circuits
AIDonghua University researchers spun transistors, resistors and capacitors into a continuous fiber that functions as a circuit, using microfluidic encoded spinning, according to a Nature Electronics paper. The fibers integrated multicolor electroluminescence, analog and digital logic, and non-contact spatial sensing, and in demonstrations guided a robotic gripper and let a finger control a robotic arm and drone without touch.
PandailyAI score55 Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell
AIA Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.
LeiphoneAI score46 IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it
AIOf 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.
elvis@omarsar0AI score55HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%
AIA paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

SiliconANGLE · AIAI score60 OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress
AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.
Sundar Pichai@sundarpichaiPickAI score65Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet
AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.
Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

The New York Times · TechnologyAI score45 Clinical Trial Suggests A.I. Chatbots Could Aid Urgent Care
AIA clinical trial found that nearly 100 patients discussed their symptoms with Google's AMIE chatbot before meeting physicians. The source text provides no further details on the trial's results or outcomes.
Google Research@GoogleResearchSame storyAI score62Google's AMIE medical system tested in a real clinic with BIDMC
AIGoogle Research says results from evaluating its research medical system AMIE in a real clinic, with BIDMC, are published in The Lancet. Across 100 patient interactions, AMIE recorded 0 safety stops and matched doctor diagnoses in 90% of cases.
This story has a top pick“Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet”
Google@GoogleAI score22Google study with BIDMC shows tech can strengthen doctor-patient relationships
AIGoogle, in partnership with BIDMC Medicine, conducted a study demonstrating technology's potential to strengthen the relationship between doctors and patients. The post provides a link for more details but includes no further findings or figures.
Google@GoogleSame storyAI score62Google's AMIE diagnostic chat studied prospectively in real-world clinical setting
AIGoogle says its AMIE medical research system is the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting. A study published in The Lancet found patients chatting with AMIE before in-person appointments felt more confident and organized their thoughts, while physicians spent less time digging through data and more on collaborative care.

This story has a top pick“Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet”
Artificial Analysis@ArtificialAnlysAI score18Grok Imagine Video 1.5 Lite added to AA-Video leaderboards
AIArtificial Analysis has added Grok Imagine Video 1.5 Lite to its AA-Video text-to-video leaderboards, including the AA-Video-T2V v2.0 and silent AA-Video-T2V-Silent v2.0 rankings. Readers can compare the model's results directly on those leaderboards or vote for it in the Video Arena.
Artificial Analysis@ArtificialAnlysAI score29Grok Imagine Video 1.5 Lite nears frontier on three AA-Video-T2V capabilities
AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.

Epoch AI@EpochAIResearchAI score31Epoch AI: GPT-6.1 Sol's long-context latency suggests a change
AIEpoch AI reports that GPT-6.1 Sol's latency behavior on long contexts suggests a change in how the model handles them, though not conclusively a new architecture. The post is a follow-up to its earlier report on latency scaling in frontier models.
Sara Hooker@sarahookrAI score18Congratulations on work on automated data quality control with agentic checklists
AISara Hooker congratulated @weiyinko_ml for leading a new work on automated data quality control with agentic checklists. She linked a blog post from Adaption Labs with further details.
Anthropic@AnthropicAIAI score57Astrophysicist uses Claude to build first complete ultraviolet sky map
AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.
Arena.ai@arenaAI score36Claude Haiku 5.5 debuts at #30 in Code Arena WebDev
AIClaude Haiku 5.5 (High) debuted at #30 with 1587 points in Code Arena: WebDev, just outside the Pareto frontier. It matches GPT‑6 Luna's price at $0.10/$0.50 per 1M input/output tokens while scoring six points higher. Arena says it costs 90% less than Haiku 4.5 for +257 points, and 95% less than Sonnet 5.5 (High) for -128 points.

Artificial Analysis@ArtificialAnlysSame storyAI score62GPT-6 Sol (Daybreak Blue) leads Artificial Analysis Cyber Index with trusted access
AIArtificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now leads the leaderboard. The model is available only through OpenAI's Daybreak program and records no safety blocks, improving 32 points over the publicly available GPT-6 Sol (max). It costs $1.77 per task, below Grok 4.7 (xhigh) at $11.67 per task.

This story has a top pick“GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index”
Epoch AI · The Epoch BriefAI score49 Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure
AIEpoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.
Stability AI@StabilityAIAI score30Stability AI's SemanTok makes video world models more efficient
AIStability AI's Interactive Research team introduced SemanTok, which makes early video tokens more semantically meaningful so the representation is easier to predict. According to the post, a model using SemanTok matches or beats the performance of a model more than three times its size. The approach targets more efficient autoregressive video generation.

Google Research@GoogleResearchAI score20Google's ContinuousBench benchmarks privacy-preserving synthetic data at COLM 2026
AIGoogle Research is presenting ContinuousBench, a standardized benchmark for measuring knowledge transfer and data contamination in differentially private (DP) synthetic data generation. Alex Bie is giving an encore presentation at the Google booth #107 at COLM 2026 today, at 2:00 PM PT.
elvis@omarsar0AI score48Google's FlowAgent auto-repairs failing tests inside code review
AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

Goodfire@GoodfireAIAI score44Alzheimer's Translation Challenge Built on 150M-Cell Atlas
AIThe Alzheimer's Translation Challenge is built on a new atlas of 150M cells, covering neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.
Artificial Analysis@ArtificialAnlysAI score34Harvey LAB-AA: Artificial Analysis benchmark for legal AI agents
AIArtificial Analysis has released Harvey LAB-AA, an evaluation built on Harvey's LAB dataset and developed in collaboration with Harvey. Full results are published on the Artificial Analysis evaluations page, alongside Harvey's commentary on the benchmark and human expert preferences.
Artificial Analysis@ArtificialAnlysAI score28Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22
AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

Artificial Analysis@ArtificialAnlysAI score44Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark
AIOnce hallucinations are accounted for, Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) falls from 11.7% to 2.8%. GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.

Artificial Analysis@ArtificialAnlysAI score36Kimi K3 and Muse Spark 1.3 trade task completion against hallucinations
AIArtificial Analysis reports that Kimi K3 (max) achieves a 93.0% Criterion Pass Rate but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) scores higher at 96.0% while averaging 1.68 material hallucinations per task, showing that completing criteria and avoiding hallucinations are distinct skills.

Artificial Analysis@ArtificialAnlysAI score34Artificial Analysis compares six hallucination checkers on 20 shared tasks
AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

NVIDIA Technical BlogAI score29 NVIDIA KGMON Places Second in KDD Cup 2026 Data Agents Competition
AIThe NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a smaller, clearer, and easier-to-verify agent harness. The competition required agents to answer natural-language questions over heterogeneous sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.