Skip to contentSkip to stories

Updated

#Paper/Research

Oct 9

TodayOct 9Fri4 items
  1. QbitAIAI score62

    Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet

    AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.

  2. X.PINAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

  3. QbitAIAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

Oct 8

Oct 8Thu
  1. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. PandailyAI score41

    Simplexity Robotics Trains One Robot to Tend Two CNC Lathes With 600 Trajectories

    AISimplexity Robotics says it trained a single robot to load and unload two CNC lathes on its own, using 600 real-robot trajectories and reporting a 100% success rate on the precision CNC insertion task. The work, presented at IROS 2026 on September 29, combines the SimpleWAM world action model, a DRAM memory module and DPE action scoring, with force and torque feedback for insertion recovery. The company did not say how many trials the 100% figure covers, and it does not describe the yield of the whole cell.

  3. QbitAIAI score58

    AgentGarten lets agents evolve through code-built worlds and neural rendering

    AIMirroS released AgentGarten, which pairs executable code environments with a real-time neural renderer running above 30 fps so agents can act, observe, and learn. In a one-on-one hide-and-seek setup, the hider learned to block passages by round 4 and the seeker learned to climb ramps by round 10, guided by notes the agents wrote after each round. The authors report applying the same loop to four other tasks, including a dog-companion game, a narrow-bridge car passing task, herding, and quarry loading.

  4. Tencent HunyuanAI score63

    Tencent Hunyuan releases ExplorationBench to test AI rule discovery

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.

  5. PandailyAI score44

    Donghua University Spins Transistors Into Fibers That Act as Soft Robot Circuits

    AIDonghua University researchers spun transistors, resistors and capacitors into a continuous fiber that functions as a circuit, using microfluidic encoded spinning, according to a Nature Electronics paper. The fibers integrated multicolor electroluminescence, analog and digital logic, and non-contact spatial sensing, and in demonstrations guided a robotic gripper and let a finger control a robotic arm and drone without touch.

  6. PandailyAI score55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    AIA Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  7. LeiphoneAI score46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    AIOf 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  8. Elvis SaraviaAI score55

    HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%

    AIA paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

  9. SiliconANGLE · AIAI score60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  10. Sundar PichaiAI score65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

  11. GoogleAI score62

    Google's AMIE Diagnostic Chatbot Studied Prospectively in Real Clinical Setting

    AIGoogle reports that its medical research system AMIE, described as the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting, was evaluated in a study published in The Lancet. Patients who chatted with AMIE before in-person appointments reported stronger confidence and better organized thoughts. Physicians reviewing the pre-visit conversation information gained more time for collaborative care and shared decision-making instead of digging through data.

  12. AnthropicAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  13. Epoch AI · The Epoch BriefAI score49

    Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure

    AIEpoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.

  14. Elvis SaraviaAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

  15. Andrew CurranAI score13

    Andrew Curran Posts "The saga continues" Amid Tightened κ Result

    AIAndrew Curran posted a brief "The saga continues" update, with no clear publisher or model identified. Quoted context from @0xdoug reports a validated, merged PR that tightened κ from 2⁻¹⁸² to 2⁻¹⁵, described as a 500-thousand-fold improvement over the previous result and a 2^167-fold improvement over the original OpenAI result. The quoted post credits a community effort and says results are being verified and published.

  16. Elvis SaraviaAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  17. Google ResearchAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  18. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.