Skip to content
TodayOct 8Thu63 items
  1. Pandaily44

    Donghua University Spins Transistors Into Fibers That Act as Soft Robot Circuits

    Donghua University researchers spun transistors, resistors and capacitors into a continuous fiber that functions as a circuit, using microfluidic encoded spinning, according to a Nature Electronics paper. The fibers integrated multicolor electroluminescence, analog and digital logic, and non-contact spatial sensing, and in demonstrations guided a robotic gripper and let a finger control a robotic arm and drone without touch.

  2. Pandaily55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    A Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  3. 雷峰网 Leiphone46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    Of 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  4. Elvis Saravia55

    HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%

    A paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

  5. SiliconANGLE · AI60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    OpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  6. Sundar Pichai65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    Google published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

  7. Google62

    Google's AMIE Diagnostic Chatbot Studied Prospectively in Real Clinical Setting

    Google reports that its medical research system AMIE, described as the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting, was evaluated in a study published in The Lancet. Patients who chatted with AMIE before in-person appointments reported stronger confidence and better organized thoughts. Physicians reviewing the pre-visit conversation information gained more time for collaborative care and shared decision-making instead of digging through data.

  8. Artificial Analysis18

    Check out Grok Imagine Video 1.5 Lite for yourself on the AA-Video Leaderboards: AA-Video-T2V v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video AA-Video-T2V-Silent v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video?audio-output=false Or vote in the Video Arena: https://artificialanalysis.ai/video/arena

    Check out Grok Imagine Video 1.5 Lite for yourself on the AA-Video Leaderboards: AA-Video-T2V v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video AA-Video-T2V-Silent v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video?audio-output=false Or vote in the Video Arena: https://artificialanalysis.ai/video/arena

  9. Artificial Analysis29

    Grok Imagine Video 1.5 Lite nears frontier on three AA-Video-T2V capabilities

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.

  10. Epoch AI31

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

  11. Anthropic57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    An astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  12. Epoch AI · The Epoch Brief49

    Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure

    Epoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.

  13. Google Research20

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

  14. Elvis Saravia48

    Google's FlowAgent auto-repairs failing tests inside code review

    Google proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

  15. Sherwin Wu62

    Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

    Artificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

  16. Goodfire44

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  17. Artificial Analysis34

    Harvey LAB-AA uses @harvey's LAB dataset and was built in collaboration with Harvey. Explore the full results: https://artificialanalysis.ai/evaluations/harvey-lab-aa See Harvey's commentary on the evaluation and human expert preferences: https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark https://www.harvey.ai/blog/augmenting-human-preference-in-complex-domains

    Harvey LAB-AA uses @harvey's LAB dataset and was built in collaboration with Harvey. Explore the full results: https://artificialanalysis.ai/evaluations/harvey-lab-aa See Harvey's commentary on the evaluation and human expert preferences: https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark https://www.harvey.ai/blog/augmenting-human-preference-in-complex-domains

  18. Artificial Analysis28

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  19. Artificial Analysis36

    Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.

    Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.

  20. Artificial Analysis34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    Artificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  21. Epoch AI14

    For example, GPT-6 Astra ran an experiment exploring why AI agents fail to learn with practice. It set AI agents’ token budgets too low. Instead of treating this as a mistake, it reported “sensitivity to the acquisition budget” as a key finding.

    For example, GPT-6 Astra ran an experiment exploring why AI agents fail to learn with practice. It set AI agents’ token budgets too low. Instead of treating this as a mistake, it reported “sensitivity to the acquisition budget” as a key finding.

  22. Epoch AI34

    Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.

    Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.

  23. Elvis Saravia46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    RSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  24. Goodfire28

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

  25. Google Research14

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.