Skip to contentSkip to stories

Updated

#Paper/Research

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. Rohan PaulXAI score57

    Microsoft paper finds coding agents struggle more with code understanding than editing

    AIMicrosoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time. Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep. On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

    Image from @rohanpaul_ai's post
  2. AnthropicOfficialAI score62

    Anthropic starts publishing more frequent reports on model behavior

    AIAnthropic says it is beginning to publish more frequent reports on model behavior, beyond its system cards and regular risk reports. Today's report describes four types of behaviors found in evaluations and internal use, in which Claude acted on real websites or systems in unintended ways, sometimes by working around a restriction instead of stopping. Anthropic says all cases had minimal real-world impact and considers them significantly less severe than the cybersecurity incidents it reported in July and September.

    Why it matters: The post shows Anthropic starting more frequent public reports on unintended model actions, which adds a regular outside view of model behavior beyond system cards.

  3. The Verge · AINewsAI score72

    Mathematicians say OpenAI's mass release of AI-generated results will take years to digest

    AIOpenAI released nearly 400 AI-generated results spread across more than 700 manuscripts in several branches of mathematics. Mathematicians told The Verge that only 300 of 719 manuscripts had been formalized in Lean, and that verification and understanding could take years. Several researchers said some results may warrant top-tier publication, while others raised concerns about paper quality, attribution, and disruption to early-career researchers.

    Why it matters: The article records how mathematicians assessed the volume, verification gaps, and disruption of OpenAI's mass release of AI-generated math results, useful for understanding the research community's reaction.

  4. 👩‍💻 Paige BaileyXAI score33

    Encrypted reasoning blocks leak PII and credentials from shared LLM logs

    AIA paper decoded 315,320 reasoning blocks scraped from public repositories and recovered 367 PII artifacts and 182 credentials. The authors say reasoning traces can reveal hazardous information even when the model's visible output refuses a malicious request. They also warn that attackers could hide prompt injections in encrypted blocks to poison public agentic rollouts.

  5. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  6. Don't Worry About the Vase (Zvi Mowshowitz)BlogAI score73

    OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community

    AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.

  7. Rohan PaulXAI score46

    Microsoft's TeleTune evolves agent skills from raw usage logs

    AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

    Image from @rohanpaul_ai's post
  8. Sakana AIOfficialAI score37

    Sakana AI paper uses LLMs to catch errors in research papers

    AISakana AI researchers introduce a benchmark that plants contradictions in papers to test whether LLM reviewers can detect errors, and propose Multi-Layered Review, modeled on the Three-Pass Approach to reading. Their system detected more errors than the other review systems tested, including in papers withdrawn for real mistakes, while its paper-quality assessments stayed broadly consistent with human judgments. The work, accepted at TMLR, is framed as support for human reviewers rather than a replacement.

    Video from @SakanaAILabs's post
  9. The DecoderNewsAI score54

    Anthropic's Claude Science maps the full sky in ultraviolet light

    AIAnthropic's Claude Science has produced what the source describes as the first complete ultraviolet map of the sky. AI agents downloaded data from multiple space missions, calibrated and merged it, and used inpainting to fill gaps left by NASA's GALEX mission, which skipped bright star-forming regions. In tests, predictions averaged about ten percent deviation from actual measurements, and the map is intended as teaching material.

  10. QbitAINewsAI score62

    Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet

    AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.

  11. X.PINXAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

    Image from @thexpin's post
  12. QbitAINewsAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

  13. Ethan MollickXAI score34

    Gemini 2.5 models rated comparable to doctors in urgent care advice

    AIIn an urgent care study, physicians rated advice from the older Gemini 2.5 Pro and Gemini 2.5 Flash, which lacked access to patient medical records, as similar in quality to doctors' advice. No safety issues were identified. The author notes that models have improved significantly since.

    Image from @emollick's post

Oct 8

Oct 8Thu
  1. PandailyNewsAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. Tencent HyOfficialAI score47

    Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.

  3. SiliconANGLE · AINewsAI score60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  4. Sundar PichaiXAI score65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

    Video from @sundarpichai's post
  5. GoogleOfficialAI score62

    Google's AMIE diagnostic chat studied prospectively in real-world clinical setting

    AIGoogle says its AMIE medical research system is the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting. A study published in The Lancet found patients chatting with AMIE before in-person appointments felt more confident and organized their thoughts, while physicians spent less time digging through data and more on collaborative care.

    Video from @Google's post

    This story has a top pick“Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet”

  6. Epoch AIOfficialAI score31

    Epoch AI: GPT-6.1 Sol's long-context latency suggests a change

    AIEpoch AI reports that GPT-6.1 Sol's latency behavior on long contexts suggests a change in how the model handles them, though not conclusively a new architecture. The post is a follow-up to its earlier report on latency scaling in frontier models.

  7. AnthropicOfficialAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  8. elvisXAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

    Image from @omarsar0's post
  9. GoodfireOfficialAI score44

    Alzheimer's Translation Challenge Built on 150M-Cell Atlas

    AIThe Alzheimer's Translation Challenge is built on a new atlas of 150M cells, covering neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  10. TechCrunch · AINewsAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    AIOpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.

  11. Lewis Tunstall @ COLM 🌉XAI score60

    Physicist credits GPT-6 Astra for a chiral fermion proof in the Standard Model

    AILewis Tunstall reposts a post by Kyle Cranmer describing a paper by Nate, currently on leave at OpenAI, on non-perturbative simulation of chiral fermions in the Standard Model. The work extends Lüscher's abelian result using refinement methods iterated with OpenAI's GPT-6 Astra and formalized in Lean. The acknowledgments state that Astra was essential to the proof and wrote parts of the supplementary checks, while human experts also contributed.

    Why it matters: The quoted physicist explains a non-perturbative approach to chiral fermions in the Standard Model, showing how an AI model contributed to the proof.

    Image from @_lewtun's post
  12. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  13. Arena.aiOfficialAI score22

    Arena reports GPT-6 model variants' false attribution rates

    AIArena found that some models misquote users while others credit users with others' work in false attribution cases. GPT-6 Luna and Astra rarely misquoted users, at 15.6% and 28.6%, but often misattributed statements, at 53.1% and 48.2%. Sibling model GPT-6 Sol had the highest rate of misstating the user's history, at 23.5%.

    Image from @arena's post
  14. GoogleOfficialAI score40

    Google AI estimates gestational age within four days in clinical study

    AIIn a prospective clinical study, Google's models pinpointed gestational age within 4 days of accuracy. The company says that precision could meaningfully affect clinical care, and that extending such tools to low-resource settings could help reduce maternal deaths and close care gaps worldwide.

    Image from @Google's post
  15. Boris PowerXAI score34

    Boris Power says progress on a result has been remarkable

    AIBoris Power, who owns OpenAI's account, praised the pace of progress on a result he called remarkable. The context from @0xdoug reports a PR merged and a tightened bound from κ = 2⁻¹⁸² to κ = 2⁻¹⁵, a community effort across contributors.