Skip to contentSkip to stories

Updated

#Paper/Research

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. AnthropicAI score38

    Harvard physicist builds toolkit to match Claude with science calculations

    AIHarvard physicist Matthew Schwartz argues that LLMs are poorly matched to science when used as human-style collaborators, so he built a toolkit for exact quantitative calculations. Working with Claude, the approach surfaced connections to ecology, population genetics, and a dozen other fields, with domain experts steering it toward interesting questions.

  2. Google GemmaAI score54

    Google Gemma credits StudentBench study comparing AI and expert human GRE tutors

    AIGoogle Gemma relays a StudentBench study reporting that AI tutors matched expert human tutors on immediate GRE learning gains. The author reports 2,383 students and a cost of 7 cents per AI tutor hour versus $75 for an expert human hour. The post also says the top AI tutor beat expert human tutors on average in 5 of 7 academic topics, and that the data and paper are publicly available.

  3. Goodfire ResearchAI score60

    Goodfire proposes protein embedding monitors for biosecurity risks in AI agents

    AIGoodfire Research developed sequence-aware monitors using protein language model embeddings to flag concerning biological sequences in dual-use AI agent tasks. On a custom benchmark, the monitors outperformed frontier model safeguards with fewer refusals on benign requests, and they held up better against paraphrasing and fragmentation attacks. The paraphrase results rely on in-silico estimates and do not establish whether the redesigned proteins keep biological activity, and the monitors run in milliseconds per sequence.

    Why it matters: The post gives a concrete benchmark setup and fragmentation results, showing how sequence embeddings can separate dual-use biology requests that task-based safeguards handle poorly.

  4. Prime IntellectAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

    Image from @PrimeIntellect's post
  5. Alexander DoriaAI score54

    SYNTH paper proposes fully synthetic single-stage training for reasoning models

    AIThe SYNTH paper, titled It's All Training, presents a fully synthetic single-stage pipeline for training workable reasoning models with high data efficiency. The authors argue this approach does not separate training into pretraining, mid-training, or post-training stages. The image shows the paper's abstract, which describes a pipeline built from a 58,000-article Wikipedia-based synthetic corpus and models named Baguettotron-600M and Baguettotron-MoE.

    Image from @Dorialexander's post
  6. Amazon ScienceAI score34

    Amazon Science Explains Graph-Centric Agentic AI for Network Root Cause Analysis

    AIAmazon Science describes a graph-centric approach in which a network digital twin graph and cascaded graph algorithms, orchestrated by an agentic AI layer, identify root causes in complex network failures. The approach was demonstrated with NTT DOCOMO at the Mobile World Conference, achieving root cause analysis in minutes on commercial networks. The article traces how graphs evolved from topology models to active reasoning substrates for agents.

Sep 30

Sep 30Wed
  1. Apple Machine Learning ResearchAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  2. Apple Machine Learning ResearchAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  3. Google · Innovation & AIAI score46

    Google AI Flu Model Ranks First in CDC FluSight Hospitalization Forecasts

    AIA flu forecasting model built with Google AI ranked first among 39 eligible models in the CDC's FluSight 2025-26 season evaluation for predicting U.S. flu-related hospital admissions. The model was developed using Empirical Research Assistance (ERA), an AI tool that generates optimization algorithms, and ERA's underlying technology is now available to trusted testers.

  4. Microsoft ResearchAI score46

    Machine learning system forecasts space-weather grid risk for 66,935 U.S. substations

    AIMicrosoft Research intern-developed machine learning pipeline forecasts location-specific geomagnetic risk for 66,935 substations in the continental United States. It combines solar-wind observations, AE and Dst forecasts, geological conductivity and grid data to estimate risk 30 to 60 minutes ahead. The pipeline detected nearly 80% of major space-weather events during the evaluation period.

  5. Liquid AIAI score42

    LongevityBench: Liquid AI's compact LFMs beat frontier models on aging tasks

    AILiquid AI and InSilicoMeds released LongevityBench, an aging benchmark with 17 tasks spanning clinical records, DNA methylation, transcriptomics, proteomics, and genetics. On several tasks, Liquid AI's compact LFMs outperformed every frontier model the team evaluated. The team plans to present the work to the longevity research community at ARDD this week.

    Video from @liquidai's post
  6. Tencent HyAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    AIResearchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

    Image from @TencentHunyuan's post
  7. OpenBMBAI score42

    Diffusion Reward Models learn full human preference distributions, not single scores

    AIOpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score. The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs. DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.

    Image from @OpenBMB's post
  8. Anthropic ResearchAI score62

    Anthropic study finds robots can do most physical tasks but rarely cost-effectively

    AIAnthropic's research rates how well present-day robots can perform US job tasks, finding they can do 74% of physical tasks, or 34% of working hours, mostly in limited settings. Robots are cost-competitive for only 0.3% of job tasks, and at a 3% annual price decline it would take about 40 years to reach 10%. The report also finds robot-exposed jobs tend to pay less and be more physically demanding than LLM-exposed jobs.

    Why it matters: The report separates current robot capability from cost, showing that physical automation is technically broad but economically narrow for now.

Sep 29

Sep 29Tue
  1. Fireworks AI BlogAI score51

    Fireworks explains how numerical mismatch and MoE routing can derail RL training

    AINumerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

  2. Apple Machine Learning ResearchAI score38

    LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness

    AIApple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.

  3. Google Developers BlogAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.

  4. Google ResearchAI score35

    Google Research unveils Diffusion Controller for steering AI image generation

    AIGoogle Research introduced Diffusion Controller, a framework that treats image generation as a continuous control problem rather than separate inference-time guidance and fine-tuning fixes. Its lightweight add-on "steering damper" network keeps the base model frozen and works on black-box or gray-box models, and it outperformed the industry standard on human preference matching. In a Stable Diffusion v1.4 test, the fully unlocked version achieved a 90% win rate over the baseline.

  5. Microsoft ResearchAI score34

    Microsoft Research unveils Quine, an early multimodal world model of biology

    AIMicrosoft Research has introduced Quine, an early-stage research effort to build a multimodal world model of biology that connects insights across biological scales and modalities. The system is designed to help scientists computationally search a space far larger than intuition allows and prioritize hypotheses before lab testing. Experimental results are meant to feed back into the model and sharpen future research directions.

    Video from @MSFTResearch's post
  6. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

    Image from @OpenBMB's post