Skip to contentSkip to stories

Updated

#Paper/Research

Sep 23

Sep 23Wed
  1. ModelScopeAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

  2. Anthropic NewsroomAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

Sep 22

Sep 22Tue
  1. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. Tencent HunyuanAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

  2. Amazon ScienceAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  3. Microsoft ResearchAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 19

Sep 19Sat
  1. Sebastian RaschkaAI score42

    Muon reduces memorization compared with AdamW in nanoGPT training experiments

    AIMuon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds. At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon. The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Sep 18

Sep 18Fri
  1. Google ResearchAI score22

    Google Research releases MilleMiglia, a public middle-mile logistics benchmark

    AIGoogle Research has introduced MilleMiglia, a standardized benchmark for optimizing middle-mile logistics, the segment that moves goods across hundreds of miles overnight. The benchmark uses spatial clustering and gravity models to simulate realistic middle-mile delivery scenarios. It addresses the difficulty of optimizing these networks without public data.

  2. SemiAnalysisAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

Sep 17

Sep 17Thu
  1. Google ResearchAI score52

    Google Research enables teachers to create generative UI learning interactives

    AIGoogle Research is sharing an experiment that lets educators generate custom, guided STEM simulations tailored to their curriculum using generative UI. It is releasing a sample library of over 30 English interactives for physics, chemistry, biology, and math, all AI-generated and reviewed by teachers. Schools using Google Workspace for Education can sign up through the Google for Education Pilot Program to give feedback.

  2. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

Sep 16

Sep 16Wed
  1. Matei ZahariaAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

Sep 15

Sep 15Tue
  1. Lewis TunstallAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  2. Tencent HunyuanAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

Sep 14

Sep 14Mon

Sep 12

Sep 12Sat
  1. Epoch AI · The Epoch BriefAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

Sep 11

Sep 11Fri
  1. Redwood Research BlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

Sep 10

Sep 10Thu
  1. Ai2 · new models on Hugging FaceAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  2. Redwood Research BlogAI score62

    Redwood Research proposes NLS depth to measure opaque serial reasoning in AI models

    AIRedwood Research defines NLS depth, a measure of how much unverbalized serial computation an AI system can perform, building on Brown-Cohen et al.'s opaque serial depth. The proposal counts only nodes that output natural language initialized from a pre-training prior as interpretable bottlenecks. Standard transformers scale with their layer count, while latent reasoning architectures would raise NLS depth sharply.

  3. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

Sep 9

Sep 9Wed
  1. Cognition Blog (Devin, Windsurf)AI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

  2. Ahead of AI (Sebastian Raschka)AI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

  3. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Dwarkesh PodcastAI score62

    Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025

    AIDwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.