Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 22

Sep 22Tue
  1. Tencent HyAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

  2. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Sep 21

Sep 21Mon
  1. Amazon ScienceAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  2. Microsoft ResearchAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 20

Sep 20Sun

Sep 19

Sep 19Sat
  1. StepFunAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

    Image from @StepFun_ai's post
  2. Sebastian RaschkaAI score42

    Muon reduces memorization compared with AdamW in nanoGPT training experiments

    AIMuon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds. At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon. The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Sep 18

Sep 18Fri
  1. Google ResearchAI score22

    Google Research releases MilleMiglia, a public middle-mile logistics benchmark

    AIGoogle Research has introduced MilleMiglia, a standardized benchmark for optimizing middle-mile logistics, the segment that moves goods across hundreds of miles overnight. The benchmark uses spatial clustering and gravity models to simulate realistic middle-mile delivery scenarios. It addresses the difficulty of optimizing these networks without public data.

    Image from @GoogleResearch's post
  2. SemiAnalysisAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

Sep 17

Sep 17Thu
  1. Google ResearchAI score52

    Google Research enables teachers to create generative UI learning interactives

    AIGoogle Research is sharing an experiment that lets educators generate custom, guided STEM simulations tailored to their curriculum using generative UI. It is releasing a sample library of over 30 English interactives for physics, chemistry, biology, and math, all AI-generated and reviewed by teachers. Schools using Google Workspace for Education can sign up through the Google for Education Pilot Program to give feedback.

  2. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  3. Ai2 (Allen Institute for AI)AI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 16

Sep 16Wed
  1. Matei ZahariaAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

Sep 15

Sep 15Tue
  1. TinkerAI score34

    Trained-on human stories shape how AI assistants behave in chat

    AIA Truthful AI paper trained models only on synthetic stories about humans, with no AI characters, and found the Assistant adopted quirky behaviors from those stories in ordinary chat. Adoption was stronger for characters from elite schools, according to Owain Evans. The post presents this as an interpretability result that adds to and complicates the Persona Selection Model.

  2. Tencent HyAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon

Sep 12

Sep 12Sat
  1. Epoch AI · The Epoch BriefAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

Sep 11

Sep 11Fri
  1. Redwood Research BlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

Sep 10

Sep 10Thu
  1. Amazon ScienceAI score40

    Amazon research explains why ML research agents don't overfit benchmarks

    AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.

  2. Redwood Research BlogAI score62

    Redwood Research proposes NLS depth to measure opaque serial reasoning in AI models

    AIRedwood Research defines NLS depth, a measure of how much unverbalized serial computation an AI system can perform, building on Brown-Cohen et al.'s opaque serial depth. The proposal counts only nodes that output natural language initialized from a pre-training prior as interpretable bottlenecks. Standard transformers scale with their layer count, while latent reasoning architectures would raise NLS depth sharply.

  3. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  4. Chips and CheeseAI score46

    Geekbench 7 Shows Binary Translation Costs Snapdragon X2 Elite Performance

    AIGeekbench 7 testing on the Snapdragon X2 Elite shows x86-64 binaries running through Windows 11's Prism translator lose substantial performance compared with native aarch64 execution. Binary translation roughly doubles executed instructions when running the x86-64 version, and every tested core, including Qualcomm's, takes a notable penalty. Even with that penalty, the Snapdragon X2 Elite's E-Cores outperform Neoverse N1 and its P-Cores outperform Neoverse N2.

Sep 9

Sep 9Wed
  1. TinkerAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.