Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Feb 4

Feb 4Wed
  1. Anthropic EngineeringOfficialAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Feb 2

Feb 2Mon
  1. Quoc LeXAI score60

    Gemini Helps Address 13 Open Erdős Problems in Math Case Study

    AIQuoc Le announced a case study using Gemini to systematically evaluate 700 conjectures labeled open in the Erdős Problems database. The team addressed 13 problems, finding 5 novel autonomous solutions and identifying 8 existing solutions missed by previous literature.

    Why it matters: The case study shows how a systematic AI sweep found new solutions and missed prior literature across 700 open Erdős conjectures, offering a concrete look at AI-assisted math research.

    Image from @quocleix's post
  2. BAAIOfficialAI score43

    BAAI's Emu3 published in Nature, first Chinese-led large model paper there

    AIThe Beijing Academy of Artificial Intelligence (BAAI) published its Emu3 multimodal large model research in Nature, which the post describes as the first large-model achievement led by a Chinese research institution in that journal. Emu3 learns from text, image, and video at scale using next-token prediction alone, reaching generation and perception performance comparable to task-specific methods. The authors frame this as a step toward scalable, unified multimodal intelligence systems.

    Video from @BAAIBeijing's post

Jan 27

Jan 27Tue
  1. Tim DettmersBlogAI score72

    Tim Dettmers describes how SERA, an open coding agent, was built

    AITim Dettmers describes building SERA, Ai2's first Open Coding Agents release, using 32 GPUs and synthetic data. The method uses soft verification, which accepts generated patches that overlap at least 50% with the target patch, and fine-tunes a 32B model on a private codebase in about 19 GPU days. The post says the resulting model can match its teacher, GLM 4.5-Air, on that private data.

    Why it matters: The post explains how a small team built an open coding agent with cheap synthetic data and soft verification, a reusable recipe for specializing models on private code.

Jan 10

Jan 10Sat
  1. Berkeley AI ResearchOfficialAI score36

    Information-Driven Design Framework Evaluates Imaging Systems by Mutual Information

    AIBerkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.

Jan 9

Jan 9Fri
  1. BAAIOfficialAI score47

    DrugCLIP screens 10 trillion protein-molecule pairs per day for drug discovery

    AITsinghua AIR and BAAI's DrugCLIP screened 10,000 proteins against 500 million molecules, identifying over 2 million drug candidates. The post claims a 1-million-fold speedup, reaching 10 trillion protein-molecule pairs per day, and positions DrugCLIP as bridging AlphaFold structures to drug candidates. The work is published in Science, with a platform available at drugclip.com.

    Image from @BAAIBeijing's post

Dec 12, 2025

Dec 12, 2025Fri
  1. Apple · new models on Hugging FaceOfficialAI score46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Oct 27, 2025

Oct 27, 2025Mon
  1. Mira MuratiXAI score54

    Thinking Machines explores on-policy distillation for training small models

    AIThinking Machines published a post on on-policy distillation, a training approach combining the error-correcting relevance of RL with the reward density of SFT. The quoted post reports that in math reasoning and an internal chat assistant, on-policy distillation can outperform other approaches at a fraction of the cost.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabOfficialAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogOfficialAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Aug 11, 2021

Aug 11, 2021Wed