Skip to contentSkip to stories

Updated

#Paper/Research

Showing low-relevance items too. Hide low-relevance items

Sep 10

Sep 10Thu
  1. Amazon ScienceAI score40

    Amazon research explains why ML research agents don't overfit benchmarks

    AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.

  2. Redwood Research BlogAI score62

    Redwood Research proposes NLS depth to measure opaque serial reasoning in AI models

    AIRedwood Research defines NLS depth, a measure of how much unverbalized serial computation an AI system can perform, building on Brown-Cohen et al.'s opaque serial depth. The proposal counts only nodes that output natural language initialized from a pre-training prior as interpretable bottlenecks. Standard transformers scale with their layer count, while latent reasoning architectures would raise NLS depth sharply.

  3. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

Sep 9

Sep 9Wed
  1. TinkerAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  2. Cognition Blog (Devin, Windsurf)AI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

  3. Ahead of AI (Sebastian Raschka)AI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

  4. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Dwarkesh PodcastAI score62

    Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025

    AIDwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.

  2. Google DeepMind · YouTubeAI score78

    DeepMind releases AlphaGenome Atlas, a predictive map of every possible DNA letter change

    AIGoogle DeepMind has used AlphaGenome to predict the molecular impact of every possible single-letter change in the human genome, around nine billion variants. The resulting AlphaGenome Atlas is a 1PB dataset that assigns each variant an AlphaGenome Variant Impact (AVI) score, covering both coding and non-coding variations, and is available to researchers worldwide. The video notes that AlphaGenome has not been validated or approved for any clinical use.

    Why it matters: The release supplies a precomputed impact score for every possible single-letter genome change, which lets researchers look up variants without running the model themselves.

Sep 5

Sep 5Sat
  1. AI at MetaAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. John SchulmanAI score34

    Schulman praises metric and dataset for training models to explain behavior

    AIJohn Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline. He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen. That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

  2. BAAI · new models on Hugging FaceAI score26

    ConsiSpace: BAAI and Peking University release geometry-consistent video spatial reasoning model

    AIBAAI and Peking University researchers released official weights for ConsiSpace, a geometry-consistent multimodal framework for spatial reasoning in long-form visual observations. The model is described in the paper "ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning" (arXiv:2607.17599).

  3. Lewis Tunstall @ COLM 🌉AI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  4. Lewis Tunstall @ COLM 🌉AI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

    Image from @_lewtun's post
  5. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  6. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. TinkerAI score25

    Tinker highlights training objectives for legible chain-of-thought and interpretability evals

    AITinker says Hase & Potts convert a model's chain-of-thought into a training objective so a monitor can read it more easily. Karvonen et al. use tested counterfactual outputs to build an interpretability eval. The post notes that counterfactuals do not explain the underlying mechanism, but their predictability is a useful foundation.

  2. TinkerAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  3. TinkerAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  4. Google DeepMind · YouTubeAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

Sep 2

Sep 2Wed
  1. TinkerAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

Aug 29

Aug 29Sat
  1. Chips and CheeseAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    AISamsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Meituan LongCatAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.