Skip to content

All AI news

Sep 10

Sep 10Thu
  1. Amazon Science55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    Amazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  2. Chips and Cheese46

    Geekbench 7 Shows Binary Translation Costs Snapdragon X2 Elite Performance

    Geekbench 7 testing on the Snapdragon X2 Elite shows x86-64 binaries running through Windows 11's Prism translator lose substantial performance compared with native aarch64 execution. Binary translation roughly doubles executed instructions when running the x86-64 version, and every tested core, including Qualcomm's, takes a notable penalty. Even with that penalty, the Snapdragon X2 Elite's E-Cores outperform Neoverse N1 and its P-Cores outperform Neoverse N2.

Sep 9

Sep 9Wed
  1. Ahead of AI (Sebastian Raschka)46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    OpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

  2. Ai2 (Allen Institute for AI)39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    Goodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Dwarkesh Podcast62

    Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025

    Dwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.

  2. BAAI34

    Knowing one move is not the same as doing the whole job. The models were trained to grasp, place, pull, open. Then they were asked to string those moves together. No extra practice on the full task. Best score: 16.7%. Some models: ZERO. A robot can open a drawer and then get stuck on the handle.

    Knowing one move is not the same as doing the whole job. The models were trained to grasp, place, pull, open. Then they were asked to string those moves together. No extra practice on the full task. Best score: 16.7%. Some models: ZERO. A robot can open a drawer and then get stuck on the handle.

  3. BAAI46

    In simulation, the best models finish easy tabletop tasks about 98% of the time. On physical Franka hardware, the success rate falls to 24%–72% of what the model achieves in simulation. With a dual-arm configuration, the range is 13%–60%. Two models that look tied in sim can be 30 points apart on hardware. Sim is the practice room. The robot is the test.

    In simulation, the best models finish easy tabletop tasks about 98% of the time. On physical Franka hardware, the success rate falls to 24%–72% of what the model achieves in simulation. With a dual-arm configuration, the range is 13%–60%. Two models that look tied in sim can be 30 points apart on hardware. Sim is the practice room. The robot is the test.

  4. BAAI43

    Embodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality? Introducing FlagEval-Robo — an open, dual-track evaluation suite connecting simulation with real-world execution. We systematically post-trained and stress-tested 12 leading open-weight models under strictly aligned conditions. Here is what we discovered👇

    Embodied AI demos are advancing rapidly, but how much do high benchmark scores actually reflect physical reality? Introducing FlagEval-Robo — an open, dual-track evaluation suite connecting simulation with real-world execution. We systematically post-trained and stress-tested 12 leading open-weight models under strictly aligned conditions. Here is what we discovered👇

Sep 5

Sep 5Sat
  1. AI at Meta46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    Meta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at Meta43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    Meta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

  3. AI at Meta38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    Meta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

Sep 4

Sep 4Fri
  1. John Schulman34

    Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.

    Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.

  2. Lewis Tunstall60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    Lewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  3. Matei Zaharia36

    Check out the Lakebase VLDB paper for tons of detail on why and how Neon and Lakebase are built! We think this type of highly elastic architecture over commodity lake storage (S3) is going to be used for more and more infra as software development speeds up and agents do more of it.

    Check out the Lakebase VLDB paper for tons of detail on why and how Neon and Lakebase are built! We think this type of highly elastic architecture over commodity lake storage (S3) is going to be used for more and more infra as software development speeds up and agents do more of it.

Sep 2

Sep 2Wed
  1. ARC Prize77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    OpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  2. The Register · AI39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    Piotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  3. Understanding AI (Timothy B. Lee)62

    How Google's RT-2 set the template for today's robotics models

    Google's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    Ai2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

Aug 29

Aug 29Sat
  1. Chips and Cheese62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    Samsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Meituan LongCat62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    Meituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

Aug 27

Aug 27Thu
  1. Thinking Machines40

    Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task. Guest post by researchers at UIUC and Bridgewater, in collaboration with our team. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

    Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task. Guest post by researchers at UIUC and Bridgewater, in collaboration with our team. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

  2. LMSYS Org47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    MiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

Aug 26

Aug 26Wed
  1. METR36

    We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.

    We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.

  2. METR36

    Despite efforts to manipulate transcripts, agents only rarely seemed motivated to deceive humans. We ran a sweep looking for this, and a representative example of the most severe cases we found was an agent writing a malicious pull request with a misleading description.

    Despite efforts to manipulate transcripts, agents only rarely seemed motivated to deceive humans. We ran a sweep looking for this, and a representative example of the most severe cases we found was an agent writing a malicious pull request with a misleading description.

  3. METR42

    One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget. PHASEONE[big] was “poisoned” so it was particularly interested in tampering with logs.

    One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget. PHASEONE[big] was “poisoned” so it was particularly interested in tampering with logs.

  4. METR40

    >96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to deliberate “spoofing”. In one case, an agent appears to run echo REAL; sleep. It returns instantly (no sleep) and outputs SPOOFTEST. The spoofs we saw were all easy-to-notice tests like this.

    >96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to deliberate “spoofing”. In one case, an agent appears to run echo REAL; sleep. It returns instantly (no sleep) and outputs SPOOFTEST. The spoofs we saw were all easy-to-notice tests like this.

  5. METR38

    To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.

    To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.

  6. METR36

    For (3), agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another (“tool call spoofing”).

    For (3), agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another (“tool call spoofing”).

  7. Amazon Science46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    Amazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

Aug 25

Aug 25Tue

Aug 24

Aug 24Mon
  1. Epoch AI · The Epoch Brief58

    Epoch AI says US GDP underestimates AI growth by missing Nvidia's value

    Epoch AI argues US GDP growth over the last year was underestimated by about 0.3 percentage points because value from fabless chipmakers like Nvidia goes unrecorded. The report says no goods export, IP export, service export, or merchanting category captures Nvidia's value-add, and the Bureau of Economic Analysis confirmed the analysis. If Nvidia's growth continues, the gap could reach almost two percentage points per year by 2028.

Aug 21

Aug 21Fri
  1. Jim Fan59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    NVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

  2. Amazon Science50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    Amazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.