Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 9

Sep 9Wed
  1. Ahead of AI (Sebastian Raschka)AI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

  2. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Dwarkesh PodcastAI score62

    Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025

    AIDwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.

  2. BAAIAI score43

    FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots

    AIBAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.

    Image from @BAAIBeijing's post

Sep 5

Sep 5Sat
  1. AI at MetaAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

    Image from @AIatMeta's post
  3. AI at MetaAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. John SchulmanAI score34

    Schulman praises metric and dataset for training models to explain behavior

    AIJohn Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline. He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen. That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

  2. Lewis Tunstall @ COLM 🌉AI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

Sep 3

Sep 3Thu
  1. TinkerAI score25

    Tinker highlights training objectives for legible chain-of-thought and interpretability evals

    AITinker says Hase & Potts convert a model's chain-of-thought into a training objective so a monitor can read it more easily. Karvonen et al. use tested counterfactual outputs to build an interpretability eval. The post notes that counterfactuals do not explain the underlying mechanism, but their predictability is a useful foundation.

  2. TinkerAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  3. TinkerAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post

Sep 2

Sep 2Wed
  1. TinkerAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

  2. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  3. The Register · AIAI score39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    AIPiotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  4. Understanding AI (Timothy B. Lee)AI score62

    How Google's RT-2 set the template for today's robotics models

    AIGoogle's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

Aug 29

Aug 29Sat
  1. Chips and CheeseAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    AISamsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Meituan LongCatAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. LMSYS OrgAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

    Image from @lmsysorg's post

Aug 26

Aug 26Wed
  1. Amazon ScienceAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    AIAmazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.