Skip to content

#Model release

Oct 8

TodayOct 8Thu1 item

Oct 7

Oct 7Wed
  1. Epoch AIAI score37

    Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.

    Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.

  2. Ai2AI score36

    Our recipe adapts existing models to bytes with a short additional training run. It keeps the model's core, adding components that group bytes into variable-length patches for it to process, then expand its outputs back to byte-level representations to predict the next byte.

    Our recipe adapts existing models to bytes with a short additional training run. It keeps the model's core, adding components that group bytes into variable-length patches for it to process, then expand its outputs back to byte-level representations to predict the next byte.

  3. Ai2AI score38

    Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 https://www.nature.com/articles/s41586-026-11111-4

    Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 https://www.nature.com/articles/s41586-026-11111-4

  4. Ai2 (Allen Institute for AI)AI score57

    Ai2's Bolmo byte-level language models are published in Nature

    Ai2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

Oct 6

Oct 6Tue
  1. ARC PrizeAI score28

    @SpaceXAI On ARC-AGI-3, Grok 4.7 scores 1.8% (vs Grok 4.7's 2.1%) in the standard harness, which lets models carry forward notes between turns, and 10.0% in a new provider adapter harness, which preserves opaque reasoning and enables auto compaction.

    @SpaceXAI On ARC-AGI-3, Grok 4.7 scores 1.8% (vs Grok 4.7's 2.1%) in the standard harness, which lets models carry forward notes between turns, and 10.0% in a new provider adapter harness, which preserves opaque reasoning and enables auto compaction.

  2. ARC PrizeAI score46

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

Oct 3

Oct 3Sat

Oct 1

Oct 1Thu
  1. ARC PrizeAI score22

    @Alibaba_Qwen @baseten - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Qwen3.8-27B results: https://arcprize.org/results/alibaba-qwen3-8-27b

    @Alibaba_Qwen @baseten - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Qwen3.8-27B results: https://arcprize.org/results/alibaba-qwen3-8-27b

Sep 30

Sep 30Wed
  1. Ant LingAI score38

    Over ~17 hours, Ling-3.1-flash built a Lua-to-x86-64 ELF compiler from scratch and passed 178/182 tests (97.8%). That’s within 1.2 percentage points of the Claude Fable 5 result shown (99.0%) and above GLM-5.3-Flash (95.9%).

    Over ~17 hours, Ling-3.1-flash built a Lua-to-x86-64 ELF compiler from scratch and passed 178/182 tests (97.8%). That’s within 1.2 percentage points of the Claude Fable 5 result shown (99.0%) and above GLM-5.3-Flash (95.9%).

Sep 28

Sep 28Mon
  1. Ali GhodsiAI score62

    Databricks finds Opus 5.5 cheaper and better, GPT-6 Luna 20x cheaper per task

    Databricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    Google Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    AIWhy it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  2. ModelScopeAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    Shanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

Sep 22

Sep 22Tue
  1. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    METR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    AIWhy it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Sep 19

Sep 19Sat
  1. StepFunAI score20

    Finance is a particular focus for Step 5 Preview. Financial work has to stand up to scrutiny. We evaluate Step 5 Preview on its ability to identify and verify reliable information, reconcile differences across reports, make assumptions explicit, and produce internally consistent forecasts and reproducible valuations. FinStepBench tests these capabilities across LiveSearch, CorporateValuation, and DeepResearch. We also evaluate Step 5 Preview on FrontierFinance across six investment use cases.

    Finance is a particular focus for Step 5 Preview. Financial work has to stand up to scrutiny. We evaluate Step 5 Preview on its ability to identify and verify reliable information, reconcile differences across reports, make assumptions explicit, and produce internally consistent forecasts and reproducible valuations. FinStepBench tests these capabilities across LiveSearch, CorporateValuation, and DeepResearch. We also evaluate Step 5 Preview on FrontierFinance across six investment use cases.

Sep 17

Sep 17Thu
  1. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    SenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

Sep 10

Sep 10Thu

Sep 2

Sep 2Wed
  1. Understanding AI (Timothy B. Lee)AI score62

    How Google's RT-2 set the template for today's robotics models

    Google's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

Jul 28

Jul 28Tue
  1. Intern Large ModelsAI score62

    Intern Large Models introduces Visual Pretraining learned from visual documents

    Intern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents. The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

Jul 20

Jul 20Mon
  1. Ahmad Al-DahleAI score31

    Still waiting on the tech report but a few innovations: 1/ KDA swaps Gated DeltaNet's single scalar decay for learned per-dimension forgetting. 2/ AttnRes retrieves selectively across depth. 3/ LatentMoE activates 16 of 896 experts, balanced from router score quantiles.

    Still waiting on the tech report but a few innovations: 1/ KDA swaps Gated DeltaNet's single scalar decay for learned per-dimension forgetting. 2/ AttnRes retrieves selectively across depth. 3/ LatentMoE activates 16 of 896 experts, balanced from router score quantiles.

Mar 26

Mar 26Thu