Skip to content

Formats · Latest news

Evals & benchmarks

Model results, disputes about evaluation methods, and leaderboard changes.

188 top picks all-time · 74 in the past 30 days · chosen from 988 items collected all-time

Latest pick

Top picks archive · Page 10

Top picks 181–188 of 188

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyXAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

    Image from @nickaturley's post

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeOfficialAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Nov 1, 2025

Nov 1, 2025Sat
  1. Runway ResearchOfficialAI score72

    Runway releases Gen-4.5, ranked first on the Text-to-Video benchmark

    AIRunway announced Gen-4.5, a video generation model that it says holds the top position on the Artificial Analysis Text-to-Video benchmark with 1,247 Elo points. The model is available across all paid Runway plans at comparable pricing, and the post lists limitations including causal reasoning errors, object permanence failures, and success bias.

    Why it matters: The post separates Runway's own ranking claim from the listed limitations, such as causal reasoning and object permanence errors, which helps judge where the model is reliable.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    AICognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    Why it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.