Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Dec 14, 2025

Dec 14, 2025Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score36

    Fun-ASR-Nano-2512 Speech Recognition Model Released by Tongyi Lab on Hugging Face

    AITongyi Lab has released Fun-ASR-Nano-2512, an end-to-end speech recognition large model trained on tens of millions of hours of real speech, supporting low-latency real-time transcription across 31 languages. The model, which has 800M parameters, targets industry use such as education and finance and claims 93% accuracy in far-field, high-noise conditions. It is available on Hugging Face and works with the FunASR toolkit.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyXAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

    Image from @nickaturley's post

Dec 10, 2025

Dec 10, 2025Wed
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score42

    Fun-CosyVoice3-0.5B-2512 Released as Open-Source Multilingual Text-to-Speech Model

    AIAlibaba's FunAudioLLM has released Fun-CosyVoice3-0.5B-2512, a 0.5B-parameter LLM-based text-to-speech model on Hugging Face, with an RL variant also published. The model supports zero-shot voice cloning across 9 languages and 18+ Chinese dialects and accents, with streaming output at latency as low as 150ms. On the source's test-en benchmark, it reports a 2.24% WER and 71.8% speaker similarity, and the RL version reports 1.68% WER.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeOfficialAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

Nov 18, 2025

Nov 18, 2025Tue
  1. Quoc LeXAI score38

    Gemini 3 Deep Think scores 45.1% on ARC-AGI-2

    AIGoogle's Gemini 3 Deep Think (Preview) reaches 45.14% on ARC-AGI-2's semi-private eval at $77.16 per task. ARC Prize says this doubles the prior state of the art, with Gemini 3 Pro scoring 31.11% at $0.81 per task.

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Nov 3, 2025

Nov 3, 2025Mon
  1. ARC PrizeOfficialAI score38

    ARC Prize Launches Verified Program to Certify ARC-AGI Benchmark Scores

    AIARC Prize Foundation announced ARC Prize Verified, a program that certifies frontier model scores on the ARC-AGI benchmark using hidden test sets and adds a third-party academic panel to audit and open-source its testing process. Five AI labs, including Google and xAI, are sponsoring ARC-AGI-3 development, and the foundation says donations do not influence verification scoring. Models that pass verification will appear on the official leaderboard with a verification badge.

Nov 1, 2025

Nov 1, 2025Sat
  1. Runway ResearchOfficialAI score72

    Runway releases Gen-4.5, ranked first on the Text-to-Video benchmark

    AIRunway announced Gen-4.5, a video generation model that it says holds the top position on the Artificial Analysis Text-to-Video benchmark with 1,247 Elo points. The model is available across all paid Runway plans at comparable pricing, and the post lists limitations including causal reasoning errors, object permanence failures, and success bias.

    Why it matters: The post separates Runway's own ranking claim from the listed limitations, such as causal reasoning and object permanence errors, which helps judge where the model is reliable.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)OfficialAI score50

    Devin Adds Claude Sonnet 4.5 as Agent Preview Model

    AICognition says Claude Sonnet 4.5 is available in Devin starting today, improving its planning performance by 18% and end-to-end eval scores by 12%. Cognition says the model's testing of its own code lets Devin run longer and handle harder tasks.

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    AICognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    Why it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.