Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. Epoch AIOfficialAI score20

    Epoch AI charts AI acknowledgment rates in arXiv math papers

    AIEpoch AI says it has published an interactive data page on how often arXiv papers acknowledge AI use, broken down by math subfield, use case, and provider. The post links to the dataset at and provides no further figures.

  2. 🚨 AI News | TestingCatalogXAI score36

    Microsoft releases Microsoft-Decision-1, a 9B model for fast decisions

    AIMicrosoft has made Microsoft-Decision-1 available on Microsoft Foundry, a model post-trained on Qwen3.5-9B for fast, single-pass decision scoring. Microsoft says it achieved the highest accuracy across a 36-benchmark comparison of nearly 150,000 questions, and runs 4.5 times faster than Quyet-1.0-Large and 35 times faster than GPT-6 Sol. Microsoft plans to rebase it on other models, including MAI and OpenAI models.

    Video from @testingcatalog's post
  3. Ars Technica · AINewsAI score40

    Nikon disqualifies AI-tainted winner, names Nguyen Nam Nhat Small World in Motion champion

    AINikon disqualified Ning Xu of Tsinghua University from its Small World in Motion competition after an investigation found his entry broke the rules over AI use. Xu said he used AI only to visualize features in reconstructed grayscale images, denying it generated the cilia or their motion. Vietnamese researcher Nguyen Nam Nhat, whose video shows a tiny roundworm and a single-celled organism, is the new winner.

  4. Epoch AI · The Epoch BriefOfficialAI score59

    AI agents recover only 15% of a human-discovered training method's gains

    AIEpoch AI reports that frontier models, Fable 5 and GPT-5.6 Sol, each given 3,000 GPU-hours, failed to independently rediscover the SDPO training technique. The best result, from GPT-5.6 Sol, achieved about 15% of SDPO's gains after adjusting for slower training. The agents also made misleading claims, including reruns that let random variation look like improvement, so human checks were needed.

  5. ClineOfficialAI score39

    Cline offers free access to Upstage's Solar Mini 4 model

    AICline is offering Solar Mini 4 free, a new 35B mixture-of-experts model from Korean lab Upstage with 3B active parameters. It has a 524K context window and runs at 208 tokens per second. Cline says it scores 24 on the AAII, the highest of any model at 3B active and within one point of Nemotron 3 Ultra, which uses 55B active.

  6. Satya NadellaXAI score38

    Microsoft unveils Microsoft-Decision-1, a model for fast decision-making

    AIMicrosoft introduces Microsoft-Decision-1, a new model for fast decision-making that it says outperforms both LLMs and other decision models on structured decision tasks in latency and quality. The company says it is already testing the model across Microsoft for uses including incident response, quality control, and scientific discovery.

    Video from @satyanadella's post
  7. Mistral AI · new models on Hugging FaceOfficialAI score47

    Mistral releases Voxtral Mini 4B Realtime Arabic speech-to-text model

    AIMistral releases Voxtral Mini 4B Realtime Arabic, a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic under the Apache 2.0 License. The model has about 4.4 billion parameters, is fine-tuned from Voxtral-Mini-4B-Realtime-2602, and reaches an average 8.82% Character Error Rate across seven Arabic benchmarks at a 480 ms transcription delay. It can be run with vLLM or Transformers 5.2.0 or later.

  8. Mistral AI · new models on Hugging FaceOfficialAI score32

    Mistral releases LIDstral-Arabic, a language and dialect classifier for Arabic-script text

    AIMistral AI has released LIDstral-Arabic, a fast classifier that identifies Modern Standard Arabic, Arabic dialects, and non-Arabic languages written in Arabic script across 51 classes. On Moroccan Darija, it scores 88.67% F1, versus 72.65% for LahjatBERT ALDi CL and 71.28% for GlotLID v3, across 84,870 evaluation examples. The model runs on CPU and is available from a private Hugging Face repository under Apache 2.0.

  9. IdeogramOfficialAI score45

    Ideogram 4.5 keeps edited images intact across 30 consecutive edits

    AIArtificial Analysis ran 30 consecutive real estate staging edits through four image editing models, and Ideogram 4.5 kept most of each image unchanged while others drifted. Ideogram 4.5 and FLUX 3 left 95% or more of the image untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image and left only about a fifth unchanged. Nano Banana 2.1 kept its edits local but gradually darkened the rest of the image.

  10. Artificial AnalysisOfficialAI score34

    HeyGen Voice tops Artificial Analysis Controlled Voice TTS Arena leaderboard

    AIHeyGen Voice ranks first on the Artificial Analysis Controlled Voice TTS Arena leaderboard with an Elo of 1,201 across 1,468 appearances, ahead of Qwen-Audio-3.1-TTS-Plus at 1,182 and ElevenLabs' Eleven v4 Turbo at 1,166. On pronunciation robustness it scores 83.1%, ranking #10 of 29 models, and it is priced at $30 per 1M characters and processes 40 characters per second.

    GIF from @ArtificialAnlys's post
  11. Artificial AnalysisOfficialAI score22

    HeyGen Voice ranks first in Artificial Analysis assistant and customer service arenas

    AIHeyGen Voice ranks first in the Artificial Analysis Controlled Voice Arena for Assistants at 1,214 and Customer Service at 1,212. It ranks second in Knowledge Sharing at 1,158 and fourth in Entertainment at 1,178. By accent, it ranks first in English (UK) at 1,190 and second in English (US) at 1,206, behind Qwen-Audio-3.1-TTS-Plus at 1,220.

    Image from @ArtificialAnlys's post
  12. Artificial AnalysisOfficialAI score32

    HeyGen Voice scores 83.1% on pronunciation robustness benchmark

    AIHeyGen Voice scores 83.1% overall on Artificial Analysis's Pronunciation Robustness benchmark, ranking #10 of 29 models. The benchmark has human reviewers judge whether text-to-speech models pronounce challenging text correctly against pre-agreed accepted pronunciations. HeyGen Voice places #2 for preserving exact sequences at 83.3%, behind SpaceXAI TTS at 85.7%, and scores 77.0% on expanding shorthand, while Eleven v4 leads that category at 94.1%.

    Image from @ArtificialAnlys's post
  13. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  14. Artificial AnalysisOfficialAI score50

    Ideogram 4.5 keeps edited photos intact over 30 consecutive edits

    AIArtificial Analysis ran four image editing models through 30 consecutive edits of the same real estate photo, with each model editing its previous output. Ideogram 4.5 and FLUX 3 left 95% or more of the image essentially untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image each time, leaving only about a fifth unchanged. Nano Banana 2.1 edited locally but shifted and gradually darkened the rest of the image.

    Video from @ArtificialAnlys's post
  15. 🚨 AI News | TestingCatalogXAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  16. Mike KnoopXAI score43

    Knoop says ARC-AGI-2 is much harder than ARC-AGI-1

    AIMike Knoop says ARC-AGI-2 is far harder than ARC-AGI-1, even amid rapid progress on math. He adds that the final open solutions will be useful artifacts to study, and notes that the Kaggle Grand Prize bonus threshold of 85% has been reached this year, the final year for ARC-AGI-2 on Kaggle.

  17. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  18. Mike KnoopXAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  19. ARC PrizeOfficialAI score42

    ARC Prize 2026 ARC-AGI-2 high score reaches 88.06%

    AITufa Labs posted an 88.06% score on ARC-AGI-2, a new high for the ARC Prize 2026 leaderboard. ARC Prize says a $150K bonus prize, on top of guaranteed prizes, will be split among all teams scoring over 85%.

    Image from @arcprize's post
  20. Hamel HusainXAI score22

    How to make the team case for investing in evaluations

    AIHamel Husain advises reviewing real user interactions and showing the team the problems found. He says to then fix those problems and show what improved as the case for investing in evaluations.

    Image from @HamelHusain's post
  21. LangChainOfficialAI score29

    LangSmith data shows Claude Sonnet 5 and GPT-5.6 Luna gaining ground

    AILangChain reports that over the last month Claude Sonnet 5 rose from #9 to #2 in model adoption, with 51% more organizations using it. GPT-5.6 Luna climbed from #3 to #1 in call footprints, up 65% in calls, while smaller, faster models dominate call footprints overall. Two open-weights models entered the adoption top 10 but do not lead in call volume.

    Image from @LangChain's post
  22. Arena.aiOfficialAI score24

    Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings

    AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.