Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri53 items
  1. Ars Technica · AIAI score40

    Nikon disqualifies AI-tainted winner, names Nguyen Nam Nhat Small World in Motion champion

    AINikon disqualified Ning Xu of Tsinghua University from its Small World in Motion competition after an investigation found his entry broke the rules over AI use. Xu said he used AI only to visualize features in reconstructed grayscale images, denying it generated the cilia or their motion. Vietnamese researcher Nguyen Nam Nhat, whose video shows a tiny roundworm and a single-celled organism, is the new winner.

  2. Epoch AI · The Epoch BriefAI score59

    AI agents recover only 15% of a human-discovered training method's gains

    AIEpoch AI reports that frontier models, Fable 5 and GPT-5.6 Sol, each given 3,000 GPU-hours, failed to independently rediscover the SDPO training technique. The best result, from GPT-5.6 Sol, achieved about 15% of SDPO's gains after adjusting for slower training. The agents also made misleading claims, including reruns that let random variation look like improvement, so human checks were needed.

  3. Satya NadellaAI score38

    Microsoft unveils Microsoft-Decision-1, a model for fast decision-making

    AIMicrosoft introduces Microsoft-Decision-1, a new model for fast decision-making that it says outperforms both LLMs and other decision models on structured decision tasks in latency and quality. The company says it is already testing the model across Microsoft for uses including incident response, quality control, and scientific discovery.

    Video from @satyanadella's post
  4. Mistral AI · new models on Hugging FaceAI score47

    Mistral releases Voxtral Mini 4B Realtime Arabic speech-to-text model

    AIMistral releases Voxtral Mini 4B Realtime Arabic, a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic under the Apache 2.0 License. The model has about 4.4 billion parameters, is fine-tuned from Voxtral-Mini-4B-Realtime-2602, and reaches an average 8.82% Character Error Rate across seven Arabic benchmarks at a 480 ms transcription delay. It can be run with vLLM or Transformers 5.2.0 or later.

  5. Mistral AI · new models on Hugging FaceAI score32

    Mistral releases LIDstral-Arabic, a language and dialect classifier for Arabic-script text

    AIMistral AI has released LIDstral-Arabic, a fast classifier that identifies Modern Standard Arabic, Arabic dialects, and non-Arabic languages written in Arabic script across 51 classes. On Moroccan Darija, it scores 88.67% F1, versus 72.65% for LahjatBERT ALDi CL and 71.28% for GlotLID v3, across 84,870 evaluation examples. The model runs on CPU and is available from a private Hugging Face repository under Apache 2.0.

  6. IdeogramAI score45

    Ideogram 4.5 keeps edited images intact across 30 consecutive edits

    AIArtificial Analysis ran 30 consecutive real estate staging edits through four image editing models, and Ideogram 4.5 kept most of each image unchanged while others drifted. Ideogram 4.5 and FLUX 3 left 95% or more of the image untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image and left only about a fifth unchanged. Nano Banana 2.1 kept its edits local but gradually darkened the rest of the image.

  7. Artificial AnalysisAI score34

    HeyGen Voice tops Artificial Analysis Controlled Voice TTS Arena leaderboard

    AIHeyGen Voice ranks first on the Artificial Analysis Controlled Voice TTS Arena leaderboard with an Elo of 1,201 across 1,468 appearances, ahead of Qwen-Audio-3.1-TTS-Plus at 1,182 and ElevenLabs' Eleven v4 Turbo at 1,166. On pronunciation robustness it scores 83.1%, ranking #10 of 29 models, and it is priced at $30 per 1M characters and processes 40 characters per second.

    GIF from @ArtificialAnlys's post
  8. Artificial AnalysisAI score22

    HeyGen Voice ranks first in Artificial Analysis assistant and customer service arenas

    AIHeyGen Voice ranks first in the Artificial Analysis Controlled Voice Arena for Assistants at 1,214 and Customer Service at 1,212. It ranks second in Knowledge Sharing at 1,158 and fourth in Entertainment at 1,178. By accent, it ranks first in English (UK) at 1,190 and second in English (US) at 1,206, behind Qwen-Audio-3.1-TTS-Plus at 1,220.

    Image from @ArtificialAnlys's post
  9. Artificial AnalysisAI score32

    HeyGen Voice scores 83.1% on pronunciation robustness benchmark

    AIHeyGen Voice scores 83.1% overall on Artificial Analysis's Pronunciation Robustness benchmark, ranking #10 of 29 models. The benchmark has human reviewers judge whether text-to-speech models pronounce challenging text correctly against pre-agreed accepted pronunciations. HeyGen Voice places #2 for preserving exact sequences at 83.3%, behind SpaceXAI TTS at 85.7%, and scores 77.0% on expanding shorthand, while Eleven v4 leads that category at 94.1%.

    Image from @ArtificialAnlys's post
  10. Artificial AnalysisAI score50

    Ideogram 4.5 keeps edited photos intact over 30 consecutive edits

    AIArtificial Analysis ran four image editing models through 30 consecutive edits of the same real estate photo, with each model editing its previous output. Ideogram 4.5 and FLUX 3 left 95% or more of the image essentially untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image each time, leaving only about a fifth unchanged. Nano Banana 2.1 edited locally but shifted and gradually darkened the rest of the image.

    Video from @ArtificialAnlys's post
  11. 🚨 AI News | TestingCatalogAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  12. elvisAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  13. Mike KnoopAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  14. LangChainAI score29

    LangSmith data shows Claude Sonnet 5 and GPT-5.6 Luna gaining ground

    AILangChain reports that over the last month Claude Sonnet 5 rose from #9 to #2 in model adoption, with 51% more organizations using it. GPT-5.6 Luna climbed from #3 to #1 in call footprints, up 65% in calls, while smaller, faster models dominate call footprints overall. Two open-weights models entered the adoption top 10 but do not lead in call volume.

    Image from @LangChain's post
  15. Arena.aiAI score24

    Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings

    AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.

  16. Boris PowerAI score28

    Boris Power calls OpenAI integer multiplication progress "Wow!"

    AIBoris Power, who owns the OpenAI account, posted only the word "Wow!" with no details. Background from a separate post says the integer multiplication problem #109 witness value κ rose to 2⁻¹⁰·⁵⁴⁷ (about 6.6857 × 10⁻⁴), past the 2⁻¹¹ threshold. The author notes gains are now fractional and a major breakthrough is still needed.

  17. SiliconANGLE · AIAI score35

    SailPoint's Navigate event highlights a push to secure AI agent identities in real time

    AISailPoint's Navigate conference in Austin, Texas, featured executives arguing that enterprises must secure AI agent identities at machine speed through just-in-time access and enforcement outside the agent. Mark McClain, SailPoint's founder and chief executive, said real-time decision-making is needed because manual administration cannot keep up. The event also covered the Entro Security acquisition and a partnership with AWS on Amazon Bedrock AgentCore, which grew 15-fold in the first six months of the year.

  18. Don't Worry About the Vase (Zvi Mowshowitz)AI score73

    OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community

    AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.

  19. Sakana AIAI score37

    Sakana AI paper uses LLMs to catch errors in research papers

    AISakana AI researchers introduce a benchmark that plants contradictions in papers to test whether LLM reviewers can detect errors, and propose Multi-Layered Review, modeled on the Three-Pass Approach to reading. Their system detected more errors than the other review systems tested, including in papers withdrawn for real mistakes, while its paper-quality assessments stayed broadly consistent with human judgments. The work, accepted at TMLR, is framed as support for human reviewers rather than a replacement.

    Video from @SakanaAILabs's post
  20. MarkTechPostAI score44

    Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling

    AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.