Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. Rohan PaulXAI score57

    Microsoft paper finds coding agents struggle more with code understanding than editing

    AIMicrosoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time. Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep. On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

    Image from @rohanpaul_ai's post
  2. OpenRouterOfficialAI score40

    Microsoft-Decision-1 is live on OpenRouter at $0.042 per million input tokens

    AIOpenRouter has released Microsoft-Decision-1, a model Microsoft says posts the highest accuracy across 36 blind benchmarks of about 150K questions. Microsoft says it runs 4.5x faster than the runner-up and 35x faster than GPT-6 Sol, with decisions flipping on only 1.3% of perturbed inputs. The model is post-trained from Qwen3.5-9B, costs $0.042 per million input tokens, has free output and a 32K context window.

  3. StepFunOfficialAI score34

    StepFun's Step 5 Preview free on Nous Portal this week

    AIStepFun's Step 5 Preview, a 600B-total, 27B-active MoE model with 1M context and vision, is free to try on Nous Portal for one week. Nous Research says it scored 33.89 on the Hermes Index, the same score as GPT-6 Luna.

  4. 🚨 AI News | TestingCatalogXAI score36

    Microsoft releases Microsoft-Decision-1, a 9B model for fast decisions

    AIMicrosoft has made Microsoft-Decision-1 available on Microsoft Foundry, a model post-trained on Qwen3.5-9B for fast, single-pass decision scoring. Microsoft says it achieved the highest accuracy across a 36-benchmark comparison of nearly 150,000 questions, and runs 4.5 times faster than Quyet-1.0-Large and 35 times faster than GPT-6 Sol. Microsoft plans to rebase it on other models, including MAI and OpenAI models.

    Video from @testingcatalog's post
  5. elvisXAI score40

    Microsoft releases Microsoft-Decision-1, a fast model for decision-making tasks

    AIMicrosoft releases Microsoft-Decision-1, a model for fast decision-making, according to Satya Nadella. Nadella says it outperforms both LLMs and other decision models on structured decision tasks in latency and quality. Microsoft is testing it internally for incident response, quality control, and scientific discovery.

    Image from @omarsar0's post
  6. Nous ResearchOfficialAI score49

    StepFun's Step 5 Preview is free on Nous Portal for a week

    AINous Research says StepFun's Step 5 Preview is free on Nous Portal for the next week. The model is a 600B total, 27B active MoE with a 1M context window and vision support. It scored 33.89 on the Hermes Index, the same score as GPT-6 Luna.

    Video from @NousResearch's post
  7. Satya NadellaXAI score38

    Microsoft unveils Microsoft-Decision-1, a model for fast decision-making

    AIMicrosoft introduces Microsoft-Decision-1, a new model for fast decision-making that it says outperforms both LLMs and other decision models on structured decision tasks in latency and quality. The company says it is already testing the model across Microsoft for uses including incident response, quality control, and scientific discovery.

    Video from @satyanadella's post
  8. Mistral AI · new models on Hugging FaceOfficialAI score47

    Mistral releases Voxtral Mini 4B Realtime Arabic speech-to-text model

    AIMistral releases Voxtral Mini 4B Realtime Arabic, a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic under the Apache 2.0 License. The model has about 4.4 billion parameters, is fine-tuned from Voxtral-Mini-4B-Realtime-2602, and reaches an average 8.82% Character Error Rate across seven Arabic benchmarks at a 480 ms transcription delay. It can be run with vLLM or Transformers 5.2.0 or later.

  9. Mistral AI · new models on Hugging FaceOfficialAI score32

    Mistral releases LIDstral-Arabic, a language and dialect classifier for Arabic-script text

    AIMistral AI has released LIDstral-Arabic, a fast classifier that identifies Modern Standard Arabic, Arabic dialects, and non-Arabic languages written in Arabic script across 51 classes. On Moroccan Darija, it scores 88.67% F1, versus 72.65% for LahjatBERT ALDi CL and 71.28% for GlotLID v3, across 84,870 evaluation examples. The model runs on CPU and is available from a private Hugging Face repository under Apache 2.0.

  10. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  11. IThome · AINewsAI score46

    JetBrains Releases Mellum2.1 Coding Model With Near-Double Qwen3.5-9B Throughput

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts coding model with 2.5B active parameters under Apache 2.0, emphasizing agentic programming. Under high load, its inference throughput in tokens is nearly twice that of Qwen3.5-9B in JetBrains' comparison, and multi-token prediction (MTP) speeds single-request responses by about 1.6x. The model is available on Hugging Face for local or private-infrastructure deployment, with GGUF and vLLM MTP support announced for later.

Oct 8

Oct 8Thu
  1. QbitAINewsAI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    AIAnthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  2. Xiaomi MiMoOfficialAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    AIXiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  3. MiniMax (official)OfficialAI score34

    MiniMax H3 nears closed-source SOTA on physics in open video world models

    AIMiniMax says its open-source H3 model is almost on par with closed-source state-of-the-art video world models on physics. The claim is supported by a quoted benchmark, World Models' Last Exam in Physics, where eight leading models scored at most 57.76/100 across 40 physics tasks, and free-fall videos averaged only 26.61/100 on composite scores.

  4. Artificial AnalysisOfficialAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

    Image from @ArtificialAnlys's post
  5. Artificial AnalysisOfficialAI score31

    Grok Imagine Video 1.5 Lite leads on quality and speed benchmark

    AIAmong 12 models on AA-Video-T2V-Silent v2.0, Grok Imagine Video 1.5 Lite is the only one that is both fastest and highest quality, with no model beating it on both measures. It generates a 10-second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher but takes 94 seconds for a 5-second clip, while Vidu Q3 Turbo is 9 seconds faster on a 5-second 720p clip yet scores well below it.

    Image from @ArtificialAnlys's post
  6. Artificial AnalysisOfficialAI score46

    Grok Imagine Video 1.5 Lite outranks Veo 3.1 at a third of the cost

    AIGrok Imagine Video 1.5 Lite ranks #17 on AA-Video-T2V v2.0, two places above Google's Veo 3.1. At 1080p with audio, it costs $0.14 per second versus $0.40 per second for Veo 3.1. Compared with Grok Imagine Video 1.5, Lite is 44% cheaper at 1080p but ranks six places lower.

    Image from @ArtificialAnlys's post
  7. Artificial AnalysisOfficialAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    AISpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

    Video from @ArtificialAnlys's post
  8. Epoch AIOfficialAI score40

    OpenAI halves GPT-6.1 Sol cached-input price, speeds long prompts

    AIEpoch AI reports that OpenAI halved GPT-6.1 Sol's cached-input price compared with GPT-6 Sol. Its measurements also show the model handles long prompts faster. The post raises, but does not confirm, a new architecture for GPT-6.1 Sol.

    Image from @EpochAIResearch's post
  9. Sherwin WuXAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  10. The DecoderNewsAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  11. Aravind SrinivasXAI score22

    Perplexity Decider ranks first on DecisionBench at lowest cost

    AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.

  12. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  13. OpenRouter · New modelsBlogAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  14. JetBrains AI BlogOfficialAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

Oct 7

Oct 7Wed
  1. Google Developers BlogOfficialAI score62

    Google's AQuA agent diagnoses production failures in a multi-agent travel concierge

    AIGoogle Developers Blog introduces AQuA, an ambient quality agent that runs in a customer's Google Cloud project and samples production sessions to find recurring agent failures. In a 32-session travel-concierge sweep, it verified six issues and traced two of them to specific prompt lines, and a replay after the fixes raised full-session passes from 5/32 to 13/32. The post notes that verification and diagnosis are model-based, and that the tool proposes edits without applying them.

    Why it matters: The post walks through a concrete production workflow, from sweep and verification to a code-anchored fix and replay, that shows how to diagnose silent agent failures.

  2. ChatGPTOfficialAI score85

    GPT-6 with Intelligent UI rolls out globally in ChatGPT, Free and Go tiers next day

    AIGPT-6 with Intelligent UI begins rolling out globally in the ChatGPT Chat tab for Plus, Pro, Business, and Enterprise users today. The rollout expands to Free and Go tiers starting tomorrow. Plus, Pro, Business, and Enterprise get GPT-6 Sol, while Free and Go get GPT-6 Luna.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  3. Hugging Face BlogOfficialAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  4. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  5. Artificial Analysis ArticlesOfficialAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    Why it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.