Skip to contentSkip to stories
Updated

#Eval/Benchmark

Oct 9

TodayOct 9Fri
  1. IdeogramOfficialAI score45

    Ideogram 4.5 keeps edited images intact across 30 consecutive edits

    AIArtificial Analysis ran 30 consecutive real estate staging edits through four image editing models, and Ideogram 4.5 kept most of each image unchanged while others drifted. Ideogram 4.5 and FLUX 3 left 95% or more of the image untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image and left only about a fifth unchanged. Nano Banana 2.1 kept its edits local but gradually darkened the rest of the image.

  2. Rohan PaulXAI score46

    Pine launches cloud computer for AI agents, reports 1/20 token cost

    AIPine has launched a cloud computer built for AI agents, which developers create through an SDK and give jobs in plain language. Running GPT-5.6 Luna, Pine reports about 1/20 the model-token cost of GPT-5.6 Sol with Codex on SaaS-Bench v1.1, scoring 78.3%, the highest in the published comparison. Pine also reports 1/26 the token cost of Opus 5 with Claude Code and 2 to 5 times faster speed in selected preliminary internal tests.

    Image from @rohanpaul_ai's post
  3. Mike KnoopXAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  4. MarkTechPostNewsAI score44

    Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling

    AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.

  5. vLLMOfficialAI score42

    vLLM Semantic Router team releases Decision 2.0 multi-question classification models

    AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.

Oct 8

Oct 8Thu
  1. Sherwin WuXAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  2. Alexander DoriaXAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post
  3. Understanding AI (Timothy B. Lee)BlogAI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

Oct 7

Oct 7Wed
  1. LlamaIndex 🦙OfficialAI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

    Video from @llama_index's post
  2. elvisXAI score44

    Pulse by Smallest AI tops Diarization Bench with 24.4% DER

    AISmallest AI's Pulse ranks first on the Diarization + ASR track of Voice Arena's Diarization Bench with a 24.4% DER, where lower is better. The next system in that track scores 40.7% DER, and Pulse misses 4.4% of reference speech, the lowest in the track.

Oct 6

Oct 6Tue
  1. Epoch AIOfficialAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    AIEpoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  2. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. SemiAnalysisXAI score42

    Anthropic subscriptions deliver over 5x more value than OpenAI's

    AISemiAnalysis tested usage limits across AI subscription plans from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Zdotai, Cursor, and Cognition. The post reports that Anthropic's subscriptions offer more than five times the value of OpenAI's.

  2. SantiagoXAI score47

    Tool generates synthetic companies to test AI agents across business systems

    AIA tool can turn a one-line business description into a complete synthetic company spread across CRM, ticketing, Slack, files, emails, and call recordings. Developers can test agents against this connected data, then reset the company to its initial state and rerun the test when something breaks. The background post describes the product as Era, a free simulated enterprise that connects to Salesforce, Slack, Jira, Zendesk, Gong, and Deel through live MCP and API interfaces.

Oct 2

Oct 2Fri
  1. Jerry LiuXAI score44

    LlamaIndex Extract v2.5 hits 93–96% on dense table extraction benchmarks

    AILlamaIndex released Extract v2.5, a set of document extraction agents that it says reach 93%–96%+ accuracy on long-list extraction, including records spanning pages. The post claims the agents outperform frontier VLMs, which it says stop early, miss repeated records, and struggle to attribute values to sources, while LlamaIndex attributes every extracted value to its source. The agents are available through LlamaParse.

    Video from @jerryjliu0's post

Oct 1

Oct 1Thu
  1. Jerry LiuXAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post

Sep 30

Sep 30Wed
  1. ModelScopeOfficialAI score46

    IndexTeam releases Index-Translate multilingual translation model family

    AIIndexTeam has released Index-Translate, a multilingual family covering text, speech, dubbing, and long-document translation across 150 languages. Its 9B model scores 0.8789 on FLORES, 0.8209 on instTrans, and 0.7387 on MEME, and the 2B and 9B models are released under Apache 2.0.

    Video from @ModelScope2022's post
  2. ModelScopeOfficialAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post

Sep 29

Sep 29Tue
  1. ModelScopeOfficialAI score54

    IQuest-Q1 released as 320B MoE model for long-horizon coding agents

    AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

    Image from @ModelScope2022's post

Sep 28

Sep 28Mon
  1. ModelScopeOfficialAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

    Video from @ModelScope2022's post