Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Arena.aiAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  2. SantiagoAI score40

    Seedance 2.5 tops evaluation of world models for physical consistency

    AISantiago says physical consistency is the most important and hardest feature of a world model, and that many generated videos show objects defying gravity. He reports that Seedance 2.5 is currently the best among the evaluated world models. The post links to a physics evaluation benchmark in which eight video world models reached a top score of 57.76/100.

  3. LeiphoneAI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    AIAnthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  4. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

    Image from @OpenBMB's post
  5. Understanding AI (Timothy B. Lee)AI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  6. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  7. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  8. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

  9. QbitAIAI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  10. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  11. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

  12. Artificial Analysis ArticlesAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    AIArtificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

  13. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. François CholletAI score44

    Chollet: Programming and math training don't boost general intelligence

    AIFrançois Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  3. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    AIWaymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  4. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. Google Developers BlogAI score62

    Google's AQuA agent diagnoses production failures in a multi-agent travel concierge

    AIGoogle Developers Blog introduces AQuA, an ambient quality agent that runs in a customer's Google Cloud project and samples production sessions to find recurring agent failures. In a 32-session travel-concierge sweep, it verified six issues and traced two of them to specific prompt lines, and a replay after the fixes raised full-session passes from 5/32 to 13/32. The post notes that verification and diagnosis are model-based, and that the tool proposes edits without applying them.

    Why it matters: The post walks through a concrete production workflow, from sweep and verification to a code-anchored fix and replay, that shows how to diagnose silent agent failures.

  6. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

    Image from @GoogleResearch's post
  7. IThome · AIAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  8. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.