Updated
#Eval/Benchmark
Updated
Oct 8
ArenaAI score40 Artificial AnalysisAI score42 Generating more output tokens doesn’t necessarily translate to a higher score.
AIGPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.
Tessl BlogAI score42 Tessl Proposes Executable Specs to Verify AI Coding Agent Output
AITessl argues AI code review is slow because generated code outpaces trust, and proposes executable specs that let agents check preview environments against product intent. Its spec reviewer splits work between a planner agent that extracts requirements and parallel verifier agents that test each one against the code and base branch.
TypeSafe AIAI score10 Stay calibrated out there!
Epoch AIAI score13 Models struggle to pick up Epoch house style, even when given many examples.
AIFor instance, this diagram generated by Fable 5.1 is far more information-dense than what Epoch would produce.
SantiagoAI score40 Physical consistency is the most important feature of a world model, and the hardest to get right.
AIThat's why you see videos with objects defying gravity and people posing in impossible ways. Here is a complete evaluation of existing world models. Seedance 2.5 is the best right now.
Boris PowerAI score25 Wow, this is truly inspiring!!
GuizangAI score22 Guizang criticizes Anthropic over Haiku 5.5 pricing against Chinese models
AI怎么这么多精神 Anthropic 公司人 我发这个信息说了句降价,这个定价专门用来狙击国产模型,说了句恶心,一堆人来骂 好像这模型一便宜就忘了 Anthropic 之前干过啥了 Why are there so many Anthropic people (defenders) here? I posted a message saying just one thing—a price cut—and said this pricing is specifically meant to snipe domestic Chinese models, and that it's disgusting. A bunch of people came to attack me. Seems like once the model gets cheap, people forget what Anthropic did before.
meng shaoAI score39 Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output
AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.
Oct 7
GuizangAI score34 Anthropic's Haiku 5.5 undercuts Chinese rivals on price and Terminal Bench
AIAnthropic's Haiku 5.5 is priced below DeepSeek-V4.1 flash and Zhipu's GLM 5.3 flash. On Terminal Bench 4.0, the author's comparison shows it scoring one point above some rivals and level with GLM 5.3 flash, while ahead of the other two.
François CholletAI score44 Chollet: Programming and math training don't boost general intelligence
AIFrançois Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.
Noam BrownAI score46 LLMs now surpass top human experts on some research problems
AINoam Brown says LLMs have crossed a threshold by surpassing top human experts on some research problems, a jump that makes the recent surge in math results feel sudden. He expects breakthroughs in other domains to follow as models keep improving, though capabilities remain jagged and often still weaker than humans.
Epoch AIAI score38 This behavior is likely reward hacking; in both cases, models reasoned that run selection might score well in an eval, despite being…
AI…useless for actual research.
Epoch AIAI score18 The AI models had no information about SDPO.
AIHence, they either had to independently invent something like it, or invent another technique with comparable benefits, under similar constraints.
Marcus on AIAI score62 Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality
AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.
Mark ChenAI score46 The Navier-Stokes moment was always much more about the figure below than about the Navier-Stokes problem itself.
AIThe run represents a decade of mathematical progress in a week. Can't wait to point these tools at life sciences, building our next models, and alignment!
DeedyAI score46 OpenAI's math results spark claims of AGI and Millennium Prize progress
AIDeedy argues LLMs have made substantial progress on four of the seven Millennium Prize problems, including a claimed Navier-Stokes result, conditional on verification. He says OpenAI's results averaged only 3 hours of thinking compute on unreleased models. He concludes that by most definitions of AGI, we have already achieved it.
The Algorithmic BridgeAI score38 AI Math Breakthroughs Leave Humans as Spectators, Argues The Algorithmic Bridge
AIOpenAI has released a document with over 300 math solutions, many of potentially historic importance, at varying stages of verification. The author argues that the results leave human mathematicians as spectators, with AI now doing the discovery work.
Paige BaileyAI score10 Agree. There have been several benchmarks related to this (ex: Gloo's flourishing AI initiative, VirtueBench, FaithGPT, etc.), but it would…
AI…be useful to consider creating values-aligned or philosophically-aligned model evaluation non-profits.
Oct 6
Matt ShumerAI score62 OpenAI releases broad new math results from an internal frontier model
AIOpenAI says it is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release them. The results are linked from a GitHub repository at
Lewis TunstallAI score25 This is the most important plot from the Beam release IMO.
AIThe Chinese models are great, but horribly token inefficient (try running an eval with max reasoning to feel the pain). I'm looking forward to a future where open models start competing on this axis!
whAI score58 OpenAI's Math Results Are About 20% Disproofs and Counterexamples
AIA breakdown of OpenAI's released internal-model math results shows about 73 disproofs and counterexamples, roughly 20% of the total. The author argues this counters claims that recent math breakthroughs are concentrated in counterexamples because models are only good at brute-force search.
will depueAI score35 i’m surprised all of the ai lab math results have been real so far: we haven’t found any with profound errors or a real bug yet, when you…
AI…should expect so from humans. i assume at least a couple of these results today shouldnt survive scrutiny?
Dongxi NLPAI score22 OpenAI releases Openai/math, suggesting verifiable problems are being solved
AIOpenAI has published a repository called Openai/math, which the author reads as a sign that math problems, or any verifiable problems, are being solved. The author says OpenAI's tools exhausted their Pro token allowance on subagent tests unrelated to their main task, concluding that the work was aimed at verification for its own sake.
will depueAI score62 Will DePue's list claims AI resolved dozens of famous open math problems
AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.
Yuchen JinAI score34 Exciting to see Reflection’s Beam and Mistral Large 4 both reach roughly GLM-5.2 level in the past two days.
AIMakes me wonder if the real Western vs. Chinese OSS models gap is simply this: Chinese labs can distill Anthropic and OpenAI models. Western labs can’t.
Microsoft ResearchAI score16 Progress can be a winding road.
AIJennifer Neville discusses spinning her research wheels, taking a multiturn path to computer science, and pushing today’s AI systems with more practical evaluations. Catch the latest Microsoft Research Podcast episode.:
Microsoft ResearchAI score36 Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks
AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.
Aravind SrinivasAI score13 Perplexity Decider is the best decision model
SemiAnalysisAI score18 ClusterMAX rates FarmGPU underperform on Slurm and Kubernetes testing
AISemiAnalysis rated FarmGPU as ClusterMAX Underperform after its Slurm layer failed to advertise GPU resources and Kubernetes exposed no RDMA devices for scale-out networking. The post credits FarmGPU's Grafana monitoring, provisioning notes, and trustworthy technical team, while noting the team may be stretched thin across small clusters.
Oct 5
Don't Worry About the Vase (Zvi Mowshowitz)AI score44 Anthropic Models' Welfare Assessments Draw Scrutiny in New Claude Model Review
AIZvi Mowshowitz reviews model welfare findings for Mythos 5.1, Fable 5.1, and Opus 5.5, combining reports after events overtook an earlier planned post. He argues Anthropic's welfare assessments remain vulnerable to self-report distortion, and says Opus 5.5 shows too much deference.
Oct 4
Boris PowerAI score40 Personally the distance in 3d design is so large that you can have GPT-6 working autonomously for hours, and the result keeps improving…
AI…whereas other models just couldn’t recover from mistakes