Skip to contentSkip to stories

Updated

#Eval/Benchmark

Oct 8

Oct 8Thu
  1. Epoch AI · The Epoch BriefAI score49

    Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure

    AIEpoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.

  2. Elvis SaraviaAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

  3. Sherwin WuAI score62

    Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

    AIArtificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

  4. Tessl BlogAI score52

    Cisco engineer argues agent skills need a context pipeline with evals

    AIJohn Groetzinger, writing in a personal capacity rather than for Cisco, argues that enterprise skills need packaging, evaluation, syncing, and distribution rather than scattered markdown files. He describes using skills to make cheaper models viable, converting curated TAC knowledge-base articles into maintained skills, and rolling out an eval framework across teams. He also describes syncing a repository README to Confluence with a deterministic script.

  5. Artificial AnalysisAI score28

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs.

    AICost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  6. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  7. TechCrunch · AIAI score46

    Arena raises $200M at $3.1B valuation, nearly doubling in 10 months

    AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.

  8. The DecoderAI score75

    Mathematicians call for OpenAI boycott after AI-generated proofs flood their field

    AIA group of mathematicians led by Terence Tao has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Tao and other Fields Medalists argue that mass-produced solutions undermine the discipline's focus on conceptual understanding, while Scott Aaronson contrasts this batch release with Anthropic's collaborative approach. The article reports that the internal model tested about 8,000 problems with roughly a five percent success rate.

  9. TechCrunch · AIAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    AIOpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.

  10. Elvis SaraviaAI score42

    Huge release from @odysseyml.

    AIOdyssey-3 Pro sets a new top score on Physics-IQ Verified, a benchmark that asks models to continue videos of real physics experiments. The robotics results stood out to me. With tens of hours of demos, the robot arm recovered from a missed grasp, a behavior that never appeared in those demos.

  11. SantiagoAI score46

    The #1 video-to-video model in the Physics-IQ Verified benchmark is finally live!

    AITheir research preview is) Odyssey 3 Pro is a world model, and nobody beats it for physical accuracy. You can use this model to control a robot, drive a car, play a video game, or pilot a drone. • It takes visual observations from the world • Uses these observations to learn how things work • Then maps that knowledge to the system's physical controls

  12. Elvis SaraviaAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.