Evaluations we read
What each evaluation measures, who runs it, and whether it counts toward the score. Only results an evaluator ran itself, the same way for every model, are scored. How it's scored
Updated Oct 8, 9:18 PM ET
Coding
18 evaluationsEngineering work in a terminal, building and extending apps, migrating code, patching security flaws.
- SciCodeScored
Turns physics, chemistry and biology problems into working scientific code.
Artificial Analysis148 modelsUpdated Oct 9, 2026
The same terminal tasks, re-run in Artificial Analysis's own harness; shares one vote with Vals's run.
Artificial Analysis141 modelsUpdated Oct 9, 2026
- Vibe Code BenchScored
Builds a web app from a written spec; the app's features are tested.
Vals AI100 modelsUpdated Oct 7, 2026
- Code MigrationScored
Moves a working program from one language or framework to another.
Vals AI74 modelsUpdated Oct 8, 2026
- LiveBench CodingScored
Fresh code-generation, completion and agentic repository tasks, averaged.
LiveBench65 modelsUpdated Oct 7, 2026
- ProgramBenchScored
Rebuilds a program from its executable alone, matching its behavior.
Vals AI61 modelsUpdated Oct 8, 2026
- APEX-SWEScored
Ships realistic software engineering work end to end, pass@1.
Mercor54 modelsUpdated Oct 9, 2026
- Terminal-Bench 4.0Scored
Finishes whole engineering tasks alone in a command-line sandbox.
Vals AI44 modelsUpdated Oct 7, 2026
- CyberBench PatchScored
Patches real security flaws in open-source projects.
Vals AI42 modelsUpdated Oct 7, 2026
- Vibe Code Bench 1–100Scored
Keeps extending an existing app without breaking what already works.
Vals AI22 modelsUpdated Oct 8, 2026
- SkillsBenchWatching
Agent tasks with and without reusable skills; run conditions still being checked.
Vals AI34 modelsUpdated Sep 27, 2026
- SRE BenchWatching
Site-reliability incidents in live systems; too new to score.
Vals AI32 modelsUpdated Oct 8, 2026
- MirrorCodeWatching
Reproduces software over long sessions; too few models so far.
Epoch AI9 modelsUpdated Sep 22, 2026
- LiveCodeBench (AA run)Reference only
Artificial Analysis's run of the contest problems.
Artificial Analysis277 modelsUpdated Oct 9, 2026
- LiveCodeBenchReference only
Contest-style coding problems; most frontier models now solve nearly all.
Vals AI130 modelsUpdated Sep 1, 2026
- SWE-bench VerifiedReference only
Fixes GitHub issues; widely trained on, so shown for reference.
Vals AI84 modelsUpdated Sep 1, 2026
- IOIReference only
Olympiad algorithm problems; shown for reference.
Vals AI41 modelsUpdated Oct 7, 2026
- SWE-bench VerifiedReference only
Epoch's own run of the same issue-fixing set, for cross-checking.
Epoch AI31 modelsUpdated Jun 25, 2026
Research and reasoning
12 evaluationsResearch math, logic and game puzzles, proofs, experiments and scientific workflows.
Expert questions across fields, answered without tools.
Artificial Analysis467 modelsUpdated Oct 9, 2026
- Chess PuzzlesScored
Finds the winning move from a position written out as text.
Epoch AI141 modelsUpdated Sep 29, 2026
- FrontierMath Tiers 1–3Scored
Unpublished research-level math problems with checkable answers.
Epoch AI83 modelsUpdated Sep 29, 2026
- Mystery Game PuzzlesScored
Works out the right move in games whose rules it has to infer.
Epoch AI73 modelsUpdated Sep 29, 2026
- LiveBench ReasoningScored
Fresh logic puzzles and competition math, averaged.
LiveBench65 modelsUpdated Oct 7, 2026
- FrontierMath Tier 4Scored
The hardest FrontierMath tier; shares one vote with Tiers 1–3.
Epoch AI63 modelsUpdated Sep 29, 2026
- ProofBenchScored
Writes complete mathematical proofs that are graded line by line.
Vals AI49 modelsUpdated Oct 7, 2026
- Terminal-Bench ScienceScored
Runs scientific workflows and analyses in a terminal.
Vals AI38 modelsUpdated Oct 8, 2026
- MysteryMechanismScored
Designs its own experiments to uncover a hidden mechanism.
Vals AI25 modelsUpdated Oct 7, 2026
- BioMysteryBenchWatching
Biology research puzzles; still checking how it is graded.
Vals AI25 modelsUpdated Oct 8, 2026
- AIME (AA run)Reference only
Competition math; saturated at the top.
Artificial Analysis207 modelsUpdated Oct 9, 2026
- OTIS Mock AIMEReference only
Competition math that frontier models have nearly saturated.
Epoch AI192 modelsUpdated Sep 29, 2026
Professional work
16 evaluationsFinance, law, tax, medical coding, spreadsheets and long agent assignments.
Resolves telecom support cases with tools and policies, alongside a simulated customer.
Artificial Analysis322 modelsUpdated Oct 9, 2026
- τ-Bench Banking (AA run)Scored
Handles bank customer service: policies, products and account changes; shares one vote with the telecom set.
Artificial Analysis154 modelsUpdated Oct 9, 2026
- MedCodeScored
Assigns diagnosis codes from hospital records.
Vals AI96 modelsUpdated Oct 8, 2026
- Finance AgentScored
An analyst's daily work: reading filings, adjusting numbers, building models.
Vals AI75 modelsUpdated Oct 7, 2026
Legal work delivered as documents, spreadsheets and slides a lawyer can review.
Vals AI75 modelsUpdated Oct 7, 2026
- Legal Research BenchScored
Finds and applies the law across practice areas.
Vals AI74 modelsUpdated Oct 7, 2026
- EMBScored
Builds and repairs financial models in spreadsheets: DCF, LBO, M&A.
Vals AI71 modelsUpdated Oct 7, 2026
- Tax Agent BenchScored
Answers professional tax questions from facts and current rules.
Vals AI66 modelsUpdated Oct 8, 2026
- LiveBench Data AnalysisScored
Reshapes tables, predicts joins and reads event sequences.
LiveBench65 modelsUpdated Oct 7, 2026
- APEX-AgentsScored
Long banking, consulting and legal assignments done as an agent, pass@1.
Mercor52 modelsUpdated Oct 9, 2026
- MedScribeWatching
Writes clinical notes; checking how the notes are graded.
Vals AI98 modelsUpdated Oct 8, 2026
- APEX-AccountingWatching
Accounting assignments; pass rates are still near the floor.
Mercor30 modelsUpdated Oct 9, 2026
- EBR-benchWatching
Learns new rules during a long task and keeps using them.
Epoch AI24 modelsUpdated Sep 29, 2026
- LegalBenchReference only
Short legal-reasoning tasks; close to saturated.
Vals AI139 modelsUpdated Oct 1, 2026
- TaxEvalReference only
Tax questions without tools; close to saturated.
Vals AI133 modelsUpdated Sep 1, 2026
- CorpFinReference only
Questions over long credit agreements; older and close to saturated.
Vals AI122 modelsUpdated Aug 12, 2026
Knowledge and accuracy
9 evaluationsFactual recall, instruction following, long documents and spotting a false premise.
- AA-LCRScored
Reasons over very long documents.
Artificial Analysis388 modelsUpdated Oct 9, 2026
- IFBench (AA run)Scored
Follows new, precisely checkable output instructions.
Artificial Analysis330 modelsUpdated Oct 9, 2026
- BullshitBenchScored
Notices a false premise instead of answering along with it.
BullshitBench115 modelsUpdated Sep 25, 2026
- SimpleQA VerifiedScored
Short factual questions answered from memory.
Epoch AI78 modelsUpdated Sep 29, 2026
- LiveBench LanguageScored
Word puzzles, typo fixing and plot reconstruction.
LiveBench65 modelsUpdated Oct 7, 2026
Rewrites and summaries under exact formatting instructions.
LiveBench65 modelsUpdated Oct 7, 2026
- GPQA DiamondReference only
Graduate-level science questions; saturated at the top.
Epoch AI214 modelsUpdated Sep 29, 2026
- GPQA DiamondReference only
Vals's run of the same science questions.
Vals AI125 modelsUpdated Sep 1, 2026
- MMLU-ProReference only
Broad multiple-choice knowledge; saturated at the top.
Vals AI125 modelsUpdated Sep 1, 2026
Human preference
1 evaluationPeople comparing two anonymous answers side by side.
- Arena TextScored
People compare two anonymous answers and pick the one they prefer (style-controlled).
LMArena385 modelsUpdated Oct 2, 2026
Composite indices
3 evaluationsEach evaluator's own summary of its evaluations: a cross-check, never scored.
- AA Intelligence IndexReference only
Artificial Analysis's composite of its own evaluations; a cross-check, not scored.
Artificial Analysis492 modelsUpdated Oct 9, 2026
- LiveBench averageReference only
LiveBench's average across all its categories; a cross-check.
LiveBench65 modelsUpdated Oct 7, 2026
- Vals IndexReference only
Vals's composite of its own benchmarks; a cross-check.
Vals AI44 modelsUpdated Oct 7, 2026
Vision
2 evaluationsUnderstanding images; shown for reference, not scored yet.
- Arena VisionReference only
People compare answers about images.
LMArena145 modelsUpdated Oct 2, 2026
- MMMU ProReference only
College-level questions about charts, diagrams and photos.
Vals AI83 modelsUpdated Sep 1, 2026
Writing and design
1 evaluationCreative work and design; shown for reference, not scored yet.
- Arena WebDevReference only
People pick the better of two generated web apps.
LMArena124 modelsUpdated Oct 6, 2026
Multilingual
1 evaluationOther languages; shown for reference, not scored yet.
- MGSMReference only
Grade-school math in ten languages.
Vals AI64 modelsUpdated Jan 9, 2026
Evaluators
Read once a day, politely: one request at a time per site, cached, under our own crawler name. When an evaluator cannot be read, its last results stay and it is marked here.
- LiveBenchRead
Results published openly with the benchmark (code under Apache 2.0).
Last read Oct 9, 2026
- Epoch AIRead
Benchmarking Hub data under CC BY 4.0; credited here.
Last read Oct 9, 2026
- LMArenaRead
Leaderboard dataset on Hugging Face under CC BY 4.0.
Last read Oct 9, 2026
- Vals AIRead
Public benchmark pages; results credited and linked.
Last read Oct 9, 2026
- MercorRead
Public APEX leaderboards; results credited and linked.
Last read Oct 9, 2026
- BullshitBenchRead
Published in its GitHub repository under the MIT license.
Last read Oct 9, 2026
Free data API with attribution; read only when ARTIFICIAL_ANALYSIS_API_KEY is set.
Last read Oct 9, 2026
Model names, release dates, prices, context windows and open-weight flags come from OpenRouter's public model list.