Gemini 2.5 Pro
Google·Released Jun 17, 2025
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
3 results · 26.7- SciCodeArtificial Analysis46.3%Reading 39.5
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor0.0%Reading at most 34.5
- Vibe Code BenchVals AINear the floor0.4%Reading at most 10.6
Research and reasoning
4 results · 27.2- Chess PuzzlesEpoch AI20.0%Reading 30.5
- Humanity's Last Exam (AA run)Artificial Analysis22.5%Reading 29.6
- FrontierMath Tier 4Epoch AINear the floor0.0%Reading at most 22.2
- FrontierMath Tiers 1–3Epoch AI24.6%Reading 21.8
Professional work
3 results · 35.6- MedCodeVals AI50.6%Reading 69.8
- τ-Bench Banking (AA run)Artificial Analysis9.7%Reading 6.4
- τ²-Bench Telecom (AA run)Artificial Analysis54.1%Reading 5.6
Knowledge and accuracy
3 results · 15.5- AA-LCRArtificial Analysis69.0%Reading 33.2
- BullshitBenchBullshitBench · default effort23.6%Reading 6.9
- IFBench (AA run)Artificial Analysis48.7%Reading 0.8
Human preference
1 result · 35.0- Arena TextLMArena1,445.6Reading 39.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBench (AA run)Artificial Analysis · CodingReference only80.1%
- SWE-bench VerifiedEpoch AI · CodingReference only57.6%
- SWE-bench VerifiedVals AI · CodingReference only54.4%
- AIME (AA run)Artificial Analysis · Research and reasoningReference only87.7%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only84.2%
- CorpFinVals AI · Professional workReference only60.8%
- MedScribeVals AI · Professional workWatching73.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only85.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only16.1
- Arena VisionLMArena · VisionReference only1,247.5
- Arena WebDevLMArena · Writing and designReference only1,227.7
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- Code Migration · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- LiveBench Reasoning · LiveBench
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- APEX-Agents · Mercor
- Harvey Legal Agent Benchmark · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- LiveBench Data Analysis · LiveBench
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench