Gemini 3.7 Flash
Google·Released Aug 13, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 60.8- CyberBench PatchVals AI · high effort87.5%Reading 75.2
- SciCodeArtificial Analysis · high effort57.2%Reading 69.3
- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench · high effort68.6%Reading 63.5
- Vibe Code BenchVals AI · high effort70.4%Reading 57.5
- Code MigrationVals AI · high effort34.8%Reading 56.1
- APEX-SWEMercor · high effort42.0%Reading 55.3
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort13.6%Reading 51.9
- Terminal-Bench 4.0Vals AI · high effort12.1%Reading 47.7
Research and reasoning
8 results · 61.4- Chess PuzzlesEpoch AI · high effort47.0%Reading 68.6
- Humanity's Last Exam (AA run)Artificial Analysis · high effort47.9%Reading 67.1
- LiveBench ReasoningLiveBench · high effort90.6%Reading 65.2
- ProofBenchVals AI · high effort58.0%Reading 59.7
- Mystery Game PuzzlesEpoch AI · high effort37.0%Reading 58.9
- FrontierMath Tiers 1–3Epoch AI · high effort71.6%Reading 58.5
- FrontierMath Tier 4Epoch AI · high effort36.6%Reading 54.2
- Terminal-Bench ScienceVals AI · high effort7.1%Reading 53.9
Professional work
9 results · 61.6- MedCodeVals AI · high effort53.4%Reading 79.2
- APEX-AgentsMercor · high effort67.8%Reading 79.1
- Finance AgentVals AI · high effort59.0%Reading 71.5
- EMBVals AI · high effort71.3%Reading 70.7
- Legal Research BenchVals AI · high effort34.6%Reading 55.7
- Harvey Legal Agent BenchmarkVals AI · high effort8.8%Reading 55.5
- τ-Bench Banking (AA run)Artificial Analysis · high effort32.8%Reading 53.3
- Tax Agent BenchVals AI · high effort22.8%Reading 51.9
- LiveBench Data AnalysisLiveBench · high effort68.0%Reading 37.3
Knowledge and accuracy
5 results · 65.3- LiveBench Instruction FollowingLiveBench · high effortCapped79.9%Reading 108.9
- SimpleQA VerifiedEpoch AI · high effort69.2%Reading 83.2
- LiveBench LanguageLiveBench · high effort85.5%Reading 75.4
- AA-LCRArtificial Analysis · high effort81.7%Reading 52.0
- BullshitBenchBullshitBench · max effortCapped23.6%Reading 6.9
Human preference
1 result · 55.0- Arena TextLMArena · high effort1,487.7Reading 51.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only67.8%
- LiveCodeBenchVals AI · CodingReference only88.7%
- SkillsBenchVals AI · CodingWatching65.9%
- SRE BenchVals AI · CodingWatching4.6%
- SWE-bench VerifiedVals AI · CodingReference only80.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only97.2%
- LegalBenchVals AI · Professional workReference only87.3%
- MedScribeVals AI · Professional workWatching83.9%
- TaxEvalVals AI · Professional workReference only74.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only94.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only90.1%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only39.1
- LiveBench averageLiveBench · Composite indicesReference only78.8%
- Vals IndexVals AI · Composite indicesReference only51.3%
- Arena VisionLMArena · VisionReference only1,296.0
- MMMU ProVals AI · VisionReference only89.0%
- Arena WebDevLMArena · Writing and designReference only1,591.9
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- MysteryMechanism · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis