Gemini 3.8 Flash
Google·Released Sep 2, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 60.7- CyberBench PatchVals AI · high effort87.5%Reading 75.2
- SciCodeArtificial Analysis · high effort56.6%Reading 67.6
- ProgramBenchVals AI · high effortNear the floor1.0%Reading at most 65.7
- Vibe Code BenchVals AI · high effort78.7%Reading 62.9
- Vibe Code Bench 1–100Vals AI · high effort18.8%Reading 62.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort19.7%Reading 58.9
- Code MigrationVals AI · high effort36.5%Reading 57.6
- Terminal-Bench 4.0Vals AI · high effort19.2%Reading 56.0
- LiveBench CodingLiveBench · high effort63.4%Reading 49.8
- APEX-SWEMercor · high effort36.3%Reading 46.8
Research and reasoning
9 results · 65.0- Chess PuzzlesEpoch AI · high effort61.0%Reading 85.6
- MysteryMechanismVals AI · high effort36.5%Reading 72.2
- Mystery Game PuzzlesEpoch AI · high effort47.0%Reading 69.1
- Humanity's Last Exam (AA run)Artificial Analysis · high effort47.8%Reading 67.0
- LiveBench ReasoningLiveBench · high effort90.4%Reading 64.5
- Terminal-Bench ScienceVals AI · high effort8.6%Reading 56.5
- FrontierMath Tiers 1–3Epoch AI · high effort68.4%Reading 55.8
- ProofBenchVals AI · high effort48.0%Reading 55.2
- FrontierMath Tier 4Epoch AI · high effort22.0%Reading 44.6
Professional work
9 results · 63.1- Finance AgentVals AI · high effort61.4%Reading 76.3
- APEX-AgentsMercor · high effort64.3%Reading 74.5
- EMBVals AI · high effort72.2%Reading 71.9
- τ-Bench Banking (AA run)Artificial Analysis · high effort44.9%Reading 69.2
- Tax Agent BenchVals AI · high effort32.1%Reading 66.1
- Harvey Legal Agent BenchmarkVals AI · high effort10.0%Reading 63.0
- MedCodeVals AI · high effort48.1%Reading 61.6
- Legal Research BenchVals AI · high effort38.9%Reading 60.8
- LiveBench Data AnalysisLiveBench · high effortCapped54.0%Reading 2.6
Knowledge and accuracy
4 results · 79.2- LiveBench Instruction FollowingLiveBench · high effortCapped81.4%Reading 117.2
- LiveBench LanguageLiveBench · high effort87.8%Reading 84.6
- SimpleQA VerifiedEpoch AI · high effort69.7%Reading 84.0
- AA-LCRArtificial Analysis · high effort81.3%Reading 51.4
Human preference
1 result · 57.5- Arena TextLMArena · high effort1,494.8Reading 53.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only56.9%
- LiveCodeBenchVals AI · CodingReference only89.5%
- SkillsBenchVals AI · CodingWatching58.0%
- SRE BenchVals AI · CodingWatching11.1%
- SWE-bench VerifiedVals AI · CodingReference only80.0%
- BioMysteryBenchVals AI · Research and reasoningWatching62.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.9%
- APEX-AccountingMercor · Professional workWatching8.7%
- LegalBenchVals AI · Professional workReference only87.0%
- MedScribeVals AI · Professional workWatching84.5%
- TaxEvalVals AI · Professional workReference only74.4%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only95.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only94.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only90.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only40.9
- LiveBench averageLiveBench · Composite indicesReference only75.8%
- Vals IndexVals AI · Composite indicesReference only54.8%
- Arena VisionLMArena · VisionReference only1,290.3
- MMMU ProVals AI · VisionReference only89.1%
- Arena WebDevLMArena · Writing and designReference only1,583.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis