Gemini 3 Flash Preview
Google·Released Dec 17, 2025
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
3 results · 30.0- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- Vibe Code BenchVals AI · high effort20.2%Reading 30.0
- Code MigrationVals AI · high effort6.4%Reading 16.8
Research and reasoning
5 results · 46.9- Chess PuzzlesEpoch AI · high effort40.0%Reading 60.0
- Humanity's Last Exam (AA run)Artificial Analysis · thinking effort36.6%Reading 52.0
- FrontierMath Tiers 1–3Epoch AI51.2%Reading 42.8
- FrontierMath Tier 4Epoch AI17.1%Reading 40.4
- Mystery Game PuzzlesEpoch AI · high effort20.0%Reading 37.8
Professional work
8 results · 36.2- MedCodeVals AI · high effortCapped55.9%Reading 87.7
- Finance AgentVals AI · high effort42.6%Reading 39.1
- τ-Bench Banking (AA run)Artificial Analysis · thinking effort20.8%Reading 34.1
- Legal Research BenchVals AI · high effort18.3%Reading 32.0
- τ²-Bench Telecom (AA run)Artificial Analysis · thinking effort80.4%Reading 30.0
- Tax Agent BenchVals AI · high effort11.8%Reading 28.2
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor0.0%Reading at most 24.8
- EMBVals AI · high effort26.0%Reading 17.1
Knowledge and accuracy
4 results · 44.8- SimpleQA VerifiedEpoch AI · high effortCapped66.8%Reading 79.3
- IFBench (AA run)Artificial Analysis · thinking effort78.0%Reading 59.9
- AA-LCRArtificial Analysis · thinking effort78.0%Reading 45.8
- BullshitBenchBullshitBench · high effortCapped10.9%Reading -52.5
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only85.6%
- LiveCodeBench (AA run)Artificial Analysis · CodingReference only90.8%
- SWE-bench VerifiedEpoch AI · CodingReference only75.4%
- SWE-bench VerifiedVals AI · CodingReference only75.0%
- AIME (AA run)Artificial Analysis · Research and reasoningReference only97.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only95.6%
- CorpFinVals AI · Professional workReference only66.4%
- LegalBenchVals AI · Professional workReference only86.9%
- MedScribeVals AI · Professional workWatching69.9%
- TaxEvalVals AI · Professional workReference only73.9%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only89.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only87.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only88.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only26.3
- MMMU ProVals AI · VisionReference only87.6%
- MGSMVals AI · MultilingualReference only93.3%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- LiveBench Coding · LiveBench
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- SciCode · Artificial Analysis
- LiveBench Reasoning · LiveBench
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor
- LiveBench Data Analysis · LiveBench
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- Arena Text · LMArena