Gemini 3.5 Flash Lite
Google·Released Jul 21, 2026
Updated Oct 9, 6:21 AM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 33.4- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench · high effort60.7%Reading 43.0
- Vibe Code BenchVals AI · high effort37.2%Reading 40.4
- APEX-SWEMercor · high effort31.2%Reading 38.7
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor1.0%Reading at most 34.5
- CyberBench PatchVals AI · high effort73.2%Reading 33.9
- SciCodeArtificial Analysis41.3%Reading 25.8
- Code MigrationVals AI · high effort6.1%Reading 16.0
Research and reasoning
6 results · 27.0- Mystery Game PuzzlesEpoch AI · low effort19.0%Reading 36.2
- Chess PuzzlesEpoch AI · high effort22.0%Reading 34.1
- FrontierMath Tiers 1–3Epoch AI · high effort26.0%Reading 23.1
- Humanity's Last Exam (AA run)Artificial Analysis18.8%Reading 22.3
- FrontierMath Tier 4Epoch AI · high effortNear the floor0.0%Reading at most 22.2
- LiveBench ReasoningLiveBench · high effort67.0%Reading 16.4
Professional work
9 results · 28.3- Finance AgentVals AI · high effort47.4%Reading 48.7
- MedCodeVals AI · high effort43.5%Reading 46.0
- EMBVals AI · high effort43.2%Reading 38.3
- APEX-AgentsMercor · high effort29.3%Reading 30.8
- τ-Bench Banking (AA run)Artificial Analysis17.5%Reading 27.6
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor0.0%Reading at most 24.7
- Legal Research BenchVals AI · high effort13.9%Reading 23.2
- Tax Agent BenchVals AI · high effort7.3%Reading 12.0
- LiveBench Data AnalysisLiveBench · high effort53.2%Reading 0.8
Knowledge and accuracy
4 results · 49.5- BullshitBenchBullshitBench · xhigh effortCapped49.1%Reading 79.6
- LiveBench Instruction FollowingLiveBench · high effort67.2%Reading 51.3
- AA-LCRArtificial Analysis76.0%Reading 42.8
- LiveBench LanguageLiveBench · high effort71.8%Reading 37.4
Human preference
1 result · 39.2- Arena TextLMArena1,455.1Reading 42.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only79.0%
- SWE-bench VerifiedVals AI · CodingReference only75.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only71.1%
- CorpFinVals AI · Professional workReference only60.7%
- LegalBenchVals AI · Professional workReference only84.1%
- MedScribeVals AI · Professional workWatching70.9%
- TaxEvalVals AI · Professional workReference only72.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only83.3%
- GPQA DiamondVals AI · Knowledge and accuracyReference only83.8%
- MMLU-ProVals AI · Knowledge and accuracyReference only85.8%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only22.2
- LiveBench averageLiveBench · Composite indicesReference only63.9%
- Arena VisionLMArena · VisionReference only1,264.4
- MMMU ProVals AI · VisionReference only83.6%
- Arena WebDevLMArena · Writing and designReference only1,439.7
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- SimpleQA Verified · Epoch AI
- IFBench (AA run) · Artificial Analysis