DeepSeek V4.1 Flash
DeepSeek·Released Sep 9, 2026Open weights
Updated Oct 8, 21:18 ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 66.1- LiveBench CodingLiveBench · max effort78.7%Reading 94.3
- APEX-SWEMercor · max effort51.9%Reading 69.5
- CyberBench PatchVals AI · high effort85.7%Reading 68.5
- Vibe Code BenchVals AI · high effort84.7%Reading 68.0
- ProgramBenchVals AI · max effortNear the floor0.5%Reading at most 65.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort26.8%Reading 65.2
- Code MigrationVals AI · high effort45.6%Reading 64.8
- Vibe Code Bench 1–100Vals AI · high effort16.4%Reading 57.6
- Terminal-Bench 4.0Vals AI · high effort19.7%Reading 56.5
- SciCodeArtificial Analysis · max effort51.9%Reading 54.8
Research and reasoning
5 results · 56.1- LiveBench ReasoningLiveBench · max effort90.0%Reading 62.9
- ProofBenchVals AI · high effort54.0%Reading 57.9
- Humanity's Last Exam (AA run)Artificial Analysis · max effort39.2%Reading 55.6
- MysteryMechanismVals AI · high effort21.2%Reading 53.6
- Terminal-Bench ScienceVals AI · high effortNear the floor0.0%Reading at most 49.0
Professional work
8 results · 54.1- LiveBench Data AnalysisLiveBench · max effort79.3%Reading 71.8
- Legal Research BenchVals AI · max effort41.3%Reading 63.5
- Finance AgentVals AI · high effort53.5%Reading 60.5
- Tax Agent BenchVals AI · max effort26.6%Reading 58.1
- EMBVals AI · high effort57.2%Reading 53.7
- APEX-AgentsMercor · max effort39.5%Reading 44.3
- Harvey Legal Agent BenchmarkVals AI · max effort6.7%Reading 40.4
- MedCodeVals AI · high effort41.2%Reading 38.0
Knowledge and accuracy
4 results · 60.2- LiveBench Instruction FollowingLiveBench · max effort70.0%Reading 62.6
- LiveBench LanguageLiveBench · max effort81.2%Reading 61.3
- BullshitBenchBullshitBench · max effort41.8%Reading 60.8
- AA-LCRArtificial Analysis · max effort84.0%Reading 56.5
Human preference
1 result · 51.5- Arena TextLMArena · max effort1,474.4Reading 47.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only40.3%
- SkillsBenchVals AI · CodingWatching69.8%
- SRE BenchVals AI · CodingWatching0.8%
- BioMysteryBenchVals AI · Research and reasoningWatching67.8%
- APEX-AccountingMercor · Professional workWatching7.8%
- LegalBenchVals AI · Professional workReference only83.3%
- MedScribeVals AI · Professional workWatching85.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only39.5
- LiveBench averageLiveBench · Composite indicesReference only81.1%
- Vals IndexVals AI · Composite indicesReference only51.3%
- Arena WebDevLMArena · Writing and designReference only1,618.6
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- SimpleQA Verified · Epoch AI
- IFBench (AA run) · Artificial Analysis