DeepSeek V4 Flash 0731
DeepSeek·Released Jul 31, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 56.1- CyberBench PatchVals AI · high effort87.5%Reading 75.2
- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- Vibe Code BenchVals AI · high effort74.7%Reading 60.2
- Code MigrationVals AI · high effort38.6%Reading 59.3
- Terminal-Bench 4.0Vals AI · high effort18.7%Reading 55.5
- SciCodeArtificial Analysis · max effort50.3%Reading 50.4
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort12.1%Reading 49.8
- LiveBench CodingLiveBench60.9%Reading 43.6
Research and reasoning
8 results · 52.7- ProofBenchVals AI · high effort56.0%Reading 58.8
- Mystery Game PuzzlesEpoch AI · max effort34.0%Reading 55.7
- Humanity's Last Exam (AA run)Artificial Analysis · max effort38.6%Reading 54.8
- LiveBench ReasoningLiveBench86.7%Reading 52.9
- Chess PuzzlesEpoch AI · max effort33.0%Reading 50.9
- Terminal-Bench ScienceVals AI · high effortNear the floor4.3%Reading at most 49.0
- FrontierMath Tiers 1–3Epoch AI · max effort57.5%Reading 47.4
- FrontierMath Tier 4Epoch AI · max effort24.4%Reading 46.4
Professional work
8 results · 55.2- LiveBench Data AnalysisLiveBench79.3%Reading 72.1
- τ-Bench Banking (AA run)Artificial Analysis · max effort39.4%Reading 62.1
- Tax Agent BenchVals AI · high effort28.1%Reading 60.3
- EMBVals AI · high effort57.0%Reading 53.5
- Finance AgentVals AI · high effort49.5%Reading 52.8
- Harvey Legal Agent BenchmarkVals AI · high effort8.3%Reading 52.7
- Legal Research BenchVals AI · high effort30.3%Reading 50.3
- MedCodeVals AI · high effort41.4%Reading 38.9
Knowledge and accuracy
5 results · 47.6- BullshitBenchBullshitBench · xhigh effort40.0%Reading 56.0
- LiveBench LanguageLiveBench79.2%Reading 55.6
- AA-LCRArtificial Analysis · max effort79.7%Reading 48.5
- LiveBench Instruction FollowingLiveBench65.5%Reading 44.6
- SimpleQA VerifiedEpoch AI · max effort33.6%Reading 30.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only32.7%
- LiveCodeBenchVals AI · CodingReference only87.3%
- SkillsBenchVals AI · CodingWatching50.7%
- SWE-bench VerifiedVals AI · CodingReference only88.8%
- BioMysteryBenchVals AI · Research and reasoningWatching64.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only94.4%
- CorpFinVals AI · Professional workReference only61.8%
- LegalBenchVals AI · Professional workReference only77.7%
- MedScribeVals AI · Professional workWatching80.4%
- TaxEvalVals AI · Professional workReference only70.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only91.0%
- GPQA DiamondVals AI · Knowledge and accuracyReference only89.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only34.3
- LiveBench averageLiveBench · Composite indicesReference only74.2%
- Vals IndexVals AI · Composite indicesReference only48.0%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- APEX-SWE · Mercor
- MysteryMechanism · Vals AI
- APEX-Agents · Mercor
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis
- Arena Text · LMArena