DeepSeek V4 Pro 0813
DeepSeek·Released Aug 13, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 59.4- CyberBench PatchVals AI · max effort85.7%Reading 68.5
- Vibe Code BenchVals AI · max effort82.3%Reading 65.8
- ProgramBenchVals AI · max effortNear the floor0.0%Reading at most 65.7
- Code MigrationVals AI · max effort41.5%Reading 61.6
- APEX-SWEMercor · max effort45.9%Reading 61.0
- Vibe Code Bench 1–100Vals AI · max effort17.5%Reading 60.2
- LiveBench CodingLiveBench66.1%Reading 56.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort14.1%Reading 52.6
- SciCodeArtificial Analysis · max effort51.0%Reading 52.3
- Terminal-Bench 4.0Vals AI · max effort14.1%Reading 50.4
Research and reasoning
8 results · 58.7- Chess PuzzlesEpoch AI · max effort47.0%Reading 68.6
- Mystery Game PuzzlesEpoch AI · max effort43.0%Reading 65.1
- LiveBench ReasoningLiveBench90.5%Reading 64.6
- Humanity's Last Exam (AA run)Artificial Analysis · max effort41.0%Reading 58.0
- ProofBenchVals AI · max effort50.0%Reading 56.1
- FrontierMath Tiers 1–3Epoch AI · max effort64.6%Reading 52.7
- Terminal-Bench ScienceVals AI · max effortNear the floor4.3%Reading at most 49.0
- FrontierMath Tier 4Epoch AI · max effort26.8%Reading 48.1
Professional work
9 results · 55.8- LiveBench Data AnalysisLiveBench79.2%Reading 71.8
- Legal Research BenchVals AI · max effort40.9%Reading 63.0
- τ-Bench Banking (AA run)Artificial Analysis · max effort39.6%Reading 62.4
- Tax Agent BenchVals AI · max effort26.8%Reading 58.3
- Finance AgentVals AI · max effort50.4%Reading 54.5
- APEX-AgentsMercor · max effort47.3%Reading 53.8
- EMBVals AI · max effort52.8%Reading 48.8
- Harvey Legal Agent BenchmarkVals AI · max effort7.5%Reading 46.9
- MedCodeVals AI · max effort42.5%Reading 42.5
Knowledge and accuracy
5 results · 52.7- LiveBench LanguageLiveBench82.1%Reading 64.0
- SimpleQA VerifiedEpoch AI · max effort52.9%Reading 58.7
- LiveBench Instruction FollowingLiveBench67.7%Reading 53.1
- AA-LCRArtificial Analysis · max effort80.3%Reading 49.7
- BullshitBenchBullshitBench · max effort32.7%Reading 35.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only51.6%
- LiveCodeBenchVals AI · CodingReference only87.5%
- SkillsBenchVals AI · CodingWatching53.8%
- SRE BenchVals AI · CodingWatching1.9%
- SWE-bench VerifiedVals AI · CodingReference only96.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.6%
- CorpFinVals AI · Professional workReference only65.4%
- LegalBenchVals AI · Professional workReference only82.4%
- MedScribeVals AI · Professional workWatching80.2%
- TaxEvalVals AI · Professional workReference only73.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only91.7%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.0%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only36.0
- LiveBench averageLiveBench · Composite indicesReference only77.4%
- Vals IndexVals AI · Composite indicesReference only47.6%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- MysteryMechanism · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis
- Arena Text · LMArena