DeepSeek V4 Pro 0423
DeepSeek·Released Apr 24, 2026Open weights
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
6 results · 43.0- ProgramBenchVals AI · max effortNear the floor0.0%Reading at most 65.7
- Code MigrationVals AI · max effort26.2%Reading 48.3
- Vibe Code BenchVals AI · max effort49.9%Reading 46.9
- Terminal-Bench 4.0Vals AI · max effort11.1%Reading 46.2
- APEX-SWEMercor · max effort33.8%Reading 42.9
- LiveBench CodingLiveBench56.3%Reading 32.5
Research and reasoning
6 results · 34.2- LiveBench ReasoningLiveBench86.7%Reading 52.9
- FrontierMath Tiers 1–3Epoch AI · max effort45.3%Reading 38.5
- ProofBenchVals AI · max effort16.0%Reading 37.4
- Mystery Game PuzzlesEpoch AI · none effort17.0%Reading 32.9
- FrontierMath Tier 4Epoch AI · max effortNear the floor2.4%Reading at most 22.2
- Chess PuzzlesEpoch AI · high effort13.0%Reading 15.0
Professional work
7 results · 42.9- LiveBench Data AnalysisLiveBench74.5%Reading 56.2
- Tax Agent BenchVals AI · max effort25.2%Reading 55.8
- EMBVals AI · max effort51.6%Reading 47.5
- Finance AgentVals AI · max effort44.1%Reading 42.1
- Legal Research BenchVals AI · max effort23.1%Reading 40.1
- MedCodeVals AI · max effort40.5%Reading 35.5
- Harvey Legal Agent BenchmarkVals AI · max effortNear the floor3.8%Reading at most 24.8
Knowledge and accuracy
4 results · 35.1- LiveBench LanguageLiveBench78.1%Reading 52.7
- SimpleQA VerifiedEpoch AI · max effort47.0%Reading 50.4
- LiveBench Instruction FollowingLiveBench62.4%Reading 32.7
- BullshitBenchBullshitBench · xhigh effortCapped10.9%Reading -52.5
Human preference
1 result · 43.1- Arena TextLMArena · high effort1,464.3Reading 44.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only87.5%
- SkillsBenchVals AI · CodingWatching51.3%
- SWE-bench VerifiedEpoch AI · CodingReference only77.6%
- SWE-bench VerifiedVals AI · CodingReference only77.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only95.6%
- CorpFinVals AI · Professional workReference only61.4%
- LegalBenchVals AI · Professional workReference only80.3%
- MedScribeVals AI · Professional workWatching75.1%
- TaxEvalVals AI · Professional workReference only72.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only89.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.2%
- LiveBench averageLiveBench · Composite indicesReference only71.6%
- Vals IndexVals AI · Composite indicesReference only38.6%
- Arena WebDevLMArena · Writing and designReference only1,581.9
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- SciCode · Artificial Analysis
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- Humanity's Last Exam (AA run) · Artificial Analysis
- APEX-Agents · Mercor
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- AA-LCR · Artificial Analysis
- IFBench (AA run) · Artificial Analysis