Claude Opus 4.7
Anthropic·Released Apr 16, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
5 results · 59.5- ProgramBenchVals AI · max effortNear the floor0.0%Reading at most 65.7
- Code MigrationVals AI · max effort43.9%Reading 63.4
- APEX-SWEMercor · max effort47.6%Reading 63.4
- Vibe Code BenchVals AI · max effort71.0%Reading 57.9
- LiveBench CodingLiveBench · xhigh effort66.4%Reading 57.6
Research and reasoning
6 results · 54.7- LiveBench ReasoningLiveBench · xhigh effort90.0%Reading 63.0
- Humanity's Last Exam (AA run)Artificial Analysis · max effort42.3%Reading 59.8
- FrontierMath Tiers 1–3Epoch AI · max effort70.2%Reading 57.3
- FrontierMath Tier 4Epoch AI · max effort31.7%Reading 51.3
- Mystery Game PuzzlesEpoch AI · max effort28.0%Reading 48.7
- Chess PuzzlesEpoch AI · xhigh effort30.0%Reading 46.7
Professional work
10 results · 58.7- MedCodeVals AI · max effort54.9%Reading 84.1
- LiveBench Data AnalysisLiveBench · xhigh effort78.3%Reading 68.3
- EMBVals AI · max effort63.7%Reading 61.2
- Legal Research BenchVals AI · max effort38.5%Reading 60.2
- Finance AgentVals AI · high effort51.5%Reading 56.6
- APEX-AgentsMercor · max effort49.2%Reading 56.0
- τ-Bench Banking (AA run)Artificial Analysis · max effort34.6%Reading 55.8
- Tax Agent BenchVals AI · max effort23.7%Reading 53.5
- τ²-Bench Telecom (AA run)Artificial Analysis · max effort88.6%Reading 42.5
- Harvey Legal Agent BenchmarkVals AI · max effort6.7%Reading 40.4
Knowledge and accuracy
6 results · 53.8- BullshitBenchBullshitBench · max effortCapped60.0%Reading 107.9
- SimpleQA VerifiedEpoch AI · xhigh effort51.7%Reading 57.0
- LiveBench LanguageLiveBench · xhigh effort77.9%Reading 52.2
- LiveBench Instruction FollowingLiveBench · xhigh effort66.7%Reading 49.3
- AA-LCRArtificial Analysis · max effort78.7%Reading 46.9
- IFBench (AA run)Artificial Analysis · max effort58.6%Reading 18.8
Human preference
1 result · 56.1- Arena TextLMArena · high effort1,501.3Reading 55.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only85.1%
- MirrorCodeEpoch AI · CodingWatching31.1%
- SWE-bench VerifiedEpoch AI · CodingReference only83.5%
- SWE-bench VerifiedVals AI · CodingReference only82.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only97.8%
- CorpFinVals AI · Professional workReference only66.1%
- EBR-benchEpoch AI · Professional workWatching19.0%
- LegalBenchVals AI · Professional workReference only85.3%
- MedScribeVals AI · Professional workWatching83.0%
- TaxEvalVals AI · Professional workReference only75.3%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.2%
- GPQA DiamondVals AI · Knowledge and accuracyReference only90.2%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only40.7
- LiveBench averageLiveBench · Composite indicesReference only76.5%
- Arena VisionLMArena · VisionReference only1,297.7
- MMMU ProVals AI · VisionReference only85.5%
- Arena WebDevLMArena · Writing and designReference only1,557.5
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- SciCode · Artificial Analysis
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI