Claude Opus 4.6
Anthropic·Released Feb 5, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
3 results · 51.7- APEX-SWEMercor · max effort41.8%Reading 55.0
- Vibe Code BenchVals AI · max effort57.6%Reading 50.6
- LiveBench CodingLiveBench · high effort63.6%Reading 50.4
Research and reasoning
6 results · 46.3- LiveBench ReasoningLiveBench · high effort89.0%Reading 59.6
- Humanity's Last Exam (AA run)Artificial Analysis · max effort39.9%Reading 56.5
- FrontierMath Tiers 1–3Epoch AI · max effort66.0%Reading 53.8
- FrontierMath Tier 4Epoch AI · max effort26.8%Reading 48.1
- Mystery Game PuzzlesEpoch AI · max effort25.0%Reading 44.9
- Chess PuzzlesEpoch AI · max effort14.0%Reading 17.6
Professional work
4 results · 51.7- MedCodeVals AI · max effort48.2%Reading 61.9
- APEX-AgentsMercor · max effort46.3%Reading 52.6
- τ²-Bench Telecom (AA run)Artificial Analysis · max effort92.1%Reading 50.5
- LiveBench Data AnalysisLiveBench · high effort69.9%Reading 42.6
Knowledge and accuracy
6 results · 50.3- BullshitBenchBullshitBench · high effortCapped89.1%Reading 216.4
- LiveBench LanguageLiveBench · high effort83.3%Reading 67.8
- SimpleQA VerifiedEpoch AI · max effort47.0%Reading 50.4
- AA-LCRArtificial Analysis · max effort78.0%Reading 45.8
- LiveBench Instruction FollowingLiveBench · high effort63.3%Reading 36.2
- IFBench (AA run)Artificial Analysis · max effortCapped53.1%Reading 8.8
Human preference
1 result · 54.5- Arena TextLMArena · high effort1,504.7Reading 56.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only84.7%
- SWE-bench VerifiedEpoch AI · CodingReference only78.7%
- SWE-bench VerifiedVals AI · CodingReference only78.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only91.1%
- APEX-AccountingMercor · Professional workWatching4.4%
- CorpFinVals AI · Professional workReference only67.0%
- EBR-benchEpoch AI · Professional workWatching12.7%
- LegalBenchVals AI · Professional workReference only85.3%
- MedScribeVals AI · Professional workWatching86.7%
- TaxEvalVals AI · Professional workReference only76.0%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only89.6%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.1%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only31.9
- LiveBench averageLiveBench · Composite indicesReference only74.5%
- Arena VisionLMArena · VisionReference only1,299.4
- MMMU ProVals AI · VisionReference only83.9%
- Arena WebDevLMArena · Writing and designReference only1,546.4
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Code Migration · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- SciCode · Artificial Analysis
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- Harvey Legal Agent Benchmark · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- τ-Bench Banking (AA run) · Artificial Analysis