Claude Sonnet 4.6
Anthropic·Released Feb 17, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 46.8- ProgramBenchVals AI · max effortNear the floor0.5%Reading at most 65.7
- Code MigrationVals AI · max effort39.9%Reading 60.3
- SciCodeArtificial Analysis · max effort50.1%Reading 49.9
- Vibe Code BenchVals AI · max effort51.5%Reading 47.6
- APEX-SWEMercor · high effort34.5%Reading 44.0
- LiveBench CodingLiveBench · medium effort60.9%Reading 43.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effortNear the floor3.0%Reading at most 34.5
Research and reasoning
4 results · 34.6- LiveBench ReasoningLiveBench · medium effort85.9%Reading 50.7
- Humanity's Last Exam (AA run)Artificial Analysis · max effort33.6%Reading 47.7
- Mystery Game PuzzlesEpoch AI14.0%Reading 27.2
- Chess PuzzlesEpoch AI · high effortCappedNear the floor5.0%Reading at most -16.4
Professional work
9 results · 51.0- LiveBench Data AnalysisLiveBench · medium effort77.9%Reading 67.2
- Legal Research BenchVals AI · max effort38.5%Reading 60.2
- EMBVals AI · max effort60.2%Reading 57.0
- Tax Agent BenchVals AI · max effort25.8%Reading 56.9
- Finance AgentVals AI · max effort51.0%Reading 55.7
- τ-Bench Banking (AA run)Artificial Analysis · max effort34.4%Reading 55.5
- APEX-AgentsMercor · high effort43.0%Reading 48.6
- Harvey Legal Agent BenchmarkVals AI · max effortNear the floor5.0%Reading at most 24.8
- τ²-Bench Telecom (AA run)Artificial Analysis · max effort75.7%Reading 24.7
Knowledge and accuracy
6 results · 44.9- BullshitBenchBullshitBench · high effortCapped92.7%Reading 245.0
- AA-LCRArtificial Analysis · max effort80.0%Reading 49.1
- LiveBench LanguageLiveBench · medium effort76.1%Reading 47.5
- LiveBench Instruction FollowingLiveBench · medium effort63.2%Reading 35.9
- SimpleQA VerifiedEpoch AI · high effort35.5%Reading 33.6
- IFBench (AA run)Artificial Analysis · max effort56.6%Reading 15.1
Human preference
1 result · 46.6- Arena TextLMArena1,472.2Reading 47.1
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only82.1%
- SkillsBenchVals AI · CodingWatching49.1%
- SWE-bench VerifiedEpoch AI · CodingReference only75.2%
- SWE-bench VerifiedVals AI · CodingReference only77.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only75.6%
- CorpFinVals AI · Professional workReference only65.3%
- LegalBenchVals AI · Professional workReference only82.1%
- TaxEvalVals AI · Professional workReference only77.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only83.3%
- GPQA DiamondVals AI · Knowledge and accuracyReference only85.6%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only30.1
- LiveBench averageLiveBench · Composite indicesReference only73.0%
- Arena VisionLMArena · VisionReference only1,275.2
- MMMU ProVals AI · VisionReference only83.6%
- Arena WebDevLMArena · Writing and designReference only1,522.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- MedCode · Vals AI