Claude Sonnet 5.5
Anthropic·Released Sep 28, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 74.3- APEX-SWEMercor · max effort66.4%Reading 91.0
- Terminal-Bench 4.0Vals AI · max effort64.1%Reading 86.7
- Code MigrationVals AI · max effort69.8%Reading 84.2
- Vibe Code BenchVals AI · max effort92.4%Reading 77.6
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort43.9%Reading 77.3
- CyberBench PatchVals AI · max effort87.5%Reading 75.2
- ProgramBenchVals AI · max effort6.5%Reading 74.1
- SciCodeArtificial Analysis · high effort53.7%Reading 59.7
- LiveBench CodingLiveBench · xhigh effort64.1%Reading 51.7
Research and reasoning
8 results · 79.5- ProofBenchVals AI · max effortNear the ceiling100.0%Reading at least 89.2
- Mystery Game PuzzlesEpoch AI · max effort65.0%Reading 87.4
- Terminal-Bench ScienceVals AI · max effort45.7%Reading 85.2
- MysteryMechanismVals AI · max effort49.1%Reading 84.8
- FrontierMath Tier 4Epoch AI · max effort80.5%Reading 80.5
- FrontierMath Tiers 1–3Epoch AI · max effort88.8%Reading 79.0
- LiveBench ReasoningLiveBench · xhigh effort91.8%Reading 69.6
- Humanity's Last Exam (AA run)Artificial Analysis · high effort45.8%Reading 64.4
Professional work
8 results · 71.1- APEX-AgentsMercor · max effort75.5%Reading 90.4
- Tax Agent BenchVals AI · max effort42.0%Reading 78.8
- MedCodeVals AI · max effort52.9%Reading 77.6
- EMBVals AI · max effort75.7%Reading 76.9
- Legal Research BenchVals AI · max effort48.1%Reading 71.0
- Finance AgentVals AI · max effort58.1%Reading 69.6
- LiveBench Data AnalysisLiveBench · xhigh effort78.6%Reading 69.6
- Harvey Legal Agent BenchmarkVals AI · max effortCappedNear the floor2.9%Reading at most 24.8
Knowledge and accuracy
4 results · 58.7- LiveBench LanguageLiveBench · xhigh effort83.4%Reading 68.3
- LiveBench Instruction FollowingLiveBench · xhigh effort70.5%Reading 64.6
- SimpleQA VerifiedEpoch AI · max effort46.5%Reading 49.7
- AA-LCRArtificial Analysis · high effort78.0%Reading 45.8
Human preference
1 result · 55.0- Arena TextLMArena · xhigh effort1,471.0Reading 46.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only83.1%
- SRE BenchVals AI · CodingWatching30.2%
- BioMysteryBenchVals AI · Research and reasoningWatching81.1%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- MedScribeVals AI · Professional workWatching91.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only95.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only46.8
- LiveBench averageLiveBench · Composite indicesReference only77.8%
- Vals IndexVals AI · Composite indicesReference only67.0%
- Arena VisionLMArena · VisionReference only1,268.3
- Arena WebDevLMArena · Writing and designReference only1,716.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- Chess Puzzles · Epoch AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis