GLM 5.1
Z.ai·Released Apr 7, 2026Open weights
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
6 results · 40.0- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- Code MigrationVals AI25.8%Reading 47.9
- APEX-SWEMercor34.5%Reading 44.0
- Vibe Code BenchVals AI31.5%Reading 37.3
- SciCodeArtificial Analysis · thinking effort44.8%Reading 35.5
- Terminal-Bench 4.0 (AA run)Artificial Analysis · thinking effortNear the floor2.0%Reading at most 34.5
Research and reasoning
3 results · 35.2- Humanity's Last Exam (AA run)Artificial Analysis · thinking effort30.1%Reading 42.4
- FrontierMath Tiers 1–3Epoch AI36.8%Reading 32.3
- Chess PuzzlesEpoch AI19.0%Reading 28.6
Professional work
7 results · 40.0- τ²-Bench Telecom (AA run)Artificial Analysis · thinking effortNear the ceiling97.7%Reading at least 60.0
- Legal Research BenchVals AI27.9%Reading 47.1
- APEX-AgentsMercor40.9%Reading 46.0
- Finance AgentVals AI44.8%Reading 43.5
- MedCodeVals AI41.6%Reading 39.5
- Harvey Legal Agent BenchmarkVals AINear the floor0.0%Reading at most 24.8
- τ-Bench Banking (AA run)Artificial Analysis · thinking effort13.6%Reading 18.3
Knowledge and accuracy
3 results · 41.8- IFBench (AA run)Artificial Analysis · thinking effort76.3%Reading 55.6
- AA-LCRArtificial Analysis · thinking effort73.7%Reading 39.4
- SimpleQA VerifiedEpoch AI34.0%Reading 31.3
Human preference
1 result · 43.1- Arena TextLMArena1,464.6Reading 44.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only81.4%
- SWE-bench VerifiedEpoch AI · CodingReference only74.2%
- SWE-bench VerifiedVals AI · CodingReference only76.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only93.3%
- APEX-AccountingMercor · Professional workWatching1.2%
- CorpFinVals AI · Professional workReference only64.5%
- LegalBenchVals AI · Professional workReference only84.4%
- MedScribeVals AI · Professional workWatching72.3%
- TaxEvalVals AI · Professional workReference only71.2%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only89.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only84.5%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only26.1
- Arena WebDevLMArena · Writing and designReference only1,508.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- LiveBench Data Analysis · LiveBench
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- BullshitBench · BullshitBench