GLM 5.2
Z.ai·Released Jun 16, 2026Open weights
Updated Oct 9, 6:21 AM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 50.1- ProgramBenchVals AI · max effortNear the floor0.5%Reading at most 65.7
- Code MigrationVals AI · max effort37.9%Reading 58.7
- LiveBench CodingLiveBench65.7%Reading 55.8
- Vibe Code BenchVals AI · max effort64.0%Reading 54.0
- SciCodeArtificial Analysis · max effort51.2%Reading 52.9
- APEX-SWEMercor · max effort37.4%Reading 48.5
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effortNear the floor1.0%Reading at most 34.5
Research and reasoning
6 results · 43.4- Humanity's Last Exam (AA run)Artificial Analysis · max effort41.1%Reading 58.2
- FrontierMath Tier 4Epoch AI · max effort29.3%Reading 49.8
- FrontierMath Tiers 1–3Epoch AI · max effort59.2%Reading 48.6
- LiveBench ReasoningLiveBench84.2%Reading 46.6
- Chess PuzzlesEpoch AI · max effort21.0%Reading 32.4
- Mystery Game PuzzlesEpoch AI · medium effort15.0%Reading 29.2
Professional work
9 results · 50.5- τ²-Bench Telecom (AA run)Artificial Analysis · max effortNear the ceiling99.1%Reading at least 60.0
- EMBVals AI61.5%Reading 58.6
- τ-Bench Banking (AA run)Artificial Analysis · max effort34.6%Reading 55.8
- LiveBench Data AnalysisLiveBench73.7%Reading 53.8
- Finance AgentVals AI49.7%Reading 53.1
- Legal Research BenchVals AI · max effort31.3%Reading 51.5
- APEX-AgentsMercor · max effort45.2%Reading 51.3
- Harvey Legal Agent BenchmarkVals AI · max effort7.1%Reading 43.7
- MedCodeVals AI40.8%Reading 36.6
Knowledge and accuracy
6 results · 36.6- IFBench (AA run)Artificial Analysis · max effort73.3%Reading 48.6
- LiveBench LanguageLiveBench76.2%Reading 47.9
- AA-LCRArtificial Analysis · max effort78.3%Reading 46.4
- LiveBench Instruction FollowingLiveBench62.3%Reading 32.5
- SimpleQA VerifiedEpoch AI · max effort34.2%Reading 31.6
- BullshitBenchBullshitBench · xhigh effortCapped18.2%Reading -14.3
Human preference
1 result · 47.4- Arena TextLMArena · max effort1,475.2Reading 48.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only69.5%
- SkillsBenchVals AI · CodingWatching45.1%
- SRE BenchVals AI · CodingWatching0.0%
- SWE-bench VerifiedEpoch AI · CodingReference only78.7%
- SWE-bench VerifiedVals AI · CodingReference only82.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only86.4%
- CorpFinVals AI · Professional workReference only66.1%
- EBR-benchEpoch AI · Professional workWatching9.5%
- LegalBenchVals AI · Professional workReference only84.1%
- MedScribeVals AI · Professional workWatching83.5%
- TaxEvalVals AI · Professional workReference only73.3%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only91.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only85.6%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.7%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only33.7
- LiveBench averageLiveBench · Composite indicesReference only73.2%
- Arena WebDevLMArena · Writing and designReference only1,603.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Tax Agent Bench · Vals AI