GLM 5.3
Z.ai·Released Aug 14, 2026Open weights
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 66.8- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort41.9%Reading 76.0
- CyberBench PatchVals AI · max effort87.5%Reading 75.2
- SciCodeArtificial Analysis · max effort59.0%Reading 74.3
- Terminal-Bench 4.0Vals AI · max effort38.9%Reading 71.0
- LiveBench CodingLiveBench69.9%Reading 67.2
- ProgramBenchVals AI · max effortNear the floor1.5%Reading at most 65.7
- Vibe Code Bench 1–100Vals AI · max effort20.0%Reading 65.3
- Code MigrationVals AI · max effort44.2%Reading 63.7
- APEX-SWEMercor · max effort47.5%Reading 63.3
- Vibe Code BenchVals AI · max effort78.1%Reading 62.6
Research and reasoning
9 results · 52.4- Humanity's Last Exam (AA run)Artificial Analysis · max effort42.3%Reading 59.8
- MysteryMechanismVals AI · max effort23.0%Reading 56.2
- FrontierMath Tiers 1–3Epoch AI · max effort68.8%Reading 56.1
- ProofBenchVals AI · max effort49.0%Reading 55.6
- Mystery Game PuzzlesEpoch AI · max effort33.0%Reading 54.6
- LiveBench ReasoningLiveBench86.9%Reading 53.3
- Terminal-Bench ScienceVals AI · max effort5.7%Reading 50.8
- FrontierMath Tier 4Epoch AI · max effort29.3%Reading 49.8
- Chess PuzzlesEpoch AI · max effort21.0%Reading 32.3
Professional work
9 results · 60.7- Tax Agent BenchVals AI · max effort39.7%Reading 76.0
- τ-Bench Banking (AA run)Artificial Analysis · max effort50.3%Reading 75.9
- Legal Research BenchVals AI · max effort49.0%Reading 72.0
- Finance AgentVals AI · max effort55.8%Reading 65.1
- APEX-AgentsMercor · max effort56.6%Reading 64.9
- EMBVals AI · max effort56.3%Reading 52.7
- Harvey Legal Agent BenchmarkVals AI · max effort8.3%Reading 52.7
- MedCodeVals AI · max effort42.9%Reading 43.8
- LiveBench Data AnalysisLiveBench70.2%Reading 43.6
Knowledge and accuracy
5 results · 58.5- BullshitBenchBullshitBench · max effort50.9%Reading 84.3
- LiveBench Instruction FollowingLiveBench69.3%Reading 59.6
- LiveBench LanguageLiveBench79.9%Reading 57.5
- AA-LCRArtificial Analysis · max effort79.7%Reading 48.5
- SimpleQA VerifiedEpoch AI · max effort41.0%Reading 41.8
Human preference
1 result · 52.5- Arena TextLMArena · max effort1,478.5Reading 49.0
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only68.4%
- LiveCodeBenchVals AI · CodingReference only80.5%
- SkillsBenchVals AI · CodingWatching47.5%
- SRE BenchVals AI · CodingWatching9.2%
- SWE-bench VerifiedVals AI · CodingReference only95.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only91.1%
- APEX-AccountingMercor · Professional workWatching9.2%
- LegalBenchVals AI · Professional workReference only84.8%
- MedScribeVals AI · Professional workWatching88.8%
- TaxEvalVals AI · Professional workReference only72.4%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only88.1%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.8%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only44.8
- LiveBench averageLiveBench · Composite indicesReference only76.1%
- Vals IndexVals AI · Composite indicesReference only53.5%
- Arena WebDevLMArena · Writing and designReference only1,622.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis