Qwen3.8 Max
Alibaba·Released Aug 2, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 53.9- APEX-SWEMercor · xhigh effort52.3%Reading 70.1
- Terminal-Bench 4.0Vals AI34.3%Reading 68.0
- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench68.8%Reading 64.0
- SciCodeArtificial Analysis53.2%Reading 58.3
- Terminal-Bench 4.0 (AA run)Artificial Analysis18.7%Reading 57.9
- Vibe Code BenchVals AI64.7%Reading 54.3
- Vibe Code Bench 1–100Vals AI12.8%Reading 48.5
- Code MigrationVals AI24.0%Reading 46.0
- CyberBench PatchVals AICapped57.1%Reading 2.3
Research and reasoning
9 results · 58.4- Terminal-Bench ScienceVals AI12.9%Reading 62.4
- LiveBench ReasoningLiveBench89.8%Reading 62.1
- FrontierMath Tiers 1–3Epoch AI · xhigh effort74.7%Reading 61.4
- Humanity's Last Exam (AA run)Artificial Analysis43.0%Reading 60.7
- Mystery Game PuzzlesEpoch AI · xhigh effort38.0%Reading 60.0
- ProofBenchVals AI58.0%Reading 59.7
- FrontierMath Tier 4Epoch AI · xhigh effort46.3%Reading 59.6
- MysteryMechanismVals AI23.9%Reading 57.4
- Chess PuzzlesEpoch AI · xhigh effort29.0%Reading 45.3
Professional work
9 results · 62.9- τ-Bench Banking (AA run)Artificial Analysis51.3%Reading 77.1
- APEX-AgentsMercor · xhigh effort63.3%Reading 73.2
- Legal Research BenchVals AI47.6%Reading 70.5
- LiveBench Data AnalysisLiveBench78.4%Reading 68.8
- Tax Agent BenchVals AI32.1%Reading 66.0
- Harvey Legal Agent BenchmarkVals AI10.4%Reading 65.3
- EMBVals AI60.1%Reading 56.9
- Finance AgentVals AI50.6%Reading 54.9
- MedCodeVals AI40.7%Reading 36.3
Knowledge and accuracy
5 results · 50.9- LiveBench Instruction FollowingLiveBench74.1%Reading 80.1
- LiveBench LanguageLiveBench79.7%Reading 57.0
- SimpleQA VerifiedEpoch AI · xhigh effort45.8%Reading 48.7
- AA-LCRArtificial Analysis78.3%Reading 46.4
- BullshitBenchBullshitBench · max effortCapped21.8%Reading 0.2
Human preference
1 result · 52.2- Arena TextLMArena1,481.6Reading 49.9
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only68.9%
- LiveCodeBenchVals AI · CodingReference only87.9%
- SkillsBenchVals AI · CodingWatching42.0%
- SWE-bench VerifiedVals AI · CodingReference only85.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only99.4%
- CorpFinVals AI · Professional workReference only65.9%
- LegalBenchVals AI · Professional workReference only83.6%
- MedScribeVals AI · Professional workWatching84.9%
- TaxEvalVals AI · Professional workReference only75.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only92.7%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only88.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only40.2
- LiveBench averageLiveBench · Composite indicesReference only78.5%
- Vals IndexVals AI · Composite indicesReference only48.3%
- Arena VisionLMArena · VisionReference only1,301.2
- MMMU ProVals AI · VisionReference only88.0%
- Arena WebDevLMArena · Writing and designReference only1,672.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis