Qwen3.6 Plus
Alibaba·Released Mar 31, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
4 results · 34.2- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench59.8%Reading 40.9
- Vibe Code BenchVals AI25.6%Reading 33.7
- Code MigrationVals AI11.1%Reading 28.3
Research and reasoning
5 results · 31.6- Humanity's Last Exam (AA run)Artificial Analysis27.8%Reading 38.8
- LiveBench ReasoningLiveBench79.8%Reading 37.2
- FrontierMath Tiers 1–3Epoch AI38.2%Reading 33.3
- Chess PuzzlesEpoch AI17.0%Reading 24.5
- Mystery Game PuzzlesEpoch AI · none effort12.0%Reading 22.8
Professional work
9 results · 30.1- τ²-Bench Telecom (AA run)Artificial AnalysisNear the ceiling97.7%Reading at least 60.0
- LiveBench Data AnalysisLiveBench69.9%Reading 42.6
- Finance AgentVals AI40.8%Reading 35.7
- τ-Bench Banking (AA run)Artificial Analysis20.8%Reading 34.1
- EMBVals AI32.9%Reading 26.1
- Legal Research BenchVals AI14.9%Reading 25.4
- Harvey Legal Agent BenchmarkVals AINear the floor1.3%Reading at most 24.8
- MedCodeVals AI36.9%Reading 23.0
- Tax Agent BenchVals AI7.7%Reading 13.8
Knowledge and accuracy
5 results · 41.0- IFBench (AA run)Artificial Analysis75.2%Reading 52.9
- AA-LCRArtificial Analysis78.3%Reading 46.4
- SimpleQA VerifiedEpoch AI44.1%Reading 46.4
- LiveBench LanguageLiveBench75.0%Reading 44.8
- LiveBench Instruction FollowingLiveBench58.3%Reading 18.1
Human preference
1 result · 37.0- Arena TextLMArena1,443.4Reading 38.5
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only86.0%
- SWE-bench VerifiedEpoch AI · CodingReference only57.9%
- SWE-bench VerifiedVals AI · CodingReference only73.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only93.3%
- CorpFinVals AI · Professional workReference only61.9%
- LegalBenchVals AI · Professional workReference only84.2%
- MedScribeVals AI · Professional workWatching77.0%
- TaxEvalVals AI · Professional workReference only74.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only87.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.7%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only27.0
- LiveBench averageLiveBench · Composite indicesReference only68.9%
- MMMU ProVals AI · VisionReference only84.2%
- Arena WebDevLMArena · Writing and designReference only1,461.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- SciCode · Artificial Analysis
- FrontierMath Tier 4 · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor
- BullshitBench · BullshitBench