Qwen3.6 27B
Alibaba·Released Apr 22, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
4 results · 28.4- Terminal-Bench 4.0 (AA run)Artificial Analysis · thinking effortNear the floor0.0%Reading at most 34.5
- LiveBench CodingLiveBench55.5%Reading 30.6
- SciCodeArtificial Analysis · thinking effort42.8%Reading 29.9
- Vibe Code BenchVals AI11.9%Reading 22.3
Research and reasoning
5 results · 26.7- Chess PuzzlesEpoch AI22.0%Reading 34.1
- FrontierMath Tiers 1–3Epoch AI35.1%Reading 30.9
- Humanity's Last Exam (AA run)Artificial Analysis · thinking effort23.1%Reading 30.7
- LiveBench ReasoningLiveBench75.1%Reading 28.8
- Mystery Game PuzzlesEpoch AI · none effort7.0%Reading 8.1
Professional work
3 results · 39.9- τ²-Bench Telecom (AA run)Artificial Analysis · thinking effort94.2%Reading 56.8
- LiveBench Data AnalysisLiveBench70.4%Reading 44.1
- τ-Bench Banking (AA run)Artificial Analysis · thinking effort16.7%Reading 25.7
Knowledge and accuracy
4 results · 25.5- AA-LCRArtificial Analysis · thinking effort77.3%Reading 44.8
- IFBench (AA run)Artificial Analysis · thinking effort67.6%Reading 36.1
- LiveBench LanguageLiveBench63.3%Reading 19.6
- LiveBench Instruction FollowingLiveBench53.2%Reading 0.1
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- SWE-bench VerifiedVals AI · CodingReference only70.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only91.1%
- CorpFinVals AI · Professional workReference only62.3%
- TaxEvalVals AI · Professional workReference only71.3%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only85.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only21.4
- LiveBench averageLiveBench · Composite indicesReference only64.0%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Code Migration · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- FrontierMath Tier 4 · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- APEX-Agents · Mercor
- Harvey Legal Agent Benchmark · Vals AI
- MedCode · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- SimpleQA Verified · Epoch AI
- BullshitBench · BullshitBench
- Arena Text · LMArena