Qwen3.7 Plus
Alibaba·Released Jun 2, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
4 results · 37.5- Vibe Code BenchVals AI46.4%Reading 45.1
- SciCodeArtificial Analysis46.1%Reading 39.0
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor1.0%Reading at most 34.5
- Code MigrationVals AI12.9%Reading 31.5
Research and reasoning
4 results · 37.8- Humanity's Last Exam (AA run)Artificial Analysis35.6%Reading 50.6
- Chess PuzzlesEpoch AI24.0%Reading 37.5
- Mystery Game PuzzlesEpoch AI17.0%Reading 32.9
- FrontierMath Tiers 1–3Epoch AI · none effort34.4%Reading 30.3
Professional work
7 results · 32.5- τ²-Bench Telecom (AA run)Artificial Analysis93.0%Reading 53.0
- EMBVals AI49.3%Reading 45.0
- Finance AgentVals AI38.2%Reading 30.4
- Legal Research BenchVals AI16.3%Reading 28.4
- τ-Bench Banking (AA run)Artificial Analysis17.5%Reading 27.5
- Harvey Legal Agent BenchmarkVals AINear the floor0.0%Reading at most 24.8
- Tax Agent BenchVals AI10.5%Reading 24.1
Knowledge and accuracy
2 results · 46.8- IFBench (AA run)Artificial Analysis78.0%Reading 59.9
- AA-LCRArtificial Analysis73.0%Reading 38.5
Human preference
1 result · 40.6- Arena TextLMArena1,455.8Reading 42.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- SkillsBenchVals AI · CodingWatching54.3%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only93.3%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only87.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.2
- Arena VisionLMArena · VisionReference only1,265.5
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor
- MedCode · Vals AI
- LiveBench Data Analysis · LiveBench
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- BullshitBench · BullshitBench