MiMo-V2.5
Xiaomi·Released Apr 22, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
4 results · 35.0- Vibe Code BenchVals AI42.2%Reading 43.0
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor0.0%Reading at most 34.5
- Code MigrationVals AI14.2%Reading 33.8
- SciCodeArtificial Analysis43.9%Reading 33.0
Research and reasoning
2 results · 35.7- Humanity's Last Exam (AA run)Artificial Analysis27.2%Reading 37.8
- ProofBenchVals AI16.0%Reading 37.4
Professional work
8 results · 22.8- EMBVals AI55.1%Reading 51.4
- τ²-Bench Telecom (AA run)Artificial Analysis90.6%Reading 46.8
- Finance AgentVals AI36.7%Reading 27.3
- Harvey Legal Agent BenchmarkVals AINear the floor1.7%Reading at most 24.8
- Tax Agent BenchVals AI7.8%Reading 14.2
- Legal Research BenchVals AI9.1%Reading 10.1
- MedCodeVals AI31.9%Reading 4.5
- τ-Bench Banking (AA run)Artificial Analysis8.7%Reading 2.5
Knowledge and accuracy
3 results · 23.2- AA-LCRArtificial Analysis73.0%Reading 38.5
- IFBench (AA run)Artificial Analysis67.1%Reading 35.3
- BullshitBenchBullshitBench · xhigh effortCapped20.0%Reading -6.8
Human preference
1 result · 33.1- Arena TextLMArena1,433.6Reading 35.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only81.5%
- SWE-bench VerifiedVals AI · CodingReference only71.0%
- CorpFinVals AI · Professional workReference only59.9%
- LegalBenchVals AI · Professional workReference only78.9%
- MedScribeVals AI · Professional workWatching72.2%
- TaxEvalVals AI · Professional workReference only71.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only81.6%
- MMLU-ProVals AI · Knowledge and accuracyReference only82.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.2
- Arena VisionLMArena · VisionReference only1,233.2
- MMMU ProVals AI · VisionReference only80.0%
- Arena WebDevLMArena · Writing and designReference only1,438.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- APEX-Agents · Mercor
- LiveBench Data Analysis · LiveBench
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench