MiniMax M3
MiniMax·Released Jun 1, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 39.1- APEX-SWEMercor · high effort35.7%Reading 45.9
- Vibe Code BenchVals AI47.6%Reading 45.7
- SciCodeArtificial Analysis47.1%Reading 41.7
- Code MigrationVals AI19.9%Reading 41.5
- CyberBench PatchVals AI75.0%Reading 38.0
- Vibe Code Bench 1–100Vals AI9.2%Reading 36.6
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor2.0%Reading at most 34.5
- Terminal-Bench 4.0Vals AINear the floor1.0%Reading at most 33.0
- LiveBench CodingLiveBench54.4%Reading 28.0
Research and reasoning
5 results · 31.1- Humanity's Last Exam (AA run)Artificial Analysis39.0%Reading 55.3
- ProofBenchVals AI18.0%Reading 39.0
- LiveBench ReasoningLiveBench75.7%Reading 29.9
- Chess PuzzlesEpoch AI14.0%Reading 17.6
- Mystery Game PuzzlesEpoch AI7.0%Reading 8.1
Professional work
10 results · 45.1- LiveBench Data AnalysisLiveBench76.2%Reading 61.3
- MedCodeVals AI46.3%Reading 55.4
- Finance AgentVals AI48.3%Reading 50.3
- Legal Research BenchVals AI29.8%Reading 49.6
- Tax Agent BenchVals AI20.6%Reading 48.0
- EMBVals AI47.8%Reading 43.3
- τ²-Bench Telecom (AA run)Artificial Analysis88.9%Reading 43.1
- APEX-AgentsMercor · high effort37.7%Reading 42.0
- Harvey Legal Agent BenchmarkVals AINear the floor4.2%Reading at most 24.8
- τ-Bench Banking (AA run)Artificial Analysis15.3%Reading 22.4
Knowledge and accuracy
5 results · 52.5- BullshitBenchBullshitBench · xhigh effort47.3%Reading 75.0
- IFBench (AA run)Artificial Analysis82.9%Reading 74.0
- AA-LCRArtificial Analysis83.0%Reading 54.5
- LiveBench LanguageLiveBench76.8%Reading 49.4
- LiveBench Instruction FollowingLiveBench57.5%Reading 15.1
Human preference
1 result · 39.0- Arena TextLMArena1,440.1Reading 37.5
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only82.2%
- SkillsBenchVals AI · CodingWatching51.5%
- SWE-bench VerifiedVals AI · CodingReference only75.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only71.1%
- APEX-AccountingMercor · Professional workWatching2.2%
- CorpFinVals AI · Professional workReference only68.1%
- LegalBenchVals AI · Professional workReference only85.4%
- MedScribeVals AI · Professional workWatching87.3%
- TaxEvalVals AI · Professional workReference only72.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only84.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only29.2
- LiveBench averageLiveBench · Composite indicesReference only67.3%
- Vals IndexVals AI · Composite indicesReference only36.5%
- Arena VisionLMArena · VisionReference only1,236.6
- MMMU ProVals AI · VisionReference only81.2%
- Arena WebDevLMArena · Writing and designReference only1,482.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- ProgramBench · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- SimpleQA Verified · Epoch AI