Mistral Medium 3.5
Mistral·Released Apr 29, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
4 results · 16.6- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor0.0%Reading at most 34.5
- SciCodeArtificial Analysis40.2%Reading 22.7
- Code MigrationVals AI · high effort5.1%Reading 12.3
- Vibe Code BenchVals AI · high effortNear the floor2.9%Reading at most 10.6
Research and reasoning
2 results · 19.9- ProofBenchVals AI · high effort9.0%Reading 30.1
- Humanity's Last Exam (AA run)Artificial Analysis13.8%Reading 10.2
Professional work
8 results · 12.9- τ²-Bench Telecom (AA run)Artificial Analysis94.2%Reading 56.8
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor0.4%Reading at most 24.8
- τ-Bench Banking (AA run)Artificial Analysis15.1%Reading 21.9
- Finance AgentVals AI · high effort32.1%Reading 17.3
- MedCodeVals AI · high effort33.8%Reading 11.5
- Legal Research BenchVals AI · high effort9.1%Reading 10.1
- Tax Agent BenchVals AI · high effortNear the floor4.5%Reading at most 0.1
- EMBVals AI · high effort11.2%Reading -11.1
Knowledge and accuracy
2 results · 32.6- IFBench (AA run)Artificial Analysis68.8%Reading 38.6
- AA-LCRArtificial Analysis69.3%Reading 33.6
Human preference
1 result · 28.7- Arena TextLMArena1,426.9Reading 33.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- SWE-bench VerifiedVals AI · CodingReference only66.4%
- CorpFinVals AI · Professional workReference only58.8%
- MedScribeVals AI · Professional workWatching67.7%
- TaxEvalVals AI · Professional workReference only68.0%
- GPQA DiamondVals AI · Knowledge and accuracyReference only34.8%
- MMLU-ProVals AI · Knowledge and accuracyReference only75.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only14.2
- Arena VisionLMArena · VisionReference only1,199.5
- Arena WebDevLMArena · Writing and designReference only1,263.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- APEX-Agents · Mercor
- LiveBench Data Analysis · LiveBench
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- BullshitBench · BullshitBench