Skip to content

Mistral Large 4

Mistral·Released Oct 6, 2026

Updated Oct 8, 9:18 PM ET

Overall score
51.1
Rank #34 overall
Evidence
17 results
52% of scored kinds covered
Context window
1M tokens
From OpenRouter
API price, USD per 1M tokens
$0.68 / $2.09
Input / output · cached input $0.07

The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.

By domain

Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.

Coding#27
56.9
5 results · 56% coverage
Research and reasoning#42
44.0
3 results · 38% coverage
Professional work#20
59.5
7 results · 78% coverage
Knowledge and accuracy#57
33.4
2 results · 33% coverage
Human preference
52.4
Not tested yet: shows the overall level

Every result behind the score

The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".

Coding

5 results · 56.9

Research and reasoning

3 results · 44.0

Professional work

7 results · 59.5

Knowledge and accuracy

2 results · 33.4

Shown, not scored

Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.

  • IOIVals AI · CodingReference only45.3%
  • BioMysteryBenchVals AI · Research and reasoningWatching67.0%
  • MedScribeVals AI · Professional workWatching80.4%
  • LiveBench averageLiveBench · Composite indicesReference only71.8%
  • Vals IndexVals AI · Composite indicesReference only48.0%

No result yet

Scored evaluations this model has not taken. A missing result neither adds nor subtracts.

  • Terminal-Bench 4.0 (AA run) · Artificial Analysis
  • ProgramBench · Vals AI
  • CyberBench Patch · Vals AI
  • APEX-SWE · Mercor
  • SciCode · Artificial Analysis
  • FrontierMath Tiers 1–3 · Epoch AI
  • FrontierMath Tier 4 · Epoch AI
  • Chess Puzzles · Epoch AI
  • Mystery Game Puzzles · Epoch AI
  • Terminal-Bench Science · Vals AI
  • Humanity's Last Exam (AA run) · Artificial Analysis
  • APEX-Agents · Mercor
  • τ²-Bench Telecom (AA run) · Artificial Analysis
  • τ-Bench Banking (AA run) · Artificial Analysis
  • SimpleQA Verified · Epoch AI
  • BullshitBench · BullshitBench
  • AA-LCR · Artificial Analysis
  • IFBench (AA run) · Artificial Analysis
  • Arena Text · LMArena
How readings and scores are computed