Skip to content

Muse Spark

Meta·Released Apr 8, 2026

Updated Oct 8, 21:18 ET

Overall score
47.8
Not ranked yet: fewer than 3 kinds of coding evaluation
Evidence
7 results
21% of scored kinds covered
Context window
—
From OpenRouter
API price, USD per 1M tokens
— / —
Input / output · cached input —

The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.

By domain

Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.

Coding
36.9
1 result · 11% coverage
Research and reasoning
55.6
1 result · 13% coverage
Professional work
58.8
2 results · 22% coverage
Knowledge and accuracy
50.5
2 results · 33% coverage
Human preference
52.0
1 result · 100% coverage

Every result behind the score

The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".

Coding

1 result · 36.9

Research and reasoning

1 result · 55.6

Professional work

2 results · 58.8

Knowledge and accuracy

2 results · 50.5

Human preference

1 result · 52.0

Shown, not scored

Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.

  • SWE-bench VerifiedVals AI · CodingReference only74.4%
  • OTIS Mock AIMEEpoch AI · Research and reasoningReference only88.9%
  • CorpFinVals AI · Professional workReference only65.1%
  • LegalBenchVals AI · Professional workReference only84.2%
  • MedScribeVals AI · Professional workWatching85.9%
  • TaxEvalVals AI · Professional workReference only77.7%
  • GPQA DiamondEpoch AI · Knowledge and accuracyReference only89.8%
  • GPQA DiamondVals AI · Knowledge and accuracyReference only89.6%
  • MMLU-ProVals AI · Knowledge and accuracyReference only87.3%
  • AA Intelligence IndexArtificial Analysis · Composite indicesReference only31.3
  • Arena VisionLMArena · VisionReference only1,293.6
  • MMMU ProVals AI · VisionReference only87.4%

No result yet

Scored evaluations this model has not taken. A missing result neither adds nor subtracts.

  • Terminal-Bench 4.0 · Vals AI
  • Terminal-Bench 4.0 (AA run) · Artificial Analysis
  • LiveBench Coding · LiveBench
  • Code Migration · Vals AI
  • ProgramBench · Vals AI
  • Vibe Code Bench 1–100 · Vals AI
  • CyberBench Patch · Vals AI
  • APEX-SWE · Mercor
  • SciCode · Artificial Analysis
  • FrontierMath Tiers 1–3 · Epoch AI
  • FrontierMath Tier 4 · Epoch AI
  • LiveBench Reasoning · LiveBench
  • Chess Puzzles · Epoch AI
  • Mystery Game Puzzles · Epoch AI
  • MysteryMechanism · Vals AI
  • Terminal-Bench Science · Vals AI
  • ProofBench · Vals AI
  • Finance Agent · Vals AI
  • APEX-Agents · Mercor
  • Harvey Legal Agent Benchmark · Vals AI
  • Legal Research Bench · Vals AI
  • Tax Agent Bench · Vals AI
  • EMB · Vals AI
  • LiveBench Data Analysis · LiveBench
  • τ-Bench Banking (AA run) · Artificial Analysis
  • SimpleQA Verified · Epoch AI
  • LiveBench Language · LiveBench
  • LiveBench Instruction Following · LiveBench
  • BullshitBench · BullshitBench
How readings and scores are computed