Muse Spark 1.3
Meta·Released Sep 2, 2026
Updated Oct 8, 21:18 ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 61.6- SciCodeArtificial Analysis · xhigh effort59.7%Reading 76.3
- LiveBench CodingLiveBench · xhigh effort72.6%Reading 74.8
- Vibe Code BenchVals AI · xhigh effort82.9%Reading 66.3
- Vibe Code Bench 1–100Vals AI · max effort20.5%Reading 66.3
- ProgramBenchVals AI · xhigh effortNear the floor0.5%Reading at most 65.7
- CyberBench PatchVals AI · xhigh effort82.1%Reading 56.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effort16.7%Reading 55.7
- Code MigrationVals AI · xhigh effort27.6%Reading 49.6
- APEX-SWEMercor · xhigh effort36.5%Reading 47.1
- Terminal-Bench 4.0Vals AI · xhigh effort10.6%Reading 45.4
Research and reasoning
9 results · 57.9- LiveBench ReasoningLiveBench · xhigh effort92.8%Reading 74.2
- MysteryMechanismVals AI · max effort36.0%Reading 71.7
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort47.5%Reading 66.6
- FrontierMath Tiers 1–3Epoch AI · xhigh effort74.4%Reading 61.0
- ProofBenchVals AI · xhigh effort55.0%Reading 58.3
- FrontierMath Tier 4Epoch AI · xhigh effort41.5%Reading 57.0
- Chess PuzzlesEpoch AI · xhigh effort35.0%Reading 53.6
- Terminal-Bench ScienceVals AI · xhigh effortNear the floor4.3%Reading at most 49.0
- Mystery Game PuzzlesEpoch AI · xhigh effort14.1%Reading 27.5
Professional work
8 results · 72.5- Harvey Legal Agent BenchmarkVals AI · xhigh effortCapped22.9%Reading 113.4
- LiveBench Data AnalysisLiveBench · xhigh effort79.6%Reading 72.9
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort47.2%Reading 72.0
- Finance AgentVals AI · xhigh effort58.9%Reading 71.2
- Tax Agent BenchVals AI · xhigh effort36.0%Reading 71.2
- APEX-AgentsMercor · xhigh effort58.6%Reading 67.3
- Legal Research BenchVals AI · xhigh effort40.9%Reading 63.0
- EMBVals AI · xhigh effort62.7%Reading 60.0
Knowledge and accuracy
3 results · 72.0- LiveBench Instruction FollowingLiveBench · xhigh effort78.0%Reading 98.8
- LiveBench LanguageLiveBench · xhigh effort82.8%Reading 66.3
- AA-LCRArtificial Analysis · xhigh effort83.0%Reading 54.5
Human preference
1 result · 57.3- Arena TextLMArena · max effort1,494.3Reading 53.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only43.9%
- SRE BenchVals AI · CodingWatching3.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only99.2%
- APEX-AccountingMercor · Professional workWatching10.8%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only45.1
- LiveBench averageLiveBench · Composite indicesReference only81.6%
- Vals IndexVals AI · Composite indicesReference only53.2%
- Arena VisionLMArena · VisionReference only1,290.2
- Arena WebDevLMArena · Writing and designReference only1,626.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- MedCode · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- SimpleQA Verified · Epoch AI
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis