Leaderboard
AI model leaderboard
Current AI models ranked on research and reasoning: research math, puzzles, proofs and scientific workflows. Same scale as the overall score.
36 scored evaluations · 7 evaluatorsUpdated Oct 8, 9:18 PM ET
Research board 30 models
Sort| 01 | GPT-6 Astra OpenAI | Sep 3, 2026 | 34 results Coverage 97% | $1 | $10 | $50 | 91.5 |
| 02 | GPT-6.1 Sol OpenAI | Sep 29, 2026 | 31 results Coverage 88% | $0.10 | $2 | $10 | 86.6 |
| 03 | Claude Opus 5.5 Anthropic | Sep 22, 2026 | 32 results Coverage 91% | $0.20 | $4 | $20 | 85.6 |
| 04 | Claude Fable 5.1 Anthropic | Sep 1, 2026 | 34 results Coverage 97% | $0.25 | $10 | $50 | 80.8 |
| 05 | Claude Sonnet 5.5 Anthropic | Sep 28, 2026 | 30 results Coverage 85% | $0.10 | $2 | $10 | 79.5 |
| 06 | Claude Opus 5 Anthropic | Jul 24, 2026 | 34 results Coverage 97% | $0.50 | $5 | $25 | 75.6 |
| 07 | Claude Fable 5 Anthropic | Jun 9, 2026 | 32 results Coverage 88% | $1 | $10 | $50 | 74.9 |
| 08 | GPT-5.6 Sol OpenAI | Jul 9, 2026 | 35 results Coverage 97% | $0.20 | $2 | $10 | 74.0 |
| 09 | GPT-6 Sol OpenAI | Sep 22, 2026 | 31 results Coverage 88% | $0.20 | $2 | $10 | 72.9 |
| 10 | GPT-5.6 Terra OpenAI | Jul 9, 2026 | 35 results Coverage 97% | $0.20 | $2 | $12 | 65.5 |
| 11 | Gemini 3.8 Flash Google | Sep 2, 2026 | 33 results Coverage 94% | $0.075 | $0.75 | $3.75 | 65.0 |
| 12 | GPT-5.5 OpenAI | Apr 23, 2026 | 30 results Coverage 85% | $0.50 | $5 | $30 | 64.6 |
| 13 | Claude Haiku 5.5 Anthropic | Oct 7, 2026 | 26 results Coverage 76% | $0.01 | $0.10 | $0.50 | 62.2 |
| 14 | Gemini 3.7 Flash Google | Aug 13, 2026 | 32 results Coverage 91% | $0.075 | $0.75 | $3.75 | 61.4 |
| 15 | GPT-5.4 OpenAI | Mar 5, 2026 | 24 results Coverage 67% | $0.25 | $2.50 | $15 | 60.9 |
| 16 | Muse Spark 1.2 Meta | Aug 5, 2026 | 27 results Coverage 79% | $0.15 | $1.25 | $4.25 | 60.7 |
| 17 | Grok 4.6 xAI | Aug 12, 2026 | 34 results Coverage 97% | $0.50 | $2 | $6 | 60.2 |
| 18 | Claude Opus 4.8 Anthropic | May 28, 2026 | 32 results Coverage 88% | $0.50 | $5 | $25 | 59.5 |
| 19 | Muse Spark 1.1 Meta | Jul 9, 2026 | 21 results Coverage 64% | $0.15 | $1.25 | $4.25 | 59.3 |
| 20 | DeepSeek V4 Pro 0813Open weights DeepSeek | Aug 13, 2026 | 32 results Coverage 91% | $0.044 | $1.32 | $3.96 | 58.7 |
| 21 | Qwen3.8 Max Alibaba | Aug 2, 2026 | 34 results Coverage 97% | — | — | — | 58.4 |
| 22 | Muse Spark 1.3 Meta | Sep 2, 2026 | 31 results Coverage 88% | $0.15 | $1.25 | $4.25 | 57.9 |
| 23 | Claude Sonnet 5 Anthropic | Jun 30, 2026 | 33 results Coverage 94% | $0.20 | $2 | $10 | 56.9 |
| 24 | Kimi K3Open weights Moonshot AI | Jul 16, 2026 | 33 results Coverage 94% | $0.31 | $0.76 | $13 | 56.2 |
| 25 | DeepSeek V4.1 FlashOpen weights DeepSeek | Sep 9, 2026 | 28 results Coverage 82% | $0.006 | $0.30 | $1.20 | 56.1 |
| 26 | Gemini 3.1 Pro Preview Google | Feb 19, 2026 | 35 results Coverage 97% | $0.20 | $2 | $12 | 55.0 |
| 27 | Claude Opus 4.7 Anthropic | Apr 16, 2026 | 28 results Coverage 79% | $0.50 | $5 | $25 | 54.7 |
| 28 | Qwen3.8 FlashOpen weights Alibaba | Aug 26, 2026 | 10 results Coverage 30% | $0.016 | $0.15 | $0.47 | 54.5 |
| 29 | Gemini 3.5 Flash Google | May 19, 2026 | 33 results Coverage 91% | $0.15 | $1.50 | $9 | 54.5 |
| 30 | Grok 4.7 xAI | Sep 21, 2026 | 31 results Coverage 88% | $0.50 | $2 | $6 | 54.4 |
- 01GPT-6 AstraOpenAI · Sep 3, 202634 results · 97% coverage · $10 in / $50 out per 1M tokens91.5
- 02GPT-6.1 SolOpenAI · Sep 29, 202631 results · 88% coverage · $2 in / $10 out per 1M tokens86.6
- 03Claude Opus 5.5Anthropic · Sep 22, 202632 results · 91% coverage · $4 in / $20 out per 1M tokens85.6
- 04Claude Fable 5.1Anthropic · Sep 1, 202634 results · 97% coverage · $10 in / $50 out per 1M tokens80.8
- 05Claude Sonnet 5.5Anthropic · Sep 28, 202630 results · 85% coverage · $2 in / $10 out per 1M tokens79.5
- 06Claude Opus 5Anthropic · Jul 24, 202634 results · 97% coverage · $5 in / $25 out per 1M tokens75.6
- 07Claude Fable 5Anthropic · Jun 9, 202632 results · 88% coverage · $10 in / $50 out per 1M tokens74.9
- 08GPT-5.6 SolOpenAI · Jul 9, 202635 results · 97% coverage · $2 in / $10 out per 1M tokens74.0
- 09GPT-6 SolOpenAI · Sep 22, 202631 results · 88% coverage · $2 in / $10 out per 1M tokens72.9
- 10GPT-5.6 TerraOpenAI · Jul 9, 202635 results · 97% coverage · $2 in / $12 out per 1M tokens65.5
- 11Gemini 3.8 FlashGoogle · Sep 2, 202633 results · 94% coverage · $0.75 in / $3.75 out per 1M tokens65.0
- 12GPT-5.5OpenAI · Apr 23, 202630 results · 85% coverage · $5 in / $30 out per 1M tokens64.6
- 13Claude Haiku 5.5Anthropic · Oct 7, 202626 results · 76% coverage · $0.10 in / $0.50 out per 1M tokens62.2
- 14Gemini 3.7 FlashGoogle · Aug 13, 202632 results · 91% coverage · $0.75 in / $3.75 out per 1M tokens61.4
- 15GPT-5.4OpenAI · Mar 5, 202624 results · 67% coverage · $2.50 in / $15 out per 1M tokens60.9
- 16Muse Spark 1.2Meta · Aug 5, 202627 results · 79% coverage · $1.25 in / $4.25 out per 1M tokens60.7
- 17Grok 4.6xAI · Aug 12, 202634 results · 97% coverage · $2 in / $6 out per 1M tokens60.2
- 18Claude Opus 4.8Anthropic · May 28, 202632 results · 88% coverage · $5 in / $25 out per 1M tokens59.5
- 19Muse Spark 1.1Meta · Jul 9, 202621 results · 64% coverage · $1.25 in / $4.25 out per 1M tokens59.3
- 20DeepSeek V4 Pro 0813Open weightsDeepSeek · Aug 13, 202632 results · 91% coverage · $1.32 in / $3.96 out per 1M tokens58.7
- 21Qwen3.8 MaxAlibaba · Aug 2, 202634 results · 97% coverage · — in / — out per 1M tokens58.4
- 22Muse Spark 1.3Meta · Sep 2, 202631 results · 88% coverage · $1.25 in / $4.25 out per 1M tokens57.9
- 23Claude Sonnet 5Anthropic · Jun 30, 202633 results · 94% coverage · $2 in / $10 out per 1M tokens56.9
- 24Kimi K3Open weightsMoonshot AI · Jul 16, 202633 results · 94% coverage · $0.76 in / $13 out per 1M tokens56.2
- 25DeepSeek V4.1 FlashOpen weightsDeepSeek · Sep 9, 202628 results · 82% coverage · $0.30 in / $1.20 out per 1M tokens56.1
- 26Gemini 3.1 Pro PreviewGoogle · Feb 19, 202635 results · 97% coverage · $2 in / $12 out per 1M tokens55.0
- 27Claude Opus 4.7Anthropic · Apr 16, 202628 results · 79% coverage · $5 in / $25 out per 1M tokens54.7
- 28Qwen3.8 FlashOpen weightsAlibaba · Aug 26, 202610 results · 30% coverage · $0.15 in / $0.47 out per 1M tokens54.5
- 29Gemini 3.5 FlashGoogle · May 19, 202633 results · 91% coverage · $1.50 in / $9 out per 1M tokens54.5
- 30Grok 4.7xAI · Sep 21, 202631 results · 88% coverage · $2 in / $6 out per 1M tokens54.4
Score: reference-group average 50, about 15 points per standard deviation; not a percentage. Ranks have no ties and stay as on the full board when filtered. At most 30 models are shown.
Not ranked yet
20 modelsThese models have a score, but their evidence is too narrow to rank: one bad result could move them a long way. They get a rank once the missing evaluations come in.
- Gemini 4 ArgonGoogleFewer than 2 kinds of knowledge evaluation · 21 results76.8
- MiMo-V2.6-ProXiaomiFewer than 2 kinds of knowledge evaluation · 21 results64.1
- Qwen3.8 Max (0902)AlibabaFewer than 3 kinds of coding evaluation · 9 results61.6
- DeepSeek V4 Flash Vision ExpDeepSeekResults from only one evaluator so far · 5 results61.2
- Ember-1FireworksResults from only one evaluator so far · 12 results60.2
- Step 5 PreviewStepFunFewer than 2 kinds of knowledge evaluation · 16 results60.0
- Hy4 previewTencentResults from only one evaluator so far · 13 results59.6
- Qwen3.8 2.4T A95BAlibabaResults from only one evaluator so far · 5 results58.7
- MiMo-V2.6-FlashXiaomiFewer than 2 kinds of knowledge evaluation · 20 results56.5
- Deepseek V4 Flash VisionDeepSeekResults from only one evaluator so far · 5 results52.5
- GPT-5.2-CodexOpenAIFewer than 3 kinds of coding evaluation · 11 results50.0
- Deepseek V4 Pro 0424DeepSeekResults from only one evaluator so far · 7 results48.8
- GPT-5.2OpenAIFewer than 3 kinds of coding evaluation · 19 results48.4
- Qwen3.6 Max PreviewAlibabaNo coding evaluations yet · 8 results48.0
- Muse SparkMetaFewer than 3 kinds of coding evaluation · 7 results47.8
- GPT-5.3-CodexOpenAIFewer than 3 kinds of coding evaluation · 6 results47.7
- Mimo V2 ProXiaomiNo coding evaluations yet · 5 results44.1
- Qwen3.5-27BAlibabaNo coding evaluations yet · 5 results43.6
- Hy3TencentFewer than 3 kinds of coding evaluation · 6 results42.3
- K2 Horizon 375b A23bInstitute of Foundation ModelsResults from only one evaluator so far · 5 results42.1
How to read this board
Every evaluation is converted to one ability scale. The reference group's average is 50 and each 15 points is about one standard deviation, so a score is not a percentage. Prices never affect the ranking, and an evaluation a model has not taken neither adds nor subtracts.
How the score worksAbout prices
Prices are list prices in US dollars per million tokens, from OpenRouter's public model list. Cached input is the price of input served from the prompt cache; cache writes and subscriptions are not included. A dash means no price is listed.
Scored evaluations by LiveBench, Epoch AI, LMArena, Vals AI, Mercor, BullshitBench, Artificial Analysis. Every evaluation and its evaluator is listed on the evaluations page; prices come from OpenRouter. The board is also available as JSON at /api/v1/leaderboard.