Skip to content

Leaderboard

Updated Oct 8, 20:31 ET

One ability score from results that independent evaluators ran themselves, the same way for every model. Weights: coding 40%, research and reasoning 20%, professional work 20%, knowledge and accuracy 15%, human preference 5%.

36 scored evaluations · 7 evaluators

Overall board 30 models

  1. 01Claude Opus 5.5Anthropic · Sep 22, 202632 results · 91% coverage · $4/$20 per 1M78.6
  2. 02Claude Fable 5.1Anthropic · Sep 1, 202634 results · 97% coverage · $10/$50 per 1M77.5
  3. 03GPT-6 AstraOpenAI · Sep 3, 202634 results · 97% coverage · $10/$50 per 1M74.1
  4. 04Claude Fable 5Anthropic · Jun 9, 202632 results · 88% coverage · $10/$50 per 1M73.6
  5. 05Claude Opus 5Anthropic · Jul 24, 202634 results · 97% coverage · $5/$25 per 1M72.9
  6. 06Claude Sonnet 5.5Anthropic · Sep 28, 202630 results · 85% coverage · $2/$10 per 1M71.4
  7. 07GPT-6.1 SolOpenAI · Sep 29, 202631 results · 88% coverage · $2/$10 per 1M70.7
  8. 08GPT-5.6 SolOpenAI · Jul 9, 202635 results · 97% coverage · $2/$10 per 1M65.5
  9. 09GPT-6 SolOpenAI · Sep 22, 202631 results · 88% coverage · $2/$10 per 1M64.7
  10. 10Gemini 3.8 FlashGoogle · Sep 2, 202633 results · 94% coverage · $0.75/$3.75 per 1M64.6
  11. 11Muse Spark 1.3Meta · Sep 2, 202631 results · 88% coverage · $1.25/$4.25 per 1M64.4
  12. 12Grok 4.6xAI · Aug 12, 202634 results · 97% coverage · $2/$6 per 1M62.7
  13. 13Kimi K3Open weightsMoonshot AI · Jul 16, 202633 results · 94% coverage · $0.76/$13 per 1M61.9
  14. 14Grok 4.7xAI · Sep 21, 202631 results · 88% coverage · $2/$6 per 1M61.7
  15. 15Gemini 3.7 FlashGoogle · Aug 13, 202632 results · 91% coverage · $0.75/$3.75 per 1M61.5
  16. 16GLM 5.3Open weightsZ.ai · Aug 14, 202634 results · 97% coverage · $0.039/$3.39 per 1M60.7
  17. 17Claude Opus 4.8Anthropic · May 28, 202632 results · 88% coverage · $5/$25 per 1M60.1
  18. 18DeepSeek V4.1 FlashOpen weightsDeepSeek · Sep 9, 202628 results · 82% coverage · $0.30/$1.20 per 1M60.1
  19. 19Muse Spark 1.2Meta · Aug 5, 202627 results · 79% coverage · $1.25/$4.25 per 1M60.1
  20. 20Claude Haiku 5.5Anthropic · Oct 7, 202626 results · 76% coverage · $0.10/$0.50 per 1M59.2
  21. 21GPT-5.5OpenAI · Apr 23, 202630 results · 85% coverage · $5/$30 per 1M59.0
  22. 22Qwen3.8 FlashOpen weightsAlibaba · Aug 26, 202610 results · 30% coverage · $0.15/$0.47 per 1M58.9
  23. 23DeepSeek V4 Pro 0813Open weightsDeepSeek · Aug 13, 202632 results · 91% coverage · $0.66/$1.98 per 1M57.4
  24. 24Claude Opus 4.7Anthropic · Apr 16, 202628 results · 79% coverage · $5/$25 per 1M57.4
  25. 25Grok 4.5xAI · Jul 8, 202630 results · 85% coverage · $2/$6 per 1M57.1
  26. 26Muse Spark 1.1Meta · Jul 9, 202621 results · 64% coverage · $1.25/$4.25 per 1M56.6
  27. 27GPT-5.6 TerraOpenAI · Jul 9, 202635 results · 97% coverage · $2/$12 per 1M56.3
  28. 28Claude Sonnet 5Anthropic · Jun 30, 202633 results · 94% coverage · $2/$10 per 1M56.0
  29. 29Qwen3.8 MaxAlibaba · Aug 2, 202634 results · 97% coverage · —/— per 1M56.0
  30. 30DeepSeek V4 Flash 0731Open weightsDeepSeek · Jul 31, 202629 results · 82% coverage · $0.018/$1.28 per 1M53.9

Score: reference-group average 50, about 15 points per standard deviation; not a percentage. Ranks have no ties and stay as on the full board when filtered. At most 30 models are shown.

Not ranked yet

20 models

These models have a score, but their evidence is too narrow to rank: one bad result could move them a long way. They get a rank once the missing evaluations come in.

How to read this board

Every evaluation is converted to one ability scale. The reference group's average is 50 and each 15 points is about one standard deviation, so a score is not a percentage. Prices never affect the ranking, and an evaluation a model has not taken neither adds nor subtracts.

How the score works

About prices

Prices are list prices in US dollars per million tokens, from OpenRouter's public model list. Cached input is the price of input served from the prompt cache; cache writes and subscriptions are not included. A dash means no price is listed.

Scored evaluations by LiveBench, Epoch AI, LMArena, Vals AI, Mercor, BullshitBench, Artificial Analysis. Every evaluation and its evaluator is listed on the evaluations page; prices come from OpenRouter. The board is also available as JSON at /api/v1/leaderboard.