Kimi K2.6
Moonshot AI·Released Apr 20, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
6 results · 44.7- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- SciCodeArtificial Analysis · thinking effort51.5%Reading 53.7
- Code MigrationVals AI27.8%Reading 49.8
- LiveBench CodingLiveBench · thinking effort62.7%Reading 48.2
- Vibe Code BenchVals AI37.9%Reading 40.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · thinking effortNear the floor0.5%Reading at most 34.5
Research and reasoning
6 results · 43.3- Humanity's Last Exam (AA run)Artificial Analysis · thinking effort37.5%Reading 53.2
- FrontierMath Tier 4Epoch AI25.6%Reading 47.3
- FrontierMath Tiers 1–3Epoch AI57.2%Reading 47.1
- LiveBench ReasoningLiveBench · thinking effort81.8%Reading 41.3
- Chess PuzzlesEpoch AI26.0%Reading 40.7
- Mystery Game PuzzlesEpoch AI18.0%Reading 34.6
Professional work
9 results · 37.0- τ²-Bench Telecom (AA run)Artificial Analysis · thinking effortNear the ceiling95.9%Reading at least 60.0
- EMBVals AI57.9%Reading 54.4
- Finance AgentVals AI44.9%Reading 43.7
- τ-Bench Banking (AA run)Artificial Analysis · thinking effort23.3%Reading 38.6
- MedCodeVals AI40.1%Reading 34.5
- LiveBench Data AnalysisLiveBench · thinking effort65.1%Reading 29.8
- Tax Agent BenchVals AI12.4%Reading 29.8
- Legal Research BenchVals AI15.9%Reading 27.4
- Harvey Legal Agent BenchmarkVals AINear the floor1.7%Reading at most 24.8
Knowledge and accuracy
5 results · 44.5- IFBench (AA run)Artificial Analysis · thinking effort76.0%Reading 54.9
- AA-LCRArtificial Analysis · thinking effort81.0%Reading 50.8
- LiveBench LanguageLiveBench · thinking effort75.1%Reading 45.1
- LiveBench Instruction FollowingLiveBench · thinking effort64.4%Reading 40.2
- SimpleQA VerifiedEpoch AI34.9%Reading 32.7
Human preference
1 result · 43.1- Arena TextLMArena1,461.0Reading 43.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only86.8%
- SWE-bench VerifiedEpoch AI · CodingReference only76.7%
- SWE-bench VerifiedVals AI · CodingReference only76.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only96.1%
- CorpFinVals AI · Professional workReference only66.7%
- EBR-benchEpoch AI · Professional workWatching2.4%
- LegalBenchVals AI · Professional workReference only84.7%
- MedScribeVals AI · Professional workWatching78.1%
- TaxEvalVals AI · Professional workReference only74.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only89.1%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only27.0
- LiveBench averageLiveBench · Composite indicesReference only70.5%
- Arena VisionLMArena · VisionReference only1,264.9
- MMMU ProVals AI · VisionReference only86.3%
- Arena WebDevLMArena · Writing and designReference only1,509.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor
- BullshitBench · BullshitBench