Kimi K2.7 Code
Moonshot AI·Released Jun 12, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 43.4- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- APEX-SWEMercor · high effort40.3%Reading 52.8
- Code MigrationVals AI25.4%Reading 47.5
- Vibe Code BenchVals AI47.2%Reading 45.5
- SciCodeArtificial Analysis47.8%Reading 43.6
- LiveBench CodingLiveBench59.8%Reading 41.0
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor1.0%Reading at most 34.5
Research and reasoning
5 results · 40.4- Humanity's Last Exam (AA run)Artificial Analysis35.0%Reading 49.7
- FrontierMath Tiers 1–3Epoch AI54.0%Reading 44.8
- LiveBench ReasoningLiveBench81.2%Reading 40.1
- FrontierMath Tier 4Epoch AI12.2%Reading 35.1
- Chess PuzzlesEpoch AI21.0%Reading 32.3
Professional work
4 results · 35.5- τ²-Bench Telecom (AA run)Artificial Analysis90.1%Reading 45.5
- APEX-AgentsMercor · high effort37.6%Reading 41.9
- τ-Bench Banking (AA run)Artificial Analysis20.2%Reading 33.0
- LiveBench Data AnalysisLiveBench62.7%Reading 23.5
Knowledge and accuracy
5 results · 35.1- LiveBench LanguageLiveBench77.9%Reading 52.1
- AA-LCRArtificial Analysis79.3%Reading 48.0
- SimpleQA VerifiedEpoch AI36.5%Reading 35.1
- IFBench (AA run)Artificial Analysis63.1%Reading 27.3
- LiveBench Instruction FollowingLiveBench56.3%Reading 10.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only82.0%
- SkillsBenchVals AI · CodingWatching50.0%
- SWE-bench VerifiedVals AI · CodingReference only78.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only95.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only87.9%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.8
- LiveBench averageLiveBench · Composite indicesReference only68.4%
- Arena WebDevLMArena · Writing and designReference only1,473.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- Harvey Legal Agent Benchmark · Vals AI
- MedCode · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- BullshitBench · BullshitBench
- Arena Text · LMArena