Grok 4.7
xAI·Released Sep 21, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 64.8- APEX-SWEMercor · xhigh effort53.6%Reading 71.9
- SciCodeArtificial Analysis · high effort57.8%Reading 71.0
- Vibe Code BenchVals AI · xhigh effort86.2%Reading 69.4
- ProgramBenchVals AI · xhigh effortNear the floor0.5%Reading at most 65.7
- Code MigrationVals AI · xhigh effort44.8%Reading 64.1
- Terminal-Bench 4.0Vals AI · xhigh effort28.8%Reading 64.1
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort24.7%Reading 63.5
- CyberBench PatchVals AI · xhigh effort83.9%Reading 62.4
- LiveBench CodingLiveBench · xhigh effort65.6%Reading 55.5
Research and reasoning
9 results · 54.4- LiveBench ReasoningLiveBench · xhigh effort89.2%Reading 60.2
- Humanity's Last Exam (AA run)Artificial Analysis · high effort42.3%Reading 59.8
- MysteryMechanismVals AI · xhigh effort25.2%Reading 59.2
- Terminal-Bench ScienceVals AI · xhigh effort10.0%Reading 58.7
- Chess PuzzlesEpoch AI · xhigh effort38.0%Reading 57.5
- Mystery Game PuzzlesEpoch AI · xhigh effort29.0%Reading 49.9
- ProofBenchVals AI · xhigh effort26.0%Reading 44.3
- FrontierMath Tiers 1–3Epoch AI · xhigh effort53.0%Reading 44.1
- FrontierMath Tier 4Epoch AI · xhigh effort17.1%Reading 40.4
Professional work
8 results · 66.0- Harvey Legal Agent BenchmarkVals AI · xhigh effort12.5%Reading 75.9
- Legal Research BenchVals AI · xhigh effort47.1%Reading 69.9
- Tax Agent BenchVals AI · xhigh effort34.2%Reading 68.9
- MedCodeVals AI · xhigh effort49.6%Reading 66.3
- EMBVals AI · xhigh effort67.0%Reading 65.2
- LiveBench Data AnalysisLiveBench · xhigh effort76.9%Reading 63.7
- APEX-AgentsMercor · xhigh effort54.6%Reading 62.5
- Finance AgentVals AI · xhigh effort52.3%Reading 58.1
Knowledge and accuracy
4 results · 62.6- LiveBench Instruction FollowingLiveBench · xhigh effort75.3%Reading 85.6
- SimpleQA VerifiedEpoch AI · xhigh effort56.0%Reading 63.1
- LiveBench LanguageLiveBench · xhigh effort80.1%Reading 58.3
- AA-LCRArtificial Analysis · high effort77.0%Reading 44.3
Human preference
1 result · 45.9- Arena TextLMArena · xhigh effort1,442.5Reading 38.3
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only57.7%
- SRE BenchVals AI · CodingWatching3.8%
- BioMysteryBenchVals AI · Research and reasoningWatching69.3%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.1%
- APEX-AccountingMercor · Professional workWatching4.2%
- LegalBenchVals AI · Professional workReference only84.4%
- MedScribeVals AI · Professional workWatching89.4%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only92.7%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only46.3
- LiveBench averageLiveBench · Composite indicesReference only77.4%
- Vals IndexVals AI · Composite indicesReference only54.9%
- Arena WebDevLMArena · Writing and designReference only1,638.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis