GPT-6 Sol
OpenAI·Released Sep 22, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 67.9- CyberBench PatchVals AI · max effort89.3%Reading 82.9
- Terminal-Bench 4.0Vals AI · max effort44.4%Reading 74.5
- Code MigrationVals AI · max effort57.2%Reading 73.7
- Vibe Code BenchVals AI · max effort87.8%Reading 71.2
- ProgramBenchVals AI · max effortNear the floor2.0%Reading at most 65.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort26.3%Reading 64.8
- SciCodeArtificial Analysis · high effort54.9%Reading 63.0
- LiveBench CodingLiveBench · max effort67.3%Reading 60.1
- APEX-SWEMercor · max effort45.0%Reading 59.7
Research and reasoning
8 results · 72.9- FrontierMath Tier 4Epoch AI · max effort90.0%Reading 91.0
- FrontierMath Tiers 1–3Epoch AI · max effort89.8%Reading 81.0
- Mystery Game PuzzlesEpoch AI · max effort56.0%Reading 78.0
- Terminal-Bench ScienceVals AI · max effort30.0%Reading 76.4
- ProofBenchVals AI · max effort83.0%Reading 73.9
- LiveBench ReasoningLiveBench · max effort92.5%Reading 72.8
- MysteryMechanismVals AI · max effort30.2%Reading 65.2
- Humanity's Last Exam (AA run)Artificial Analysis · high effort44.1%Reading 62.1
Professional work
8 results · 54.5- LiveBench Data AnalysisLiveBench · max effort81.2%Reading 79.0
- EMBVals AI · max effort71.5%Reading 71.0
- APEX-AgentsMercor · max effort54.3%Reading 62.1
- MedCodeVals AI · max effort47.1%Reading 58.0
- Finance AgentVals AI · max effort49.0%Reading 51.9
- Legal Research BenchVals AI · max effort28.8%Reading 48.4
- Tax Agent BenchVals AI · max effort15.0%Reading 36.4
- Harvey Legal Agent BenchmarkVals AI · max effortNear the floor1.7%Reading at most 24.8
Knowledge and accuracy
5 results · 63.6- LiveBench LanguageLiveBench · max effort85.3%Reading 74.8
- SimpleQA VerifiedEpoch AI · max effort60.7%Reading 70.0
- BullshitBenchBullshitBench · max effort41.8%Reading 60.8
- LiveBench Instruction FollowingLiveBench · max effort68.6%Reading 56.6
- AA-LCRArtificial Analysis · high effort83.7%Reading 55.8
Human preference
1 result · 49.7- Arena TextLMArena · max effort1,457.2Reading 42.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only82.6%
- SRE BenchVals AI · CodingWatching30.5%
- BioMysteryBenchVals AI · Research and reasoningWatching74.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- EBR-benchEpoch AI · Professional workWatching53.3%
- MedScribeVals AI · Professional workWatching82.0%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only94.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only42.4
- LiveBench averageLiveBench · Composite indicesReference only79.2%
- Vals IndexVals AI · Composite indicesReference only57.5%
- Arena WebDevLMArena · Writing and designReference only1,687.5
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- Chess Puzzles · Epoch AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis