GPT-6 Luna
OpenAI·Released Sep 22, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 55.8- CyberBench PatchVals AI · max effort87.5%Reading 75.2
- ProgramBenchVals AI · max effortNear the floor0.5%Reading at most 65.7
- Vibe Code BenchVals AI · max effort81.6%Reading 65.3
- Code MigrationVals AI · max effort42.6%Reading 62.4
- LiveBench CodingLiveBench · max effort65.1%Reading 54.2
- APEX-SWEMercor · max effort38.8%Reading 50.6
- SciCodeArtificial Analysis · high effort50.3%Reading 50.4
- Terminal-Bench 4.0Vals AI · max effort13.6%Reading 49.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effortNear the floor4.5%Reading at most 34.5
Research and reasoning
9 results · 47.8- FrontierMath Tiers 1–3Epoch AI · max effort78.9%Reading 65.6
- FrontierMath Tier 4Epoch AI · max effort56.1%Reading 64.9
- ProofBenchVals AI · max effort64.0%Reading 62.6
- MysteryMechanismVals AI · max effort19.4%Reading 50.9
- LiveBench ReasoningLiveBench · max effort85.4%Reading 49.6
- Terminal-Bench ScienceVals AI · max effortNear the floor4.3%Reading at most 49.0
- Chess PuzzlesEpoch AI · max effort31.0%Reading 48.1
- Humanity's Last Exam (AA run)Artificial Analysis · high effort32.9%Reading 46.7
- Mystery Game PuzzlesEpoch AI · max effortCapped7.0%Reading 8.1
Professional work
8 results · 49.3- EMBVals AI · max effort68.5%Reading 67.1
- Finance AgentVals AI · max effort49.9%Reading 53.5
- LiveBench Data AnalysisLiveBench · max effort73.4%Reading 52.6
- Legal Research BenchVals AI · max effort30.3%Reading 50.3
- APEX-AgentsMercor · max effort44.3%Reading 50.2
- MedCodeVals AI · max effort44.7%Reading 50.0
- Tax Agent BenchVals AI · max effort19.9%Reading 46.8
- Harvey Legal Agent BenchmarkVals AI · max effortNear the floor2.9%Reading at most 24.8
Knowledge and accuracy
5 results · 37.7- AA-LCRArtificial Analysis · high effort79.3%Reading 48.0
- SimpleQA VerifiedEpoch AI · max effort41.4%Reading 42.4
- LiveBench LanguageLiveBench · max effort73.8%Reading 42.0
- BullshitBenchBullshitBench · max effort34.5%Reading 41.0
- LiveBench Instruction FollowingLiveBench · max effortCapped55.9%Reading 9.5
Human preference
1 result · 41.7- Arena TextLMArena · max effort1,442.9Reading 38.4
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only55.6%
- SRE BenchVals AI · CodingWatching2.7%
- BioMysteryBenchVals AI · Research and reasoningWatching61.5%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.9%
- MedScribeVals AI · Professional workWatching83.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only32.9
- LiveBench averageLiveBench · Composite indicesReference only72.0%
- Vals IndexVals AI · Composite indicesReference only51.2%
- Arena WebDevLMArena · Writing and designReference only1,582.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis