gpt-oss-120b
OpenAI·Released Aug 5, 2025Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
3 results · -1.0- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effortNear the floor0.0%Reading at most 34.5
- SciCodeArtificial Analysis · high effort34.0%Reading 4.6
- APEX-SWEMercor · high effortCappedNear the floor3.3%Reading at most -37.8
Research and reasoning
3 results · 16.2- Chess PuzzlesEpoch AI · high effort20.0%Reading 30.5
- Humanity's Last Exam (AA run)Artificial Analysis · high effort19.6%Reading 23.9
- Mystery Game PuzzlesEpoch AI · high effortNear the floor0.0%Reading at most -0.8
Professional work
3 results · -3.2- τ-Bench Banking (AA run)Artificial Analysis · high effort12.8%Reading 16.0
- τ²-Bench Telecom (AA run)Artificial Analysis · high effort65.8%Reading 15.2
- APEX-AgentsMercor · high effortCappedNear the floor4.4%Reading at most -30.5
Knowledge and accuracy
3 results · 8.3- IFBench (AA run)Artificial Analysis · high effort69.0%Reading 39.1
- AA-LCRArtificial Analysis · high effort52.0%Reading 13.7
- BullshitBenchBullshitBench · high effortCappedNear the floor3.6%Reading at most -106.6
Human preference
1 result · 9.6- Arena TextLMArena1,351.6Reading 11.3
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only83.2%
- LiveCodeBench (AA run)Artificial Analysis · CodingReference only87.8%
- SWE-bench VerifiedVals AI · CodingReference only33.6%
- AIME (AA run)Artificial Analysis · Research and reasoningReference only93.4%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only88.9%
- APEX-AccountingMercor · Professional workWatching0.0%
- CorpFinVals AI · Professional workReference only58.2%
- LegalBenchVals AI · Professional workReference only75.9%
- TaxEvalVals AI · Professional workReference only71.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only75.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only78.5%
- MMLU-ProVals AI · Knowledge and accuracyReference only79.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only11.6
- MGSMVals AI · MultilingualReference only92.0%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- LiveBench Coding · LiveBench
- Code Migration · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- Harvey Legal Agent Benchmark · Vals AI
- MedCode · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI
- LiveBench Data Analysis · LiveBench
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench