GPT-5.4 Mini
OpenAI·Released Mar 17, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 38.9- ProgramBenchVals AI · xhigh effortNear the floor0.0%Reading at most 65.7
- SciCodeArtificial Analysis · xhigh effort52.1%Reading 55.3
- Vibe Code BenchVals AI · xhigh effort48.0%Reading 45.9
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effortNear the floor2.0%Reading at most 34.5
- LiveBench CodingLiveBench · xhigh effort56.6%Reading 33.3
- Terminal-Bench 4.0Vals AI · xhigh effortNear the floor2.5%Reading at most 33.0
- Code MigrationVals AI · xhigh effort12.9%Reading 31.6
Research and reasoning
6 results · 28.4- FrontierMath Tiers 1–3Epoch AI · xhigh effort51.2%Reading 42.8
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort28.1%Reading 39.3
- FrontierMath Tier 4Epoch AI · xhigh effort9.8%Reading 31.8
- LiveBench ReasoningLiveBench · xhigh effort74.9%Reading 28.5
- Chess PuzzlesEpoch AI · high effort18.0%Reading 26.6
- Mystery Game PuzzlesEpoch AI · medium effort7.0%Reading 8.1
Professional work
8 results · 33.1- LiveBench Data AnalysisLiveBench · xhigh effort70.8%Reading 45.1
- Finance AgentVals AI · xhigh effort45.4%Reading 44.6
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort25.6%Reading 42.4
- EMBVals AI · xhigh effort45.4%Reading 40.7
- τ²-Bench Telecom (AA run)Artificial Analysis · xhigh effort83.3%Reading 33.9
- Harvey Legal Agent BenchmarkVals AI · xhigh effortNear the floor0.0%Reading at most 24.8
- Legal Research BenchVals AI · xhigh effort12.5%Reading 19.8
- Tax Agent BenchVals AI · xhigh effort8.8%Reading 18.3
Knowledge and accuracy
6 results · 31.6- IFBench (AA run)Artificial Analysis · xhigh effort73.3%Reading 48.5
- AA-LCRArtificial Analysis · xhigh effort77.0%Reading 44.3
- LiveBench LanguageLiveBench · xhigh effort71.0%Reading 35.4
- SimpleQA VerifiedEpoch AI · high effort29.4%Reading 23.8
- LiveBench Instruction FollowingLiveBench · xhigh effort59.8%Reading 23.3
- BullshitBenchBullshitBench · high effort25.4%Reading 13.1
Human preference
1 result · 37.6- Arena TextLMArena · high effort1,447.1Reading 39.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only81.5%
- SWE-bench VerifiedVals AI · CodingReference only73.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only87.2%
- CorpFinVals AI · Professional workReference only60.9%
- TaxEvalVals AI · Professional workReference only71.2%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only83.6%
- GPQA DiamondVals AI · Knowledge and accuracyReference only83.1%
- MMLU-ProVals AI · Knowledge and accuracyReference only84.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only24.1
- LiveBench averageLiveBench · Composite indicesReference only66.4%
- Vals IndexVals AI · Composite indicesReference only33.2%
- Arena VisionLMArena · VisionReference only1,250.8
- MMMU ProVals AI · VisionReference only79.2%
- Arena WebDevLMArena · Writing and designReference only1,397.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor
- MedCode · Vals AI