GPT-5.4 Nano
OpenAI·Released Mar 17, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
5 results · 36.0- SciCodeArtificial Analysis · xhigh effort47.2%Reading 42.0
- LiveBench CodingLiveBench · xhigh effort58.8%Reading 38.5
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effortNear the floor0.5%Reading at most 34.5
- Code MigrationVals AI · high effort14.5%Reading 34.1
- Vibe Code BenchVals AI · high effort26.1%Reading 34.1
Research and reasoning
6 results · 34.2- LiveBench ReasoningLiveBench · xhigh effort86.0%Reading 51.1
- Chess PuzzlesEpoch AI · high effort30.0%Reading 46.7
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort28.3%Reading 39.6
- FrontierMath Tiers 1–3Epoch AI · high effort44.9%Reading 38.3
- FrontierMath Tier 4Epoch AI · high effort12.2%Reading 35.1
- Mystery Game PuzzlesEpoch AI · high effortNear the floor5.0%Reading at most -0.8
Professional work
9 results · 25.6- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort27.4%Reading 45.4
- EMBVals AI · high effort44.8%Reading 40.0
- MedCodeVals AI · high effort41.0%Reading 37.5
- LiveBench Data AnalysisLiveBench · xhigh effort67.6%Reading 36.4
- Finance AgentVals AI · high effort38.2%Reading 30.3
- τ²-Bench Telecom (AA run)Artificial Analysis · xhigh effort76.0%Reading 25.0
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor0.0%Reading at most 24.8
- Tax Agent BenchVals AI · high effortNear the floor5.0%Reading at most 0.1
- Legal Research BenchVals AI · high effort6.3%Reading -1.1
Knowledge and accuracy
6 results · 26.6- IFBench (AA run)Artificial Analysis · xhigh effort75.9%Reading 54.8
- LiveBench Instruction FollowingLiveBench · xhigh effort67.2%Reading 51.2
- AA-LCRArtificial Analysis · xhigh effort76.7%Reading 43.8
- LiveBench LanguageLiveBench · xhigh effort62.5%Reading 18.1
- SimpleQA VerifiedEpoch AI · high effortCapped11.7%Reading -16.5
- BullshitBenchBullshitBench · high effortCapped10.9%Reading -52.5
Human preference
1 result · 27.3- Arena TextLMArena · high effort1,401.3Reading 26.1
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only84.0%
- SWE-bench VerifiedVals AI · CodingReference only69.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only87.8%
- CorpFinVals AI · Professional workReference only61.2%
- LegalBenchVals AI · Professional workReference only77.9%
- MedScribeVals AI · Professional workWatching77.1%
- TaxEvalVals AI · Professional workReference only67.4%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only78.5%
- GPQA DiamondVals AI · Knowledge and accuracyReference only77.5%
- MMLU-ProVals AI · Knowledge and accuracyReference only77.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only20.7
- LiveBench averageLiveBench · Composite indicesReference only69.6%
- Arena VisionLMArena · VisionReference only1,199.3
- MMMU ProVals AI · VisionReference only73.6%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- APEX-Agents · Mercor