Nemotron 3 Ultra
NVIDIA·Released Jun 4, 2026Open weights
Updated Oct 9, 6:21 AM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 21.5- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · thinking effortNear the floor0.5%Reading at most 34.5
- LiveBench CodingLiveBench54.7%Reading 28.7
- SciCodeArtificial Analysis · thinking effort40.3%Reading 23.0
- Vibe Code BenchVals AI7.6%Reading 16.2
- Code MigrationVals AINear the floor4.9%Reading at most 11.8
- APEX-SWEMercor · high effort16.4%Reading 8.9
Research and reasoning
4 results · 32.4- LiveBench ReasoningLiveBench81.7%Reading 41.0
- Humanity's Last Exam (AA run)Artificial Analysis · thinking effort28.4%Reading 39.8
- Mystery Game PuzzlesEpoch AI20.0%Reading 37.8
- Chess PuzzlesEpoch AI12.0%Reading 12.3
Professional work
9 results · 23.6- τ²-Bench Telecom (AA run)Artificial Analysis · thinking effort83.3%Reading 33.9
- Finance AgentVals AI37.7%Reading 29.2
- MedCodeVals AI38.6%Reading 29.1
- Legal Research BenchVals AI15.4%Reading 26.4
- EMBVals AI31.7%Reading 24.7
- Harvey Legal Agent BenchmarkVals AINear the floor0.4%Reading at most 24.7
- APEX-AgentsMercor · high effort22.7%Reading 20.6
- τ-Bench Banking (AA run)Artificial Analysis · thinking effort14.2%Reading 19.9
- LiveBench Data AnalysisLiveBench54.5%Reading 3.6
Knowledge and accuracy
5 results · 48.5- LiveBench Instruction FollowingLiveBenchCapped73.4%Reading 77.2
- IFBench (AA run)Artificial Analysis · thinking effortCapped81.4%Reading 69.4
- AA-LCRArtificial Analysis · thinking effort79.3%Reading 48.0
- LiveBench LanguageLiveBench70.8%Reading 35.2
- BullshitBenchBullshitBench · xhigh effort30.9%Reading 30.5
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only86.0%
- SWE-bench VerifiedVals AI · CodingReference only69.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only86.7%
- APEX-AccountingMercor · Professional workWatching0.3%
- CorpFinVals AI · Professional workReference only65.5%
- LegalBenchVals AI · Professional workReference only82.1%
- TaxEvalVals AI · Professional workReference only73.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only85.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only86.1%
- MMLU-ProVals AI · Knowledge and accuracyReference only85.8%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only22.9
- LiveBench averageLiveBench · Composite indicesReference only67.4%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Tax Agent Bench · Vals AI
- SimpleQA Verified · Epoch AI
- Arena Text · LMArena