Rankings / NVIDIA
Nemotron 3 Ultra 550B A55b
released 2026-06-04
BenchAtlas Index
as of 2026-09-11
thinking
rank #196
7 families · 4 categories · medium
base configuration
rank #200
6 families · 3 categories · medium
base configuration
rank #364
6 families · 4 categories · medium
Benchmark evidence
22 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 38.7 % | independent | LiveBench | |
| τ²-Bench | thinking | 83.3 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | thinking | 14.2 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | 70.7 % | independent | LiveBench | |
| SciCode | thinking | 40.3 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 54.5 % | independent | LiveBench |
external indices
| AA Coding Index | thinking | 49.3 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | thinking | 23.4 points | independent | Artificial Analysis |
| Vals Index | 27.4 points | independent | Vals AI |
knowledge science
| GPQA Diamond implementation=artificial-analysis | thinking | 86.7 % | independent | Artificial Analysis |
| GPQA Diamond implementation=vals-ai | 86.1 % | independent | Vals AI | |
| Humanity's Last Exam implementation=artificial-analysis | thinking | 28.4 % | independent | Artificial Analysis |
| MMLU-Pro implementation=vals-ai | 85.8 % | independent | Vals AI |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 70.8 % | independent | LiveBench |
long context instruction
| AA-LCR | thinking | 79.3 % | independent | Artificial Analysis |
| IFBench | thinking | 81.4 % | independent | Artificial Analysis |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 73.5 % | independent | LiveBench |
professional
| CorpFin | 65.5 % | independent | Vals AI | |
| LegalBench | 82.1 % | independent | Vals AI | |
| TaxEval | 73.1 % | independent | Vals AI |
reasoning math
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 88.7 % | independent | LiveBench | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 74.7 % | independent | LiveBench |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + Nemotron 3 Ultra 550B A55b | Terminal-Bench 2.1 | 53.9 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Nemotron 3 Ultra 550B A55b | Terminal-Bench Hard | 36.4 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
