/ Models

Language models

NVIDIA: Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Maker
Nvidia
Context window
512K
Max output
Input / 1M
$0.6
Output / 1M
$3.60
1M in + 1M out
$4.20

Everything else

Max output
Input / 1M$0.6
Output / 1M$3.60
Cache read / 1M$0.2
Cache write / 1M
1M in + 1M out$4.20
Takestext
Returnstext
Toolsyes
Reasoningoptional · high, medium
Open weightsyes
Released2026-06-04
Knowledge cutoff
Retires
Catalog idnvidia/nemotron-3-ultra-550b-a55b

Using it

In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.

← all models