/ Models

Language models

NVIDIA: Nemotron 3 Ultra (batch)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Maker
Nvidia
Context window
512K
Max output
Input / 1M
$0.3
Output / 1M
$1.80
1M in + 1M out
$2.10

Everything else

Max output
Input / 1M$0.3
Output / 1M$1.80
Cache read / 1M$0.1
Cache write / 1M
1M in + 1M out$2.10
Takestext
Returnstext
Toolsyes
Reasoningoptional · high, medium
Open weightsyes
Released2026-06-04
Knowledge cutoff
Retires
Catalog idnvidia/nemotron-3-ultra-550b-a55b:batch

Using it

In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.

← all models