Nemotron Ultra
Summary
Nemotron Ultra is an open-weight language model from NVIDIA, released on 4 Jun 2026. It accepts text and generates text, with a 203K-token context window and up to 203K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.50 per million input tokens and $2.20 per million output tokens (OpenRouter), across 7 providers we track.
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | NVIDIA |
|---|---|
| Release date | 4 Jun 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [Hugging Face] |
| License | openmdw-1.1 [Hugging Face] |
| Parameters | 561B [Hugging Face] |
| Context window | 203K (202,800 tokens) [models.dev] |
| Max output | 203K tokens [models.dev] |
| Knowledge cutoff | Feb 2026 [models.dev] |
| Input | Text [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
| Structured output | Yes [models.dev] |
| Hugging Face | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 [Hugging Face] |
Nemotron Ultra API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| NVIDIA NIM nvidia/nemotron-3-ultra-550b-a55b | $0.50 | $2.50 | $0.15 | 1M | models.dev, 4 Oct 2026 |
| OpenRouter nvidia/nemotron-3-ultra-550b-a55b | $0.50 | $2.20 | $0.10 | 262K | OpenRouter, 4 Oct 2026 |
| Fireworks AI accounts/fireworks/models/nemotron-3-ultra-nvfp4 | $0.60 | $2.40 | $0.12 | 262K | models.dev, 4 Oct 2026 |
| Baseten nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B | $0.60 | $2.40 | $0.12 | 203K | models.dev, 4 Oct 2026 |
| Together AI nvidia/nemotron-3-ultra-550b-a55b | $0.60 | $3.60 | $0.20 | 512K | models.dev, 4 Oct 2026 |
| Vercel AI Gateway nvidia/nemotron-3-ultra-550b-a55b | $0.60 | $2.40 | $0.12 | 1M | models.dev, 4 Oct 2026 |
| Nebius AI Studio nvidia/Nemotron-3-Ultra-550b-a55b | $1 | $3 | $1 | 1M | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of Nemotron Ultra?
Nemotron Ultra has a context window of 202,800 tokens (203K) and can generate up to 202,800 tokens in a single response, according to models.dev.
How much does the Nemotron Ultra API cost?
Nemotron Ultra is listed from $0.50 per million input tokens and $2.20 per million output tokens (OpenRouter). Prices last verified 4 Oct 2026.
Is Nemotron Ultra open source?
Nemotron Ultra is an open-weight model: its weights are published on Hugging Face as nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 under the openmdw-1.1 license. Check the license terms before commercial use.
When was Nemotron Ultra released?
NVIDIA released Nemotron Ultra on 4 Jun 2026, according to models.dev.
Which providers offer Nemotron Ultra?
We track Nemotron Ultra on 7 providers: NVIDIA NIM, OpenRouter, Fireworks AI, Baseten, Together AI, Vercel AI Gateway, Nebius AI Studio.
Change history
No changes detected since we started tracking this model. We re-check every source daily.
Sources
- https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b
- https://github.com/sst/models.dev/blob/dev/providers/baseten/models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B.toml
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/nemotron-3-ultra-550b-a55b.toml
- https://github.com/sst/models.dev/blob/dev/providers/fireworks-ai/models/accounts/fireworks/models/nemotron-3-ultra-nvfp4.toml
- https://github.com/sst/models.dev/blob/dev/providers/togetherai/models/nvidia/nemotron-3-ultra-550b-a55b.toml
- https://github.com/sst/models.dev/blob/dev/providers/vercel/models/nvidia/nemotron-3-ultra-550b-a55b.toml
- https://github.com/sst/models.dev/blob/dev/providers/nebius/models/nvidia/Nemotron-3-Ultra-550b-a55b.toml