NVIDIA

Llama 3.1 Nemotron Ultra 253B

Generally availableOpen weights128K contextReasoning

Summary

Llama 3.1 Nemotron Ultra 253B is an open-weight language model from NVIDIA, released on 7 Apr 2025. It accepts text and generates text, with a 128K-token context window and up to 16K output tokens per response.

Flagship Nemotron model for high-throughput reasoning and complex agents

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperNVIDIA
Release date7 Apr 2025 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [models.dev]
Context window128K (128,000 tokens) [models.dev]
Max output16K tokens [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]

Llama 3.1 Nemotron Ultra 253B API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
NVIDIA NIM
nvidia/llama-3.1-nemotron-ultra-253b-v1
FreeFree–128Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of Llama 3.1 Nemotron Ultra 253B?

Llama 3.1 Nemotron Ultra 253B has a context window of 128,000 tokens (128K) and can generate up to 16,384 tokens in a single response, according to models.dev.

Is Llama 3.1 Nemotron Ultra 253B open source?

Llama 3.1 Nemotron Ultra 253B is an open-weight model. Check the license terms before commercial use.

When was Llama 3.1 Nemotron Ultra 253B released?

NVIDIA released Llama 3.1 Nemotron Ultra 253B on 7 Apr 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources