NVIDIA

Llama 3.3 Nemotron Super 49B v1.5

DeprecatedOpen weights131K contextReasoning

Summary

Llama 3.3 Nemotron Super 49B v1.5 is an open-weight language model from NVIDIA, released on 25 Jul 2025. It accepts text and generates text, with a 131K-token context window and up to 131K output tokens per response. NVIDIA has deprecated this model.

Nemotron model for efficient reasoning, coding, and specialized AI agents

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperNVIDIA
Release date25 Jul 2025 [models.dev]
StatusDeprecated [models.dev]
WeightsOpen weights [models.dev]
Context window131K (131,072 tokens) [models.dev]
Max output131K tokens [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]
Structured outputYes [models.dev]

Llama 3.3 Nemotron Super 49B v1.5 API pricing by provider

USD per million tokens. Sorted by input price.

We do not currently track a paid API listing for this model.

Frequently asked questions

What is the context window of Llama 3.3 Nemotron Super 49B v1.5?

Llama 3.3 Nemotron Super 49B v1.5 has a context window of 131,072 tokens (131K) and can generate up to 131,072 tokens in a single response, according to models.dev.

Is Llama 3.3 Nemotron Super 49B v1.5 open source?

Llama 3.3 Nemotron Super 49B v1.5 is an open-weight model. Check the license terms before commercial use.

When was Llama 3.3 Nemotron Super 49B v1.5 released?

NVIDIA released Llama 3.3 Nemotron Super 49B v1.5 on 25 Jul 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources