Llama 3.3 Nemotron Super 49B v1.5
Summary
Llama 3.3 Nemotron Super 49B v1.5 is an open-weight language model from NVIDIA, released on 25 Jul 2025. It accepts text and generates text, with a 131K-token context window and up to 131K output tokens per response. NVIDIA has deprecated this model.
Nemotron model for efficient reasoning, coding, and specialized AI agents
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | NVIDIA |
|---|---|
| Release date | 25 Jul 2025 [models.dev] |
| Status | Deprecated [models.dev] |
| Weights | Open weights [models.dev] |
| Context window | 131K (131,072 tokens) [models.dev] |
| Max output | 131K tokens [models.dev] |
| Input | Text [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
| Structured output | Yes [models.dev] |
Llama 3.3 Nemotron Super 49B v1.5 API pricing by provider
USD per million tokens. Sorted by input price.
We do not currently track a paid API listing for this model.
Frequently asked questions
What is the context window of Llama 3.3 Nemotron Super 49B v1.5?
Llama 3.3 Nemotron Super 49B v1.5 has a context window of 131,072 tokens (131K) and can generate up to 131,072 tokens in a single response, according to models.dev.
Is Llama 3.3 Nemotron Super 49B v1.5 open source?
Llama 3.3 Nemotron Super 49B v1.5 is an open-weight model. Check the license terms before commercial use.
When was Llama 3.3 Nemotron Super 49B v1.5 released?
NVIDIA released Llama 3.3 Nemotron Super 49B v1.5 on 25 Jul 2025, according to models.dev.
Change history
No changes detected since we started tracking this model. We re-check every source daily.