# Llama 3.3 Nemotron Super 49B v1.5

> Llama 3.3 Nemotron Super 49B v1.5 is an open-weight language model from NVIDIA, released on 25 Jul 2025. It accepts text and generates text, with a 131K-token context window and up to 131K output tokens per response. NVIDIA has deprecated this model.

Source page: https://www.aimodel.directory/models/llama-3-3-nemotron-super-49b-v1-5
Last verified: 4 Oct 2026

## Specifications

- **Developer:** NVIDIA
- **Release date:** 25 Jul 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Status:** Deprecated (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Weights:** Open weights (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Context window:** 131K (131,072 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Max output:** 131K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Input:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)
- **Structured output:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5.toml)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|

## FAQ

### What is the context window of Llama 3.3 Nemotron Super 49B v1.5?

Llama 3.3 Nemotron Super 49B v1.5 has a context window of 131,072 tokens (131K) and can generate up to 131,072 tokens in a single response, according to models.dev.

### Is Llama 3.3 Nemotron Super 49B v1.5 open source?

Llama 3.3 Nemotron Super 49B v1.5 is an open-weight model. Check the license terms before commercial use.

### When was Llama 3.3 Nemotron Super 49B v1.5 released?

NVIDIA released Llama 3.3 Nemotron Super 49B v1.5 on 25 Jul 2025, according to models.dev.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Llama 3.3 Nemotron Super 49B v1.5", https://www.aimodel.directory/models/llama-3-3-nemotron-super-49b-v1-5