Llama 3.1 Nemotron 70B Instruct
Summary
Llama 3.1 Nemotron 70B Instruct is an open-weight language model from NVIDIA, released on 15 Apr 2025. It accepts text and generates text, with a 128K-token context window and up to 8K output tokens per response.
Nemotron model for efficient reasoning, coding, and specialized AI agents
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | NVIDIA |
|---|---|
| Release date | 15 Apr 2025 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [models.dev] |
| Context window | 128K (128,000 tokens) [models.dev] |
| Max output | 8K tokens [models.dev] |
| Input | Text [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | No [models.dev] |
| Tool calling | Yes [models.dev] |
Llama 3.1 Nemotron 70B Instruct API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| NVIDIA NIM nvidia/llama-3.1-nemotron-70b-instruct | Free | Free | – | 128K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of Llama 3.1 Nemotron 70B Instruct?
Llama 3.1 Nemotron 70B Instruct has a context window of 128,000 tokens (128K) and can generate up to 8,192 tokens in a single response, according to models.dev.
Is Llama 3.1 Nemotron 70B Instruct open source?
Llama 3.1 Nemotron 70B Instruct is an open-weight model. Check the license terms before commercial use.
When was Llama 3.1 Nemotron 70B Instruct released?
NVIDIA released Llama 3.1 Nemotron 70B Instruct on 15 Apr 2025, according to models.dev.
Change history
No changes detected since we started tracking this model. We re-check every source daily.