NVIDIA

Llama 3.1 Nemotron 70B Instruct

Generally availableOpen weights128K context

Summary

Llama 3.1 Nemotron 70B Instruct is an open-weight language model from NVIDIA, released on 15 Apr 2025. It accepts text and generates text, with a 128K-token context window and up to 8K output tokens per response.

Nemotron model for efficient reasoning, coding, and specialized AI agents

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperNVIDIA
Release date15 Apr 2025 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [models.dev]
Context window128K (128,000 tokens) [models.dev]
Max output8K tokens [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingYes [models.dev]

Llama 3.1 Nemotron 70B Instruct API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
NVIDIA NIM
nvidia/llama-3.1-nemotron-70b-instruct
FreeFree–128Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of Llama 3.1 Nemotron 70B Instruct?

Llama 3.1 Nemotron 70B Instruct has a context window of 128,000 tokens (128K) and can generate up to 8,192 tokens in a single response, according to models.dev.

Is Llama 3.1 Nemotron 70B Instruct open source?

Llama 3.1 Nemotron 70B Instruct is an open-weight model. Check the license terms before commercial use.

When was Llama 3.1 Nemotron 70B Instruct released?

NVIDIA released Llama 3.1 Nemotron 70B Instruct on 15 Apr 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources