NVIDIA

nemotron-mini-4b-instruct

DeprecatedOpen weights128K context

Summary

nemotron-mini-4b-instruct is an open-weight language model from NVIDIA, released on 21 Aug 2024. It accepts text and generates text, with a 128K-token context window and up to 8K output tokens per response. NVIDIA has deprecated this model.

Compact Nemotron model for efficient reasoning and deployable AI agents

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperNVIDIA
Release date21 Aug 2024 [models.dev]
StatusDeprecated [models.dev]
WeightsOpen weights [models.dev]
Context window128K (128,000 tokens) [models.dev]
Max output8K tokens [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingYes [models.dev]

nemotron-mini-4b-instruct API pricing by provider

USD per million tokens. Sorted by input price.

We do not currently track a paid API listing for this model.

Frequently asked questions

What is the context window of nemotron-mini-4b-instruct?

nemotron-mini-4b-instruct has a context window of 128,000 tokens (128K) and can generate up to 8,192 tokens in a single response, according to models.dev.

Is nemotron-mini-4b-instruct open source?

nemotron-mini-4b-instruct is an open-weight model. Check the license terms before commercial use.

When was nemotron-mini-4b-instruct released?

NVIDIA released nemotron-mini-4b-instruct on 21 Aug 2024, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources